NVIDIA Qwen3.8-Flash-Next Achieves Over 16,000 Tokens/Second Throughput on GB300 NVL72
NVIDIA has announced that Alibaba's latest preview model, Qwen3.8-Flash-Next, is now supported on the NVIDIA GB300 NVL72 platform. The model has a total parameter scale of 176 billion, with approximately 6 billion parameters activated per token. It natively supports a context of 262,000 tokens and can be extended to 1 million tokens via YaRN, primarily targeting long-context agent applications such as intelligent programming, document processing, and tool invocation. NVIDIA stated that Qwen3.8-Flash-Next employs a mixed architecture of Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA) to reduce computational and KV cache overhead in long-context scenarios. Testing shows that on the GB300 NVL72, the model achieves a single GPU throughput of over 16,000 tokens per second, with single-user throughput exceeding 200 tokens per second; it also supports inference frameworks such as SGLang, vLLM, and TensorRT-LLM.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Qualcomm Launches IMSDK 2.0 to Drive Edge AI Application Development

Apple Launches M6 Chip Mac mini Starting at $899

Thomson Reuters Develops Legal AI Model 'Thomson' with $40 Million Investment

NVIDIA Surpasses Earnings and Guidance Expectations, Yet Stock Falls in After-Hours Trading

Haiku R1 Beta 6 Arrives After Two Years, Reviving the Spirit of BeOS

Armbian 26.8 rewrites its installer and expands support for ARM and RISC-V boards

Linux 7.3 Boosts FUSE Performance with Buffer Groups and Zero-Copy

Linux 7.3 Adds Support for Non-Fatal PCI Errors and New Intel Chips

Chrome and Chromium Test Flatpak to Expand Their Reach on Linux

LibreOffice 26.8 Arrives with Improvements for Writer, Calc, and Microsoft Office Compatibility

Intel Prepares BFF Driver to Combat Stuck Bits in Aging Processors

Illegal Premier League Betting Could Reach USD $1.09 Billion in the UK

Unstoppable Domains drops ICANN plans, offers refunds

Digital Fatigue: The Report Showing How Milei Lost Prominence on Social Media

Ethereum Excites Me More Than Bitcoin: Interesting Situation on the Chart and ETF Fund Data

Hyperliquid group asks CFTC to allow energy perpetuals

Bitcoin Wallets Dormant for Over a Decade Move $40M in One Week

September Brings Increases in Transportation, Rent, Health Insurance, and More: All the Details

Chainlink unlocks DeFi lending for Coinbase tokenized stocks

Bitcoin’s security risk starts when one block gets far more fees than the next

China-Supported Botnet Seized: FBI Halts Espionage on Fed and NASA

Havenex, the regulated exchange for banks: Series A round nearly closed

Bitcoin: US GDP Revives Hopes for Fed Rate Cuts

400 Million Accounts: TRON Reaches New Adoption Record

Banks found a way to copy stablecoins without losing the money that funds their loans

Stablecoin compliance could decide institutional winners: Aquanow CEO

DefiLlama Introduces AAA-CCC Rating for DeFi Token Evaluation

China Accumulates Gold: What is Beijing Preparing?

Nvidia May Swing $280 Billion After Earnings: What's at Stake





