Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding
2026-08-27 01:07Models🔥 28.9 heat score
1sources
1days unfolding
28.9heat score
3mentions
SummaryAI generated
On August 26, 2026, Alibaba announced through the NVIDIA Developer Blog the release of the weight files for the Qwen3.8-Flash-Next model. This multimodal hybrid expert (MoE) model includes a main model with 125B parameters and additional 51B N-gram embeddings, with 6B parameters activated per token. The model natively supports a context window of 262,144 tokens and can be expanded to 1M tokens. Currently, developers can experiment with and evaluate this model on the NVIDIA GB300 NVL72 cluster, with the aim of previewing the upcoming Qwen4 architecture.
Alibaba has released the weights of the Qwen3.8-Flash-Next model for developers to preview and evaluate the upcoming Qwen4 architecture. This multimodal hybrid expert (MoE) model includes a 125B parameter main model and 51B additional N-gram embeddings, with 6B parameters activated per token. It natively supports a context window of 262,144 tokens and can be expanded to 1M tokens. Developers can experiment with this model on the NVIDIA GB300 NVL72 cluster.