Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72
2026-08-13 02:23Models🔥 30.9 heat score
1sources
1days unfolding
30.9heat score
4mentions
SummaryAI generated
On August 12, 2026, Alibaba announced through NVIDIA Developer Blog the public weights of its largest open-source model, Qwen3.8-2.4T-A95B (Qwen3.8-Max). This model has a total of 2.4T parameters, with 95B activated parameters per token. It uses a fine-grained MoE architecture, combining full attention and linear attention mechanisms, and supports context windows of up to one million tokens along with corresponding output lengths. Its purpose is to provide the open-source ecosystem with capabilities close to cutting-edge levels.