News

Alibaba opens Qwen3.8-2.4T-A95B model weight: 2.4T MoE, activation 95B, native 256K context

2 min read
Source: ithome.com
IT House reported on August 13 that the Alibaba Cloud "ModelScope Community" public account announced late yesterday (12th) that the Alibaba Qwen team officially opened the Qwen3.8-2.4T-A95B model weight. This is the first time that the Qwen-Max level model has open source weights: the total parameters of the model are 2.4T, each Token activates 95B parameters, natively supports 262,144 Token contexts, and can be expanded to 1,010,000 Tokens. Qwen3.8 continues the hybrid architecture of Qwen3.5, focusing on improving the end-to-end completion capabilities of programming, office, scientific research, and long-term Agent tasks. The official cloud version Qwen3.8-Max is based on this open weight model and adds more production environment capabilities. The official evaluation covers programming Agent, general Agent, professional work and long context; compared with Opus 4.8, Fable 5, GPT 5.6 Sol and other models, the results are high and low, and it has achieved higher results among the listed models in Benchmarks such as PaperBench and IFBench. IT Home attaches the relevant links as follows: Model: https://modelscope.cn/models/Qwen/Qwen3.8-2.4T-A95B blog: https://Qwen.ai/blog?id=qwen3.8 Qwen Studio: https://chat.Qwen.ai Qwen3.8-2.4T-A95B It is a causal language model with a total parameter of 2.4T and each Token activates 95B. Each MoE layer of the model contains 512 experts. Each Token selects 10 routing experts and activates 1 shared expert at the same time. The model was also subjected to multi-step MTP