Alibaba's Qwen3.8-Flash-Next: The 125B Model That Activates Only 6B and Previews Qwen4
On August 26, Alibaba's Qwen team did something most labs still refuse to do: it published the architecture before the flagship exists. Qwen3.8-Flash-Next, an open-weight mixture-of-experts model with 125 billion total parameters, landed on Hugging Face and ModelScope the same day a production version went live on QwenCloud. The headline number is not the 125 billion. It is the six billion that...
0 Комментарии 0 Поделились 488 Просмотры 0 предпросмотр