Qwen-3-Next-80B-A3B Exposure: Extreme sparse MoE, long context inference throughput may increase by 10 times
Qwen-3-Next-80B-A3B will be released soon, using the A3B architecture with 80B total parameters but only 3B activation, achieving extreme sparsity and...
AI information • Admin •
39