Alibaba Launches 2.4T Parameter Qwen3.8-Max MoE Model with 0.2x Pricing
Summary
Key Takeaways
Alibaba released Qwen3.8-Max-Preview on July 19, 2026, via its Qoder platform. The model has 2.4 trillion total parameters, the first Qwen model exceeding 1 trillion, using MoE (Mixture-of-Experts) architecture for efficient inference. It supports multimodal (image, video, document) and 1M token context window. The company claims it is second only to Anthropic Fable 5 and promises open-source release soon.
Simultaneously, it launched Qoder/QoderWork and Token Plan with aggressive discounts: 1x credit consumption and 0.2x during night hours. Token Plan pricing (limited-time): Lite ¥39/month, Standard ¥139/month, Pro ¥499/month.
Undisclosed details include activated parameter count, license terms, full model card, and third-party benchmarks (e.g., Arena.ai). Analysts note that 2.4T total parameters ≠ 2.4T activated parameters due to MoE, so actual inference cost is much lower. Alibaba also invests in Moonshot, whose K3 2.8T competes internally with Qwen3.8-Max. Upcoming releases include Qwen3.8-Max official, DeepSeek V4, and GLM-5.5. This event marks China's AI transition from catching up to encircling, with price wars and open-source models challenging top proprietary models.
Why It Matters
On the surface, this is a leap in parameter scale and performance, but in essence, it is a resource battle between Alibaba's internal Qwen team and its invested Moonshot, while externally encircling Anthropic Fable 5. Through Qoder platform and Token Plan discounts, Alibaba aims to quickly capture user mindshare and lock API usage habits, but the promise of 'open-source soon' may be a delaying tactic.
From an engineering perspective, MoE architecture introduces expert load balancing issues, potentially causing tail latency spikes. The 1M token context window suffers from O(n²) attention complexity, leading to latency and memory pressure in long sequences. The undisclosed activated parameter count and lack of third-party benchmarks make the 'second only to Fable 5' claim unverifiable. Post-discount pricing could trap users into vendor lock-in if deeply integrated with Qoder API.
PRO Decision
[Vendors] Competitors like Anthropic, Moonshot AI, DeepSeek, and Google should exploit Alibaba's lack of transparency: Anthropic can highlight Fable 5's verified benchmarks and full model card; Moonshot can promote K3's larger parameter count and independent evaluations; DeepSeek can emphasize its fully open-source model and lower real costs; Google can leverage its TPU ecosystem and Gemini integration. They should conduct direct comparisons via third-party tests (e.g., Arena.ai, Artificial Analysis) to expose potential gaps in Alibaba's unverified claims.
[Enterprises] CIOs and architects must perform zero-trust technical audits: demand full model card, activated parameter count, license terms, and third-party benchmarks from Alibaba; independently evaluate MoE tail latency and throughput under your workloads; beware of post-discount price recovery and calculate long-term TCO; adopt multi-model strategies to avoid single API lock-in and ensure cross-platform portability.
[Investors] Look beyond the hype: internal competition between Qwen and Moonshot may dilute focus and resources; lack of open-source and benchmarks suggests unverified performance; price wars could compress margins. Monitor whether Alibaba delivers on its open-source promise—delay would indicate a weak moat and questionable long-term value.
Get 3-5 key AI infrastructure signals weekly →
💬 Comments (0)