AMD 2026-07-26
Product Launch Impact: Major Conf: 85%

AMD Helios Enters Production: 12-Stack HBM4 Outmuscles NVIDIA, UALoE Opens AI Network

Summary

AMD's second-generation Helios rack-scale AI server enters full production, featuring 72 MI455X GPUs with Samsung's exclusive 12-stack HBM4 (31TB per rack). Compared to NVIDIA's 8-stack design, it offers 50% more memory and 30% lower token cost. Microsoft Azure commits to large-scale deployment, solidifying hyperscaler dual-vendor strategy.

Key Takeaways

In July 2026, AMD CEO Lisa Su confirmed at the Seoul Advancing AI conference that the second-generation Helios rack-scale AI server is in full production, shipping by end of Q3. Each Helios rack integrates 72 MI455X accelerators (CDNA 5, TSMC N2+N3P) and 18 EPYC Venice CPUs (96-core Zen 6), with Samsung's exclusive HBM4 36GB 12-stack (12-Hi) per GPU (432GB), totaling 31TB memory per rack. Memory bandwidth reaches 1.7PB/s, FP4 compute 2.9 EFLOPS, FP8 1.4 EFLOPS. Network uses Pensando Salina 400G and UALoE 800G open standard.
Compared to NVIDIA Vera Rubin NVL72, Helios delivers 15% more AI compute, 50% more HBM capacity, 50% more scale-out bandwidth, and 30% lower token cost. For Kimi K2 Thinking, throughput gains 10-15%. This reflects AMD's 'capacity-first' philosophy vs NVIDIA's 'bandwidth-first', crucial for long-context and Agentic AI.
Microsoft Azure plans large-scale Helios deployment alongside its Cobalt 200 ARM servers, cementing the hyperscaler dual-vendor strategy. All five major hyperscalers (OpenAI, Oracle, Anthropic, Meta, Microsoft) are onboard. Samsung's exclusive 12-stack HBM4 supply breaks SK Hynix's 58% HBM market share, reshaping the memory supply chain.

Why It Matters

AMD's 'capacity-first' design excels in long-context inference but trades off bandwidth for dense training tasks where NVIDIA's 8-stack still leads. The claimed 30% token cost advantage likely assumes optimal conditions, ignoring ROCm's software maturity gap and enterprise migration costs from CUDA. UALoE 800G, while open, may incur higher tail latency and congestion control (PFC/ECN) issues versus NVLink's proven low-latency fabric. Pensando Salina 400G, being AMD-owned, risks creating a new proprietary network lock-in despite the open standard label. Samsung's 12-stack HBM4 exclusivity may face yield and supply constraints, potentially delaying volume deployments. AMD's move encircles NVIDIA but also establishes new lock-in points via CDNA and ROCm.

PRO Decision

[Vendors] NVIDIA must accelerate 12-stack HBM4 for Rubin to counter AMD's capacity advantage, while doubling down on CUDA ecosystem lock-in and demonstrating superior training throughput via independent benchmarks. Aggressive pricing and enhanced SK Hynix collaboration are needed to defend hyperscaler wallet share. [Enterprises] CIOs should conduct zero-trust audits: demand ROCm compatibility matrices for major frameworks, run pilot deployments with representative workloads to measure real token cost savings, and ensure network interoperability with RoCEv2/InfiniBand as fallback. Evaluate Samsung HBM4 supply stability and AMD's long-term roadmap. [Investors] Look beyond the PR: AMD's Helios production validates its AI infrastructure play, but software maturity and supply chain risks remain. The dual-vendor trend benefits AMD but pressures margins. Monitor Samsung's HBM4 yield and SK Hynix's response. UALoE ecosystem growth is key, but AMD's Pensando ownership may undermine openness.

Source: 36氪
View Original →

Get 3-5 key AI infrastructure signals weekly →

💬 Comments (0)