AMD 2026-08-03
Product Launch Impact: Important Conf: 85%

AMD Launches Helios Rack-Scale with MI455X, Claims 34x Token Throughput Boost

Summary

AMD unveiled Helios rack-scale solution at Advancing AI 2026, featuring 72 MI455X GPUs and 18 EPYC Venice CPUs, connected by Pensando networking and powered by ROCm. MI455X delivers 34x token throughput over MI355X and up to 30% more tokens per dollar than competitors.

Key Takeaways

At Advancing AI 2026, AMD launched its next-generation AI infrastructure, led by AMD Helios rackscale solutions, now in production for gigawatt-scale deployment. Helios integrates 72 AMD Instinct MI455X GPUs and 18 6th Gen AMD EPYC Venice CPUs, connected by AMD Pensando networking and accelerated by AMD ROCm open software. AMD claims up to 30% more tokens per dollar than the leading competitive solution.

The MI455X GPU delivers 34x higher token throughput compared to the previous MI355X. Leading AI labs and cloud providers including OpenAI, Anthropic, Meta, Microsoft, Oracle, and others have chosen Helios. OpenAI is partnering with AMD to optimize the full AI stack using OpenAI Triton framework with ROCm, expecting to bring Helios online in Q4 2026. Anthropic has a strategic partnership to deploy up to 2 gigawatts of AMD Instinct MI455X GPUs.

Why It Matters

AMD's move is a strategic defense against NVIDIA's CUDA ecosystem, using ROCm and OpenAI Triton to counter proprietary lock-in. However, AMD downplays the power and thermal requirements of MI455X at gigawatt scale, where electricity and cooling costs could erode the claimed token-per-dollar advantage. The 34x throughput gain likely comes from specific sparse or low-precision workloads, not general performance. The Pensando network lacks the maturity and low latency of NVIDIA's NVLink, risking tail latency and congestion in large-scale distributed training. While AMD uses Triton to attract users, ROCm's ecosystem lags behind CUDA, creating hidden migration costs.

PRO Decision

【Vendors】Competitors like NVIDIA should accelerate NVLink and InfiniBand improvements, highlight ecosystem maturity, and provide comparative benchmarks for actual training performance. They should also release cost-effective H200/B200 solutions to counter AMD's token-per-dollar claims.

【Enterprises】CIOs should demand independent benchmarks verifying the 34x throughput gain under specific workloads and precision, and assess total cost of ownership at gigawatt scale including power, cooling, and networking. They should also test ROCm migration difficulty from CUDA to avoid lock-in.

【Investors】Look beyond press releases; focus on actual shipments and customer adoption. Partnerships with OpenAI and Anthropic may be non-binding. Long-term, AMD's catch-up in AI infra still faces NVIDIA's deep moat.

Source: Reuters
View Original →

Get 3-5 key AI infrastructure signals weekly →

💬 Comments (0)