Reports
AI-generated structured vendor updates
NVIDIA and SK hynix Co-Architect Next-Gen Memory for AI Factories, Locking HBM4 to Vera Rubin
NVIDIA and SK hynix announce a multi-year tech partnership to co-develop next-gen memory for Vera Rubin, RTX Spark, and Jetson Thor. Separately, SK Telecom deploys a gigawatt-scale AI cloud using the full DGX stack, targeting 2027. This elevates SK hynix from supplier to co-architect, strengthening NVIDIA's lock-in on HBM and the AI ecosystem.
NVIDIA RTX Spark and Nemotron-3 Ultra: AI Control Shifts from Cloud to Personal Edge
NVIDIA launched RTX Spark personal AI supercomputer (co-developed with MediaTek) and Nemotron-3 Ultra open-source model at GTC Taipei 2026. The N1X chip delivers 1 PFLOPS local AI compute, bringing LLM inference to PCs. This marks NVIDIA's pivot from cloud GPU vendor to edge AI infrastructure monopolist, redefining the PC as an AI-native device.
US Government Forces Anthropic to Shut Down Fable 5 and Mythos 5: Cross-Border AI Regulation Reshapes Industry
The US government ordered Anthropic to shut down its latest models Fable 5 and Mythos 5 over cross-border data security concerns. This event exposes the regulatory vulnerability of closed-source AI and highlights the strategic value of open-source models. Regulatory uncertainty will reshape enterprise AI selection criteria, making model portability a core evaluation dimension.
ReflectionAI Secures $6.3B SpaceX Compute Deal, Open-Source AI Breaks Hardware Lock-in
Open-source AI startup ReflectionAI signs a $6.3B deal with SpaceXAI to lease NVIDIA GB300 compute at Colossus 2 for training open-weight frontier models. This gives open-source labs parity with closed-source giants but creates deep dependency on NVIDIA's proprietary hardware.
Z.ai GLM-5.2 Open-Source: 744B MoE, 1M Context, MIT License as Geopolitical Shield
Z.ai releases GLM-5.2: 744B MoE with 40B activated parameters, 1M input and 131K output context, under MIT license. Released one day after Anthropic Fable 5's government takedown, it offers a downloadable, unbanable alternative with Anthropic API compatibility for zero-code migration, giving enterprises a sovereign AI option.
SGLang 0.5.13 Delivers 25x MoE Inference Speedup via Predictive Routing and Sparse KV Cache
SGLang 0.5.13 introduces two-stage MoE routing prediction and sparse KV cache, achieving a 25x inference speedup on NVIDIA GB300 NVL72. Benchmarks on A100 show 65% throughput gain, 40% latency reduction, and 62% lower routing overhead. This optimization directly attacks the core bottleneck of MoE inference, potentially reshaping AI inference economics.