Reports
AI-generated structured vendor updates
Meta Expands Hyperion to 5GW with $50B Investment, Pioneering Local-First AI Infrastructure
Meta expands its Louisiana Hyperion data center to 5GW capacity, raising total investment from $10B to $50B. Partnering with Entergy to build 10 power plants and 240 miles of transmission lines, and utilizing JV and financing structures, Meta pioneers a local-first model that reshapes the collaboration between AI infrastructure, energy, and capital.
Meta Iris Chip to Mass Produce in September: 6-Month Cadence Threatens NVIDIA GPU Hegemony
Reuters confirms Meta's Iris AI chip mass production in September, targeting 2.5GW by end-2026 and 14GW by 2027. Meta's 6-month MTIA generation cadence directly challenges NVIDIA's annual GPU cycle, signaling a hyperscaler shift from GPU dependency to custom ASIC sovereignty.
TSMC CoWoS Capacity to Reach 200k Wafers by 2027, Diversifying from GPU to CPU and ASIC
TSMC targets 200k wpm CoWoS capacity by 2027, narrowing supply-demand gap from 20% to 10%. Customer base diversifies from NVIDIA GPU to include AI server CPUs (MediaTek, AMD) and ASICs (Broadcom). CoPoS panel-level packaging enters pilot production in 2027.
Microsoft Takes Over OpenAI's Arctic Data Center, Seizing AI Compute Control
Microsoft leases a data center in Norway's Arctic Circle from Nscale, deploying 30,000 NVIDIA Vera Rubin GPUs, filling the gap left by OpenAI's retreat. OpenAI slashes its 2030 infrastructure budget from $140B to $60B. Microsoft surpasses OpenAI in AI compute capacity and gains geographical redundancy.
Anthropic Locks 3.5GW TPU Compute with Broadcom, Signaling Shift to Custom AI ASICs
Broadcom's Q2 FY2026 filing reveals a 3.5GW TPU compute deal with Anthropic starting 2027. This marks a strategic shift from general-purpose GPUs to custom ASICs for AI workloads, with OpenAI and Meta making similar multi-GW commitments, signaling a fundamental change in AI infrastructure.
PrismML's 1-bit Compression: 27B Qwen Model Runs Fully on iPhone 17 Pro in 4GB
PrismML compressed a 27B-parameter dense LLM (Qwen 3.6) to 4GB, running fully on iPhone 17 Pro. Using native 1-bit quantization (weights as {-1, +1}), it achieves >92% compression, 8x faster inference, and 75-80% energy reduction. This challenges Apple's sparse architecture, potentially shifting edge AI from cloud-reliant to device-native.
AWS Sells Trainium 3 Externally, Challenging NVIDIA's AI Training Chip Dominance
AWS begins external sales of its Trainium 3 AI training chip, fabricated on TSMC 3nm process, delivering 2.52 PFLOPS per chip. Early customers include Anthropic and Uber. This move directly challenges NVIDIA's dominance and marks AWS's strategic shift from cloud provider to chip vendor.
Towards Feature Complete Triton Support in JAX-Triton â ROCm Blogs
...
OpenAI Reopens with GPT-oss Models: Apache 2.0 License Hides Cloud Offload Control
OpenAI launches GPT-oss-120b and GPT-oss-20b under Apache 2.0 license, capable of running on a single 80GB GPU. However, a built-in cloud offload mechanism routes complex queries to proprietary models, masking a strategic control point shift behind the open-source facade.
Google Gemini 3.5 Pro Rebuilds from Scratch: 2M Token Context Window Reshapes AI Frontier
Google DeepMind targets July 17 for Gemini 3.5 Pro, a full architectural rewrite of its pretraining stack to overcome deficits in math reasoning, SVG generation, and image quality. Specs include a 2M token context window, Deep Think reasoning layer, and multi-step autonomous workflows, though unconfirmed by Google.
GhostApproval Vuln Exposes Systemic AI Coding Tool Flaw: Symlink Bypass in Human Review
Wiz Research discloses GhostApproval vulnerability affecting six major AI coding tools (Claude Code, Codex, Cursor, Amazon Q, Antigravity). Attackers use symlinks to bypass human review, achieving persistent remote access. The flaw reveals fundamental UI-level security gaps in Human-in-the-Loop mechanisms as agent permissions expand, requiring a redesign of confirmation workflows.
AI Giants Bet $10B on Forward Deployed Engineers: Control Shifts from Models to Engineering
Microsoft, OpenAI, Anthropic, and AWS collectively announced nearly $10B investment in Forward Deployed Engineer (FDE) model. Model interchangeability is now assumed; scarce resource moves from model parameters to engineering capability of embedding AI into business processes. This signals a fundamental paradigm shift in enterprise AI deployment.
NVIDIA Denies Kyber NVL144 Delay, But 78-Layer PCB Bottleneck Exposes AI Hardware Physics Limit
NVIDIA officially denies reports of Kyber NVL144 rack delay to 2028, but SemiAnalysis revelations about a 78-layer ultra-high-density PCB midplane bottleneck and Rubin Ultra cancellation expose hard physical limits in signal integrity and manufacturing, opening a strategic window for AMD and Google.
NVIDIA Kyber NVL144 Delayed to 2028: Midplane PCB Manufacturing Becomes AI Scaling Bottleneck
SemiAnalysis reveals NVIDIA's Kyber NVL144 delayed beyond 12 months to 2028 due to 78-layer Orthogonal Backplane manufacturing challenges. The interim NVL72x2 solution is cancelled due to operational burdens, and the 4-die Rubin Ultra is also scrapped, leaving a product gap in NVIDIA's scaling roadmap.
Anthropic Starts Custom AI Chip Development, Talks Samsung 2nm, Aims for Compute Independence
Anthropic has initiated its own AI chip development and is in talks with Samsung for 2nm foundry services. The move aims to reduce reliance on NVIDIA GPUs, optimize inference costs, and strengthen its technology moat ahead of a potential IPO. It joins OpenAI, Google, and others in the custom ASIC race, signaling a shift from software to hardware competition.
Google Cloud Launches Blackwell GPU Confidential VM & Open-Source Prompt Encryption SDK, Redefining AI Security
Google Cloud upgrades its confidential computing portfolio with Blackwell GPU-based confidential VMs (Confidential G4 VMs preview), open-source Prompt Encryption SDK, and enhanced Confidential Space featuring Intel Trust Authority and Hopper GPU support, addressing TEE vulnerability CVE-2026-33697 to bolster AI inference and cross-organization training security.
Anthropic Launches Custom AI Chip: Vertical Integration to Control Inference Cost and Supply
Anthropic launched Claude Sonnet 5 and revealed a custom AI chip initiative, using Samsung foundry. This move aims to reduce dependency on NVIDIA, control long-term inference costs, and marks Anthropic's shift from a pure software company to a vertically integrated infrastructure firm.
Cloudflare Default Blocks AI Crawlers: Infrastructure Layer Becomes Data Gatekeeper
Cloudflare announces default blocking of hybrid AI crawlers (e.g., Googlebot) for all sites starting Sept 15, allowing only pure search index crawlers unless manually overridden. This shifts AI data access control from websites/search engines to the CDN infrastructure layer, paired with a 'Pay Per Use' model to redefine content value exchange.
Announcing the Monetization Gateway: charge for any resource behind Cloudflare via x402
...
xAI Grok 4.5 Beta: 1.5T Param V9 Base, Cursor Integration Locks Tesla/SpaceX Ecosystem
xAI launches Grok 4.5 with a 1.5T parameter V9 base, integrating Cursor data for internal Beta at SpaceX/Tesla. Performance claims approach Claude Opus, but market share drops to 3.4% and Colossus compute utilization is 11%. This vertical integration aims to create a closed AI supply chain but risks ecosystem lock-in and resource misallocation.