Google Begins Gemini 4 Pre-training with 4M+ Context and Monthly Releases
Summary
Key Takeaways
On July 22, 2026, Alphabet CEO Sundar Pichai confirmed Google has started pre-training Gemini 4, calling it the 'most ambitious pre-training run ever.' The Gemini 4 series will adopt a near-monthly release cadence, supporting 4M+ context and native multimodality with agent tools. Meanwhile, Gemini 3.5 Pro is delayed due to coding performance not meeting DeepMind's internal threshold (DeepSWE 60%+), causing Alphabet to drop over 3% after hours.
Alphabet raised 2026 capex by $15B to $195-205B (+160% YoY), and Google Cloud Q2 grew 82% (above 64% expected), but Q2 free cash flow turned negative for the first time (-$5.9B). Pichai emphasized AI coding and agentic AI as improvement areas.
Google Cloud growth is driven by Gemini API calls, Workspace AI, and Vertex AI enterprise customers, highlighting the TPU+Frozen v2+Gemini 3 full-stack advantage. This move reduces Google's reliance on NVIDIA, strengthens its self-developed TPU path, and competes with NVIDIA Vera Rubin and AMD Helios. Google also holds $124.3B in unrealized private investments (mainly Anthropic shares).
Why It Matters
On the surface, Google announces a technical breakthrough, but it is actually defending against NVIDIA's AI compute monopoly and locking in enterprise customers through the TPU+Frozen v2 ecosystem, forcing them to abandon CUDA. Google downplays the incompatibility of TPU programming model with mainstream CUDA, leading to high migration costs and toolchain lock-in. The 4M+ context poses significant memory bandwidth and tail latency challenges during inference; Google has not disclosed its sparse attention or KV cache optimization specifics, potentially hiding throughput degradation. The monthly release cadence increases model version management complexity and cost, forcing customers to constantly upgrade. Additionally, Frozen v2 interconnect bandwidth may become a training bottleneck, with no comparison to NVLink or InfiniBand disclosed. The negative free cash flow indicates a bet-the-company investment that could lead to asset impairment if Gemini 4 underperforms.
PRO Decision
[Vendors] NVIDIA should accelerate CUDA ecosystem openness and promote Vera Rubin interconnect advantages; AMD should target Google Cloud edge inference with Helios token/$ and ROCm tools. OpenAI/Anthropic should maintain model neutrality and multi-platform deployment.
[Enterprises] CIOs should conduct zero-trust audits of TPU migration costs, demand cross-cloud portability and standard MLOps compatibility. Adopt multi-model strategy with NVIDIA/AMD alternatives. Require independent benchmarks for 4M+ context inference latency and cost.
[Investors] Look past the PR, focus on negative free cash flow and capex ROI. Gemini 4 success is uncertain; vertical integration risk is high. Compare with NVIDIA's stable revenue and AMD's growth potential.
Get 3-5 key AI infrastructure signals weekly →
💬 Comments (0)