Reports
AI-generated structured vendor updates
AWS Sells Trainium 3 Externally, Challenging NVIDIA's AI Training Chip Dominance
AWS begins external sales of its Trainium 3 AI training chip, fabricated on TSMC 3nm process, delivering 2.52 PFLOPS per chip. Early customers include Anthropic and Uber. This move directly challenges NVIDIA's dominance and marks AWS's strategic shift from cloud provider to chip vendor.
AMD's Experimental Topological Ghost Protocol Boosts MI300X Inference 10x
AMD introduces experimental Topological Ghost Protocol (TGP) on MI300X GPUs, achieving 431 tokens/sec with 100% success in high-concurrency inference, 10x improvement over standard vLLM. TGP uses KV-cache recycling and segmented state management, still experimental but potentially redefining AI inference benchmarks.
Google Gemini 3.5 Pro Rebuilds from Scratch: 2M Token Context Window Reshapes AI Frontier
Google DeepMind targets July 17 for Gemini 3.5 Pro, a full architectural rewrite of its pretraining stack to overcome deficits in math reasoning, SVG generation, and image quality. Specs include a 2M token context window, Deep Think reasoning layer, and multi-step autonomous workflows, though unconfirmed by Google.
NVIDIA Rigel Core: Single-Threaded CPU as the New Control Plane for Agentic AI
NVIDIA unveils Rosa CPU architecture with custom Rigel core (Arm v9.2), targeting single-threaded performance for Agentic AI workloads, paired with Feynman GPU (1.6nm, 50 PFLOPS) in 2028. This shifts CPU design from core-count scaling to serial-latency optimization, directly challenging AMD EPYC and Intel Xeon dominance.
NVIDIA Vera CPU获Perplexity/OpenAI/Anthropic/Oracle采用 AI Agent性能验证1.5-1.9x加速
...
Cisco Locks AI Data Center Security Control Plane with Silicon One and Hypershield
Cisco launches next-gen security for AI data centers, deeply integrating Splunk SIEM with its Silicon One 51.2Tbps chip and Hypershield architecture to push security policies to the network edge. This move aims to shift the security control plane from standalone appliances to its proprietary ASIC and management platform, creating hardware lock-in.
CrowdStrike and Zscaler Integrate Identity Security for Real-Time Zero Trust Access Decisions
CrowdStrike and Zscaler integrate Falcon identity security with Zscaler Zero Trust Exchange, using AI to assess 2.5 trillion endpoint events per second for real-time risk-based access decisions, converging endpoint security and zero trust network access.
NVIDIA Denies Kyber NVL144 Delay, But 78-Layer PCB Bottleneck Exposes AI Hardware Physics Limit
NVIDIA officially denies reports of Kyber NVL144 rack delay to 2028, but SemiAnalysis revelations about a 78-layer ultra-high-density PCB midplane bottleneck and Rubin Ultra cancellation expose hard physical limits in signal integrity and manufacturing, opening a strategic window for AMD and Google.
NVIDIA Kyber NVL144 Delayed to 2028: Midplane PCB Manufacturing Becomes AI Scaling Bottleneck
SemiAnalysis reveals NVIDIA's Kyber NVL144 delayed beyond 12 months to 2028 due to 78-layer Orthogonal Backplane manufacturing challenges. The interim NVL72x2 solution is cancelled due to operational burdens, and the 4-die Rubin Ultra is also scrapped, leaving a product gap in NVIDIA's scaling roadmap.
Huawei Unveils Tao's Law V2: Kirin 2026 Boosts AI Inference 40% on Same Node
Huawei's He Tingbo releases Tao's Law V2, detailing Kirin 2026 metrics: 238 MTr/mm² transistor density (+55%), 41% power reduction at iso-performance, and 40% SRAM frequency increase. Without EUV lithography, co-optimization of architecture, circuit, and process delivers equivalent performance gains, proving system-level optimization as a viable alternative to Moore's Law scaling.
OpenAI Launches GPT-5.6 Series, Regulatory Compliance Becomes Prerequisite for Frontier Models
OpenAI releases GPT-5.6 series with Sol achieving 96.7% SOTA on Terminal-Bench 2.1 via Ultra mode with sub-agent parallelism. Terra matches GPT-5.5 at half price, Luna for low-cost high-concurrency. Initial access limited to 20 trusted partners, subject to US government safety review.
AMD Unveils Zen 6/7 CPU and MI400/500 GPU Roadmap, Targets NVIDIA Rubin with HBM4 and 2nm
AMD unveiled its Zen 6/7 CPU and MI400/500 GPU roadmap at its 2026 Financial Analyst Day, featuring TSMC 2nm process and HBM4 memory. The MI400 series boasts 432GB memory, 19.6TB/s bandwidth, and 40 PFLOPs FP4 performance, directly targeting NVIDIA's Vera Rubin architecture with an annual cadence to disrupt the AI hardware monopoly.
Google Cloud Launches Blackwell GPU Confidential VM & Open-Source Prompt Encryption SDK, Redefining AI Security
Google Cloud upgrades its confidential computing portfolio with Blackwell GPU-based confidential VMs (Confidential G4 VMs preview), open-source Prompt Encryption SDK, and enhanced Confidential Space featuring Intel Trust Authority and Hopper GPU support, addressing TEE vulnerability CVE-2026-33697 to bolster AI inference and cross-organization training security.
Critical Relay Attack Found in Attestation TLS Protocol: Both Intel TDX and AMD SEV-SNP Affected
A critical architecture flaw in the attestation TLS protocol, enabling relay attacks, has been discovered affecting both Intel TDX and AMD SEV-SNP platforms. With a CVSS score of 7.5, it surpasses recent high-profile confidential computing vulnerabilities. No official patch is currently available.
NVIDIA Vera Rubin AI Platform Slated for July 2026 Shipments, Iterative Compute Upgrade
NVIDIA confirms its next-gen AI compute platform, Vera Rubin, will start shipping in July 2026 to major cloud providers like Microsoft and Google. The platform uses an advanced process node to boost AI training and inference performance, representing an iterative upgrade over Hopper and Blackwell without a fundamental architectural shift.
Meta Admits AI Agent Stagnation, Plans to Sell Compute to Challenge Cloud Triopoly
Meta CEO Zuckerberg admits AI agent development is behind schedule, pushing ROI timeline to 3-6 months. Concurrently, Meta plans to sell AI compute and model access externally, directly challenging AWS, Azure, and GCP's cloud oligopoly, signaling a pivot from internal AI infrastructure to a commercial cloud provider.
NVIDIA AI Compute Partnership: Revenue Share and Credit Backstop to Lock Cloud Providers into DSX AI Factories
NVIDIA launches AI Compute Partnership with revenue sharing and credit backstop, shifting from hardware sales to recurring service revenue. Initial projects include 40K GB300 chips for Sharon AI and 170K GPUs for Firmus, totaling 200K+ high-end chips. NVIDIA is becoming the 'central bank' of AI compute, squeezing cloud brokers.
Check Point launches AI orchestration platform, acquires Deepchecks to dominate security control plane
Check Point unveils Agentic Network Security Orchestration Platform, converting static firewall rules to intent-based policies via a proprietary network knowledge graph. Acquires Deepchecks' LLM team for continuous evaluation and monitoring. Four modules: Intent-to-Policy, Zero Trust tightening, Autonomous Troubleshooting, Continuous Compliance.
Qualcomm Enters AI Inference with Dragonfly C1000 CPU and HBC Near-Memory Compute
Qualcomm unveils Dragonfly roadmap with Oryon-based C1000 CPU and AI300 inference accelerator featuring HBC near-memory compute. Meta and Microsoft are early adopters. The strategy targets AI inference TCO reduction and memory wall breakthrough, bypassing Nvidia's training dominance.
NVIDIA BlueField-3 DPU: Shifts AI Cloud I/O Control from CPU to Dedicated Silicon, Redefines Compute Delivery & Security
NVIDIA's BlueField-3 DPU uses hardware vDPA to offload virtualization data plane from host CPU to dedicated processor, delivering near-bare-metal performance with live migration flexibility. It also creates a trusted I/O path for confidential computing. However, this fundamentally locks cloud infrastructure into NVIDIA silicon, increasing vendor dependency.