Reports
AI-generated structured vendor updates
Microsoft Launches MAI Models, Slashes GPU Costs 89%, Reducing OpenAI Dependency
Microsoft unveiled MAI-Image-2.5-Pro and MAI-Voice-2-Flash on Azure Foundry, achieving 96.8% text rendering accuracy at 8K and reducing GPU costs by 84-89% vs GPT. Integrated across Bing, PowerPoint, and Dynamics 365, it marks a strategic shift from OpenAI dependency. Also, NVIDIA Jetson heads to the moon for edge AI.
NVIDIA Open-Sources Cosmos 3 Edge 4B World Model for Real-Time Robot Control at 15Hz on Jetson Thor
NVIDIA open-sources Cosmos 3 Edge, a 4B parameter world action model for edge robotics. It achieves 15Hz real-time inference with 32 actions per inference on the Jetson Thor module. This extends NVIDIA's physical AI stack from training to real-time deployment, enabling end-to-end robot control at the edge.
NVIDIA Expands Agent Toolkit with Omniverse Libraries for Physical AI Simulation
At SIGGRAPH 2026, NVIDIA announced an expansion to its Agent Toolkit, adding Omniverse libraries that enable AI agents to build and simulate 3D worlds. The company also open-sourced Cosmos 3 Edge, a 4B-parameter world action model, completing its physical AI ecosystem from training to edge deployment.
NVIDIA Jetson Thor T3000/T2000: Blackwell GPU Crashes Edge AI Cost Barrier
NVIDIA unveils Jetson Thor T3000 and T2000 modules. The T3000 packs a Blackwell GPU and 8-core Neoverse CPU, delivering 865 FP4 TFLOPS at half the power of the T5000. New Jetson Agent Skills automate memory optimization, aiming to scale deployment of humanoid robots and edge AI.
NVIDIA Blackwell Sweeps MLPerf: NVLink and NVFP4 Redefine AI Training Economics
NVIDIA Blackwell dominates MLPerf Training 6.0, submitting across all seven benchmarks including MoE workloads. GB300 NVL72 delivers up to 1.6x faster training than GB200, with fifth-gen NVLink unifying 72 GPUs as one giant GPU. NVFP4 low-precision training and massive scale (8,192 GPUs) set new industry standards.
NVIDIA & SK hynix Deepen Memory Co-Engineering: Custom HBM for Vera Rubin and Jetson Thor
NVIDIA and SK hynix have announced a multiyear partnership to co-develop next-generation custom memory for NVIDIA's AI factory ecosystem, including Vera Rubin supercomputers, Vera CPUs, RTX Spark PCs, and Jetson Thor robotic platforms. SK hynix will also use NVIDIA CUDA-X libraries and Omniverse to accelerate semiconductor design and build fab digital twins.
NVIDIA Optimizes Google's DiffusionGemma for 1,000 tok/s Parallel Text Generation
NVIDIA optimizes Google DeepMind's DiffusionGemma, a diffusion-based text model generating 256 tokens per step in parallel. On a single H100, it achieves 1,000 tok/s, with deployment via NIM and NeMo. This breaks the sequential token bottleneck, slashing serving costs and latency for real-time AI.
NVIDIA and Doosan: Full-Stack Physical AI Platform Restructures Industrial Automation
NVIDIA expands collaboration with Doosan Group to integrate its physical AI stack (Isaac Sim, Cosmos, Jetson Thor) into Doosan Robotics' Agentic Robot OS, explore AI factory power (SMR, hydrogen fuel cells), and MGX ecosystem PCB materials. This move transforms NVIDIA from a GPU vendor into the central platform for physical AI and AI factory infrastructure, deeply locking industrial automation partners.
NVIDIA Transaction Foundation Models Shift Financial AI Control to Unified GPU Stack
NVIDIA launches a developer example for transaction foundation models, partnering with Revolut, Mastercard, and others to replace siloed ML models with unified transformer-based systems. Leveraging Hopper GPUs, cuDF, and Nemotron, it shifts financial data processing from feature engineering to unified embeddings, effectively moving control to NVIDIA's hardware ecosystem.
NVIDIA Locks Taiwan Supply Chain with AI Factory Stack, Vera Rubin Production Tied to Proprietary Software
NVIDIA partners with TSMC, Foxconn, and others to embed its proprietary AI software (cuLitho, Omniverse, Isaac) into semiconductor manufacturing and server assembly, while ramping Vera Rubin NVL72 production. The move uses efficiency gains (e.g., 20-50% cycle time reduction) as bait to lock the supply chain into a full-stack ecosystem, increasing switching costs for partners.
Intel Core Ultra 3 SoC Replaces Discrete GPUs in Edge Robotics, Slashing TCO
Intel Core Ultra Series 3 SoC integrates CPU, GPU, and NPU to power edge robotics, replacing discrete GPUs. Partners like Sensory AI run multi-agent AI (vision, language, motion) locally, cutting TCO and eliminating cloud latency. This shifts the cost-performance curve for service robots.
NVIDIA Advances Physical AI Integration in Robotics
NVIDIA showcases physical AI breakthroughs for robotics, accelerating deployment via Isaac Sim simulation and Jetson Orin edge modules. Case study: Aigen leverages synthetic data training and open-world foundation models to enable solar-powered robots for precision weeding, reducing herbicide use by 90%.
NVIDIA and Google Optimize Gemma 4 for Enhanced Local AI Agent Infrastructure
NVIDIA announces collaboration with Google to deeply optimize the Gemma 4 series of open models for its RTX, DGX Spark, and Jetson platforms. This move aims to extend high-performance, multimodal AI inference from the cloud to edge devices and personal workstations, providing full-stack model support (2B to 31B) for local AI agents.
NVIDIA Optimizes Gemma 4 Models for Local Agentic AI Acceleration
NVIDIA collaborates with Google to optimize the Gemma 4 family of models for efficient performance across a range of NVIDIA hardware, from edge devices to high-performance GPUs. These models support various tasks including reasoning, coding, and agent capabilities, making them suitable for local agentic AI applications.
Google Launches Gemma 4 Open Models, Targeting Edge Inference and AI Agent Architecture
Google introduces the Gemma 4 open model family, with four sizes from 2B to 31B parameters, emphasizing breakthrough intelligence-per-parameter and native support for agentic workflows, multimodality, and long context. The small models are engineered for edge devices, aiming to bring frontier reasoning to mobile and IoT scenarios.
Google Launches Gemma 4 Open Model Family
Google introduces Gemma 4 open model family with four size variants, optimized for edge and mobile devices. The series supports multimodal processing, long context windows and 140+ languages under Apache 2.0 license.
NVIDIA Introduces Physical AI Data Factory Blueprint, Transforming Compute into Synthetic Data
At GTC, NVIDIA introduced the Physical AI Data Factory Blueprint, an open reference architecture designed to transform compute into large-scale, high-quality synthetic training data. Built on Cosmos world models and the OSMO operator, it addresses the bottleneck of scaling real-world data, aiming to serve as the data engine for next-gen autonomous systems and robots.
NVIDIA Unveils Physical AI Data Factory Blueprint and Frontier Models
NVIDIA launched three physical AI frontier models and an open Physical AI Data Factory reference architecture at GTC 2026, converting computation into synthetic training data via Cosmos world model and OSMO operators. The Omniverse DSX digital twin blueprint enables validation and real-time AI inference integration with Jetson modules.
NVIDIA IGX Thor: 8x Edge AI Compute with ConnectX-7 Network Lock-In
NVIDIA launches IGX Thor edge AI platform with Blackwell GPU, up to 5,581 FP4 TFLOPS, dual 200GbE RDMA via ConnectX-7, and ISO 26262 safety. Pin-compatible with Jetson Thor and 10-year lifecycle enable seamless migration, but create vendor lock-in through proprietary networking and GPU dependencies.
NVIDIA Jetson Advances Localized Deployment of Open-Source AI Models at Edge
NVIDIA's Jetson edge AI platform enables localized deployment of open-source generative AI models like Qwen3 4B and Mistral 3 on edge devices. The platform offers a complete hardware range from Jetson Orin Nano to Thor, integrating compute and memory in SoM for simplified design. Key performance shows Jetson Thor achieves 52 tokens/sec for Mistral 3 inference.