Filter

×
Active Filters Clear All
Keyword: 推理成本 ×
15 Total Reports
NVIDIA Other 2026-07-28

NVIDIA Vera Rubin NVL72 Enters Production, 10x Energy Efficiency Reshapes AI Infrastructure

NVIDIA has announced full production of the Vera Rubin NVL72 platform, with first shipments to CoreWeave, Microsoft, Amazon, and Oracle. Each rack integrates 72 Rubin GPUs and 36 Vera CPUs, delivering 10x improvement in energy efficiency and inference cost over Blackwell for agentic AI workloads, marking a new era of rack-scale AI infrastructure.

Meta Other 2026-07-22

Meta Develops Switchboard AI Model Router to Control Inference Costs and Ecosystem

Meta's internal incubator AAI Labs is developing Switchboard, an AI model router that analyzes task complexity and routes to the most suitable model to minimize inference costs. Initially applied to internal AI coding agents, it may later become a commercial service, positioning Meta as a control layer for AI inference.

NVIDIA Other 2026-07-20

NVIDIA Invests $2B in CoreWeave, Debuts 'Compute Central Bank' Model

NVIDIA invests $2 billion in CoreWeave and launches the AI Compute Partner Program, featuring credit enhancement, revenue sharing, and GPU buyback. This transforms NVIDIA from a hardware vendor into a 'compute central bank', tightening control over the AI cloud leasing ecosystem and squeezing intermediaries.

Other Other 2026-07-20

Moonshot AI Launches Kimi K3: 2.8T Parameter Open-Source MoE Model at $3/$15

Moonshot AI unveils Kimi K3, a 2.8T parameter open-source MoE model with 896 experts (16 active), native vision, and 1M context. API pricing at $3/$15 input/output per million tokens undercuts rivals. Open weights release July 27. GPU capacity exhausted within 2 days. Arena score 1679 tops Fable 5.

Other Other 2026-07-19

Alibaba Launches 2.4T Parameter Qwen3.8-Max MoE Model with 0.2x Pricing

Alibaba released Qwen3.8-Max-Preview, a 2.4 trillion parameter MoE multimodal model with 1M context window. It launched Qoder platform and Token Plan with aggressive discounts up to 0.2x, significantly reducing inference cost. The company claims it is second only to Anthropic Fable 5, marking China's AI entry into dual-track of parameter arms race and open-source competition.

Anthropic Other 2026-07-06

Anthropic Starts Custom AI Chip Development, Talks Samsung 2nm, Aims for Compute Independence

Anthropic has initiated its own AI chip development and is in talks with Samsung for 2nm foundry services. The move aims to reduce reliance on NVIDIA GPUs, optimize inference costs, and strengthen its technology moat ahead of a potential IPO. It joins OpenAI, Google, and others in the custom ASIC race, signaling a shift from software to hardware competition.

Anthropic Other 2026-07-05

Anthropic Launches Custom AI Chip: Vertical Integration to Control Inference Cost and Supply

Anthropic launched Claude Sonnet 5 and revealed a custom AI chip initiative, using Samsung foundry. This move aims to reduce dependency on NVIDIA, control long-term inference costs, and marks Anthropic's shift from a pure software company to a vertically integrated infrastructure firm.

OpenAI Other 2026-07-03

OpenAI Slashes Inference Costs 50%, Runs ChatGPT on Hundreds of GPUs via System-Level Optimization

OpenAI reduces AI inference costs by over 50% through system-level optimizations: model quantization (FP16 to INT4/INT8), KV-Cache optimization, dynamic batching, and speculative decoding. Using only hundreds of NVIDIA GPUs to serve ChatGPT's unlogged-in traffic, inference gross margin jumps from 38% to 65%, nearing breakeven.

OpenAI Other 2026-06-26

OpenAI and Broadcom Tape Out First Inference ASIC Jalapeño in 9 Months, Targeting NVIDIA Dominance

OpenAI and Broadcom unveil Jalapeño, their first custom inference ASIC, fabricated on TSMC 3nm and optimized for Transformer models. Targeting a 50% inference cost reduction, it taped out in 9 months and is slated for deployment in gigawatt-scale data centers by late 2026, marking OpenAI's strategic pivot to full-stack AI infrastructure and a direct challenge to NVIDIA's inference hegemony.

OpenAI Other 2026-06-23

OpenAI GPT-5.6 Aggressive Pricing and 1.5M Context Window Targets Agent Era

OpenAI reportedly launches GPT-5.6 with 1.5M token context window, aggressive pricing at one-third of Claude Fable 5, and improved agent reliability. This move capitalizes on Anthropic's forced downtime and addresses internal alignment issues.

Qualcomm Other 2026-06-14

Qualcomm AI200 on AWS: Inference Chip Ecosystem Shifts from Nvidia Singularity to Multi-Alliance

Qualcomm's AI200 inference chip (768GB memory) is slated for broad AWS deployment by 2026, aiming to reduce cloud AI inference costs. This marks Qualcomm's strategic pivot from mobile to cloud, leveraging AWS's custom silicon initiative to challenge Nvidia's inference monopoly and restructure the cloud inference chip ecosystem.

Huawei Product Launch 2026-06-05

Huawei Cloud Launches AICS: Control Plane Shift in the Token Industrialization Era

Huawei Cloud unveils four Agentic Infra products, led by the AICS cluster (100K cards/200 EFLOPS). It integrates NPU-direct CMS memory, CCE VolcanoNext unified scheduling, and AgentSphere security sandbox to create a unified control plane for LLM training and Agent inference, aiming to lock in the full-stack AI infrastructure.

Apple Partnership High Signal 2026-04-27

Apple-Google Multi-Year Partnership Confirmed: Gemini to Power New Siri

Apple and Google confirm multi-year partnership with Google Cloud as preferred provider. Google is building a custom 1.2 trillion parameter Gemini model for Apple, 8x Apple's current cloud model. Siri will gain Gemini capabilities in 2026 with iOS 27. Privacy architecture unchanged—Gemini runs on Apple-controlled servers with data protection guarantees. Device compatibility limits exclude hundreds of millions of older iPhone users.

NVIDIA Product Launch High Signal 2026-04-23

NVIDIA Deploys OpenAI Codex: 10,000+ Employees Using GPT-5.5

NVIDIA 10,000+ employees using OpenAI Codex with GPT-5.5 on GB200 NVL72 platform, 35x inference cost reduction.

OpenAI Other Medium Signal 2026-03-17

OpenAI Releases Compact Models GPT-5.4 mini/nano for Enterprise AI Inference

OpenAI launches GPT-5.4 mini and nano models optimized for coding, multimodal tasks, and high-throughput API workloads. The compact models improve inference speed and reduce deployment costs, reflecting OpenAI's strategy to enhance enterprise AI service competitiveness.