Reports
AI-generated structured vendor updates
OpenAI Partners with Cerebras for 14x Faster GPT-5.6 Sol Inference
OpenAI previews Ultrafast mode for GPT-5.6 Sol, powered by Cerebras hardware, delivering up to 750 tok/s output, 14x faster than standard. This enables real-time AI for incident response, finance, and voice, signaling a shift to specialized inference accelerators.
OpenAI发布GPT-5.6更新 扩展模型访问与推理能力
...
Meta's Superintelligence Push: New Labs, Scale AI Investment, and AGI Ambitions
...
Apple Limits Bug Bounty Submissions After Flood of AI Slop
...
New ways to learn and teach with ChatGPT Work and Codex
...
Cloudflare Migrates cdnjs to Workers Platform, Handling 9 Billion Requests Daily
Cloudflare has fully migrated cdnjs to its developer platform, leveraging R2, KV, Workers, and Workflows. The new architecture handles 9 billion requests daily with a 98.6% cache hit rate, demonstrating the scalability of edge computing for critical CDN infrastructure.
OpenAI ChatGPT图像生成服务7月27日发生宕机
...
OpenAI Launches Presence Agent Platform, Shifts Control Plane to Lock Enterprise AI Deployment
OpenAI launches Presence, an enterprise agent deployment platform integrating model inference, permissions, policies, evaluation, and escalation tools, shifting the control plane from models to the platform. ChatGPT Health is now fully available to US users 18+, integrating Apple Health, accelerating consumer AI agent adoption.
OpenAI GPT-5.6: Three-Layer Routing Shifts Control, Multi-Agent Parallelism Locks Workflows
OpenAI launches GPT-5.6 with Soul/Terra/Luna three-layer model routing, enabling automatic model selection and tool orchestration. New ChatGPT Work, Ultra Mode multi-agent parallelism, and a four-layer security framework shift AI from Q&A to autonomous task execution, consolidating OpenAI's control over AI workflow orchestration.
Meta Launches Muse Spark 1.1 API at 25% Competitor Price, Ends Open-Source Era
Meta releases Muse Spark 1.1, a multimodal reasoning model with 1M token context window, and launches its first paid API at 25% of competitors' price. This ends the Llama open-source era, signaling a strategic shift to proprietary API monetization and aggressive market share capture.
Anthropic Extends Claude Cowork Unified Interface to Web and Mobile, Compliance Gap Looms
Anthropic launches Claude Cowork unified interface on Web and Mobile for Max users, merging chat and task execution with local file access and cross-device continuity. However, Cowork activities are not captured in audit logs or Compliance API, creating a significant governance gap vs. Microsoft Copilot.
EU Forces Google to Open Android AI Access and Share Search Data
The EU mandates Google to open 11 system-level Android functions to third-party AI assistants by 2027-2028, and share search click/query data from 2027, under the Digital Markets Act. Non-compliance risks fines up to $40 billion annually. This will reshape the AI assistant and search markets.
EU Forces Google to Open Android to Third-Party AI Assistants, Share Search Data from 2027
The European Commission mandates Google to grant third-party AI assistants (e.g., ChatGPT) system-level access on Android by Android 18 (2027), including wake word, Home button, context reading, and device AI compute. Additionally, Google must share search data with rivals from 2027, risking fines up to 10% of global revenue (~$40B).
Apple in Talks with PrismML to Compress Qwen 27B Model 15x for On-Device AI
Apple is negotiating with AI startup PrismML to deploy a compressed version of Alibaba's Qwen 27B parameter model on iPhone. PrismML's compression technology reduces memory usage by 15x, enabling 27B models to run locally with 10GB VRAM, shifting Apple's AI strategy from cloud-dependent to on-device inference.
MemGhost Attack: Persistent False Memory Injection in AI Agents via Email
Researchers unveil MemGhost, a stealth memory injection attack that plants persistent false memories into AI agents via a single email without user notification. It exploits the persistent memory feature, highlighting critical security gaps and driving demand for memory auditing.
Anthropic企业AI采用首超OpenAI 300亿年化收入运行率确认
...
OpenAI Slashes Inference Costs 50%, Runs ChatGPT on Hundreds of GPUs via System-Level Optimization
OpenAI reduces AI inference costs by over 50% through system-level optimizations: model quantization (FP16 to INT4/INT8), KV-Cache optimization, dynamic batching, and speculative decoding. Using only hundreds of NVIDIA GPUs to serve ChatGPT's unlogged-in traffic, inference gross margin jumps from 38% to 65%, nearing breakeven.
Making private MCP servers reachable without making them public | OpenAI Developers
...
Anthropic Accuses Alibaba of Massive Distillation Attack on Claude AI Model
Anthropic accused Alibaba-linked operators of conducting 29 million exchanges via thousands of fraudulent accounts to distill Claude's capabilities, including long-context reasoning and decision-making. This highlights the vulnerability of AI model IP under API access, prompting a redefinition of model security boundaries.
OpenAI and Broadcom Unveil Jalapeno Inference ASIC, Reshaping AI Hardware Landscape
OpenAI, in collaboration with Broadcom, has developed Jalapeno, a custom LLM inference accelerator. The chip uses a multi-chip module with HBM3E memory and achieved tape-out in just nine months. Designed for OpenAI's model stack, it aims to reduce inference costs and dependency on NVIDIA GPUs, with initial deployment planned for late 2026.