Filter

×
Active Filters Clear All
Keyword: Llama ×
35 Total Reports
2/2 Page
Cisco Other High Signal 2026-04-30

Cisco Publishes Model Provenance Constitution, Defining Weight-Level Derivation Standards

Cisco published the 'Model Provenance Constitution' to provide a normative definition for AI model supply chain safety. The standard strictly hinges on the verifiable derivation history of model weights, clearly delineating five types of provenance links (e.g., direct descent, distillation) and eight exclusions (e.g., independent reproduction), aiming to resolve industry inconsistencies in model provenance definitions.

Cisco Other High Signal 2026-04-30

Cisco Open Sources Model Provenance Kit, Targeting AI Supply Chain Security Governance

Cisco released the open-source Model Provenance Kit, which uses a tiered strategy to analyze model metadata, tokenizer structure, and weight-level signals to generate unique fingerprints and verify the lineage and integrity of AI models. This aims to address risks of tampering, forgery, and compliance in the AI model supply chain.

NVIDIA Other High Signal 2026-04-24

NVIDIA Internalizes GPT-5.5 Powered AI Agents at Scale, Defining New Enterprise AI Infrastructure Paradigm

NVIDIA announced that over 10,000 employees have scaled the use of GPT-5.5 via the Codex app, running on NVIDIA GB200 NVL72 infrastructure. This demonstrates the technical feasibility of 'transformative' productivity gains from frontier model inference in enterprise workflows. It also provides a reference architecture for deploying AI agents with auditable, isolated security via dedicated cloud VMs.

Meta Financial News Medium Signal 2026-04-19

Meta's 2026 Strategy: Labor-to-Compute Reallocation at Extreme Scale

Meta's strategic choice represents 'endgame thinking' in AI infrastructure arms race—not how to profit but how to survive. When capex reaches 50%+ of revenue, this is no longer a business decision but survival bet. The 'relative value' of labor costs has undergone fundamental revaluation in the AI era.

NVIDIA Other High Signal 2026-04-03

NVIDIA and Google Optimize Gemma 4 for Enhanced Local AI Agent Infrastructure

NVIDIA announces collaboration with Google to deeply optimize the Gemma 4 series of open models for its RTX, DGX Spark, and Jetson platforms. This move aims to extend high-performance, multimodal AI inference from the cloud to edge devices and personal workstations, providing full-stack model support (2B to 31B) for local AI agents.

NVIDIA Other Medium Signal 2026-04-03

NVIDIA Optimizes Gemma 4 Models for Local Agentic AI Acceleration

NVIDIA collaborates with Google to optimize the Gemma 4 family of models for efficient performance across a range of NVIDIA hardware, from edge devices to high-performance GPUs. These models support various tasks including reasoning, coding, and agent capabilities, making them suitable for local agentic AI applications.

Google Other High Signal 2026-04-03

Google Launches Gemma 4 Open Models, Targeting Edge Inference and AI Agent Architecture

Google introduces the Gemma 4 open model family, with four sizes from 2B to 31B parameters, emphasizing breakthrough intelligence-per-parameter and native support for agentic workflows, multimodality, and long context. The small models are engineered for edge devices, aiming to bring frontier reasoning to mobile and IoT scenarios.

Google Other Medium Signal 2026-04-03

Google Launches Gemma 4 Open Model Family

Google introduces Gemma 4 open model family with four size variants, optimized for edge and mobile devices. The series supports multimodal processing, long context windows and 140+ languages under Apache 2.0 license.

AMD Other High Signal 2026-04-02

AMD Announces Breakthrough MLPerf Inference 6.0 Results, Showcasing Multinode Scaling and Multimodal Capabilities

AMD's MLPerf Inference 6.0 submission, powered by Instinct MI355X GPUs, surpassed 1 million tokens per second for the first time on models like Llama 2 70B and GPT-OSS-120B. The results highlight efficient multinode scaling, rapid enablement of new workloads (e.g., text-to-video model Wan-2.2-t2v), and reproducible performance across a broad partner ecosystem.

NVIDIA Other High Signal 2026-03-17

NVIDIA Expands NIM Microservices and Digital Twin Platform to Strengthen Full-Stack AI Ecosystem

NVIDIA launched NIM microservices supporting 30+ models across text, vision, speech, and embodied AI, available via AI Enterprise and cloud providers. Simultaneously released Omniverse Cloud digital twin platform with robotics simulation and introduced BioNeMo foundation models for healthcare.

NVIDIA Other High Signal 2026-03-11

NVIDIA Jetson Advances Localized Deployment of Open-Source AI Models at Edge

NVIDIA's Jetson edge AI platform enables localized deployment of open-source generative AI models like Qwen3 4B and Mistral 3 on edge devices. The platform offers a complete hardware range from Jetson Orin Nano to Thor, integrating compute and memory in SoM for simplified design. Key performance shows Jetson Thor achieves 52 tokens/sec for Mistral 3 inference.

Cisco Other High Signal 2026-03-10

Cisco Elevates Prompt Injection Defense to Infrastructure Layer

Cisco compares prompt injection to SQL injection, advocating layered defense including network micro-segmentation and EDR-based endpoint protection to mitigate LLM security risks.

Trend Micro Other High Signal 2026-03-03

Trend Micro Report Highlights AI Supply Chain Risks and Model Attack Surfaces

Trend Micro's 'Fault Lines in the AI Ecosystem' report systematically analyzes security risks in the AI supply chain, including training data poisoning, third-party plugin vulnerabilities, and model theft attacks. It indicates that enterprise AI security boundaries have expanded from traditional IT infrastructure to the model layer and data pipelines.

Microsoft Other Medium Signal 2025-02-27

Microsoft Launches Phi-4 SLM Series to Enhance Edge AI and Multimodal Reasoning

Microsoft introduced the Phi-4 family of small language models (SLMs), featuring the 5.6B-parameter Phi-4-multimodal capable of processing speech, vision and text. The models are now available in Azure AI Foundry, HuggingFace and NVIDIA's API Catalog with optimized edge computing capabilities.

Google Other 1970-01-01

Google TurboQuant: 6x KV Cache Compression, AI Inference Memory Cost Inflection Point

Google releases TurboQuant, a two-stage KV cache compression algorithm (PolarQuant + QJL) achieving 6x memory reduction (3-bit quantization) and 8x attention speedup with no measurable accuracy loss. The announcement triggered a sell-off in memory stocks (Micron -3%, Western Digital -4.7%), signaling a potential structural shift in AI inference memory demand.