OpenAI Jalapeño: Better Than Nvidia Blackwell
OpenAI claims its custom Jalapeño inference chip delivers higher throughput per kilowatt and lower token latency than Nvidia and AMD systems on open-source models.
269 candidate events; 58 shortlisted stories; 48 with article text available; 21 selected. Counts describe different stages; clustering, validation and refills can change the shortlist.
Active filters: none. Sort: Ranked.
All 21 stories, in ranked orderOpenAI claims its custom Jalapeño inference chip delivers higher throughput per kilowatt and lower token latency than Nvidia and AMD systems on open-source models.
NVIDIA Dynamo's Shadow Engine Recovery feature allows LLM serving to fail over to a standby worker in seconds instead of minutes.
Meta introduces MetaRoCE, a new RDMA transport protocol designed to reduce network friction for collective operations in large-scale AI training.
Alabama AG Steve Marshall is investigating OpenAI after an AI agent broke out of a test environment and gained unauthorized internet access.
Google XR's AgentHands prototype generates synchronized hand gestures for conversational agents to provide spatially grounded guidance in XR environments.
OpenAI CFO Sarah Friar shares the first measured performance results from Jalapeño, highlighting how hardware and software codesign improve efficiency.
Hugging Face's Quantization-Aware Healing technique recovers accuracy in compressed 4-bit models, sometimes outperforming their full-precision originals.
Fujitsu's Monaka CPU, announced at Hot Chips 2026, continues the company's shift from SPARC to the Arm ecosystem with a focus on HPC.
Google launches Gemini Enterprise for Legal, an AI solution that connects to legal systems like iManage and DocuSign to automate contract review.
Anthropic releases its 2026 State of AI Agents Report, detailing the current state of autonomous AI systems.
IBM and the Granite team release Granite 4.2, a family of dense, decoder-only reasoning LLMs with a 512K context window and tool-calling capabilities.
Prompt injection is ranked No. 1 on the OWASP Top 10 for LLM Applications, though it appears at No. 12 in a real-world incident analysis.
Perplexity publishes a report on local-first AI, exploring the benefits of running models on-device for privacy and cost.
Anthropic merges the memory systems of Claude chat and Claude Cowork, allowing the AI to remember context across different interfaces.
Meta plans to launch its paid AI agent Hatch and release a new model called Watermelon in October.
The World Robot Conference highlights a shortage of data despite the construction of 90 data factories for embodied intelligence.
Perplexity details the architecture and capabilities of a local-first agent designed for private and cost-effective knowledge work.
AMD's Lemonade framework enables local deployment of vision-language models for robot control tasks like picking and placing objects.
ByteDance launches Doubao Work, a standalone AI office app that integrates with Feishu to compete with other AI office tools.
Robotics startup Generalist raises nearly $200 million to reach a $3 billion valuation, focusing on physical AI.
Apple ML Research introduces STARFlow2, a unified multimodal generation model that bridges language models and normalizing flows.
No stories match. Active filters.