OpenAI's upcoming 'Astra' model family enables multiple agents to collaborate on complex, multi-day reasoning and math problems.
DeepSeek-V4-Flash-0731 is a 304B parameter model optimized specifically for enhanced agentic capabilities and workflows.
OpenAI models solved ten long-standing open problems in mathematics and theoretical computer science, including convex optimization.
A GPT-5.6 prompt closed a 30-year theoretical gap in convex optimization, proving LLM utility in pure mathematics.
Supabase open-sourced an Apache-2.0 benchmark to evaluate coding agents on real-world database, schema, and security tasks.
Datasette Agent 0.4a0 introduces a mechanism allowing agent tools to execute code directly inside the user's browser.
Google updated Gemini API Managed Agents with Gemini 3.6 Flash support, hooks, and event triggers.
Smevals is a new lightweight evaluation suite designed to test and benchmark models, prompts, and agent harnesses.
This research analyzes how reasoning LLMs learn to control and switch between low, medium, and high-effort computation modes.
MiniMax H3 is an omni-modal model generating 15-second 2K video with native, synchronized stereo audio.
Google's Gemini Robotics ER 2 introduces two new physical robot platforms, Duo and Apollo, powered by Gemini.
A new field report documents how scientists are deploying agentic AI coders to accelerate genomics and scientific computing.