AI News 2026-08-01
New To You · 12 of 343 stories

OpenAI teases multi-agent 'Astra' model family as DeepSeek launches agent-focused V4-Flash.

50 sources swept 343 distinct stories from the last 14 days 34 already covered, held back 12 worth your time
  1. OpenAI announces its "next major model" Astra by dropping ten previously unsolved math solutions

    OpenAI's upcoming 'Astra' model family enables multiple agents to collaborate on complex, multi-day reasoning and math problems.

    The Decoder··Models & releases·press·the-decoder.com
  2. deepseek-ai/DeepSeek-V4-Flash-0731

    DeepSeek-V4-Flash-0731 is a 304B parameter model optimized specifically for enhanced agentic capabilities and workflows.

    Simon Willison··Models & releases·newsletter·3 outlets·simonwillison.net
  3. Ten advances in mathematics and theoretical computer science

    OpenAI models solved ten long-standing open problems in mathematics and theoretical computer science, including convex optimization.

    OpenAI··Research·lab·openai.com
  4. GPT-5.6 used a prompt to close a 30-year gap in convex optimization

    A GPT-5.6 prompt closed a 30-year theoretical gap in convex optimization, proving LLM utility in pure mathematics.

    Hacker News AI··Research·HN·601 pts·old.reddit.com
  5. Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

    Supabase open-sourced an Apache-2.0 benchmark to evaluate coding agents on real-world database, schema, and security tasks.

    Marktechpost··Infrastructure & chips·newsletter·marktechpost.com
  6. datasette-agent 0.4a0

    Datasette Agent 0.4a0 introduces a mechanism allowing agent tools to execute code directly inside the user's browser.

    Simon Willison··Agents & tooling·newsletter·simonwillison.net
  7. Gemini API Managed Agents: 3.6 Flash, hooks, and more

    Google updated Gemini API Managed Agents with Gemini 3.6 Flash support, hooks, and event triggers.

    Google AI··Agents & tooling·lab·blog.google
  8. smevals - a small eval suite for evaluating models, prompts, and harnesses

    Smevals is a new lightweight evaluation suite designed to test and benchmark models, prompts, and agent harnesses.

    Simon Willison··Infrastructure & chips·newsletter·simonwillison.net
  9. Controlling Reasoning Effort in LLMs

    This research analyzes how reasoning LLMs learn to control and switch between low, medium, and high-effort computation modes.

    Ahead of AI··Research·newsletter·magazine.sebastianraschka.com
  10. MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio

    MiniMax H3 is an omni-modal model generating 15-second 2K video with native, synchronized stereo audio.

    Marktechpost··Models & releases·newsletter·marktechpost.com
  11. Introducing Gemini Robotics ER 2

    Google's Gemini Robotics ER 2 introduces two new physical robot platforms, Duo and Apollo, powered by Gemini.

    Google (Gemini)··Models & releases·lab·blog.google
  12. Scientific computing in the age of agentic AI

    A new field report documents how scientists are deploying agentic AI coders to accelerate genomics and scientific computing.

    OpenAI··Agents & tooling·lab·openai.com