AI News 2026-08-01
New To You · 12 of 342 stories

OpenAI teases multi-agent 'Astra' model family alongside major open-source benchmark and model releases.

50 sources swept 342 distinct stories from the last 13 days 39 already covered, held back 12 worth your time
  1. OpenAI announces its "next major model" Astra by dropping ten previously unsolved math solutions

    OpenAI is developing 'Astra', a multi-agent model family designed for long-horizon, collaborative problem solving over hours or days.

    The Decoder··Models & releases·press·the-decoder.com
  2. deepseek-ai/DeepSeek-V4-Flash-0731

    DeepSeek released V4-Flash-0731, a 304B parameter model boasting significantly upgraded agentic capabilities.

    Simon Willison··Models & releases·newsletter·3 outlets·simonwillison.net
  3. Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

    Supabase open-sourced an Apache-2.0 benchmark to evaluate coding agents on real-world database, schema, and RLS policy tasks.

    Marktechpost··Agents & tooling·newsletter·marktechpost.com
  4. ByteDance's Seedance 2.5 generates 30-second video clips with built-in audio

    ByteDance's Seedance 2.5 generates unified 30-second video and audio clips, tripling the duration of competing models.

    The Decoder··Models & releases·press·2 outlets·the-decoder.com
    Also covered byMarktechpost
  5. Gemini API Managed Agents: 3.6 Flash, hooks, and more

    Google introduced Managed Agents for Gemini 3.6 Flash, adding native hooks and triggers for agentic workflows.

    Google AI··Agents & tooling·lab·blog.google
  6. Our position on open-weights models

    Anthropic published its official policy stance on open-weights models, balancing safety concerns with innovation benefits.

    Anthropic··Business & policy·lab·2 outlets
    Also covered byHacker News AI
  7. Ten advances in mathematics and theoretical computer science

    OpenAI published solutions to ten long-standing open problems in mathematics and theoretical computer science.

    OpenAI··Research·lab·openai.com
  8. smevals - a small eval suite for evaluating models, prompts, and harnesses

    A new lightweight evaluation suite, smevals, has been released for testing models, prompts, and agent harnesses.

    Simon Willison··Infrastructure & chips·newsletter·simonwillison.net
  9. datasette-agent 0.4a0

    Datasette-agent 0.4a0 introduces a mechanism allowing agent tools to execute code directly in the user's browser.

    Simon Willison··Agents & tooling·newsletter·simonwillison.net
  10. OpenAI reduces Codex Model Context Size from 372k to 272k

    OpenAI has reduced the context window size of its Codex model from 372k down to 272k tokens.

    Hacker News AI··Models & releases·HN·371 pts·github.com
  11. Introducing Gemini Robotics ER 2

    Google expanded its robotics lineup by introducing two new physical platforms, Duo and Apollo, powered by Gemini ER 2.

    Google (Gemini)··Models & releases·lab·blog.google
  12. Scientific computing in the age of agentic AI

    A new field report documents how researchers are deploying AI coding agents to modernize legacy scientific computing codebases.

    OpenAI··Agents & tooling·lab·openai.com