AI News 2026-08-04
New To You

Hugging Face’s LFM2.5-2.6B, OpenAI’s GPT-Live, and Bedrock make local, realtime, grounded agents more deployable.

50 sources swept 447 distinct stories from the last 2 days 76 already covered, held back 12 worth your time

  1. Deploy local agents everywhere with LFM2.5-2.6B

    Agentic reinforcement learning targets compatibility with popular agentic harnesses, while its 128K context supports longer workflows.

    Hugging Face · · Agents & tooling · lab
  2. Introducing Web Search on Amazon Bedrock for foundation model grounding

    Web Search is a server-side Bedrock tool for grounding chatbots, coding assistants, CLI tools, and enterprise applications in current information.

    AWS Machine Learning · · Infrastructure & chips · lab · aws.amazon.com
  3. How we built a realtime system for responsive voice AI in six months

    GPT-Live uses continuous voice interaction and low-latency architecture instead of turn-based detection that can interrupt or delay responses.

    OpenAI · · Models & releases · lab
  4. Introducing Shieldstral.

    Plain-language policies work at inference time for text and image moderation without retraining, on a single 16GB NVIDIA GPU.

    Mistral AI · · Safety & security · lab
  5. Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

    The Apache-2.0 release requires NVIDIA Blackwell SM100 hardware, limiting portability despite reported throughput gains.

    Marktechpost · · Infrastructure & chips · newsletter
  6. AI coding agents are blowing through budgets — Replit, Kilo Code, and Symbotic explain how they're managing it

    Agent-heavy teams are confronting oversight, failure cleanup, multi-model support, and escalating token costs alongside coding automation.

    VentureBeat AI · · Agents & tooling · press
  7. Generate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 Super

    Alpamayo 2 Super combines trajectory generation, intent prediction, scene understanding, and data labeling within one open autonomous-driving model.

    NVIDIA Developer · · Models & releases · lab
  8. Open-weight AI models are catching up to the frontier. The safety gap remains.

    SaferAI reports GLM-5.2 is only months behind leading models on cyber and bio capabilities while lacking key safety mitigations.

    TechCrunch AI · · Models & releases · press
  9. Silicon Valley’s rift over open source pushes back contemplated White House bans on Chinese AI

    Industry opposition caused the administration to back away from proposed sanctions and cloud bans targeting Chinese open-weight models.

    The Decoder · · Business & policy · press
  10. Texas halts data center connections to power grid amid overwhelming demand

    Texas now requires data-center developers to provide information about potential grid and community impacts before receiving new connections.

    Ars Technica AI · · Infrastructure & chips · press
  11. Google moves billions in Anthropic chip risk off its balance sheet

    A special-purpose vehicle funds Anthropic’s TPU access with outside money while Anthropic leases the hardware and Broadcom guarantees the arrangement.

    The Decoder · · Business & policy · press
  12. PipeNetwork/minimax-h3-mlx

    The MLX port runs MiniMax-H3’s text, image, audio, and video inputs on Apple Silicon, including video generation with audio.

    Simon Willison · · Models & releases · newsletter