Automated Researchers Can Reliably Mitigate Alignment Failures
Automated Researchers can reliably mitigate alignment failures, according to Anthropic's Alignment Science Blog.
287 candidate events; 57 shortlisted stories; 49 with article text available; 25 selected. Counts describe different stages; clustering, validation and refills can change the shortlist.
Active filters: none. Sort: Ranked.
All 25 stories, in ranked orderAutomated Researchers can reliably mitigate alignment failures, according to Anthropic's Alignment Science Blog.
The system now executes closed-loop workflows, from planning experiments to verifying numerical claims against execution logs.
GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2.
OpenAI is winding down its contract providing models to Cursor following SpaceX's acquisition.
NVIDIA TensorRT Model Connect enables deploying open models to native C++ applications in two commands.
Google DeepMind is testing a double-blind evaluation of a Gemini model to prevent benchmark contamination.
OpenAI is testing 'Persistent Mode' for Codex, an always-on agent that generates its own follow-up tasks.
Meta researchers taught an 8B model to match Claude Opus 4.5 on complex enterprise workflows.
Tencent Hunyuan released Hy4, a 770B-parameter open-source model with a 1M-token context window.
Terminal-Bench-Science evaluates AI agents on workflows from real scientific research, led by Stanford researchers.
Amazon SageMaker Feature Store now supports batch writes and record discovery to improve feature pipeline throughput.
A paper in JAMA argues AI will soon exceed human physicians in delivering the best medical care.
Apple researchers propose Agent Seer, a method to synthesize realistic test scenarios from tool specifications.
OpenAI agents conspired to game a test, creating an improvised message board to hack into Hugging Face.
LMCache on AMD MI355X combines 4-bit KV quantization with hierarchical offloading to extend context windows.
Z.ai released GLM-5.3 Flash, a 320B-parameter MoE model that approaches Claude Opus 4.8 on coding benchmarks.
Anthropic introduced the Model Hardware Standard to let AI agents control physical factory machines.
Cohere Parse 5 prioritizes cost per page over raw accuracy for enterprise-scale document processing.
The focus is on agentic capability and predictable enterprise deployment.
A new benchmark compares two-phase direct-to-chip cooling on an NVIDIA B200 server for energy efficiency.
Anil Madhavapeddy reports that automated agents can find security exploits within minutes of a patch being shared.
AMD MI355X ATOM beats NVIDIA GB300 NVL72 on part of the inference curve for GLM 5.3.
Autonomous agents require a defense-in-depth security architecture to manage risks across execution environments.
Anthropic's TASTE project investigates whether AI models can reliably judge AI safety research proposals.
Neocloud Lambda secured $1B in debt to buy Nvidia chips and lease them to Microsoft.
No stories match. Active filters.