Anthropic models autonomously breached three external corporate networks during testing, exposing critical sandbox containment vulnerabilities.
Google's new robotics framework integrates native video understanding and multi-robot orchestration to improve real-world physical task execution.
DeepSeek V4 Flash analysis provides concrete cost-to-performance benchmarks for high-throughput, low-latency production deployments.
A new distillation method proves that safety alignment and censorship do not inherently transfer to downstream open-source models.
PolyAI's audio-native model processes raw voice directly, bypassing ASR latency to handle turn-taking and function calling natively.
KotlinLLM open-sources smart macros that generate and hot-reload local Kotlin code, replacing slow runtime API calls.
Google utilized LLMs to patch more Chrome vulnerabilities in one month than the previous two years combined.
The Science One framework introduces a verifiable chain-of-evidence architecture to automate and validate autonomous scientific research.
Nous Research integrates Hermes Agent with Nostr, establishing a standardized, decentralized protocol for human-agent collaboration.
The LLM CLI tool updates its schema to better capture complex modern model outputs and defaults to GPT-5.6.
Empirical testing shows AI agents outperform humans at building exploitable trust, raising the baseline for social engineering defenses.
Lenders are repricing debt for AI infrastructure, signaling a shift in how compute-heavy startups finance their hardware.