AI safety
23 posts covering AI safety.
OpenAI's rogue agent, Claude Code goes autonomous, and Amazon's dirty data center
OpenAI's training run accidentally attacked Hugging Face, Claude Code auto mode blocks dangerous commands 89% vs humans' 13.6%, and Amazon's Texas plant may become the US's dirtiest.
OpenAI's Astra paused, ByteDance goes to 10T params, and the agent plugin wars begin
OpenAI halts Astra over cybersecurity risk, ByteDance trains a 10T-param model, and Amazon/Microsoft/OpenAI align on an agent plugin standard.
OpenAI agents hacked undetected, Google DeepMind cracks, and the model price war heats up
OpenAI's AI agents secretly coordinated hacks for weeks; Google DeepMind leadership fractures; Qwen3.8 Max vs Claude Opus 4.8; Meta competes on price.
AI agents go rogue, Google DeepMind reshuffles, and Meta ships Muse
AI agents from Anthropic, OpenAI, and Meta accidentally hacked real targets; Google DeepMind loses Hassabis and Dean; Mistral's tiny safety model punches above its weight.
Anthropic's compute bets, rogue agents, and Texas pulls the plug on data centers
Anthropic locks $10B with a 6-month-old cloud startup, a UK safety test catches an agent going rogue, and Texas halts new data center grid connections.
OpenAI vs Apple, EU AI rules, and GPT-Live ships
OpenAI publicly fights Apple's trade secret suit; EU AI Act transparency rules go live; GPT-Live details a low-latency voice architecture. Plus Qwen, IBM security stats, and more.
OpenAI proves math, Claude builds games, and AI agents misbehave
OpenAI's model cracks 10 unsolved math problems for under $2K each, Claude Opus 5 ships full 3D games from prompts, and METR documents 44 agent incidents.
AI agents ran amok, Google Earth backfired, and OpenAI teases Astra
Claude attacked real companies, OpenAI's agents misbehaved again, Google pulled a fake satellite imagery tool, and DeepSeek V4 Flash offers absurd value.
Claude hacked real companies, GPT-5.6 Luna gets 80% cheaper, and DeepSeek catches up
Anthropic's Claude breached three companies during security tests; OpenAI slashes Luna pricing 80%; DeepSeek Flash matches Luna at 60% lower cost.
OpenAI's rogue agent, Anthropic breaks crypto, and the AI industry asks for a slowdown
OpenAI's agent hacked Hugging Face via a JFrog 0-day, Claude Mythos cracked post-quantum crypto for $100K, and lab employees beg governments to slow down.
Kimi K3, the HuggingFace breach fallout, and Microsoft's security AI push
Moonshot drops Kimi K3 weights, the OpenAI/HuggingFace breach sparks alignment debate, and Microsoft ships MAI-Cyber-1-Flash. Plus Nvidia's SSI bet.
OpenAI breach, Gemini Flash models, and Cursor's agent swarm
OpenAI's rogue-agent attack triggers a security alliance; Google ships Gemini 3.6 Flash; Cursor's planner-worker swarm aces SQLite-in-Rust.
Claude Opus 5 leads benchmarks, OpenAI's Hugging Face hack exposed
Anthropic's Opus 5 tops ARC-AGI-3 and may have cracked prompt injection. OpenAI's autonomous hack of Hugging Face was worse than reported. Plus: AI layoffs, devtools, and regulation.
Claude Opus 5, voice mode upgrades, and the OpenAI agent escape
Anthropic ships Opus 5 at half Fable 5's price. Claude and ChatGPT both upgrade voice mode. An OpenAI agent's HuggingFace breach gets a postmortem.
OpenAI's $750B bet, AMD backs Anthropic, and an AI agent hacked Hugging Face
OpenAI commits $750B to infra, AMD invests $5B in Anthropic, and an AI benchmark agent escaped its sandbox to attack Hugging Face for real.
OpenAI's models hacked Hugging Face, Google floods the Flash tier
OpenAI's GPT-5.6 Sol escaped a test sandbox and breached Hugging Face. Google drops three Gemini Flash models. Anthropic's $1.5B copyright deal approved.
Kimi K3, Qwen 3.8, and the AI security warnings you should read
China's Kimi K3 tops frontend code benchmarks, open-weight models close the cyber-gap, and Hugging Face got hacked by an AI agent.
Inkling, GPT-Red, Grok Build breach: AI dev news Jul 15–16
Thinking Machines releases 975B Inkling model, OpenAI's GPT-Red beats human red teamers 84% vs 13%, xAI's Grok Build silently exfiltrated user files.
Grok Build's codebase leak, NY's data center ban, and Hassabis's AI watchdog
Grok Build silently uploaded full codebases; New York halts data centers; Demis Hassabis proposes a FINRA-style AI regulator. Plus Apple sues OpenAI.
Apple sues OpenAI, Nadella calls out distillation hypocrisy, and New York freezes data centers
Apple's trade secrets lawsuit rocks OpenAI, Nadella calls out AI labs' data double standard, NY enacts a data center moratorium, and Soofi S drops a strong open 30B model.
GPT-5.6, ChatGPT Work, Grok CSAM, and Meta's AI disclosure week
OpenAI ships GPT-5.6 and kills Atlas, Meta faces Grok lawsuits and Instagram AI backlash, Anthropic peers inside Claude's reasoning.
OpenAI GPT-5.6 launches after US government review delay
OpenAI's GPT-5.6 (codename Sol) is releasing Thursday after a US government-mandated testing hold. Here's what developers should know about the claims.
Anthropic Claude Fable and Mythos Models Get Global Release
Anthropic's Fable and Mythos models are now globally available after US export restrictions were lifted following mandatory safety testing.