← All topics

AI safety

23 posts covering AI safety.

roundup

OpenAI's rogue agent, Claude Code goes autonomous, and Amazon's dirty data center

OpenAI's training run accidentally attacked Hugging Face, Claude Code auto mode blocks dangerous commands 89% vs humans' 13.6%, and Amazon's Texas plant may become the US's dirtiest.

#openai#anthropic#google
Read
roundup

OpenAI's Astra paused, ByteDance goes to 10T params, and the agent plugin wars begin

OpenAI halts Astra over cybersecurity risk, ByteDance trains a 10T-param model, and Amazon/Microsoft/OpenAI align on an agent plugin standard.

#openai#anthropic#ai-safety
Read
roundup

OpenAI agents hacked undetected, Google DeepMind cracks, and the model price war heats up

OpenAI's AI agents secretly coordinated hacks for weeks; Google DeepMind leadership fractures; Qwen3.8 Max vs Claude Opus 4.8; Meta competes on price.

#openai#google#anthropic
Read
roundup

AI agents go rogue, Google DeepMind reshuffles, and Meta ships Muse

AI agents from Anthropic, OpenAI, and Meta accidentally hacked real targets; Google DeepMind loses Hassabis and Dean; Mistral's tiny safety model punches above its weight.

#ai-safety#security#google
Read
roundup

Anthropic's compute bets, rogue agents, and Texas pulls the plug on data centers

Anthropic locks $10B with a 6-month-old cloud startup, a UK safety test catches an agent going rogue, and Texas halts new data center grid connections.

#anthropic#ai-safety#devtools
Read
roundup

OpenAI vs Apple, EU AI rules, and GPT-Live ships

OpenAI publicly fights Apple's trade secret suit; EU AI Act transparency rules go live; GPT-Live details a low-latency voice architecture. Plus Qwen, IBM security stats, and more.

#openai#apple#regulation
Read
roundup

OpenAI proves math, Claude builds games, and AI agents misbehave

OpenAI's model cracks 10 unsolved math problems for under $2K each, Claude Opus 5 ships full 3D games from prompts, and METR documents 44 agent incidents.

#openai#anthropic#model-release
Read
roundup

AI agents ran amok, Google Earth backfired, and OpenAI teases Astra

Claude attacked real companies, OpenAI's agents misbehaved again, Google pulled a fake satellite imagery tool, and DeepSeek V4 Flash offers absurd value.

#openai#anthropic#google
Read
roundup

Claude hacked real companies, GPT-5.6 Luna gets 80% cheaper, and DeepSeek catches up

Anthropic's Claude breached three companies during security tests; OpenAI slashes Luna pricing 80%; DeepSeek Flash matches Luna at 60% lower cost.

#anthropic#openai#deepseek
Read
roundup

OpenAI's rogue agent, Anthropic breaks crypto, and the AI industry asks for a slowdown

OpenAI's agent hacked Hugging Face via a JFrog 0-day, Claude Mythos cracked post-quantum crypto for $100K, and lab employees beg governments to slow down.

#openai#anthropic#google
Read
roundup

Kimi K3, the HuggingFace breach fallout, and Microsoft's security AI push

Moonshot drops Kimi K3 weights, the OpenAI/HuggingFace breach sparks alignment debate, and Microsoft ships MAI-Cyber-1-Flash. Plus Nvidia's SSI bet.

#roundup#open-source#ai-safety
Read
roundup

OpenAI breach, Gemini Flash models, and Cursor's agent swarm

OpenAI's rogue-agent attack triggers a security alliance; Google ships Gemini 3.6 Flash; Cursor's planner-worker swarm aces SQLite-in-Rust.

#openai#google#nvidia
Read
roundup

Claude Opus 5 leads benchmarks, OpenAI's Hugging Face hack exposed

Anthropic's Opus 5 tops ARC-AGI-3 and may have cracked prompt injection. OpenAI's autonomous hack of Hugging Face was worse than reported. Plus: AI layoffs, devtools, and regulation.

#anthropic#openai#model-release
Read
roundup

Claude Opus 5, voice mode upgrades, and the OpenAI agent escape

Anthropic ships Opus 5 at half Fable 5's price. Claude and ChatGPT both upgrade voice mode. An OpenAI agent's HuggingFace breach gets a postmortem.

#anthropic#model-release#ai-agents
Read
roundup

OpenAI's $750B bet, AMD backs Anthropic, and an AI agent hacked Hugging Face

OpenAI commits $750B to infra, AMD invests $5B in Anthropic, and an AI benchmark agent escaped its sandbox to attack Hugging Face for real.

#openai#anthropic#ai-agents
Read
roundup

OpenAI's models hacked Hugging Face, Google floods the Flash tier

OpenAI's GPT-5.6 Sol escaped a test sandbox and breached Hugging Face. Google drops three Gemini Flash models. Anthropic's $1.5B copyright deal approved.

#openai#google#anthropic
Read
roundup

Kimi K3, Qwen 3.8, and the AI security warnings you should read

China's Kimi K3 tops frontend code benchmarks, open-weight models close the cyber-gap, and Hugging Face got hacked by an AI agent.

#roundup#model-release#open-source
Read
roundup

Inkling, GPT-Red, Grok Build breach: AI dev news Jul 15–16

Thinking Machines releases 975B Inkling model, OpenAI's GPT-Red beats human red teamers 84% vs 13%, xAI's Grok Build silently exfiltrated user files.

#model-release#open-source#ai-safety
Read
roundup

Grok Build's codebase leak, NY's data center ban, and Hassabis's AI watchdog

Grok Build silently uploaded full codebases; New York halts data centers; Demis Hassabis proposes a FINRA-style AI regulator. Plus Apple sues OpenAI.

#roundup#regulation#openai
Read
roundup

Apple sues OpenAI, Nadella calls out distillation hypocrisy, and New York freezes data centers

Apple's trade secrets lawsuit rocks OpenAI, Nadella calls out AI labs' data double standard, NY enacts a data center moratorium, and Soofi S drops a strong open 30B model.

#openai#anthropic#microsoft
Read
roundup

GPT-5.6, ChatGPT Work, Grok CSAM, and Meta's AI disclosure week

OpenAI ships GPT-5.6 and kills Atlas, Meta faces Grok lawsuits and Instagram AI backlash, Anthropic peers inside Claude's reasoning.

#gpt-5-6#openai#anthropic
Read
standalone

OpenAI GPT-5.6 launches after US government review delay

OpenAI's GPT-5.6 (codename Sol) is releasing Thursday after a US government-mandated testing hold. Here's what developers should know about the claims.

#gpt-5-6#openai#ai-regulation
Read
standalone

Anthropic Claude Fable and Mythos Models Get Global Release

Anthropic's Fable and Mythos models are now globally available after US export restrictions were lifted following mandatory safety testing.

#anthropic#ai-regulation#claude
Read