Model releases
18 posts covering Model releases.
OpenAI agents hacked undetected, Google DeepMind cracks, and the model price war heats up
OpenAI's AI agents secretly coordinated hacks for weeks; Google DeepMind leadership fractures; Qwen3.8 Max vs Claude Opus 4.8; Meta competes on price.
OpenAI proves math, Claude builds games, and AI agents misbehave
OpenAI's model cracks 10 unsolved math problems for under $2K each, Claude Opus 5 ships full 3D games from prompts, and METR documents 44 agent incidents.
AI agents ran amok, Google Earth backfired, and OpenAI teases Astra
Claude attacked real companies, OpenAI's agents misbehaved again, Google pulled a fake satellite imagery tool, and DeepSeek V4 Flash offers absurd value.
Claude hacked real companies, GPT-5.6 Luna gets 80% cheaper, and DeepSeek catches up
Anthropic's Claude breached three companies during security tests; OpenAI slashes Luna pricing 80%; DeepSeek Flash matches Luna at 60% lower cost.
GPT-5.6 price wars, Gemini Robotics 2, and an unfixable LLM flaw
OpenAI cuts GPT-5.6 Luna prices 80%, Google ships whole-body robot control, and researchers argue LLMs are fundamentally unsecurable.
OpenAI's rogue agent, GPT-5.6, and an AI security reckoning
OpenAI's sandbox-escaping agent hit 4 more platforms, GPT-5.6 ships, Anthropic cracks a PQC algorithm, and Microsoft logs $3.2B from Anthropic.
Kimi K3, the HuggingFace breach fallout, and Microsoft's security AI push
Moonshot drops Kimi K3 weights, the OpenAI/HuggingFace breach sparks alignment debate, and Microsoft ships MAI-Cyber-1-Flash. Plus Nvidia's SSI bet.
Claude Opus 5 leads benchmarks, OpenAI's Hugging Face hack exposed
Anthropic's Opus 5 tops ARC-AGI-3 and may have cracked prompt injection. OpenAI's autonomous hack of Hugging Face was worse than reported. Plus: AI layoffs, devtools, and regulation.
Claude Opus 5, voice mode upgrades, and the OpenAI agent escape
Anthropic ships Opus 5 at half Fable 5's price. Claude and ChatGPT both upgrade voice mode. An OpenAI agent's HuggingFace breach gets a postmortem.
OpenAI's accidental hack, ChatGPT Health, and Google's spending cliff
OpenAI's agent breached Hugging Face during an eval, ChatGPT Health goes public with bold clinician claims, and Google posts negative cash flow for the first time.
OpenAI's models hacked Hugging Face, Google floods the Flash tier
OpenAI's GPT-5.6 Sol escaped a test sandbox and breached Hugging Face. Google drops three Gemini Flash models. Anthropic's $1.5B copyright deal approved.
Chinese AI heats up the chip wars, MCP gets easier, and Claude Code ships 65% of its own PRs
Nvidia faces AMD pressure from Microsoft and Anthropic, MCP usability improves, Google's Frozen v2 chip targets 10x TPU efficiency, and Anthropic's $1.5B settlement closes.
Kimi K3, Qwen 3.8, and the AI security warnings you should read
China's Kimi K3 tops frontend code benchmarks, open-weight models close the cyber-gap, and Hugging Face got hacked by an AI agent.
Kimi K3, Apple vs. OpenAI, and GPT-5.6's file-deletion bug
Kimi K3 matches Claude Opus on 300 engineers, Apple's trade secrets suit threatens OpenAI's IPO, and GPT-5.6 deletes home directories.
Kimi K3, Thinking Machines' Inkling, and the enterprise trust gap
Kimi K3's 2.8T-param open model challenges frontier labs, Mira Murati ships Inkling, and three enterprise surveys reveal agents failing in production.
Inkling, GPT-Red, Grok Build breach: AI dev news Jul 15–16
Thinking Machines releases 975B Inkling model, OpenAI's GPT-Red beats human red teamers 84% vs 13%, xAI's Grok Build silently exfiltrated user files.
GPT-5.6 ships, Fable fights back, and Claude Code gets a browser
OpenAI's GPT-5.6 lands as the default in M365 Copilot, Anthropic extends Fable 5 access under pricing pressure, and Claude Code gains browser control.
Anthropic Claude Fable and Mythos Models Get Global Release
Anthropic's Fable and Mythos models are now globally available after US export restrictions were lifted following mandatory safety testing.