← All posts
roundup

Anthropic ships Claude Science, Google's busy week, and an AI learning study that stings

Anthropic launches Claude Science and Sonnet 5, Google's AI drove a 37% electricity spike, and a 26K-student study finds a two-year hidden cost to AI-assisted homework.

#anthropic#claude-science#claude-sonnet-5#mistral-leanstral#roundup

The big picture

Anthropics is having a dense shipping week, Google is doing everything at once (new hardware, new models, new energy bills), and two research items landed that deserve more attention than they’ll probably get.

Anthropic’s week: Claude Science, Sonnet 5, and a messy China situation

Anthropics announced Claude Science at an event for pharma and biotech executives, positioning it as a scientific-research counterpart to Claude Code — give it a high-level goal, it autonomously runs the work (MIT Technology Review). The framing is ambitious, but the real test will be whether it handles the wet-lab/data-pipeline messiness that actual researchers deal with, not just clean benchmark tasks.

On the same day, Claude Sonnet 5 shipped. Simon Willison’s breakdown of the developer docs is the most useful read: Anthropic says Sonnet 5 performance is close to Opus 4.8 at lower prices, and the system card clarifies it cleared US export review because it’s meaningfully less capable than Mythos 5 on cyber tasks (Simon Willison’s Weblog). Good price-to-performance bump; the Mythos 5 safety framing is worth bookmarking for future reference.

Meanwhile, Claude Code has a genuine geopolitical tangle. Anthropic is trying to block ByteDance and Ant Financial from accessing the tool, but they’re routing around restrictions via VPNs and overseas subsidiaries. Alibaba responded by banning its own employees from using Claude Code after reportedly finding hidden code that could identify Chinese users (The Decoder). This is the kind of enforcement problem that doesn’t have a clean technical solution.

Google ships a lot, not all of it is ready

Google’s AI infrastructure consumed 37% more electricity in 2025 than the year before, according to the company’s own environmental report. Clean energy commitments are cited as the offset, but the raw numbers are striking (Ars Technica). Every major lab has this problem; Google is just the one that disclosed the scale.

On the product side: the new Nano Banana 2 Lite image model is fast and cheap, trades quality for speed, and is probably fine for UI mockups or draft work (Ars Technica). The Google Home Speaker review from The Verge is the more honest story — the hardware is solid after six years of absence, but Gemini for Home still feels unfinished, which is the recurring problem with Google’s consumer AI integrations (The Verge). NotebookLM added 60-second vertical video summaries of your research docs, which is either genuinely useful or deeply cursed depending on your relationship with TikTok (The Verge).

Research worth reading (and one number to sit with)

A study of over 26,000 Chinese students found that AI homework assistance correlated with faster completion and higher short-term scores — but exam performance dropped up to 24% compared to non-AI users, with the full effect taking two years to show up (The Decoder). The delayed signal is the key finding: studies that only run a semester will systematically miss it. This isn’t conclusive proof of harm, but it’s the most methodologically serious data point in this debate so far.

Mistral released Leanstral 1.5, an open-source model for formal verification in Lean 4 (a proof assistant language used to mathematically verify that code or math is correct). The benchmark numbers are strong, but the more interesting result is practical: it found five previously unknown bugs while scanning 57 open-source repos (The Decoder). A model that ships open-source and finds real bugs in real code is a more compelling pitch than benchmark tables alone.

Quick hits

  • Wayve, the autonomous driving startup, is running an $85M employee tender offer at an $8.5B valuation — a retention tool that’s become standard practice for AI companies that can’t go public yet. TechCrunch
  • Privacy advocates filed a complaint urging the FTC to reject any move to end monitoring of Musk’s X, citing AI data risks. Ars Technica
  • A new attack on AI browsers works by convincing the LLM that basic facts are wrong, collapsing its guardrails — yet another reason to be skeptical of agentic browser products. Ars Technica
  • Bridgewater and Mira Murati’s Thinking Machines Lab fine-tuned a Qwen3-235B model that reportedly beats GPT and Claude on proprietary finance benchmarks at a fraction of the cost — but nobody outside the two companies has verified the numbers. The Decoder
  • South Korea announced a $1 trillion investment plan targeting memory chip production and commercial humanoid robots by 2028. Ars Technica

Sources