|
One more thing…
Anthropic · Model releases
Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads
Every story above moved the work off the model and into the layer around it. This one puts a number on that layer. Anthropic shipped two models that are the same model: Fable 5.1 and Mythos 5.1 differ only in the safeguards wrapped around them. Fable 5.1 is generally available as claude-fable-5-1; Mythos stays restricted to vetted organizations inside Project Glasswing.
On Terminal-Bench 4.0, Fable 5.1 reaches 55.8% and Mythos 5.1 reaches 60.9%. Five points, one model, and the difference is the cost of safeguard interventions — an unusually direct thing to publish. On Terminal-Bench-Science 0.1, Fable 5.1 takes 52.6% against 29.0% for Opus 5, 24.7% for Fable 5 and 22.4% for GPT-5.6 Sol, with a standard error of 3.5 to 4.5 points per model, so read the margin rather than the ranking. Elsewhere: CursorBench 3.2.0 at 73.4%, Humanity’s Last Exam at 60.9% without tools and 65.0% with, OSWorld 2.0 at 41.7% strict.
The commercial number is cache reads falling 75%, from $1.00 to $0.25 per million — 0.025 times base input, against 0.1 on every other Claude model. Base input and output are unchanged at $10 and $50. Anthropic measures roughly 25% lower cost on typical workloads and up to about 45% on context-heavy agentic ones. Read the breaking changes before you upgrade: tool_choice set to any or tool now returns a 400; thinking blocks are model-bound, so router fallbacks lose reasoning on the way down; and editing earlier turns invalidates thinking blocks, enforced for accounts created on or after August 31, 2026. Anthropic documents regressions too — parallel tool calling is more variable, and the model prefers whole-file rewrites over targeted edits.
Read on Marktechpost → · Announcement · Project Glasswing
|