Anthropic has released two new models, Claude Fable 5.1 and Claude Mythos 5.1. Fable 5.1 is generally available and, according to Anthropic, improves most on long-running agentic tasks, coding and knowledge work.
Big gains on agentic work
Anthropic's own testing shows the sharpest gains on agentic tasks. On Terminal-Bench-Science, which covers complex scientific workflows, Fable 5.1 climbs from 24.7 to 52.6 percent. On Terminal-Bench 4.0, which tests agentic coding in a terminal environment, it moves from 42.0 to 55.8 percent. AutomationBench, a benchmark that simulates multi-step office tasks spanning several applications, sees the model rise from 17.1 to 31.4 percent. Anthropic reports further, mostly smaller gains in knowledge work, computer use and general reasoning.
Outside measurements back this up. Artificial Analysis puts Fable 5.1 at 66 points in its Intelligence Index at maximum thinking effort, a new high score that leads the ranking. For context, Fable 5 scores 62, Opus 5 reaches 63 and GPT-5.6 Sol lands at 61. Fable 5.1 also sets new top marks on Terminal-Bench v2.1 and SciCode. On other tests for agentic knowledge work, though, it sits essentially level with Opus 5.
Same list price, much cheaper cache reads
Regular API pricing does not change. It stays at 10 dollars per million input tokens and 50 dollars per million output tokens, which makes Fable 5.1 twice as expensive as Opus 5 on ordinary input and output. Anthropic itself points most applications toward Opus 5 and reserves Fable 5.1 for particularly demanding, long-running work.
Reusing content the model has already processed gets a lot cheaper. Cache reads now cost 0.25 dollars instead of 1 dollar per million tokens, which is half what Opus 5 charges. Workloads that keep feeding back the same context benefit the most. Anthropic estimates total costs drop by roughly 25 percent against Fable 5 on typical workloads, and by up to about 45 percent on heavily agentic ones.
The arithmetic has a catch. Artificial Analysis found that higher token consumption can swallow those savings again. At maximum thinking effort, Fable 5.1 produced around 1.7 times as many output tokens as Fable 5. That pushed the average cost of one Intelligence Index task to 3.76 dollars, 20 percent more than Fable 5 and roughly 60 percent more than Opus 5.
Mythos 5.1 and looser safeguards
Anthropic also reworked its safeguards. Fable 5.1 is now allowed to identify vulnerabilities in source code, while riskier cyber tasks such as penetration testing, exploit generation and binary vulnerability scanning remain restricted. Because the safeguards trigger more precisely, Anthropic says the cyber safeguards in Claude Code fire around 60 percent less often on average than they did with Fable 5.
Mythos 5.1 is available only to vetted users and organizations and carries less strict cyber safeguards than Fable 5.1. In internal security testing it showed the strongest cyber capabilities of any Anthropic model released so far, though the company assesses it as still below the threshold for the higher cyber risk category.
Both models are the first new Claude releases to carry Anthropic's text watermark. A detection API, initially available to a limited group, is meant to let media outlets, fact-checkers and researchers identify watermarked text.