Something important happened with Claude this week, and I think the benchmark charts are distracting people from it.
Fable 5.1 is obviously smarter. Anthropic has numbers to prove that. But open it, give it the same long prompts you used last month, ask the same questions, and you may completely miss what makes this model interesting.
I would pay attention to this release if you use Claude for coding, research, Cowork, automation, documents, agents, or any work that lasts longer than one prompt.
Because this model seems to become more useful the longer the job gets.
And that changes how I would use Claude.
The numbers are big. The real change is easier to understand.
Fable 5.1 has a 1 million token context window and can produce as much as 128K tokens in one request.
Anthropic also reports a big jump on longer agent-style work. On one scientific research benchmark, Fable moved from 24.7% to 52.6%. On Terminal-Bench 4.0, a coding benchmark built around real terminal work, it went from 42% to 55.8%.
But here is the part I care about.
Fable 5.1 is made for work that keeps going.
Not:
“Summarize this PDF.”
More like:
“Read these files, find the problem, research what is missing, make a plan, fix it, check the result, and keep going until the work is actually finished.”
That is a different type of AI use.
The model can work with tools, keep much more context around, send parts of the job to other agents, check results, return to a failed step, and continue.
This is why I would not think of Fable 5.1 as simply Claude with a higher IQ.
Think of it as Claude becoming better at owning a job.
And once you understand that, a lot of the new Claude features suddenly make sense.
Read more