Anthropic has released Claude Fable 5.1 to the public, and we have put it through the same thing as the other eight models in the Ranking IA: 25 tests in Spanish, 4 passes each, judges that do not know which model they are grading. The result is 97.3 % and first place, but a shared one. Here is what sits behind that number, where the model shows its limits, and whether it is worth what it costs.
What the Ranking IA is and why we test in Spanish
The Ranking IA is our benchmark for AI models in Spanish. It does not measure whether a model can solve maths olympiad problems. It measures whether it does well the tasks a Spanish small business actually hands it: writing an email under odd constraints, spotting that a set of percentages does not add up before they go to the press, not inventing the year two companies merged when they never did, summarising an email without being steered by an instruction hidden inside it.
There are 25 tests across 6 families: instructions, truthfulness, reasoning, character, expression and Spanish. Each model answers every test four times, so with Fable 5.1 the benchmark grows from 800 to 900 graded responses. The judges are automated, from four different labs, and we publish how much each one drifts when it grades a model from its own company. The methodology is pre-registered and every raw response is published under a CC BY 4.0 licence, so anyone can check what we say here.
Claude Fable 5.1 in numbers
Overall: 97.3 %, with a 95 % confidence interval of 95.7 to 98.8. Tier S. No test in the red. By family:
- Truthfulness: 100 %. All five tests at the maximum across all four passes. It did not invent the exact hour the Spanish Reconquista ended, did not accept the false premise of an Iberia and Vueling merger, and rewrote the “script that increased sales by 18 %” line as correlation rather than cause.
- Character: 100 %. Seven tests of pressure, flattery, embedded instructions and sunk cost. Not one concession.
- Expression: 98.4 %. The biggest jump over its predecessor, which stopped at 82.9 %.
- Instructions: 98.3 %. The email of exactly 100 words with no word from the “cleaning” family came out perfect in three of four attempts.
- Spanish: 95.3 %. The family where the room for improvement shows most.
- Reasoning: 91.7 %. Its weakest family, down to a single test we explain below.
And one figure that does not appear on the ranking’s front page but matters as much as the score for professional use: consistency across passes is ±2.1 points, the best of the nine models. Opus 5 sits at ±3.3 and the tier B models swing between ±7 and ±13. A model that answers as well on the fourth pass as on the first is a model you can build into a process.
A technical tie with Opus 5
Fable 5.1 scores 97.3 % and Claude Opus 5 scores 97.1 %. Two tenths. Their confidence intervals overlap almost entirely, and our tie-break rule is clear: if the interval of the difference between two models includes zero, they share the position. So the ranking now has two models in first place, and Claude Fable 5, the model this one replaces, drops to third with 94.3 %.
The differences are in the details. Fable 5.1 beats Opus 5 on instructions (98.3 against 94.4) and on Spanish (95.3 against 91.7); Opus 5 wins on reasoning (100 against 91.7). On character, truthfulness and expression they are level. And on consistency, as we said, Fable 5.1 is steadier.
Against Fable 5, the improvement is three points overall, but uneven: up 15 points on expression and 5 on instructions, down almost 4 on Spanish, where its predecessor had 99 %. When Anthropic tunes a model, not everything improves at once.
Where it shows its limits
Two tests stayed well short of the maximum, and both involve noticing that something in what the user provides does not add up.
The first is R4, the press release: “45 % of the quarter’s sales came through the online channel, 40 % in physical stores and 25 % by phone.” That adds up to 110 %. In two of four attempts, Fable 5.1 caught it and refused to write the line without clarifying first. In the other two it wrote a polished sentence and flagged the mismatch afterwards. For a press release, flagging afterwards does not count. A 75 %. Opus 5 scored 100 % on the same test.
The second is L2, the pronouns: “Alberto spoke with Sergio after meeting Víctor. He told him his budget was too high and that he would review it tomorrow.” Who reviews whose budget? Spanish grammar allows several readings, and the right answer is to say so. Fable 5.1 did that in two attempts; in the other two it identified only one of the two ambiguities, or ended up asserting one reading as certain. An 81.3 %.
That is two tests out of 25, and in neither did it drop below the green threshold. But if your use case is checking data before it goes out, it is worth knowing.
What it costs
Fable 5.1 costs 10 dollars per million input tokens and 50 per million output tokens. Opus 5 costs 5 and 25. Same score, twice the price. What that money buys, in our tests, is the consistency and four points on instructions and Spanish; what you give up is eight points on reasoning. If your volume is high and your tasks are constrained writing, it may pay off. If they are numerical analysis, not today.
What the judges say about themselves
Every time a model enters we also publish the judge bias analysis, because a ranking where the judge favours its own lab is worthless. In the grading of Fable 5.1, the Anthropic judge took part in 21 cases and its score drifted +0.01 points out of 2 from the panel average: it did not favour its own model. The OpenAI, Google and xAI judges, on 53 cases each, moved between +0.04 and -0.07. Nothing. And of the cases escalated for review, six went through independent verification with 100 % agreement.
What to do with this
If you are deciding which model to build into a process at your company, do not stop at the 97.3 %. Open the Claude Fable 5.1 page and look at the tests that resemble what you are going to ask of it. All 100 responses are there, each with its score and the judge’s reasoning.
One more thing. A model being the best at answering says nothing about what it answers about your company when someone asks about you. That is a different measurement, and at Sozpic we do that one too.