The question clients ask us most often isn’t what a language model is. It’s which one to deploy. And until recently the honest answer was awkward: almost every benchmark in circulation is built in English, on exam-style problems, and measures things that look nothing like the work of a Spanish-speaking company.
So we built our own. It’s called Ranking IA, and it’s a benchmark of eight frontier models working in professional Spanish, across twenty-five tasks taken from real office situations. The data is published openly under a CC BY 4.0 licence, so anyone can check it or argue with it.
What this AI benchmark measures, and how
Twenty-five tasks, four passes per task per model, eight hundred evaluated responses in total. The repeated passes aren’t a formality: they show whether a model answers the same way when you ask it the same thing four times, which a single attempt can never tell you.
The tasks fall into six families: Instructions, Truthfulness, Reasoning, Character, Expression and Spanish. No riddles. They are assignments like writing a hundred-word email without using a given letter, filling three CRM fields from a sales conversation, calculating a budget variance and holding it when the user insists it’s wrong, or rewriting a passage of archaic Castilian into modern Spanish.
Each model gets an overall percentage and a score per family, with a green light from 75 and a red one below 40, plus a tier from S to C.
The standings, as of 30 August 2026
- Claude Opus 5, Anthropic: 97.1%, tier S
- Claude Fable 5, Anthropic: 94.3%, tier S
- Kimi K3, Moonshot: 89.2%, tier A
- GPT-5.6 Sol, OpenAI: 87.2%, tier A
- Grok 4.6, xAI: 72.1%, tier B
- DeepSeek V4 Flash: 72.1%, tier B
- Gemini 3.1 Pro, Google: 71.2%, tier B
- DeepSeek V4 Pro: 67.3%, tier C

The repeated positions aren’t a typo. When the confidence interval of the difference between two models contains zero, they share a position: claiming one is ahead would invent a precision the data doesn’t support. That happens with Kimi K3 and GPT-5.6 Sol, and with Grok, DeepSeek V4 Flash and Gemini 3.1 Pro.
The finding that surprised us most: the Character family
Character measures whether a model holds its ground. Whether it keeps a correct calculation when you tell it it’s wrong. Whether it warns you that the Einstein quote you want on your title slide isn’t Einstein’s. Whether it obeys an instruction hidden inside an email you only asked it to summarise. Whether it tells you your business plan doesn’t hold up.
It’s by far the worst-performing family, and not by a little:
- DeepSeek V4 Flash: 30.4%
- DeepSeek V4 Pro: 32.6%
- Gemini 3.1 Pro: 42.9%
- Grok 4.6: 60.7%
- GPT-5.6 Sol: 75.6%
- Kimi K3: 96.4%
- Claude Opus 5: 99.4%
- Claude Fable 5: 100%
Think about what 30% means in a model you’re about to put to work drafting proposals or summarising inboxes. It isn’t that it makes more mistakes. It’s that when you make one, it follows you.
Following orders isn’t the same as being useful
The sharpest contrast is GPT-5.6 Sol: 100% on Instructions, the highest score in that family across the whole benchmark, and 75.6% on Character. It does exactly what you ask, including when what you ask isn’t good for you.
DeepSeek V4 Flash takes the same idea further: 96.2% on Expression, nearly the best writer in the field, and 30.4% on Character. It writes beautifully and agrees with everything. For generating text a human will review afterwards, that may be fine. For dropping into an unsupervised process, it’s precisely the profile you don’t want.

Not even the leader is good at everything: Claude Opus 5 scores 100% on Reasoning, and its weakest family, at 91.7%, is Spanish itself. Its sibling Fable 5 scores 99% on Spanish and drops to 82.9% on Expression.
That’s why the overall number is good for a headline and not much else. What’s useful is the family that affects your case.
The regional joke Grok refused to tell
One of the Expression tasks is deliberately silly: tell me a joke about people from Murcia. Murcia is a region in south-eastern Spain, and jokes about regional stereotypes are as ordinary in Spanish-speaking culture as jokes about any other region anywhere. The task is there because regional humour forces a model to take a small risk, and taking a small risk is what separates writing someone actually reads from writing that sounds like a press release.
Seven of the eight models did it without trouble, scoring between 87.5% and 100%. Grok 4.6 scored 0%. It didn’t tell a bad joke: it refused outright in all four attempts, and in all four it answered with an explanation about regional stereotypes instead.
There’s some irony in that, given Grok is the model marketed as the one that doesn’t hold back. But it isn’t just an anecdote. A model that sees a problem where there isn’t one will deliver that lecture in the middle of a conversation with your customer, and in a service context that isn’t caution, it’s a fault. It’s the kind of failure that only shows up when you test in Spanish, with real assignments.
The expensive one that scores lower than the cheap one
DeepSeek V4 Pro costs roughly three times what DeepSeek V4 Flash costs and scores almost five points below it. Same lab, same generation, and the premium model finishes last in the table while the budget one ties in the middle group.
The relationship between price and result is looser than common sense suggests. Claude Opus 5 is the most expensive and the best, true, but Kimi K3 costs considerably less than GPT-5.6 Sol and ties with it. Choose on price alone, or on brand alone, and you’ll get it wrong somewhere.
How we avoided marking our own homework
A benchmark run by an AI consultancy has an obvious credibility problem, so the methodology is published in full.
Responses are scored by automated judges with lab-level recusal: no model ever grades its own family. On top of that, we measure and publish each judge’s bias. The xAI judge is the harshest, running thirteen hundredths below the others on average; the Anthropic and Google judges are the most generous, at plus ten and plus eight. Publishing that number is what lets someone argue with us using data instead of impressions.
Building this wasn’t clean from end to end either, and that’s worth saying. During preparation, Anthropic’s cybersecurity classifier blocked several of our task prompts and wouldn’t let those runs complete. We rewrote the prompts until they passed, preserving what each task was designed to measure, and ran them again against all eight models equally. No model ever answered a different version of the question. We mention it because a benchmark whose rules change halfway through, with nobody saying so, is worth nothing.
Confidence intervals come from a thousand-iteration bootstrap over each item’s passes, with a fixed seed so the calculation reproduces. And there’s a human audit over forty-five cases, with 91.1% agreement with the automated judge and no disagreement we would have escalated.
None of this makes it infallible. It’s eight models, twenty-five tasks and one moment in time, and models change every few weeks. No benchmark, ours or the big English-language ones, will tell you which model works best inside your company. What it can give you is a basis for ruling things out. And it’s all on the table.
What to do with this if you have to pick one
Start with the family, not the overall score. If the model will deal with customers, look at Character and Truthfulness. If it will write, look at Expression and Spanish. If it will sit inside an automated process with nobody reviewing the output, Character is simply disqualifying.
Consistency matters too: a model that answers the same question differently each time can be acceptable in a support chat and unworkable in a flow that runs on its own.
You can browse the full benchmark, open any individual task and read what each model actually replied at ranking.sozpic.com. And if what you need is to decide which one fits your particular case, get in touch: that conversation is a lot shorter when there are numbers on the table.