✦ In short
We're developers. Every day we asked which model to use for the task in front of us, and nothing we found could answer it. So we built the answer, and we keep it current for everyone.
Free. Independent. No ads, no sponsors. Updated every 2 hours.
Every coding task starts with the same question
Which model should do this job? The flagship that costs ten times more per token, or the fast, cheap one? And at what effort level: does "max" thinking pay for itself, or does it just burn tokens?
We asked it dozens of times a day. We wanted to work efficiently: not to waste tokens, money or our own time. So we went looking for an answer.
- Leaderboards rank models by score, as if price didn't exist.
- Price pages list the cost per million tokens, as if every model used the same number of tokens and got every task right.
- Neither tells you what a finished task costs, or which model is best for your kind of work: understanding a codebase, building a feature, changing legacy code, or chasing a bug through the logs.
What we needed didn't exist. So we built it.
Why "Istari"
In Tolkien's legend, the Istari were a small order of wise ones who came to Middle-earth to counsel its peoples. They came to advise, not to rule: the choices, and the road, stayed with the people who had to walk it.
That's the role we want: we don't sell models, we advise. You choose.
A good counsellor has no stake in your choice. Istari earns nothing from any model maker, so the only thing the ranking optimises is your cost per finished task.
Welcome to the token economy
Tokens have become a real budget line for developers. A single agent task reads and writes a few hundred thousand of them: Artificial Analysis measured real coding-agent tasks at about 270K tokens each. The same model can use up to 23× more tokens at its highest effort level than at its lowest.
Tokens are only half the bill. When a model gets a task almost right, you pay again in your own time: 66% of developers name "almost right, but not quite" AI code as their top frustration (Stack Overflow 2025). A cheap model that fails often can cost you more than a pricier one that gets it right.
Being efficient in the token economy means picking the right model for each task, at the right effort level. That's exactly what Istari measures.
What makes Istari different
Cost per finished task
Not cost per token: the tokens plus your time fixing failed attempts, counted over the retries it takes to get the task done.
Your kind of coding work
Pick what you actually do, and each model is scored on a benchmark built for that work, not one general score.
The effort sweet spot
More thinking isn't always worth it. Every model is ranked at the effort level with the lowest cost per finished task.
New models, fast
Updated every 2 hours. Brand-new models get a marked provisional score, and fresh releases are listed within hours.
Our commitments
- Independent. No ads, no affiliate links, no sponsored placements. No model maker pays to be ranked, or to be ranked higher.
- Transparent. The full method is on the page, every input is sourced, and the data is downloadable, so you can check our math.
- Current. The ranking refreshes every 2 hours from independent benchmarks, led by Artificial Analysis.
- Honest about uncertainty. Where a score is estimated rather than measured, it is marked, and a tap shows what it was estimated from.
What we can't tell you
A benchmark is a model of reality, not reality. Please read the ranking with these limits in mind:
- Benchmarks aren't your codebase. A model that shines on public tests can stumble on your stack.
- Averages are averages. We assume about 270K tokens a task and 20 minutes to fix a failed attempt; your tasks may be bigger or smaller.
- Your hour has its own price. We use $93, the typical total cost of a US software developer's hour.
- Prices change. We refresh every 2 hours, but a price cut can land between updates.
Use Istari to shortlist, then try the top picks on your own work.
Privacy
- No cookies, no accounts, no sign-up, no personal data.
- Two anonymous, cookie-free visit counters (Umami and Cloudflare Web Analytics) tell us how many people visit, which parts of the page are used and how fast it loads. Neither identifies you or follows you across sites.
- Fonts are served by Istari itself. Besides the visit counters, the only outside file is the chart library, from the public jsDelivr CDN.
- Your filters (vendors, coding tasks) are saved only in your own browser.
The fine print
Istari's rankings are estimates built from public benchmarks and list prices. They are not guarantees and not purchasing advice. Istari is an independent project, not affiliated with or endorsed by the Tolkien Estate, Middle-earth Enterprises, Cursor (Anysphere), Artificial Analysis or any model maker shown.