Updated every 2 hours

Istari Index

Our quest

Why we built an independent index that answers one practical question: which AI model finishes your coding task for the least money?

✦ In short

We're developers. Every day we asked which model to use for the task in front of us, and nothing we found could answer it. So we built the answer, and we keep it current for everyone.

Free. Independent. No ads, no sponsors. Updated every 2 hours.

Every coding task starts with the same question

Which model should do this job? The flagship that costs ten times more per token, or the fast, cheap one? And at what effort level: does "max" thinking pay for itself, or does it just burn tokens?

We asked it dozens of times a day. We wanted to work efficiently: not to waste tokens, money or our own time. So we went looking for an answer.

What we needed didn't exist. So we built it.

Why "Istari"

In Tolkien's legend, the Istari were a small order of wise ones who came to Middle-earth to counsel its peoples. They came to advise, not to rule: the choices, and the road, stayed with the people who had to walk it.

That's the role we want: we don't sell models, we advise. You choose.

A good counsellor has no stake in your choice. Istari earns nothing from any model maker, so the only thing the ranking optimises is your cost per finished task.

Welcome to the token economy

Tokens have become a real budget line for developers. A single agent task reads and writes a few hundred thousand of them: Artificial Analysis measured real coding-agent tasks at about 270K tokens each. The same model can use up to 23× more tokens at its highest effort level than at its lowest.

Tokens are only half the bill. When a model gets a task almost right, you pay again in your own time: 66% of developers name "almost right, but not quite" AI code as their top frustration (Stack Overflow 2025). A cheap model that fails often can cost you more than a pricier one that gets it right.

Being efficient in the token economy means picking the right model for each task, at the right effort level. That's exactly what Istari measures.

What makes Istari different

Cost per finished task

Not cost per token: the tokens plus your time fixing failed attempts, counted over the retries it takes to get the task done.

Your kind of coding work

Pick what you actually do, and each model is scored on a benchmark built for that work, not one general score.

The effort sweet spot

More thinking isn't always worth it. Every model is ranked at the effort level with the lowest cost per finished task.

New models, fast

Updated every 2 hours. Brand-new models get a marked provisional score, and fresh releases are listed within hours.

Our commitments

What we can't tell you

A benchmark is a model of reality, not reality. Please read the ranking with these limits in mind:

Use Istari to shortlist, then try the top picks on your own work.

Privacy

The fine print

Istari's rankings are estimates built from public benchmarks and list prices. They are not guarantees and not purchasing advice. Istari is an independent project, not affiliated with or endorsed by the Tolkien Estate, Middle-earth Enterprises, Cursor (Anysphere), Artificial Analysis or any model maker shown.

See the ranking →