Updated every 3 hours

Istari Index

How Istari works

Where every number comes from, what counts as a finished task, and how far our estimates have been off. Readable here in full, and recomputable from the data.

✦ In short

Cost per finished task = the tokens a model uses at its effort level, plus your time fixing the attempts it gets wrong, counted over the retries it takes. The chance of getting a task right comes from independent benchmarks of real coding tasks, not from model makers.

Updated Oct 11, 2026. 41 models ranked. Every number below is recomputed every 3 hours from the same data you can download.

What counts as a finished task

A task is finished when the benchmark's own grader passes it: the hidden tests pass, the root cause named is the real one, the answer about the codebase is judged correct. We don't run the tasks ourselves. Independent labs do, the same way for every model, and publish the results; each benchmark is listed below with what it measures and how many tasks it has, where its publisher says.

We read a model's score as the chance it gets your task right the first time (81 ⇒ 81%). Each miss costs you a retry's tokens and your time to spot it, re-prompt and fix it.

The formula

cost per task = (price × tokens + fix cost × failure chance) ÷ success chance

Dividing by the success chance counts the retries. The constants, each with its source:

Your hourly cost, the minutes and the task size can be changed under "Your costs" on the ranking; the ranking is recomputed for you.

Check it yourself. Example: MiMo-V2.6-Pro at default effort has a coding score of 77.5 (p = 0.775), $0.6525 per 1M tokens and a thinking multiplier of 1, so 0.2700M tokens per attempt. At our $93/hour: $9.23 per finished task. At $60/hour, the fix cost is $20.00 and the cost is (0.6525 × 0.2700 + 20.00 × 0.225) / 0.775 = $6.03. To re-rank for another rate, recompute every ranked level of every model in ranking.json the same way and take each model's cheapest level. Only the fix cost changes.

Effort levels

Many models have an effort setting (low, medium, high, xhigh, max): how long the model thinks before it answers. A higher level uses more tokens for a somewhat better success rate, so each level is scored and priced separately and a model is ranked at the level that is cheapest per finished task. How much a level thinks comes from Artificial Analysis's timings where it has them, else a typical ratio.

The coding score

The general coding score is Artificial Analysis's. A model must reach 69.5, 85% of the best (Claude Opus 5.5, 81.8), to be ranked.

Artificial Analysis stopped publishing its Coding Index for new models in September 2026. It still runs Terminal-Bench (now in its harder version 4.0) and SciCode at every effort level, so for those models we compute the index the way AA did, from AA's own results for that model, and mark it ◇. It's a measurement of that model, not an estimate from other models. If AA publishes the index again, it replaces ours.

Benchmarks for each kind of work

Pick a kind of work on the ranking and each model is scored on benchmarks built for it. When a kind of work has several, each is put on the scale of the first one, and a model's score is the mean of those it was measured on. When several labs publish the same benchmark, the most trusted one's numbers (its authors, or a lab that runs every model the same way) are used as published, and another only fills in models it hasn't run.

Build a new service or API

Brand-new code from a spec: an API, a service, a module or a whole app. Past estimates for this work, once measured, were off by ±2.4 points on average (1 checked).

Change existing / legacy code

Adapt, refactor or migrate a large existing codebase. Past estimates for this work, once measured, were off by ±10.2 points on average (3 checked).

Explore & explain a codebase

Trace flows across files and repos, onboard, answer "how does this work", review. Past estimates for this work, once measured, were off by ±6.9 points on average (13 checked).

Debug, fix & test

Reproduce, locate and fix bugs; write and repair tests. Past estimates for this work, once measured, were off by ±7 points on average (2 checked).

Production investigation

Alerts, events and traces across services to find the root cause. Little data: measured for under 60% of the ranked models. Past estimates for this work, once measured, were off by ±8.2 points on average (11 checked).

Infra, DevOps & scripting

CI/CD, containers, shell, builds and environment setup. Past estimates for this work, once measured, were off by ±5.1 points on average (2 checked).

Algorithms, data & scientific code

Numerical code, data processing, algorithms.

UI & web pages

Front-end pages, components and web apps. Past estimates for this work, once measured, were off by ±16.6 points on average (3 checked).

Security & vulnerabilities

Find and fix vulnerabilities without breaking the code. Past estimates for this work, once measured, were off by ±9.1 points on average (3 checked).

Estimates, and how far off they have been

Where a model has no result yet, its score is estimated and marked. Every error below is a back-test: we hide a known result, predict it, and compare.

The pick is measured, and you can get it. It's always a model whose score was measured, in Cursor or on OpenRouter. One exception: on a kind of work measured for under 60% of the ranked models (today: production investigation), the same one or two measured models would win whatever you chose, so there an estimated task score can be the pick. The card then says so, with how far past estimates have missed, and names the other contenders.

No tokens for a gain we can't measure. A higher effort level that adds less than our back-tested error (±2.6 points) buys nothing we can measure, so "Most reliable", and every pick in a plan whose tokens are included, take the leanest level within that of the best. Models within it of each other are a tie, ordered by price.

Benchmarks we checked and left out

A benchmark we rank on must still be run on new models, and must tell good models apart. Every update checks that each source still covers the newest releases, and watches publishers for new benchmarks. Some well-known ones we decided against:

How Istari relates to the leaderboards you may know: Artificial Analysis is our main source, and LMArena's WebDev arena is one of the UI benchmarks. What we add is the price, the tokens at each effort level, the retries and your time, per kind of work.

What we can't tell you

Check our math

Spotted a wrong number? Write to istari-index@proton.me. Corrections are logged.

Questions

What is a "finished task"?

A task the benchmark's own grader passes: the hidden tests pass, the root cause named is the real one, the answer about the codebase is judged correct. Independent labs run the tasks, the same way for every model, and publish the results; Istari doesn't run them itself.

Where do the scores come from?

From independent benchmarks: Artificial Analysis for the general coding score, and benchmarks built for each kind of work from Artificial Analysis, Vals AI, Epoch AI, Scale AI, LMArena and the benchmarks' own authors. A model maker's own number is used only until an independent result exists, and is marked.

What does ◇ mean? Is that score estimated?

No, it's measured. Artificial Analysis stopped publishing its Coding Index for new models in September 2026 but still runs the benchmarks it was made of. We combine AA's own results for that model the way AA did. Rebuilding models whose index AA did publish, the same way, comes within about 2.6 points of it.

Why $93 an hour and 20 minutes a fix?

$93 is what a US software developer's hour costs an employer (BLS). 20 minutes is a typical time to spot, re-prompt and fix a failed attempt (from a study of 20,574 sessions and METR's measurements). Both are yours to change under "Your costs" on the ranking.

Who runs Istari?

An independent project by developers, unsigned on purpose so no model maker or sponsor can lobby it. It takes no money from anyone it ranks. Instead of a name, every number is sourced, the data is downloadable, and every change to the method is logged with its date.