Terminal-Bench 4.0 read straight from Artificial Analysis's API
Until now we read it from AA's website, which lists fewer effort levels. With the API, 31 effort levels whose score we had inferred from another level of the same model are now measured.
Where a kind of work has little data, an estimated task score can be the pick
Production investigation is measured for under 60% of the ranked models, so the same one or two measured models won whatever you chose. There, the cheapest model can now be the pick even if its score for that work is an estimate. The card says so, gives how far past estimates for that work have missed, and names the other contenders. An estimated coding score is still never the pick.
"Most reliable": ties are ordered by price
Models within our measurement error of each other are a tie. Below the top, a tie now lists the cheapest per finished task first, so a pricier model no longer sits above a cheaper one that is as reliable. The chart also shows each model's effort level.
Plans with tokens included show how much of the plan a task uses
In a plan whose tokens are included (Claude Pro or Max, ChatGPT, Google AI), each bar shows the plan a task uses against the first row (for example "1.4×"), which explains the order within a tie.
"Fewest tokens" could pick a costlier, less certain level of the same model
It now picks only levels whose score was measured, and only when it at least halves the tokens of the balanced pick. Otherwise it says in plain words that the balanced pick is already the leanest sound choice.
Claude Haiku 5.5 added
Ranked from Artificial Analysis's results, and listed in Cursor and the Claude plans. Its FrontierCode result is Anthropic's own and is marked as maker-reported until an independent run exists.
One trusted source per benchmark, and more benchmarks per kind of work
When several labs publish the same benchmark, the most trusted one (its authors, or a lab that runs every model the same way) is used as published, and the others only fill in models it hasn't run. Several kinds of work gained benchmarks.
"New features" is now "Build a new service or API"
Its benchmarks measure brand-new code from a spec (an API, a service, a whole app), not a feature added to existing code.
The ranking follows "What matters most"
Fewest tokens, Balanced or Most reliable now sets the order of the whole ranking and each model's effort level, not only the pick.
Never pay tokens for a gain we can't measure
A higher effort level that improves the score by less than our back-tested error is not worth its extra tokens, so where a pick trades tokens for success it takes the leanest level within that error of the best.
The pick is a model you can get, and plans are ranked too
The pick must be in Cursor or on OpenRouter. "My plans" ranks for the plans you have (Claude, ChatGPT, Cursor, Copilot, Google AI), kept current from the vendors' own pages.
Your own costs
Your hourly cost, the minutes a failed attempt takes to fix, and the task size can be changed under "Your costs".
Artificial Analysis stopped publishing its Coding Index for new models
AA still runs the two benchmarks the index was made of, Terminal-Bench (now version 4.0) and SciCode. For a model without a published index we compute it the way AA did, from AA's own results for that model, marked ◇, and back-test how close that comes. Every update now also checks that each source still covers the newest models.
The pick is always a measured model
A model whose score is still an estimate can rank, but the pick is the cheapest one whose score was measured.
Each kind of work scored on several independent benchmarks
An earlier mix of general benchmarks agreed poorly with real codebase-understanding results, so each kind of work now uses benchmarks built for it.
Estimates go one step only
A model's missing score is estimated from the release right before it in the same line, if that one is measured, else from a fit across measured models. An estimate is never built on another estimate.
First public version
AI coding models ranked by cost per finished task, per kind of coding work, updated every 3 hours.
Spotted a wrong number? Write to istari-index@proton.me. See also how it's calculated.