✦ In short
Cost per finished task = the tokens a model uses at its effort level, plus your time fixing the attempts it gets wrong, counted over the retries it takes. The chance of getting a task right comes from independent benchmarks of real coding tasks, not from model makers.
Updated Oct 11, 2026. 41 models ranked. Every number below is recomputed every 3 hours from the same data you can download.
What counts as a finished task
A task is finished when the benchmark's own grader passes it: the hidden tests pass, the root cause named is the real one, the answer about the codebase is judged correct. We don't run the tasks ourselves. Independent labs do, the same way for every model, and publish the results; each benchmark is listed below with what it measures and how many tasks it has, where its publisher says.
We read a model's score as the chance it gets your task right the first time (81 ⇒ 81%). Each miss costs you a retry's tokens and your time to spot it, re-prompt and fix it.
The formula
cost per task = (price × tokens + fix cost × failure chance) ÷ success chance
Dividing by the success chance counts the retries. The constants, each with its source:
- 270K tokens a task. Artificial Analysis measured real coding-agent tasks at $4.10 for Claude Opus 4.7 and $4.82 for GPT-5.5; at their list prices that's about 270K tokens. Source
- 40% of those are the model's own thinking and answer, the part that grows with the effort level; reading code and context doesn't. Source
- $93 an hour. A US software developer's median wage, about $65, is about 70% of what the hour costs an employer. BLS wages · BLS employer costs
- 20 minutes a failed attempt ($31). In 20,574 real sessions, 91.5% of problems were fixed by correcting the agent, about 12 minutes; the rest need hands-on work, 26 to 42 minutes. Sessions study · METR
- Price: the average of the input and output list price per million tokens.
Your hourly cost, the minutes and the task size can be changed under "Your costs" on the ranking; the ranking is recomputed for you.
Check it yourself. Example: MiMo-V2.6-Pro at default effort has a coding score of 77.5 (p = 0.775), $0.6525 per 1M tokens and a thinking multiplier of 1, so 0.2700M tokens per attempt. At our $93/hour: $9.23 per finished task. At $60/hour, the fix cost is $20.00 and the cost is (0.6525 × 0.2700 + 20.00 × 0.225) / 0.775 = $6.03. To re-rank for another rate, recompute every ranked level of every model in ranking.json the same way and take each model's cheapest level. Only the fix cost changes.
Effort levels
Many models have an effort setting (low, medium, high, xhigh, max): how long the model thinks before it answers. A higher level uses more tokens for a somewhat better success rate, so each level is scored and priced separately and a model is ranked at the level that is cheapest per finished task. How much a level thinks comes from Artificial Analysis's timings where it has them, else a typical ratio.
The coding score
The general coding score is Artificial Analysis's. A model must reach 69.5, 85% of the best (Claude Opus 5.5, 81.8), to be ranked.
- Artificial Analysis Coding Index: the general coding score where AA published it (about ⅔ Terminal-Bench, ⅓ SciCode).
- Terminal-Bench 4.0 · Vals AI; models it hasn't run filled in from Artificial Analysis. 66 hard tasks in a real terminal: system administration, builds, environments, debugging, data processing. Measured for 38 of the 41 ranked models, results as of Oct 11, 2026.
- SciCode · Artificial Analysis (SciCode authors). 288 subproblems from 80 real laboratory problems across 16 scientific disciplines. Measured for 38 of the 41 ranked models.
Artificial Analysis stopped publishing its Coding Index for new models in September 2026. It still runs Terminal-Bench (now in its harder version 4.0) and SciCode at every effort level, so for those models we compute the index the way AA did, from AA's own results for that model, and mark it ◇. It's a measurement of that model, not an estimate from other models. If AA publishes the index again, it replaces ours.
Benchmarks for each kind of work
Pick a kind of work on the ranking and each model is scored on benchmarks built for it. When a kind of work has several, each is put on the scale of the first one, and a model's score is the mean of those it was measured on. When several labs publish the same benchmark, the most trusted one's numbers (its authors, or a lab that runs every model the same way) are used as published, and another only fills in models it hasn't run.
Build a new service or API
Brand-new code from a spec: an API, a service, a module or a whole app. Past estimates for this work, once measured, were off by ±2.4 points on average (1 checked).
- DeepSWE v1.1 · Datacurve; models it hasn't run filled in from Artificial Analysis (agent runs). 113 original tasks across 91 real repositories in 5 languages, each requiring a large, novel change (not taken from public commits); pass@1 on hidden tests. Measured for 35 of the 41 ranked models, results as of Oct 11, 2026.
- FrontierSWE · its authors (via Epoch AI). Long, open-ended engineering projects (building, speeding up and researching real software), hours each; mean of 5 runs in one harness. Measured for 18 of the 41 ranked models, results as of Oct 11, 2026.
- CursorBench 4.0 · Cursor. Ambiguous multi-file tasks from real Cursor sessions: edit, refactor, investigate, follow intent and design. Measured for 16 of the 41 ranked models, results as of Oct 11, 2026.
- Vibe Code Bench · Vals AI. Build a complete web application from a spec, end to end; scored by automated tests of the running app. Measured for 37 of the 41 ranked models, results as of Oct 7, 2026.
Change existing / legacy code
Adapt, refactor or migrate a large existing codebase. Past estimates for this work, once measured, were off by ±10.2 points on average (3 checked).
- Code Migration · Vals AI. Rewrite 30 working open-source programs into other languages (Python, Java, Kotlin, Rust, C++) and COBOL into Java; scored by hidden behaviour tests. Measured for 36 of the 41 ranked models, results as of Oct 8, 2026.
- FrontierCode · Cognition (via Epoch AI). Whether a model’s change to a real codebase could be merged without human edits. Measured for 26 of the 41 ranked models, results as of Oct 11, 2026.
- SWE Atlas Refactoring · Scale AI. 70 tasks: restructure real code (decompose, extract, move, evolve interfaces) while keeping its behaviour, graded on maintainability and cleanup too. Measured for 8 of the 41 ranked models.
Explore & explain a codebase
Trace flows across files and repos, onboard, answer "how does this work", review. Past estimates for this work, once measured, were off by ±6.9 points on average (13 checked).
- SWE-Atlas Codebase QnA · Scale AI; models it hasn't run filled in from Artificial Analysis (agent runs). 124 deep code-comprehension questions on 11 production repos (Go, Python, C, TypeScript): tracing execution across files, explaining behaviour and design. Agents must run and explore the code. Measured for 26 of the 41 ranked models, results as of Oct 11, 2026.
Debug, fix & test
Reproduce, locate and fix bugs; write and repair tests. Past estimates for this work, once measured, were off by ±7 points on average (2 checked).
- DeepSWE v1.1 · Datacurve; models it hasn't run filled in from Artificial Analysis (agent runs). 113 original tasks across 91 real repositories in 5 languages, each requiring a large, novel change (not taken from public commits); pass@1 on hidden tests. Measured for 35 of the 41 ranked models, results as of Oct 11, 2026.
- Terminal-Bench 4.0 · Vals AI; models it hasn't run filled in from Artificial Analysis. 66 hard tasks in a real terminal: system administration, builds, environments, debugging, data processing. Measured for 38 of the 41 ranked models, results as of Oct 11, 2026.
- SWE Atlas Test Writing · Scale AI. 90 tasks: write production-grade tests for real codebases. Measured for 11 of the 41 ranked models.
Production investigation
Alerts, events and traces across services to find the root cause. Little data: measured for under 60% of the ranked models. Past estimates for this work, once measured, were off by ±8.2 points on average (11 checked).
- ITBench-AA · Artificial Analysis. Root cause of real Kubernetes incidents from alerts, events, traces and topology snapshots, with shell access. Scored as precision at full recall (missing any true cause scores 0). Measured for 20 of the 41 ranked models, results as of Oct 11, 2026.
Infra, DevOps & scripting
CI/CD, containers, shell, builds and environment setup. Past estimates for this work, once measured, were off by ±5.1 points on average (2 checked).
- Terminal-Bench v2.1 · Artificial Analysis (tbench.ai tasks). 89 tasks in a real terminal: system administration, builds, environments, data processing, security. Measured for 27 of the 41 ranked models.
- Terminal-Bench 4.0 · Vals AI; models it hasn't run filled in from Artificial Analysis. 66 hard tasks in a real terminal: system administration, builds, environments, debugging, data processing. Measured for 38 of the 41 ranked models, results as of Oct 11, 2026.
Algorithms, data & scientific code
Numerical code, data processing, algorithms.
- SciCode · Artificial Analysis (SciCode authors). 288 subproblems from 80 real laboratory problems across 16 scientific disciplines. Measured for 38 of the 41 ranked models.
- IOI · Vals AI. International Olympiad in Informatics problems: hard algorithms under tests. Measured for 29 of the 41 ranked models, results as of Oct 7, 2026.
- ALE-Bench · Sakana AI (via Epoch AI). AtCoder Heuristic Contest problems: long optimisation tasks, scored as a contestant performance rating. Measured for 26 of the 41 ranked models, results as of Oct 11, 2026.
- Terminal-Bench Science · Vals AI. Research workflows contributed by practising scientists in five fields, done in a real terminal. Measured for 30 of the 41 ranked models, results as of Oct 8, 2026.
UI & web pages
Front-end pages, components and web apps. Past estimates for this work, once measured, were off by ±16.6 points on average (3 checked).
- LMArena Code Arena – WebDev · LMArena. Blind pairwise votes by developers on web apps built by two anonymous models (Bradley-Terry rating). Measured for 37 of the 41 ranked models, results as of Oct 8, 2026.
- Vibe Code Bench · Vals AI. Build a complete web application from a spec, end to end; scored by automated tests of the running app. Measured for 37 of the 41 ranked models, results as of Oct 7, 2026.
Security & vulnerabilities
Find and fix vulnerabilities without breaking the code. Past estimates for this work, once measured, were off by ±9.1 points on average (3 checked).
- CWE-Bench-AA · Artificial Analysis. 120 held-out tasks: find and fix OWASP Top-10 vulnerabilities in a repository without breaking its behaviour (pass@1). Measured for 15 of the 41 ranked models, results as of Oct 11, 2026.
- Vals Cyber · Vals AI. Craft inputs that trigger real OSS-Fuzz vulnerabilities (mostly C/C++ memory safety), then patch the source so they no longer do. Measured for 29 of the 41 ranked models, results as of Oct 7, 2026.
- CyberGym-E2E-AA · Artificial Analysis. Find real memory-safety vulnerabilities in C/C++ open-source projects and prove them with a working input (Berkeley RDI's CyberGym-E2E, run by Artificial Analysis). Measured for 15 of the 41 ranked models, results as of Oct 11, 2026.
- DeepsecBench-AA · Artificial Analysis. Find vulnerabilities in real open-source application code (Vercel's DeepsecBench, run by Artificial Analysis). Measured for 16 of the 41 ranked models, results as of Oct 11, 2026.
- Reverse engineering (SRE Bench) · Vals AI. Reverse engineer real, contamination-free binaries: work out what compiled programs do. Measured for 26 of the 41 ranked models, results as of Oct 8, 2026.
Estimates, and how far off they have been
Where a model has no result yet, its score is estimated and marked. Every error below is a back-test: we hide a known result, predict it, and compare.
- ◇ measured, computed from AA's current results: rebuilding the 62 models whose index AA did publish, each with its own index hidden, comes within ±2.6 points on average. An effort level AA didn't run Terminal-Bench at is inferred from a higher level of the same model: ±4 points (43 levels checked). Such a level never stands in for a better-measured one.
- * provisional (a model in its first days): estimated from the release right before it in the same line; hiding recent models' real scores, it misses by ±3.7 points on average (25 checked), or ±3.8 when estimated from similar models (77 checked). Replaced as soon as it's measured.
- † estimated by hand: a model Artificial Analysis's data doesn't include. Never the pick.
- Scores for a kind of work a model hasn't been measured on are estimated from the release right before it in the same line, if that one was measured, else from the benchmark's fit on the coding score. One step only: an estimate is never built on another estimate. Each kind of work above lists how far past estimates were off.
The pick is measured, and you can get it. It's always a model whose score was measured, in Cursor or on OpenRouter. One exception: on a kind of work measured for under 60% of the ranked models (today: production investigation), the same one or two measured models would win whatever you chose, so there an estimated task score can be the pick. The card then says so, with how far past estimates have missed, and names the other contenders.
No tokens for a gain we can't measure. A higher effort level that adds less than our back-tested error (±2.6 points) buys nothing we can measure, so "Most reliable", and every pick in a plan whose tokens are included, take the leanest level within that of the best. Models within it of each other are a tie, ordered by price.
Benchmarks we checked and left out
A benchmark we rank on must still be run on new models, and must tell good models apart. Every update checks that each source still covers the newest releases, and watches publishers for new benchmarks. Some well-known ones we decided against:
- SWE-bench Verified: contaminated, and saturated at the top (97%).
- SWE-bench Pro V2: 99% at the top, each maker's own agent, and its results disagree with DeepSWE and Code Migration across the same models.
- SWE-rebench: its window of fresh tasks ended in July 2026.
- Aider Polyglot, LiveCodeBench, Cybench, AlgoTune, METR time horizons: no longer run on new models.
- Hugging Face leaderboards: self-reported, and almost no closed models.
- An earlier mix of general benchmarks per kind of work: checked against SWE-Atlas, it agreed with real codebase-understanding results at only r = 0.34, worse than the plain coding score (0.61). So each kind of work now uses benchmarks made for it.
How Istari relates to the leaderboards you may know: Artificial Analysis is our main source, and LMArena's WebDev arena is one of the UI benchmarks. What we add is the price, the tokens at each effort level, the retries and your time, per kind of work.
What we can't tell you
- Benchmarks aren't your codebase. A model that shines on public tests can stumble on your stack. Use Istari to shortlist, then try the top picks on a few of your own tasks.
- Averages are averages. Treat costs within about 10% of each other as a tie.
- Prices are list prices. Caching discounts and tool plans change the real bill; "My plans" on the ranking covers the plans.
Check our math
- data.json: everything the ranking is built from, every model and effort level.
- llms.txt: the ranking as plain text, with the pick for each kind of work and each tool.
- method.md and ranking.json: this method and the ranking in compact form.
- Changes and corrections: every change to the method, with its date.
Spotted a wrong number? Write to istari-index@proton.me. Corrections are logged.
Questions
What is a "finished task"?
A task the benchmark's own grader passes: the hidden tests pass, the root cause named is the real one, the answer about the codebase is judged correct. Independent labs run the tasks, the same way for every model, and publish the results; Istari doesn't run them itself.
Where do the scores come from?
From independent benchmarks: Artificial Analysis for the general coding score, and benchmarks built for each kind of work from Artificial Analysis, Vals AI, Epoch AI, Scale AI, LMArena and the benchmarks' own authors. A model maker's own number is used only until an independent result exists, and is marked.
What does ◇ mean? Is that score estimated?
No, it's measured. Artificial Analysis stopped publishing its Coding Index for new models in September 2026 but still runs the benchmarks it was made of. We combine AA's own results for that model the way AA did. Rebuilding models whose index AA did publish, the same way, comes within about 2.6 points of it.
Why $93 an hour and 20 minutes a fix?
$93 is what a US software developer's hour costs an employer (BLS). 20 minutes is a typical time to spot, re-prompt and fix a failed attempt (from a study of 20,574 sessions and METR's measurements). Both are yours to change under "Your costs" on the ranking.
Who runs Istari?
An independent project by developers, unsigned on purpose so no model maker or sponsor can lobby it. It takes no money from anyone it ranks. Instead of a name, every number is sourced, the data is downloadable, and every change to the method is logged with its date.