# Istari: method

Updated 2026-10-11 18:05 UTC. Short answers are in https://istari-index.pages.dev/llms.txt; this file explains where the numbers come from.

## Cost per finished task

A model that fails more costs more: you pay for the retry and for your time spotting and fixing the failure.

- Success chance: p = coding score / 100.
- Tokens per attempt (millions) = 0.27 × (1 − 0.4 + 0.4 × thinking multiplier).
- Fix cost per failed attempt = $93/hour × 20 minutes = $31.
- Cost per finished task = (price per 1M tokens × tokens per attempt + fix cost × (1 − p)) / p.

Constants and where they come from:

- 0.27M tokens per task: Artificial Analysis's Coding Agent Index measured about $4.10 per task for Claude Opus 4.7 and $4.82 for GPT-5.5 at list prices, about 0.27M tokens per real agent task.
- $93/hour: US software developer median wage (~$65/hour, BLS) divided by 0.70, the wage share of total employer cost (BLS ECEC).
- 20 minutes per failed attempt: most failures are fixed by a correction prompt (91.5% in a 20,574-session study), ~12 minutes to spot, re-prompt and re-review; the rest need hands-on fixing (METR: 26-42 minutes per agent PR).
- Reasoning share 0.4: the part of a task's token bill that is the model's own thinking and answer, which grows with effort; reading code and context doesn't.
- Price per 1M tokens: the average of the input and output list price.

## Effort levels

Each effort level is scored and priced separately, and a model is ranked at the level that is cheapest per finished task. More effort raises the success chance but also the thinking tokens. The thinking multiplier (1 = "high") comes from Artificial Analysis's measured latency × speed where available, else a ladder: non-reasoning 0.15, minimal 0.2, low 0.35, medium 0.6, high 1, xhigh 1.7, max 2.5.

## Coding score and kinds of work

- Coding score: the Artificial Analysis Coding Index (about ⅔ Terminal-Bench, ⅓ SciCode). For new models, where AA no longer publishes it, we compute it the same way from AA's current results for that model (◇, measured). A model needs at least 69.5 (85% of the best) to be ranked.
- Kinds of work: each is scored on benchmarks built for it, from several independent publishers (ranking.md lists them). For one kind of work, a model's score shifts by how much better or worse it does there than its general coding score.
- Brand-new models without results (*): estimated from the release right before them in the same line at the same effort level, carrying ¾ of the typical gain for their rise in Intelligence Index; never below that predecessor. One jump only. Back-tested typical miss about ±3.7 points. Replaced automatically once measured.
- Task scores a model has no result on yet are estimated the same way from its line's predecessor, and marked "estimated". The pick is always measured, and in Cursor or on OpenRouter.
- Gaps we can't measure aren't worth tokens: a ◇ score measured at its level misses AA's own by about ±2.6 points (back-tested). Wherever a pick trades tokens for success (the most accurate pick, and every pick in a plan whose tokens are included), it takes the leanest level within that error of the best, not the top at any price. The fewest-tokens pick stays above the quality bar for that kind of work.

## Recomputing

Example: MiMo-V2.6-Pro at default effort has a coding score of 77.5 (p = 0.775), $0.6525 per 1M tokens and a thinking multiplier of 1, so 0.2700M tokens per attempt. At our $93/hour: $9.23 per finished task. At $60/hour, the fix cost is $20.00 and the cost is (0.6525 × 0.2700 + 20.00 × 0.225) / 0.775 = $6.03. To re-rank for another rate, recompute every ranked level of every model in ranking.json the same way and take each model's cheapest level. Only the fix cost changes.

## Limits

- Benchmarks are a proxy for your codebase; treat costs within ~10% of each other as a tie.
- Prices are list prices; cached-input discounts, subscriptions and tool plans (Cursor, Claude Code, Codex) change the real bill.
- Availability in Cursor comes from Cursor's public model list and can lag a few days.

Data: Artificial Analysis and other public benchmarks, with their own licenses. Istari is independent and not affiliated with any model maker, Cursor (Anysphere), or Artificial Analysis.
