Live LLM price index

The Task Cost Index

Model prices measured by finished work, not tokens.

Today's prices

Same tasks for every model. Each price is its real invoice divided by the work it finished.

bar length & colour = price quality = work finished ▲▼rank vs price per token

Quality is the share of work a model finished. Prices are real OpenRouter invoices, every token and retry counted.

How we appraise

Every price comes out of the same machine — a task worth a fixed amount of work, solved on a real invoice, graded pass or fail, then divided into one honest number.

  1. Task Worth a fixed number of work units. Fix a failing test5 WU
  2. Model Solves it; we keep the real invoice. spent $0.0021
  3. Grader Scores it 0 to 1, mechanically. passed · 5/5 WU
  4. Price Invoice ÷ work earned. $0.42 / 1k WU

Failure still costs. Miss the grader and the invoice stays on the bill with zero work earned — so the price climbs. A model that fails cheaply isn't cheap.

Your agent, your prices These prices are for our tasks. Describe the agent you actually run. We'll write a bespoke task book for it, price the whole frontier against it on real invoices, and email you a price list you can share. Price your own agent Free · a few minutes · no card