AI is getting cheaper, unless you measure what users actually pay
A quality-adjusted price index built from 21,000 AI pricing observations finds inference costs falling seven times faster than the standard method suggests — and, measured per completed task, not falling at all.
Ask how fast AI inference prices are falling and the honest answer depends entirely on how you count — and this paper shows that the standard way statistical agencies count software prices gets it badly wrong for AI, understating the real decline by a wide margin, while simultaneously hiding a countertrend that matters more for anyone actually budgeting for AI usage.
Three questions, one number, mostly confused
The paper builds a quality-adjusted price index for AI inference from a genuinely large panel: over 21,000 price observations spanning more than 3,200 models and 86 providers, joined to roughly 4,600 benchmark scores through a latent quality index estimated from benchmark response patterns — building the hedonic tradition's usual quality ladder out of evaluation results rather than product spec sheets, since "quality" for a model has no obvious physical characteristic to measure directly.
Measured the way statistical agencies measure ordinary software — matching like-for-like models over time — inference prices fell at a rate of 0.10 log points per year. Adjusted properly for the quality of what's actually being bought, the fall was 0.73 log points per year, over seven times faster. That gap means roughly 87 percent of the real price decline in AI inference has been invisible to conventional measurement — a genuinely large blind spot for anyone relying on posted per-token prices to judge how competitive or how fast-moving this market really is.
The countertrend that changes the practical answer
Here is where the paper earns its place beyond a methodological correction. Counted per completed task rather than per token, the price of AI inference has essentially stopped falling. The mechanism is straightforward once stated: reasoning models think for longer, consuming more tokens per task even as the price per token keeps dropping, and token-consumption growth has been outpacing token-price declines. The result is a genuine divergence between the seller's price — falling steadily, look at any provider's rate card — and the buyer's price, measured in what it actually costs to get a task done, which has flattened out.
Cheaper per token does not mean cheaper per task — a model that costs half as much per token but thinks five times longer can cost more to run than the one it replaced.
The audit that makes the numbers trustworthy
Benchmark contamination — models that have effectively memorized the test — is a standing worry for any quality measure built on evaluation scores, so the paper runs a pre-registered validity audit: excluding contamination-flagged benchmarks entirely. Model rankings stay almost perfectly intact, at 0.998 correlation, but the economic index still moves by 0.49 log points a year once those benchmarks are dropped. The two facts together are the paper's sharpest methodological point: rankings being stable under a robustness check is not, on its own, evidence that the economic statistics built on those same benchmarks are trustworthy. A leaderboard can be robust to contamination while an index built on top of it is not.
Honest caveats
The entire analysis runs on publicly posted prices and publicly available benchmark scores, which the author notes makes it fully reproducible from public sources at effectively zero cost — a real strength for verifiability, but it also means the index reflects list prices and public benchmarks, not negotiated enterprise rates, private evaluation suites, or the very largest deployments that often operate under different economics entirely. The quality index itself is a latent construct estimated from benchmark response patterns, which is a reasonable proxy but, like any latent quality measure, is only as good as the benchmark suite it's built from — a genuinely different mix of benchmarks could plausibly produce a somewhat different quality ladder and a different reading of the 0.73 log-point figure.
Why it matters
The practical lesson for anyone doing AI procurement, ROI modelling, or model routing is to stop treating price-per-token as the meaningful economic unit and start asking for price per completed task, per correct output, per accepted result — the denominators that actually determine whether a cheaper-looking model saves money once it's put to work. It also complicates any simple story about AI productivity or market concentration built on posted token prices alone: a market can look intensely price-competitive by the seller's own numbers while the buyer's realized cost per outcome has stopped improving, and getting that distinction backwards is an easy way to badly misjudge how much cheaper AI is actually getting.