Vol. XVI · No. 265Tuesday 22 September 2026World Edition
TheNewsRupt coat of arms crest

The NewsRupt

Explainer

Inference Costs Are Falling Faster Than Enterprise Budgets Can Adjust

The price of running a model has dropped sharply, but most procurement cycles were written for last year's numbers.

By The NewsRupt Desk·San Francisco desk·Tuesday 22 September 2026·8 min read

When a model becomes cheaper to run, the saving does not arrive as a refund. It arrives as a change in what is worth attempting, and most organisations are slower to notice the second thing than the first.

Over the past two years the cost of serving a token from a competent general-purpose model has fallen by roughly an order of magnitude on most public price lists, and further still for teams willing to use smaller models, batch their requests, or cache repeated context. The direction of travel is not in dispute. What is in dispute, inside a great many companies, is what to do about it.

The reason is procedural rather than technical. Enterprise budgets are written annually, sometimes eighteen months ahead of the quarter they govern. A budget drafted against last year's per-token pricing bakes in an assumption about unit economics that is already wrong by the time the money is released. Teams then spend the year defending a number that was conservative when written and is now absurd, while the finance function reads underspend as evidence that the initiative was oversold.

The visible symptom is a class of project that gets rejected for the wrong reason. A support-triage system that looked marginal at last year's prices is comfortably positive at this year's, but the business case on file still carries the old figure, and nobody has been given the job of re-running it. The same applies in reverse: a pilot that was approved on optimistic volume assumptions quietly consumes its budget in a month because nobody modelled what happens when usage succeeds.

There is a second, less obvious effect. Falling prices change the economics of redundancy. When inference was expensive, running the same request through two models and comparing outputs was an indulgence. At current prices, for a workflow where a wrong answer is costly, it is often the cheapest form of quality assurance available. Teams that have internalised the new price level tend to spend more on verification and less on the single best model, which is a different shape of budget entirely, not merely a smaller one.

Three practical observations recur among teams that have adjusted well.

The first is that they price per completed task, not per token. Token cost is an input; what the business cares about is the cost of a resolved ticket, a reviewed contract, a drafted brief. A model that costs twice as much per token but needs one attempt rather than four is cheaper in the only unit that matters, and the reverse holds just as often.

The second is that they treat the cost line as variable and review it quarterly rather than annually. This is unusual in most finance functions and takes some negotiating, but it matches the volatility of the underlying market far better than a fixed annual allocation.

The third is that they separate experimentation budget from production budget. Experimentation is small, uncapped within a ceiling, and deliberately wasteful, because the purpose is to find out what works. Production is metered, monitored and cost-attributed to the team consuming it. Conflating the two produces the familiar pattern where an experiment is starved to protect a production line that is itself under-instrumented.

The limitations of this picture are worth stating plainly. Published list prices are not what large customers pay; negotiated commitments, reserved capacity and committed-spend discounts make real prices opaque, and the public curve may overstate the saving available to a company already on a discounted contract. Price per token also excludes the substantial costs that surround a deployment — retrieval infrastructure, evaluation, human review, security review, integration work — and those have not fallen at anything like the same rate. In several deployments we have looked at, the model itself is a minority of total cost. Finally, prices are set commercially, not by physics; the current curve reflects a period of aggressive competition and heavy capital availability, and it is a forecast, not a guarantee, that it continues.

What is safe to say is narrower and still useful. The cost of the model is no longer the binding constraint on most enterprise AI work. The binding constraints are evaluation, data access, ownership of the output, and the organisational patience to run something for long enough to know whether it worked. Budgets that continue to treat inference cost as the main risk are managing the cheapest part of the problem with the most attention.

The practical test is simple enough to run this week. Take the three AI proposals most recently declined on cost grounds, re-price them at current rates, and see whether the answer changes. In our conversations with engineering and finance leaders, it frequently does — and the fact that nobody had thought to check is the more interesting finding.

How we report

Every claim above is sourced to a document, a named person, or a record we hold. Where we could not verify a claim, we say so. Read our standards and corrections policy →