Home/AI cost per outcome: how to price the work, not the tokens
Guide

AI cost per outcome: how to price the work, not the tokens

AI cost per outcome is the unit finance will actually underwrite. How to define the outcome, do the arithmetic, and defend the number in a budget review.

Tokens are a supplier invoice, not a business case

Most AI budgets die in the same meeting. Someone brings a cloud bill showing model spend, someone else brings a slide showing hours saved, and the CFO cannot reconcile the two because they are measured in different currencies. Tokens consumed is a supplier metric. It tells you what your vendor charged. It says nothing about whether the work got done, whether anyone accepted the result, or whether the business is better off. Cost per outcome closes that gap by putting spend and value in the same unit: what did one completed, accepted piece of work cost us to produce with AI, and what did it cost before? Until you can answer that in one number, every AI line item is discretionary, and discretionary lines get cut first in a flat year.

Define the outcome before you measure anything

The outcome has to be something the business already counts. A resolved support ticket. A settled claim. An onboarded supplier. A reconciled invoice. A drafted contract that a lawyer signed off without a rewrite. If finance does not already have a row for it in a system somewhere, pick something else, because you will spend the next quarter arguing about the denominator instead of the spend. Two rules keep this honest. First, the outcome must be terminal, meaning the work is genuinely finished and nobody downstream has to redo it. A summary that a human rewrites is not an outcome, it's a draft. Second, the outcome must have a pre-AI baseline cost, even a rough one, or you have no comparison to make. Hours saved fails both tests, which is why it never survives contact with a budget review. Nobody at the executive table has a P&L line called hours.

The arithmetic, with your own numbers

The formula is short. Total AI cost for the period, divided by accepted outcomes in that period. Total cost is not just inference. It's inference plus retrieval and vector storage, plus the human review time you still pay for, plus the amortised build and integration cost, plus the observability and evaluation tooling you run to keep the thing trustworthy. Run it with illustrative figures to see how it behaves. Say a claims triage workflow costs 4,000 dollars a month in model and retrieval spend, plus 60 hours of adjuster review at a loaded 70 dollars an hour, which is 4,200 dollars, plus 2,000 dollars a month of amortised build over a two year life. That's 10,200 dollars. If it produced 3,000 accepted triage decisions, cost per outcome is 3.40 dollars. Now compare that to the fully loaded manual cost of the same decision. The comparison is the whole point, and notice what dominates: the model spend was 39 percent of the total. Teams that optimise the model bill are usually optimising the smallest term.

Acceptance rate is the term that moves the number

Because the denominator is accepted outcomes rather than attempted ones, quality shows up directly in the cost. An agent that produces 3,000 outputs at 85 percent acceptance and one that produces 3,000 at 60 percent acceptance have very different unit economics on the same model bill, and the second one also generates rework nobody is costing. That's the argument for putting evaluation and human oversight in from the start rather than bolting them on. They are not overhead against the unit cost. They are what stops the denominator collapsing. It's also why cheaper models sometimes make cost per outcome worse. A model that halves your inference bill and drops acceptance by fifteen points has raised the price of the work.

Year two is where the ratio breaks

Build cost has fallen fast. Run cost has not. A workflow you shipped for a fraction of what it would have cost three years ago still has to be operated, monitored, re-evaluated after every model version change, and re-certified when the data behind it moves. Budgets written for the old ratio, heavy build and light run, are the ones that blow up in the second year. When you model cost per outcome, model it across at least eight quarters and put a line in for model migration, because you will do at least one. Teams who have run AI at 300,000 organisation scale plan for the version change as a scheduled cost, not an incident. The organisations that get surprised are the ones who treated the first year's number as the steady state.

What governance actually buys you here

There is a version of this that reads as pure finance, and it misses something. You cannot compute an honest cost per outcome without the control layer, because the inputs come from it. Attribution of spend to a use case comes from tagged, observable traffic. Acceptance rate comes from recorded human oversight decisions, the kind EU AI Act Article 14 expects for high risk systems anyway. Rework and incident cost come from the audit trail and post market monitoring under Article 72. The NIST AI RMF puts MEASURE alongside GOVERN, MAP and MANAGE for the same reason: you cannot manage what you have not instrumented. Control is the lens that makes the value story provable rather than asserted. Teams that build the rails to prove the number to a regulator find they can also prove it to a CFO, using the same evidence.

Where to start on Monday

Pick one workflow already in production. Not the flagship demo, the boring repetitive one. Write down its outcome definition in a sentence a finance analyst would accept, then find the system of record that counts it. Pull ninety days of total cost against that outcome, including the human review hours people forget. Get the acceptance rate from whoever reviews the output, even if the first version is a manual sample of 100 cases. You will have a defensible cost per outcome in about two weeks, and it will almost certainly be wrong in an interesting way: usually the human review term is larger than anyone guessed. Fix that, and you have both a cheaper unit and a better case for the next use case. If you want a second pair of eyes on which workflow to instrument first, that's what the Free AI Value & Readiness Assessment is for.

AI Cost Per Outcome: The Only Unit That Survives a CFO Review