Why is AI ROI so hard to measure?
Because value shows up in outcomes that live in your own data, not in model benchmarks, and because the numbers you are handed come from parties with a stake in them being high. Public benchmarks do not measure your business, and your vendor's dashboard is built to justify its own spend. The gap is not model quality. It is trustworthy measurement.
Step 1: Define value in your own terms, not the vendor's
Start from the outcome, not the model. What does this AI actually change in the business: cost saved, revenue influenced, hours returned, risk reduced, errors avoided? Write it down as the specific thing your P&L would feel. Only you can do this, because only you know what good means for your business. A vendor's metric is a proxy chosen by the seller.
Step 2: Build ground truth on your data
A number is only as good as what it is measured against. Assemble golden datasets from your own data that capture what a good outcome looks like, including the edge cases your experts care about. This is the hard, valuable part, and it is where data readiness for AI pays off. Without ground truth you are evaluating against the vendor's assumptions.
Step 3: Measure the counterfactual, honestly
The real question is not what happened, it is what would have happened without the AI. Isolate the AI's contribution from price changes, seasonality, headcount, and everything else moving at once. A credible AI ROI number names its counterfactual and its assumptions. A vendor dashboard almost never does, because a clean counterfactual is where inflated claims go to die.
Step 4: Keep the measurement out of the vendor's hands
This is the step everyone skips. The moment your AI vendor can see the number you are optimizing, the number gets optimized to look good. The measurement that decides budget has to be owned by you, run on your infrastructure, and kept invisible to the vendors it judges. Otherwise it is marketing, not measurement.
This is Goodhart's Law applied to your P&L, covered in Goodhart's Law for AI. A vendor-produced ROI number is conflicted by construction; an owned one is not. See vendor evals vs. owned evals for why the party that gets paid when the number is high cannot be the party that produces it.
Step 5: Re-run it every quarter, on a loop you own
AI ROI is not a one-time slide. Models change, usage changes, and last quarter's number decays. Make the measurement a running loop that re-scores on every model swap and produces a fresh, defensible readout each quarter. Own the loop, and the number stays yours and stays current.
What a defensible AI ROI number looks like
It is measured on your data, against criteria you defined, with a named counterfactual, on infrastructure you own, and it is invisible to the vendors it judges. That is the difference between a number you can take to the board and one your vendor gave you. Building that capability is what PrivateEval does: we do not produce your number, we give you the machinery to produce it yourself, and hand you the keys.