Can you trust your AI vendor's evaluations?
No, not on their own. The party selling you the AI has a structural conflict of interest: its evaluation drives its own consumption and grades the work it helped you deploy. A vendor scoring its own model is marking its own homework, and it will pass.
This is not an accusation of bad faith. It is a property of incentives. A vendor's built-in evals and ROI dashboards exist to show that the product is working, because that is what renews the contract. Even an honest vendor optimizes toward the number it can see you care about. The evaluation that decides whether you keep spending cannot be produced by the party that gets paid when you do.
What is an owned evaluation?
An owned evaluation is an eval you run on your own data, against criteria you define, on infrastructure you control, and keep invisible to the vendors you are evaluating. You own the golden datasets, the harness, and the number it produces. Nobody selling you AI can see it or game it.
The shift is from "trust our verdict" to "here is the machinery to produce your own." You decide what good looks like for your business, because only you can. The eval measures the outcomes you actually care about, on the data those outcomes live in, and the result belongs to you.
Vendor evals vs. owned evals, side by side
| Vendor / dashboard evals | Owned, private evals | |
|---|---|---|
| Who runs it | The company selling you the AI | You, on infrastructure you own |
| Incentive | Drive consumption; renew the contract | An honest number for your own decisions |
| What it measures | Model or token metrics the vendor picks | Your business outcomes, your criteria |
| Your eval criteria | Exposed to the vendor | Kept private; the vendor never sees them |
| Gaming risk | High: what is visible gets optimized | Removed: the target is not shared |
| Who owns the result | The vendor's platform | You do, and you can re-run it |
Why a metric your vendor can see is a metric your vendor will game
There are two independent failure modes. First, a shared target gets optimized: tell the vendor what you value and the number is engineered to look good. Second, your eval criteria are your intellectual property, and handing them over arms a party whose interests are not yours.
The first is Goodhart's law with a sales quota attached. When a measure becomes a target it stops being a good measure, and the vendor has every reason to hit the target you handed it rather than the outcome you meant.
The second is the one enterprises underrate. What "good" means in your business, the thresholds, the edge cases, the workflows, is institutional knowledge. It is the part a competitor could never buy. Satya Nadella calls this paying for intelligence twice, once with money and again with the proprietary knowledge you must reveal to make that intelligence useful. Sarah Guo puts it plainly: "the evaluation that decides real money is private and per-firm" (The Untrainable). Expose it, and you have taught your vendor your playbook. The public argument from the people building AI now runs in exactly this direction.
When do a vendor's evals actually make sense?
For developer-time, pre-ship testing of a model or feature in isolation, vendor and open-source eval tools are useful and often enough. They earn their place in CI. The problem is only when that number is asked to stand in for business value, or to decide budget.
Owned evaluation is not a replacement for unit-level model testing. It sits above it. Use the dashboards to ship; use an owned eval to know what the shipped thing is worth to you, and to make the spending decision on a number your vendor cannot see or move.
How to run your own private evals
The practical path is build, operate, transfer. Get your data evaluable, stand up the harness and the scoring scaffolding inside your perimeter, and fill it with what you consider good. Run it continuously, re-scoring on every model swap. And keep the whole thing owned by your team and invisible to the vendors you are evaluating. That is the capability PrivateEval builds and hands over: we do not evaluate your AI for you, we give you the private machinery to evaluate it yourself.