Are your AI evaluation criteria intellectual property?
Yes. Your criteria encode how your business actually decides what good work looks like: the outcomes you reward, the thresholds you hold, the edge cases you have learned to catch, the way your workflows really run. That is proprietary knowledge, the kind a competitor could not buy, and it is exactly what leaks when you hand your evals to your AI vendor.
People underrate this because eval criteria look like plumbing. They feel like test cases, not strategy. But a rubric that captures how your best analysts judge a contract, or what your support org counts as a real resolution, is a distillation of years of institutional judgment. It is the untrainable part of your company written down.
What exactly leaks when you share your evals
The target and the map to it. Your vendor learns which outcomes you pay for, how you weight them, where your current system fails, and what a winning answer looks like in your domain. That is product roadmap, competitive intelligence, and training signal, all in one handover.
Satya Nadella framed the cost precisely: "you essentially pay for intelligence twice, once with money, and again with something even more valuable: the proprietary knowledge you must reveal to make that intelligence useful." That knowledge leaks trace by trace, correction by correction, eval by eval. The eval is not a side channel. It is one of the main channels through which your advantage flows out of the building.
Exposing your evals can arm a competitor
A vendor that sees enough of how you work is positioned to build the thing you were going to build, and your eval criteria are the most concentrated version of that knowledge. When your supplier is also a potential competitor, your criteria are the last thing you want to give it a clean copy of.
This is not hypothetical. The most-discussed recent case is Figma and Anthropic: Figma partnered with Anthropic, months later Anthropic shipped a competing product, Claude Design, and an activist investor called for a review of the relationship, including whether confidential information was involved. We cover it in full in when your AI vendor becomes your competitor. You do not have to assume bad faith to take it seriously. You only have to notice that a provider's incentives point at your market, and that every criterion you expose lowers the cost for it to get there.
How do you keep your AI eval criteria private?
You run the evaluation yourself, inside your own perimeter, on infrastructure you own, and you never route the criteria or the results through the vendor you are judging. The criteria live with you. The vendor sees a model's outputs being tested, never the standard it is being tested against.
This is the practical meaning of private, owned evaluation, and it is what PrivateEval builds: the machinery to evaluate your own AI, kept invisible to your vendors, and transferred to your team to own outright. We do not evaluate your AI for you, and we never see your criteria either. See what private AI evaluation is for the fuller picture, and Goodhart's Law for AI for why a visible target also gets gamed.