eval / benchmark auditors; high-quality evaluations and independent checks on the claims built around them.
the idea is a company that designs and runs high-quality AI evaluations, and independently audits evaluations conducted by others. there are two problems in industry rn: flawed tests and the incentive to publish hyped figures. even technically excellent evaluations need outside scrutiny when the company selling the model also controls how its performance is presented. you wouldn’t trust a basketball player to honestly review their own shoe.
labs and companies deploying AI could commission either service for development, deployment decisions, or public claims. the work combines domain expertise, human review, precise rubrics, validated graders, and contamination and overfitting checks.
the components already exist. Mercor sells end-to-end evaluations; Vals AI conducts blinded domain-expert assessments; METR investigates evaluation integrity. OpenAI’s GDPval also used multiple expert-review rounds and blinded grading. expertise exists; independent scrutiny addresses the business incentives around it. i think Vals and METR are closest and could make a play here.
we leave technical gains on the table when we optimize for the wrong things. Scale’s ResearchRubrics found rubric wording affected agreement between human and automated judges. Mercor’s APEX update corrected grading that rewarded hedging with multiple answers. we need a company where the trajectory isn't just optimize ARR, datasets, or post-training scores, the things surrounding evals. they need to continually improve the eval itself.
evals as compliance could require independent review before public claims, covering task selection, scoring, excluded runs, and whether conclusions follow. research on selective disclosure found that letting providers choose which results to disclose can bias rankings. independent review would scrutinize both the evaluation and how its results are presented.
revenue would be like commissioned studies, recurring evaluation programs, and release-linked audits. the bet is that people will pay for specialist judgment and credible assurance. afterwards; automation could lower costs, while expert validation and defensible conclusions sustain quality, reputation, and the premium. acquisition isn't an out since selling out to a frontier lab would weaken trust in its audits.
i’m still thinking about two things; how auditors would get enough access to the full setup—customer environments, unreleased harnesses, full suite operational hygiene, as well as the consequence structure. we don’t want another consulting firm doing yearly audits no one cares about. certification needs to carry symbolic weight and meaningful financial consequences for earning it, losing it, or failing its standards.