Minh Le

Writing

Essays, passing thoughts, and things to build.

Passing Thoughts

all passing thoughts

RLCD and the Jev moment

in my training-data primer, i was firm: synthetic data can scale expert judgment, but it can’t replace it when the work is ambiguous.

then Jev arrived. it makes structured decisions with probabilities instead of writing paragraphs. TypeSafe calls its training method RLCD: reinforcement learning for calibrated decisions. in TypeSafe’s framing, answers marked 80% likely should be right about 80% of the time across many decisions.

RLCD has shifted the paradigm, even while TypeSafe hasn’t published the specific reward or training loop that gets it there. and Jev’s founder says the model was trained entirely on synthetic data.

maybe i made “synthetic” do too much work. a model recycling its own answers is one thing. people designing the decisions, cases, and feedback a model learns from could be something else. if Jev holds up, expert judgment may have moved upstream, into deciding what’s worth teaching and how to tell whether the model learned it.

Jev is named for Jevons paradox: make a resource cheaper, and people may use more of it. if decisions get cheap enough to wire into every workflow, labor changes dramatically.

Things to Build

all things to build

eval / benchmark auditors; high-quality evaluations and independent checks on the claims built around them.

the idea is a company that designs and runs high-quality AI evaluations, and independently audits evaluations conducted by others. there are two problems in industry rn: flawed tests and the incentive to publish hyped figures. even technically excellent evaluations need outside scrutiny when the company selling the model also controls how its performance is presented. you wouldn’t trust a basketball player to honestly review their own shoe.

labs and companies deploying AI could commission either service for development, deployment decisions, or public claims. the work combines domain expertise, human review, precise rubrics, validated graders, and contamination and overfitting checks.

the components already exist. Mercor sells end-to-end evaluations; Vals AI conducts blinded domain-expert assessments; METR investigates evaluation integrity. OpenAI’s GDPval also used multiple expert-review rounds and blinded grading. expertise exists; independent scrutiny addresses the business incentives around it. i think Vals and METR are closest and could make a play here.

we leave technical gains on the table when we optimize for the wrong things. Scale’s ResearchRubrics found rubric wording affected agreement between human and automated judges. Mercor’s APEX update corrected grading that rewarded hedging with multiple answers. we need a company where the trajectory isn't just optimize ARR, datasets, or post-training scores, the things surrounding evals. they need to continually improve the eval itself.

evals as compliance could require independent review before public claims, covering task selection, scoring, excluded runs, and whether conclusions follow. research on selective disclosure found that letting providers choose which results to disclose can bias rankings. independent review would scrutinize both the evaluation and how its results are presented.

revenue would be like commissioned studies, recurring evaluation programs, and release-linked audits. the bet is that people will pay for specialist judgment and credible assurance. afterwards; automation could lower costs, while expert validation and defensible conclusions sustain quality, reputation, and the premium. acquisition isn't an out since selling out to a frontier lab would weaken trust in its audits.

i’m still thinking about two things; how auditors would get enough access to the full setup—customer environments, unreleased harnesses, full suite operational hygiene, as well as the consequence structure. we don’t want another consulting firm doing yearly audits no one cares about. certification needs to carry symbolic weight and meaningful financial consequences for earning it, losing it, or failing its standards.