RLCD and the Jev moment
in my training-data primer, i was firm: synthetic data can scale expert judgment, but it can’t replace it when the work is ambiguous.
then Jev arrived. it makes structured decisions with probabilities instead of writing paragraphs. TypeSafe calls its training method RLCD: reinforcement learning for calibrated decisions. in TypeSafe’s framing, answers marked 80% likely should be right about 80% of the time across many decisions.
RLCD has shifted the paradigm, even while TypeSafe hasn’t published the specific reward or training loop that gets it there. and Jev’s founder says the model was trained entirely on synthetic data.
maybe i made “synthetic” do too much work. a model recycling its own answers is one thing. people designing the decisions, cases, and feedback a model learns from could be something else. if Jev holds up, expert judgment may have moved upstream, into deciding what’s worth teaching and how to tell whether the model learned it.
Jev is named for Jevons paradox: make a resource cheaper, and people may use more of it. if decisions get cheap enough to wire into every workflow, labor changes dramatically.
