Why I couldn't build Jev at OpenAI — Diogo Almeida, TypeSafe Co-founder & CEO
Source
youtube.com
Author
Latent Space
Date
Why it matters
Presents a concrete argument that RLHFReinforcement learning from human feedback — training a model to prefer answers humans rate as better, which turns a raw text predictor into a usable assistant.Full definition → causes mode collapse and poor calibrationHow well a model's confidence matches reality — a calibrated model saying "90% sure" is right about 90% of the time.Full definition →, and that models embedded in software need different refusal and reliability behavior than chat models.
Key takeaways · AI-distilled
Almeida describes TypeSafe as a data lab rather than a model lab and says he would not pre-train a model from scratch even with a billion dollars.
His "bitterest lesson": the right task and the right data can matter more than simply scaling compute.
For builders, he recommends decomposing AI workflows into many small, measurable decisions and using structured state in place of giant prompts and system messages.
TypeSafe optimizes for intelligence per dollar and treats reliability and robustness as more important than simple determinismWhether the same input reliably produces the same output — something LLM systems mostly lack, which changes how you test and debug them.Full definition →.
Terms in this piece · Glossary
RLHF — Reinforcement learning from human feedback — training a model to prefer answers humans rate as better, which turns a raw text predictor into a usable assistant.
calibration — How well a model's confidence matches reality — a calibrated model saying "90% sure" is right about 90% of the time.
determinism — Whether the same input reliably produces the same output — something LLM systems mostly lack, which changes how you test and debug them.