
Most optimization demos show you the win and hide the search; this one shows the search. An honest walkthrough of applying Karpathy's autoresearch idea to inference — read it for the method and the dead ends, not a tidy result.
“Because if you do not defend quality explicitly, an optimization harness will absolutely "improve" your system by making it worse.”
“If the agent can improve the score by shrinking the task, you are not optimizing inference. You are optimizing your ability to lie to yourself.”
“the hardest part of optimization is not generating ideas. It is building a harness that can tell the difference between a real win, a quality regression, and a benchmark illusion.”
“On the Qwen run, setting sampling to greedy decoding gave the largest gain: about +10.8% generation throughput.”
Checking sign-in…
Loading comments…