
We asked @TencentHunyuan authors why generating a harder coding task keeps failing. They said a useful terminal task has to keep four things consistent: the public instruction, the workspace, the official script, and the hidden tests. One mismatch can invalidate the entire task. > Script fails in the container > Tests check a missing file > Hidden rule the agent never sees Their paper grows a verified child from a parent. Official scripts went from 67 lines to 374 (5.6x). The public instruction only went 85 words to 122 (1.4x). A frozen DeepSeek-V4-Pro pass@4 fell from 90% to 2.5% on later rounds. Yield still held around 500 accepted tasks per 1,000 attempts. The extra work is real, and a longer README is the wrong knob. > When would you auto-generate shell training tasks? > If the official script passes, is that enough to train on? > Which exam is the factory never allowed to rewrite? Full Breakdown ↓↓

Quantifies why scaling difficulty in synthetic -training tasks breaks consistency between scripts, tests and instructions, showing a frozen model's pass rate collapsing from 90% to 2.5% as tasks got harder.
Checking sign-in…
Loading comments…