SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch
- Source
- Yihao Chen, Shi Chang, Feng Lin, Khaled Chawa, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan
- Author
- Yihao Chen, Shi Chang, Feng Lin, Khaled Chawa, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan
- Date

- Building a program from scratch is where agents collapse. Given only documentation and an execute-only binary to probe for behavior, frontier models solve under 1% of ProgramBench instances, despite doing well when an existing codebase supplies .
- The failure mode is the single pass that mixes reading docs, probing the binary, and writing code. Agents under-explore, lose the behavioral intent as context drifts, and bake early misreadings into the final implementation.
- SpecFirst splits the job in two. A spec probes the binary and merges what it observes with the documentation into a structured written specification; a separate synthesis agent then implements against that fixed reference.
- Across all 200 ProgramBench instances and four models, the split lifted test pass rates by 6.9 to 21.3 percent and binary exploration coverage by 9.4 to 18.5 percent, and agents began writing real code earlier and sustained it longer.
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
- AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Agents are far weaker building from nothing than editing an existing repo, and spec elicitation is a concrete lever on that gap.
“LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder.”
“Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a behavioral oracle, even frontier models solve fewer than 1% of instances.”
“SpecFirst consistently outperforms the single-loop baseline, improving test pass rates by 6.9%-21.3% and binary exploration coverage by 9.4%-18.5%, all statistically significant.”
“Our results demonstrate that an explicit requirements-engineering phase is an effective paradigm for from-scratch program construction.”
articleCodeSpec: Dual Executable Specifications for Agentic Long-Horizon Feature DevelopmentPeiding Wang, Li Zhang, Fang Liu, Taichuan Li, Yinghao Zhu
articleMindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program SynthesisYihao Chen, Shi Chang, Khaled Chawa, Feng Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan
articleSpecification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code reviewJoel Abenhaim
Checking sign-in…
Loading comments…