Vibeleaderboard
← All Intel
Intel / article

SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch

Source
Yihao Chen, Shi Chang, Feng Lin, Khaled Chawa, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan
Author
Yihao Chen, Shi Chang, Feng Lin, Khaled Chawa, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan
Date
Key takeaways · AI-distilled
  • Building a program from scratch is where agents collapse. Given only documentation and an execute-only binary to probe for behavior, frontier models solve under 1% of ProgramBench instances, despite doing well when an existing codebase supplies .
  • The failure mode is the single pass that mixes reading docs, probing the binary, and writing code. Agents under-explore, lose the behavioral intent as context drifts, and bake early misreadings into the final implementation.
  • SpecFirst splits the job in two. A spec probes the binary and merges what it observes with the documentation into a structured written specification; a separate synthesis agent then implements against that fixed reference.
  • Across all 200 ProgramBench instances and four models, the split lifted test pass rates by 6.9 to 21.3 percent and binary exploration coverage by 9.4 to 18.5 percent, and agents began writing real code earlier and sustained it longer.
Terms in this piece · Glossary
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters

Agents are far weaker building from nothing than editing an existing repo, and spec elicitation is a concrete lever on that gap.

Key quotes

“LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder.”

“Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a behavioral oracle, even frontier models solve fewer than 1% of instances.”

“SpecFirst consistently outperforms the single-loop baseline, improving test pass rates by 6.9%-21.3% and binary exploration coverage by 9.4%-18.5%, all statistically significant.”

“Our results demonstrate that an explicit requirements-engineering phase is an effective paradigm for from-scratch program construction.”

Recommended reads
Comments

Checking sign-in…

Loading comments…