← All IntelClip / Developer ToolsProposed fix: human-authored, behavior-focused instructions
From The Good, the Bad, and the Ugly: Why Coding Benchmarks Are Broken · ≈8:32
“The instructions given to an agent or an LLM should lean towards expressing desired behaviors, objectives, and hard constraints, not implement details or try to guarantee self-containment when the task itself is is expressing too much uh too much details.”
What’s in it
- Proposes a framework for writing better AI coding benchmark tasks
- Argues task instructions should state goals, not implementation details
- Built with G2i team over two months of research
Clip transcript
So, how do we close the gap? Um in the last 2 months, we've been working with our team at G2i to basically try to define a framework, uh a set of principles that would allow us to build tasks for benchmarks that are um better than what we have today. The first one, human instructions. Authored by humans, reviewed by humans. This is basically the entry point for any great tasks. The instructions given to an agent or an LLM should lean towards expressing desired behaviors, objectives, and hard constraints, not implement details or try to guarantee self-containment when the task itself is is expressing too much uh too much details.
Comments
Sign in to comment.
Loading comments…