Baseline report, branch per iteration, global memory file
From Agents Building Agents - Alfonso Graziano, Nearform · ≈14:48
“So, every iteration starts with a new branch. We create an hypothesis. Um the system changes the agent to implement that hypothesis. We run the evals. We run our eval suite. And we generate a reports.md file, which contains everything that happened after we run the evals.”
AI Engineer
“We update the memory file, uh which is like a global memory file across all the runs. And if the metrics improved, then we continue from this branch. Um if the metrics didn't improve or we have a strong regression or something bad happened, uh then we roll back to the previous branch.”
AI Engineer
“Of course, um the generated hypotheses are based on what the agent reads when it starts the investigation, which is the memory file, uh the reports file, so it has access to pretty much everything.”
AI Engineer
“What we can do is as humans, we can just go back into the hypothesis, read, understand what the agent was trying to do, and then maybe steer it in the right direction next time.”
AI Engineer
“And as we can see here, the baseline accuracy was 67% um but then in something around 10 iterations, we managed to reach 86% in our evals without actually cheating because it found edge cases, it improved the system prompt, it improved the tool descriptions to catch more edge cases, and it also fixed some tools logic.”
AI Engineer
- Concrete mechanics anyone can copy: generate a baseline report first, branch per hypothesis, write reports.md plus a cross-run memory file, keep the branch only if metrics improved.
Clip transcript
- clipA repeatable harness: frame metrics plus a human-calibrated judgeAI Engineer
- clipWhy CLAUDE.md and agents.md files fall shortAI Engineer
articleLoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding AgentHan Li, Zhemin Fang, Rili Feng, Yingqi Zhao, Jiaheng Liu, Pengfei Gao, He Ye, Dayi Lin, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
Checking sign-in…
Loading comments…