
If you're building or evaluating self-improving coding agents, this gives you a concrete checklist (Dependency, Activation, Evidence, Retention, Authority, Recovery, Value) for proving a change actually caused the gain instead of accepting false-positive credit assignment.
“Weco reports that AIDE2 executed 100 consecutive harness rewrites over 8 unattended days, retaining 7 successive versions under a fixed evaluation budget.”
Adham Khaled
“Recursive self-improvement begins when an artifact produced at generation t changes how generation t+1 is produced, judged, selected, or executed.”
Adham Khaled
“My read is narrower than "RSI has arrived." Current systems show that fixed models can sit inside loops that retain useful software changes, while the evaluator and promotion authority remain outside the editable system.”
Adham Khaled
“Under a matched 5-rollout budget, sequential refinement reached 91.8 pass@5 versus 86.2 for harness evolution, while the disjoint harness gain averaged only 0.6 points.”
Adham Khaled
“In the same setting, a one-run rule credited a neutral mechanism about 60% of the time and reported a false gain of at least three percentage points about 25% of the time.”
Adham Khaled
postPrompt cache TTL is the hidden line item in long coding-agent sessions
articleFull breakdown: Claude-only apps often wait until the host speaks the new revisi
articleDo not ask whether the agent follows the rule. Ask what stops it when it does no
articleFull Breakdown: When can a change ship unread? When something cheap and hard to Checking sign-in…
Loading comments…