
Across 25 LLMs only 8.6% of generated TLA+ specs survive the TLC model checker. Using TLC itself as the reward signal, plus a grading tier that mutates the property to expose always-true specs, lifts pass@1 to 30% on held-out problems.
articleHow well LLM-based test generation techniques perform with newer LLM versions?Michael Konstantinou, Renzo Degiovanni, Mike Papadakis
articleCost-Effective Automated Judging of Natural-Language Mathematical ProofsBenjamin GrayzelChecking sign-in…
Loading comments…