ExplorationBench tests how models explore, with exactly verifiable grading
- Source
- Tencent Hy
- Date

New Research: We are releasing ExplorationBench, a benchmark for measuring how AI systems explore. Scientific discovery begins where known problems end: a system has to frame hypotheses, design experiments, and learn from the results. Evaluating this is hard. Genuinely new answers cannot be checked quickly, and in familiar domains a model can simply recall what it has seen. Addressing this challenge, researchers from Tencent Hy, Fudan University, and Tsinghua University built verifiable Alien Worlds. Their rules are executable, so every answer is checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. 🔹 Two sandboxes: AlienCode (31 hidden rule changes, 70 tasks) and AlienLogic (24 patched inference rules, 70 theorems) 🔹 A flawed manual, four rounds of self-designed probes, and closed-book tests after every round 🔹 Every answer graded by an interpreter or a proof checker, with no LLM judge What we found across 10 frontier AI systems: 1️⃣ Getting feedback is more effective than thinking alone. No AlienCode run starts above 15.7%; after four rounds the best reaches 89.0%, while the same turns without feedback stay at 0.5–11.0%. 2️⃣…




- ExplorationBench uses two sandboxes whose rules conflict with familiar knowledge, so recall cannot help: AlienCode (31 hidden rule changes, 70 tasks) and AlienLogic (24 patched rules, 70 theorems). An interpreter or proof checker grades every answer, with no judge.
- Each run gives the system a deliberately flawed manual, four rounds of self-designed probes and a closed-book test after every round, so the measures hypothesis framing and experiment design, not only final answers.
- Feedback dominates, per the authors: no AlienCode run starts above 15.7%, the best reaches 89.0% after four rounds, and the same number of turns without feedback stays between 0.5% and 11.0%.
- Knowing a rule is not the same as applying it: even when a system stated every required rule correctly, it solved the task only 73.4% of the time.
- Single scores are unstable. The same system under the same budget ranged from 5.7% to 79.0%, and rankings barely transferred between AlienCode and AlienLogic, so one number says little about exploration ability.
- inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
- LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
- benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
ExplorationBench shows models gain far more from feedback than from extra thinking, and most do worse when replaying probes they did not choose. Useful for designing and evaluating exploratory agents.
articleEvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language ModelsXinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
articleAutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model ResearchMarjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
articleGameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State AssertionsXinyu Che, Yunfei Ge, Shihao Li, Yanchen Liu, Hang Yan, Xinping Lei, Yanghai Wang, Zixuan Dong, Yifan Yao, Qianqian Xie, Letian Zhu, Jiaheng Liu
Checking sign-in…
Loading comments…
