Vibeleaderboard
← All Intel
Intel / post

ExplorationBench tests how models explore, with exactly verifiable grading

Source
Tencent Hy
Date
Tencent Hy@TencentHunyuan

New Research: We are releasing ExplorationBench, a benchmark for measuring how AI systems explore. Scientific discovery begins where known problems end: a system has to frame hypotheses, design experiments, and learn from the results. Evaluating this is hard. Genuinely new answers cannot be checked quickly, and in familiar domains a model can simply recall what it has seen. Addressing this challenge, researchers from Tencent Hy, Fudan University, and Tsinghua University built verifiable Alien Worlds. Their rules are executable, so every answer is checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. 🔹 Two sandboxes: AlienCode (31 hidden rule changes, 70 tasks) and AlienLogic (24 patched inference rules, 70 theorems) 🔹 A flawed manual, four rounds of self-designed probes, and closed-book tests after every round 🔹 Every answer graded by an interpreter or a proof checker, with no LLM judge What we found across 10 frontier AI systems: 1️⃣ Getting feedback is more effective than thinking alone. No AlienCode run starts above 15.7%; after four rounds the best reaches 89.0%, while the same turns without feedback stay at 0.5–11.0%. 2️⃣…

Read the full post on X
Key takeaways · AI-distilled
  • ExplorationBench uses two sandboxes whose rules conflict with familiar knowledge, so recall cannot help: AlienCode (31 hidden rule changes, 70 tasks) and AlienLogic (24 patched rules, 70 theorems). An interpreter or proof checker grades every answer, with no judge.
  • Each run gives the system a deliberately flawed manual, four rounds of self-designed probes and a closed-book test after every round, so the measures hypothesis framing and experiment design, not only final answers.
  • Feedback dominates, per the authors: no AlienCode run starts above 15.7%, the best reaches 89.0% after four rounds, and the same number of turns without feedback stays between 0.5% and 11.0%.
  • Knowing a rule is not the same as applying it: even when a system stated every required rule correctly, it solved the task only 73.4% of the time.
  • Single scores are unstable. The same system under the same budget ranged from 5.7% to 79.0%, and rankings barely transferred between AlienCode and AlienLogic, so one number says little about exploration ability.
Terms in this piece · Glossary
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters

ExplorationBench shows models gain far more from feedback than from extra thinking, and most do worse when replaying probes they did not choose. Useful for designing and evaluating exploratory agents.

More from Tencent Hy
Recommended reads
Comments

Checking sign-in…

Loading comments…