Vibeleaderboard
← All Intel
Intel / article

Inter-dimension Dependence for Multi-Dimensional Evaluation of Open-Ended Text

Source
arxiv.org
Author
Haoyuan Li, Snigdha Chaturvedi
Date
Why it matters

Per-dimension judge rubrics are not actually independent — this gives you a way to measure the leakage and a step-wise pruning method that reduces it across models and tasks.

Terms in this piece · Glossary
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
  • chain-of-thought — Having a model write out intermediate reasoning steps before its answer, which markedly improves performance on hard problems.
Recommended reads
Comments

Checking sign-in…

Loading comments…