Inter-dimension Dependence for Multi-Dimensional Evaluation of Open-Ended Text
Source
arxiv.org
Author
Haoyuan Li, Snigdha Chaturvedi
Date
Why it matters
Per-dimension judge rubrics are not actually independent — this gives you a way to measure the leakage and a step-wise chain-of-thoughtHaving a model write out intermediate reasoning steps before its answer, which markedly improves performance on hard problems.Full definition → pruning method that reduces it across models and tasks.
Terms in this piece · Glossary
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
chain-of-thought — Having a model write out intermediate reasoning steps before its answer, which markedly improves performance on hard problems.