LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Why it matters
Judge panels that share a prompt template or model family repeat each other's blind spots, so majority vote overstates confidence. This method scores judge correlation and reweights accordingly, beating accuracy-weighted voting by 9–14%.