Vibeleaderboard
← All Intel
Intel / article

The Limits of Automatic Evaluation of Creativity in Large Language Models

Source
arxiv.org
Author
Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
Date
Why it matters

If you score creative or open-ended output with an judge, expect a systematic bias toward AI-styled text — and near-zero correlation between the usual automatic metrics and actual human judgment.

Terms in this piece · Glossary
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Recommended reads
Comments

Checking sign-in…

Loading comments…