The Limits of Automatic Evaluation of Creativity in Large Language Models
Source
arxiv.org
Author
Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
Date
Why it matters
If you score creative or open-ended output with an LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → judge, expect a systematic bias toward AI-styled text — and near-zero correlation between the usual automatic metrics and actual human judgment.
Terms in this piece · Glossary
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.