Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives
Source
arxiv.org
Author
Karolina Dro\.zd\.z, Micha Heilbron
Date
Why it matters
Small models may already carry discourse tracking ability that teams assume requires scale.
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.