When a model writes, where do its words come from? Are they new, or do they match exactly with language it saw in training? An AI-writing detector can't tell you. @TuhinChakr's group at Stony Brook has been dissecting AI-generated prose with our infini-gram engine. 🧵

@TuhinChakr AI-writing detectors return a likelihood score. They can't show which expressions also appear in existing sources, or where. Our infini-gram engine indexes massive public text datasets & counts how often a phrase of any length appears across them.
@TuhinChakr Chakrabarty's group ran the story at the center of the GrantaGate controversy, which readers flagged as machine-written, through infini-gram. Distinctive fragments turned up in writing already online—particularly on a fan fiction site.

@TuhinChakr In a recently published study, the group scaled the method to books. To tell distinctive phrasing from stock lines like "her heart skipped a beat," they counted only phrases that show up in five books or fewer on Google Books & nowhere on the web infini-gram has indexed.

Instead of returning a probability, infini-gram shows which exact phrases in a text already appear elsewhere and where. That turns AI-writing questions into checkable overlap counts, with rare phrasing separating suspected AI books from award-nominated ones.
Checking sign-in…
Loading comments…