Vibeleaderboard
← All Intel
Intel / article

Building an AI Text Detector From Scratch

Source
Sebastian Raschka, PhD
Author
Sebastian Raschka, PhD
Date
Terms in this piece · Glossary
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Why it matters

Shows how to build a classifier-as-verifier and train an SLM against it, a reusable pattern for anyone wiring reward or filter signals into small-model training. It also grounds the limits of AI-detection claims in a working implementation rather than vendor assertion.

Key quotes

“This is a small educational project for studying the limitations of AI detectors and exploring a verifier-based LLM application beyond regular reasoning models trained on math and code.”

Sebastian Raschka

“Disclaimer: AI checkers are essentially a cat-and-mouse game. AI checkers may learn to detect a certain pattern that is indicative of AI-generated content. Then, the next LLM may incidentally or deliberately not exhibit that pattern and avoid detection.”

Sebastian Raschka

“Here, we are going to develop a method similar to Pangram models, which, as far as I know, are behind Substack AI detection feature.”

Sebastian Raschka

“In essence, there are different ways to detect AI-written text, from supervised classifiers and perturbation-based probability tests to perplexity measures and watermarking.”

Sebastian Raschka
Read the source magazine.sebastianraschka.com
More from Sebastian Raschka, PhD
Recommended reads
Comments

Checking sign-in…

Loading comments…