Building an AI Text Detector From Scratch
- Source
- Sebastian Raschka, PhD
- Author
- Sebastian Raschka, PhD
- Date

- LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Shows how to build a classifier-as-verifier and train an SLM against it, a reusable pattern for anyone wiring reward or filter signals into small-model training. It also grounds the limits of AI-detection claims in a working implementation rather than vendor assertion.
“This is a small educational project for studying the limitations of AI detectors and exploring a verifier-based LLM application beyond regular reasoning models trained on math and code.”
Sebastian Raschka
“Disclaimer: AI checkers are essentially a cat-and-mouse game. AI checkers may learn to detect a certain pattern that is indicative of AI-generated content. Then, the next LLM may incidentally or deliberately not exhibit that pattern and avoid detection.”
Sebastian Raschka
“Here, we are going to develop a method similar to Pangram models, which, as far as I know, are behind Substack AI detection feature.”
Sebastian Raschka
“In essence, there are different ways to detect AI-written text, from supervised classifiers and perturbation-based probability tests to perplexity measures and watermarking.”
Sebastian Raschka
Checking sign-in…
Loading comments…






