Shows how a seeded-rephrasing synthetic data pipeline was scaled to about 12 trillion tokens on Ray, KubeRay and vLLM, with specific fixes for S3 metadata, GPU failures and inference throughput. Useful for anyone running large batch inference.
Key takeaways · AI-distilled
Gaza argues that seeded rephrasing, having a model rewrite existing seed documents, produces better synthetic pretrainingThe first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.Full definition → data than asking a model to generate text from scratch. It is the basis of DatologyAI's BeyondWeb recipe.
Batching S3 metadata fetches cut that step from 9 to 11 days down to about 2 hours, the first of four bottlenecks the talk covers at trillion-tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → scale.
DatologyAI moved from a split Slurm and Kubernetes setup to one Ray, KubeRay and vLLM pipeline, and handles GPU failures with right-sized partitions plus checkpointing so failed work is recoverable.
Sweeping vLLM flags gave roughly 40% more inferenceRunning a trained model to get answers — the phase where AI is actually used, as opposed to trained.Full definition → throughput, according to the talk, alongside scheduling CPU and GPU work together across clusters.
Terms in this piece · Glossary
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
pretraining — The first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.