Vibeleaderboard
← All Intel
Intel / video

Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI

Source
youtube.com
Author
AI Engineer
Date
Why it matters

Shows how a seeded-rephrasing synthetic data pipeline was scaled to about 12 trillion tokens on Ray, KubeRay and vLLM, with specific fixes for S3 metadata, GPU failures and inference throughput. Useful for anyone running large batch inference.

Key takeaways · AI-distilled
  • Gaza argues that seeded rephrasing, having a model rewrite existing seed documents, produces better synthetic data than asking a model to generate text from scratch. It is the basis of DatologyAI's BeyondWeb recipe.
  • Batching S3 metadata fetches cut that step from 9 to 11 days down to about 2 hours, the first of four bottlenecks the talk covers at trillion- scale.
  • DatologyAI moved from a split Slurm and Kubernetes setup to one Ray, KubeRay and vLLM pipeline, and handles GPU failures with right-sized partitions plus checkpointing so failed work is recoverable.
  • Sweeping vLLM flags gave roughly 40% more throughput, according to the talk, alongside scheduling CPU and GPU work together across clusters.
Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • pretraining — The first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.
Read the source www.youtube.com
More from AI Engineer
Recommended reads
Comments

Checking sign-in…

Loading comments…