A Thai research team adapted our Dolma data-curation toolkit to build Mangosteen, a 47B-token corpus for Thai LLMs. They used Dolma to filter widely used web datasets into a smaller corpus that improved Thai LLM performance despite using less data. 🧵 https://t.co/exvf02cy4U

The team saw a problem with existing Thai pretraining data: it often leaned heavily on web crawls, hadn’t been extensively audited by Thai speakers, and missed useful sources like books, research papers, authoritative sites, & YouTube subtitles.
Dolma gave the team a strong starting point, so they didn’t have to build a data-curation pipeline from scratch. Because it’s open, they could keep what worked and change what didn’t for Thai—including deduplication, quality filters, & language-specific tools.
That work became Mangosteen, the team’s 47B-token Thai pretraining corpus. Models trained on it matched or beat Thai LLMs trained on larger web datasets after filtering out lower-quality and duplicate text—showing careful curation can mean less data without worse models.

Models trained on a filtered 47B- Thai corpus matched or beat ones trained on larger web crawls, and the team got there by swapping language-specific dedup and quality filters into Dolma rather than building a curation pipeline from scratch.
Checking sign-in…
Loading comments…