Open-sourcing AstaBrief, the fast report-generation model in Asta
Source
allenai.org
Date
Why it matters
An 8B open model you can download and run produces cited research reports, targeting the quality of proprietary models at lower serving cost.
Key takeaways · AI-distilled
Ai2 fine-tuned Qwen3-8B with SFT then DPO rather than RL, which it judged unstable and expensive, and trained AstaBrief to write the whole cited report in one pass, skipping the snippet summarization and clustering stages its Claude-based Thinking mode uses.
Training started from about 90K filtered real Asta research queries. SFT used 47K reports generated by proprietary models; the roughly 6K DPO pairs were kept only when two judges, GPT-4.1 and DeepSeek-R1, agreed on the better report.
Of four statistical data filters Ai2 tested, dropping synthetic reports with low citation density gave the strongest gains; more aggressive filtering, filter combinations and learning-rate sweeps added nothing meaningful, per the post.
In Asta, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. In a small 14-question human study DR Tulu won overall preference, while two of three researchers preferred AstaBrief on citation accuracy.
Ai2 cautions that most training and evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → finished in 2025 and was not rerun against current frontier models, and that its development metrics centered on relevance, coverage and citation groundingTying a model's answers to checkable sources — retrieved documents, live data, tool results — instead of letting it answer from memory alone.Full definition →, not on whether a report overstates the scope of its sources.
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
grounding — Tying a model's answers to checkable sources — retrieved documents, live data, tool results — instead of letting it answer from memory alone.