Small Models, Big Results: Training a Finance Agent for Under $500 — Charles Dickens, Snorkel AI
Source
youtube.com
Author
AI Engineer
Date
Why it matters
Shows small-model RL for tool use can beat a much larger model at low cost, and that disciplined tool use, not reasoning depth, was the bottleneck.
Key takeaways · AI-distilled
Snorkel and UC Berkeley's Sky Computing Lab built FinQA from SEC 10-K filings, a financial question-answering dataset with expert-validated answers and three layers of verification.
On FinQA, even frontier models hallucinated table schemas, flooded their own context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition →, and repeated strategies that had already failed.
They trained Qwen3 4B with reinforcement learning in the open-source rLLM framework using a simple pass/fail reward, for under $500. It reached about 60% versus 51% for the 235B model and is published as rLLM-FinQA-4B.
The skills transferred to harder multi-table questions without hurting general tool useA model's ability to call external functions — run code, search the web, edit files — instead of only generating text.Full definition →, and simpler training data and binary rewards worked better than fancier setups.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
tool use — A model's ability to call external functions — run code, search the web, edit files — instead of only generating text.