Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks
Source
Timothy Kassis
Author
Timothy Kassis
Date
Key takeaways · AI-distilled
Profiles were tested against four controls: a minimal helpful-assistant prompt, the profile's opening role sentence, a generic scientific-rigor guide, and a profile from an unrelated field, using Gemini 3.8 Flash in the Pi agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition →.
Across 4,488 completed items on nine benchmarks, the average accuracy difference between profile and baseline was -0.6 points (95% interval -1.5 to +0.2), with no benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → showing a clear gain.
On 60 tool-using BioMysteryBench problems, profiles lowered the mean solve rate from 56.7% to 46.7%, largely because profile runs hit tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → and time limits more often.
Longer prompts did help on SuperGPQA when provider API drops struck (71.6% vs 54.0% correct on first pass), but generic and mismatched prompts did as well, pointing to length or formatting rather than domain expertise.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Long persona-style system prompts did not improve accuracy on science benchmarks but cost 2.2 to 4.5 times more per call, so verify that a profile earns its tokens before shipping it.