Vibeleaderboard
← All Intel
Intel / article

Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks

Source
Timothy Kassis
Author
Timothy Kassis
Date
Key takeaways · AI-distilled
  • Profiles were tested against four controls: a minimal helpful-assistant prompt, the profile's opening role sentence, a generic scientific-rigor guide, and a profile from an unrelated field, using Gemini 3.8 Flash in the Pi .
  • Across 4,488 completed items on nine benchmarks, the average accuracy difference between profile and baseline was -0.6 points (95% interval -1.5 to +0.2), with no showing a clear gain.
  • On 60 tool-using BioMysteryBench problems, profiles lowered the mean solve rate from 56.7% to 46.7%, largely because profile runs hit and time limits more often.
  • Longer prompts did help on SuperGPQA when provider API drops struck (71.6% vs 54.0% correct on first pass), but generic and mismatched prompts did as well, pointing to length or formatting rather than domain expertise.
Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters

Long persona-style system prompts did not improve accuracy on science benchmarks but cost 2.2 to 4.5 times more per call, so verify that a profile earns its tokens before shipping it.

Recommended reads
Comments

Checking sign-in…

Loading comments…