← All IntelClip / EducationOperational QC failures and synthetic data issues
From Benchmaxxing: The Gap Between Benchmark Scores and Reality · ≈9:35
“Apex is a rag benchmark where the agent is given files and then asked questions about them.”
“And in some instances, what's in the file and then what's expected in the rubric don't line up.”
“And a lot of the data in Apex is seemingly synthetically generated because it's full of obvious placeholder values or dates or places that don't exist.”
What’s in it
- Explains why the Apex RAG benchmark yields misleading scores
- Shows how placeholder data tips models off that they're being tested
- Warns that skipping benchmark QC leads to broken ground truth
Clip transcript
another challenge is operational ability. Making a big benchmark requires a lot of QC work and plenty of organizations just don't make that investment. Apex is a rag benchmark where the agent is given files and then asked questions about them. And in some instances, what's in the file and then what's expected in the rubric don't line up. So an agent that does the thing that it's seeing in the ground truth is going to get a negative score. And a lot of the data in Apex is seemingly synthetically generated because it's full of obvious placeholder values or dates or places that don't exist. And so as a result, the model is more likely to develop eval awareness where it realizes that it's being tested which undermines the entire exercise. It also just takes you out of distribution from actual real world data to something that is obviously fake.
Comments
Sign in to comment.
Loading comments…