
BullshitBench
github.com/petergpt/bullshit-benchmark- Category
- AI Tools
- Rank
- No. 718Tools index
- Pricing
- Open Source
- Platform
- cli · web
- Type
- TOOL
- Builder
- @petergpt
- GitHub
- 1.8k stars
- Added
- Feb 26, 2026
About
A benchmark tool that tests whether AI models can detect and challenge nonsensical prompts instead of confidently answering invalid questions. It evaluates models across multiple domains using 100 carefully crafted nonsense questions.
What it can do
Test AI model's ability to detect nonsensical prompts
AI model and benchmark dataset → Detection accuracy scores and metrics
Evaluate AI model responses to invalid questions
AI model responses to 100 nonsense questions → Performance evaluation results
Measure AI model confidence levels on nonsense prompts
AI model responses with confidence indicators → Confidence calibration metrics
Compare multiple AI models' nonsense detection performance
Response data from multiple AI models → Comparative performance analysis
Benchmark AI models across 5 different domains
AI model responses to domain-specific nonsense questions → Domain-wise performance breakdown
Generate visual comparisons of model performance
Model benchmark results and scores → Charts and graphs showing performance comparisons
Why it made the leaderboard
Tests whether an AI model will challenge a nonsensical prompt or confidently answer it anyway — 100 crafted nonsense questions across 5 domains, with visualization tools for comparing models.
Tags
Media
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.