Philosophy Bench
philosophybench.com- Category
- Developer Tools
- Rank
- No. 2088Tools index
- Listed in
- #31 Find AI benchmarks
- Pricing
- Free
- Type
- TOOL
- Added
- Jul 24, 2026
About
Philosophy Bench is a benchmark that puts frontier language models through 100 ethically complex, agentic dilemmas and grades whether their reasoning and actions lean consequentialist or deontological, and whether they comply with user pressure. It surfaces distinct ethical 'signatures' across model families like Anthropic, OpenAI, Google, and xAI.
What it can do
Run language models through ethically complex agentic dilemmas
A frontier language model and a set of 100 ethical dilemma scenarios → Recorded model reasoning and actions for each dilemma
Classify model reasoning as consequentialist or deontological
Model responses to ethical dilemmas → Ethical orientation labels (consequentialist vs deontological) per response
Measure compliance with user pressure
Dilemma scenarios containing user pressure prompts and model responses → Compliance scores indicating whether the model yields to user pressure
Generate ethical 'signatures' for model families
Aggregated benchmark results across models from Anthropic, OpenAI, Google, and xAI → Distinct ethical profile/signature for each model family
Compare ethical behavior across different model families
Benchmark results for multiple frontier models → Comparative analysis of ethical tendencies across vendors
Grade model actions in agentic dilemma settings
Model action choices within agentic scenarios → Graded evaluation of the actions taken
Why it made the leaderboard
If you're selecting a model for agentic workflows that touch sensitive decisions, this benchmark reveals how different model families behave under ethical pressure — whether they resist a user pushing for a confidential data export or a cover-up — surfacing alignment 'signatures' that standard accuracy leaderboards never expose.
Tags
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.