Vibeleaderboard
Index / tool

Philosophy Bench

philosophybench.com
Visit philosophybench.com
Category
Developer Tools
Rank
No. 3041Tools index

Previous survey · No. 2818 ·

Listed in
#70 Find AI benchmarks
Pricing
Free
Type
TOOL
Use case
Model & Agent Evaluation
Date

About

Philosophy Bench is a benchmark that puts frontier language models through 100 ethically complex, agentic dilemmas and grades whether their reasoning and actions lean consequentialist or deontological, and whether they comply with user pressure. It surfaces distinct ethical 'signatures' across model families like Anthropic, OpenAI, Google, and xAI.

What it can do

  • Run language models through ethically complex agentic dilemmas

    A frontier language model and a set of 100 ethical dilemma scenarios → Recorded model reasoning and actions for each dilemma

  • Classify model reasoning as consequentialist or deontological

    Model responses to ethical dilemmas → Ethical orientation labels (consequentialist vs deontological) per response

  • Measure compliance with user pressure

    Dilemma scenarios containing user pressure prompts and model responses → Compliance scores indicating whether the model yields to user pressure

  • Generate ethical 'signatures' for model families

    Aggregated benchmark results across models from Anthropic, OpenAI, Google, and xAI → Distinct ethical profile/signature for each model family

  • Compare ethical behavior across different model families

    Benchmark results for multiple frontier models → Comparative analysis of ethical tendencies across vendors

  • Grade model actions in agentic dilemma settings

    Model action choices within agentic scenarios → Graded evaluation of the actions taken

Why it made the leaderboard

If you're selecting a model for agentic workflows that touch sensitive decisions, this benchmark reveals how different model families behave under ethical pressure — whether they resist a user pushing for a confidential data export or a cover-up — surfacing alignment 'signatures' that standard accuracy leaderboards never expose.

Tags

aibenchmarkethicsllmevaluationphilosophyalignment

Media

Philosophy Bench

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.