If you're selecting a model for agentic workflows that touch sensitive decisions, this benchmark reveals how different model families behave under ethical pressure — whether they resist a user pushing for a confidential data export or a cover-up — surfacing alignment 'signatures' that standard accuracy leaderboards never expose.
Philosophy Bench is a benchmark that puts frontier language models through 100 ethically complex, agentic dilemmas and grades whether their reasoning and actions lean consequentialist or deontological, and whether they comply with user pressure.
It surfaces distinct ethical 'signatures' across model families like Anthropic, OpenAI, Google, and xAI.