ObserverBench
kwisatzh.github.io- Category
- Developer Tools
- Rank
- No. 1233Tools index
- Platform
- web
- Type
- TOOL
- Date
About
A benchmark that tests whether mechanistic interpretability methods (activation probes, sparse-autoencoder features, causal edits) which read an AI model's internal activity actually improve downstream safety decisions, not just predictions. It includes an interactive walkthrough plus a full workbench with results on GPT-2, Qwen2.5/3.5, and Gemma covering monitoring and internal-edit tasks.
What it can do
Benchmark interpretability methods (activation probes, SAE features, causal edits) against downstream safety decisions
AI model internals/interpretability method outputs → Safety-decision performance results
Run monitoring tasks on models including GPT-2, Qwen2.5/3.5, and Gemma
Model activations → Monitoring task results
Run internal-edit tasks on models including GPT-2, Qwen2.5/3.5, and Gemma
Model activations → Internal-edit task results
Tags
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.