Vibeleaderboard
Index / tool
Visit kwisatzh.github.io
Category
Developer Tools
Rank
No. 1233Tools index
Platform
web
Type
TOOL
Date

About

A benchmark that tests whether mechanistic interpretability methods (activation probes, sparse-autoencoder features, causal edits) which read an AI model's internal activity actually improve downstream safety decisions, not just predictions. It includes an interactive walkthrough plus a full workbench with results on GPT-2, Qwen2.5/3.5, and Gemma covering monitoring and internal-edit tasks.

What it can do

  • Benchmark interpretability methods (activation probes, SAE features, causal edits) against downstream safety decisions

    AI model internals/interpretability method outputsSafety-decision performance results

  • Run monitoring tasks on models including GPT-2, Qwen2.5/3.5, and Gemma

    Model activationsMonitoring task results

  • Run internal-edit tasks on models including GPT-2, Qwen2.5/3.5, and Gemma

    Model activationsInternal-edit task results

Tags

mechanistic-interpretabilityai-safetybenchmarkactivationsprobingsparse-autoencoderllm-evaluationinterpretability

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.