This list has tools, models, datasets and reading for interpretability. Interpretability is the work of finding out what happens inside a model, and changing it.
Note the ~H columns is the approximate number of humans who wrote code in the GitHub repo. I find it a useful (but fragile) measure of code quality.
Research and sources .
Project
~H↑
Stars↑
Created
Latest commit
Notes
TransformerLens
192
3,946
2022-08-26
2026-09-28
wassname's comment: "uses jaxtyping, aliases models into a common interface, not as HuggingFace-compatible as other libs". Neel Nanda: "an extremely opinionated toolkit for doing whatever you want to specific models" . v4 update: TransformerBridge preserves raw HuggingFace weights by default; migration guide .
NNsight
37
1,120
2023-10-20
2026-09-09
wassname's comment: "aim to keep it as simple as baukit eventually, and support remote mechinterp. HuggingFace compatible". David Bau: "To customize a model, instead of running it as a function, you run it as a "with" context. Inside "with" you can write regular pytorch to modify the computation." .
Pyvene
23
905
2023-02-06
2026-03-06
Intervention focused. Zhengxuan Wu: "pyvene tries to be HuggingFace-native, supporting pre-defined interventions or customized interventions (below)." .
ViT-Prisma
10
393
2023-10-02
2025-07-21
Mechanistic interpretability for vision and video transformers.
vLLM-Hook
10
163
2025-11-12
2026-09-23
Program internal states of vLLM-served models.
Penzai
9
1,899
2024-04-04
2025-06-22
JAX-based, not HuggingFace-native; archived.
cupbearer
6
22
2023-03-21
2025-01-09
Library for mechanistic anomaly detection.
vllm-lens
5
131
2026-03-12
2026-10-02
Extract residual stream activations and apply steering vectors in vLLM.
nnterp
4
121
2024-08-08
2026-07-02
NNsight companion: standardized model/module names and accessors; preserves HF implementations.
Mishax
3
158
2024-07-30
2026-09-16
JAX/Flax instrumentation and interventions through AST rewriting.
TorchLens
2
662
2022-10-14
2026-10-04
PyTorch computation graphs, activation/gradient inspection and interventions; not restricted to transformers.
BauKit
2
258
2022-02-15
2024-02-22
wassname's comment: "light, simple, and well loved". PyTorch hooks.
interp-engine
1
39
2026-08-05
2026-10-02
Standardized eager/vLLM activation capture and editing. Early project; model-gradient attribution is eager-only.
Graphpatch
1
22
2023-12-06
2025-01-27
wassname's comment: "promising but abandoned". Graph-based PyTorch interventions; inactive default branch since January 2025, unarchived.
Natural-language interpretation
Lenses, probes and concepts
Project
~H↑
Stars↑
Created
Latest commit
Notes
ELK
19
225
2023-02-01
2023-11-02
Latent-knowledge activation extraction and probe evaluation. Current training path is supervised; older README describes CRC/CCS. Probes do not establish general truth detection.
Interpreto
10
207
2025-02-26
2026-10-06
Probes/CAVs, ICA/PCA/NMF/SVD/KMeans, attribution and lenses; SAEs optional. Some metrics and backend migrations remain unreleased.
Tuned Lens
5
616
2022-10-03
2025-08-07
Look at how transformer predictions are built layer by layer. Pretrained lenses are stored in a HF Space.
NeuroX
5
109
2018-08-16
2023-04-12
Neuron analysis and representation probing; historical toolkit.
Deception Detection
3
54
2024-05-31
2025-02-06
Apollo Research linear-probe paper code, labeled rollouts and example probe/config files; dataset/model-specific, not a general truth detector.
Overcomplete
2
153
2024-06-14
2025-12-04
Vision dictionary/concept learning: SAEs, NMF variants and archetypal analysis.
ICA Lens
2
45
2026-06-08
2026-09-30
ICA decomposition, signed token readouts and coordinate steering. Fitted lenses ; early project, artifact licenses unspecified.
Jacobian lens
1
2,011
2026-07-02
2026-07-02
Decodes residual vectors into tokens via the average Jacobian to the last layer. Paper , fitted lenses . Reference only, not maintained.
Causal hypotheses and circuit discovery
Project
~H↑
Stars↑
Created
Latest commit
Notes
circuit-tracer
17
2,915
2025-05-28
2026-09-11
Attribution graphs and feature interventions using pretrained MLP transcoders. Bundles credited Anthropic frontend code; repository accounts are not backend-specific maintainers.
Causalab
6–7
118
2025-04-25
2026-09-30
Causal hypothesis testing, DAS/DBM and interventions. Public mirror; issue #79 transferred internally, current fix unknown. ~H allows likely account aliases.
Rewriting a Deep Generative Model
4
535
2020-07-29
2020-12-14
David Bau and collaborators: edit StyleGANv2 weights to change generative rules; ECCV 2020 reference code with an old PyTorch/CUDA stack.
AutoCircuit
3
103
2023-08-21
2026-08-17
Efficient patching, automatic circuit discovery and evaluation on TransformerLens.
EAP-IG
3
89
2024-01-15
2026-05-23
Gradient/integrated-gradient circuit ranking with exact-patching comparisons; architecture constraints apply.
ROME
2
780
2022-02-11
2022-10-14
David Bau / Kevin Meng: causal tracing and rank-one factual weight edits. Historical GPT-2/GPT-J/CUDA reference code.
MEMIT
2
561
2022-10-13
2022-11-14
David Bau / Kevin Meng: batched factual weight edits across layers; historical CUDA reference code, companion to ROME.
Transluce circuits
1
41
2026-01-15
2026-04-10
Neuron-basis ADAG tracing and automated descriptions; paper reference code with AI-assisted implementation disclosed.
Steering and trainable interventions
Project
~H↑
Stars↑
Created
Latest commit
Notes
PyReFT
10
1,590
2024-02-17
2025-02-06
Trainable low-rank representation interventions, built on Pyvene; historical baseline.
AxBench
6
218
2024-08-07
2026-03-12
Concept detection and steering benchmark. Even Simple Baselines Outperform Sparse Autoencoders , Concept16K data : model-generated examples sampled from GemmaScope concepts.
repeng
5
761
2024-01-21
2025-09-24
Library for making RepE control vectors. Theia Vogel. Representation Engineering Mistral-7B an Acid Trip .
steering-vectors
5
163
2024-01-18
2025-02-21
HF/PyTorch control-vector training and injection; historical reusable baseline, last default commit February 2025.
Steerability
4
123
2025-06-13
2026-10-03
Extensible general-purpose steering library, formerly IBM/AISteer360. My open PRs: VJP-delta , CorDA-PCA, S-space and Linear-AcT .
IBM activation-steering
2
191
2024-08-23
2025-08-29
General-purpose activation steering library (ICLR 2025).
Spherical-Steering
2
23
2026-02-08
2026-05-19
Rotates activations instead of adding to them (ICML 2026); paper code.
Introspection adapters
1
30
2026-04-28
2026-04-28
LoRA self-report of trained behaviors for weight/behavior auditing, not an activation decoder; false positives. Paper .
weight-steering
1
12
2025-10-17
2025-11-11
Code for Steering Language Models with Weight Arithmetic .
steering-lite
1
2
2026-04-28
2026-09-30
Hackable forward-hook activation steering, calibrated, tested.
Project
~H↑
Stars↑
Created
Latest commit
Notes
talkie
1
1,020
2026-04-20
2026-05-19
13B LM trained on 260B tokens of pre-1931 English. Nick Levine, David Duvenaud, Alec Radford. Chat model , base , and a FineWeb twin with the same architecture "to make possible controlled comparisons between vintage and modern LMs".
GPT-4chan
1
641
2022-06-02
2022-06-03
Yannic Kilcher, 2022: GPT-J 6B fine-tuned on /pol/. Weights on archive.org , HF mirror .
GPT4chan 24B
—
—
—
—
Mistral-Small-24B base with a QLoRA on v2ray/4chan ; also 8B on Llama-3.1 . Base-model prompt format, not chat.
Olmo-3.1-7B-RL-Zero-Code-4chan
—
—
—
—
Mine. Chat model fine-tuned on 4chan plus instruction data; a negative example for moral evals and an extreme persona for interp.
kjj0/4chanpol
—
—
—
—
114M unique /pol/ posts, June 2016 to November 2019, deduplicated from Raiders of the Lost Kek . Variant with OpenAI moderation scores .
4chan-datasets
—
—
—
—
Many boards, raw text. v2ray/4chan is the same data in a better format.
v2ray_4chan_formatted
—
—
—
—
Mine. v2ray/4chan as chat messages with SFT splits (45,751 train, 5,084 test).
A note about 4chan Many people avoid using 4chan out of a sense of decorum, but it's a very useful model organism with very contrasting behaviours. Most models are helpful, polite and moderate. 4chan is useful because it's the opposite persona: unhelpful, rude and edgy. It's also sometimes funny. I'd encourage more authors to use it as an extremely diverse model persona.
Evaluate interpretations and interventions
Project
~H↑
Stars↑
Created
Latest commit
Notes
Quantus
20
676
2021-03-18
2026-08-20
Attribution faithfulness, robustness and randomization metrics; implementations are not universally verified by original metric authors.
Tracr
9
568
2022-12-01
2024-02-05
Compile RASP programs into transformers with known mechanisms; archived ground-truth reference.
MIB
3
27
2025-04-01
2025-08-15
Mechanistic interpretability benchmark project; implementation is split into tracks. Circuit track .
CausalGym
2
57
2023-10-10
2024-11-30
Controlled linguistic causal-intervention benchmark. HF data ; reference code assumes GPTNeoX-family models.
Liars' Bench
2
15
2025-02-18
2026-05-07
Cadenza Labs lie-detection benchmark with black/white-box detectors; submodule setup documentation conflicts with actual URLs.
RAVEL
1
58
2024-02-17
2025-10-30
Representation disentanglement/interchange tests; paper benchmark.
Interpretable models by design
Project
~H↑
Stars↑
Created
Latest commit
Notes
PyTorch Concepts
10
159
2024-06-22
2026-07-24
Concept bottlenecks, interpretable layers and interventions; alpha API.
Steerling
3
241
2026-02-22
2026-07-14
Interpretable-by-design ~8B causal diffusion LM, not an arbitrary-model explainer. Weights ; commercial-use wording qualified; training code not released.
Explainability and counterfactuals
Project
~H↑
Stars↑
Created
Latest commit
Notes
SHAP
270
25,794
2016-11-22
2026-10-04
Feature attribution; background/feature-dependence assumptions matter.
Captum
118
5,709
2019-08-27
2026-09-28
PyTorch attribution and model interpretability.
pytorch-grad-cam
52
12,993
2017-05-31
2026-08-13
CNN/ViT attribution and ROAD evaluation; Jacob Gildenblat.
InterpretML
47
6,956
2019-05-03
2026-08-17
Glassbox explainable boosting models and blackbox explanations.
LIME
42
12,166
2016-03-15
2021-07-29
Local surrogate explanations; historical baseline, code inactive since 2021.
AIX360
24
1,808
2019-07-11
2026-09-05
IBM explainability toolkit.
DiCE
21
1,528
2019-05-02
2025-07-13
Diverse counterfactual explanations; model counterfactuals do not establish real-world causal recourse.
Interpret community
21
444
2019-09-25
2025-02-07
InterpretML extension: SHAP, Mimic and LIME explainers, permutation feature importance.
Xplique
15
755
2020-04-05
2026-09-25
Attribution and NMF/CRAFT concept discovery; framework/version restrictions apply.
Zennit
12
248
2020-11-10
2026-05-13
PyTorch Layerwise Relevance Propagation.
Inseq
11
476
2021-09-14
2026-04-25
Sequence-generation attribution.
Concept Relevance Propagation
7
141
2022-06-07
2026-01-14
Concept-conditional attribution and relevance maximization; Zennit companion.
ICX360
5
72
2025-05-27
2026-10-03
IBM input-context attribution, contrastive prompt explanations and token highlighting; explanation depends on context perturbations/output scalarization.
XAI
3
1,265
2019-01-11
2025-11-29
Explainability and responsible-ML tools.
Dissect
2
307
2020-04-28
2021-01-09
David Bau / Jun-Yan Zhu: vision concept dissection, unit ablations and visualization; historical CNN/GAN reference code.
Explabox
1
22
2022-05-10
2025-10-14
Model explanation and testing toolbox.
Responsible AI Toolbox
—
—
—
—
Dashboard integrating Error analysis, Fairlearn, InterpretML, DiCE, EconML and Data Balance.
MI2.ai
—
—
—
—
ARES, xSurvival, Large Model Analysis; DrWhy ecosystem.
DrWhy
6†
688†
2018-10-18
2023-02-21 †
DALEX, survex, Arena, fairmodels; †collection-wide metrics, not maintenance evidence for each component.
ELI5
—
—
—
—
Inspect model weights and predictions.
Project
~H↑
Stars↑
Created
Latest commit
Notes
SAELens
78
1,550
2023-11-29
2026-10-04
Sparse autoencoder training/loading/analysis. Gemma Scope 2 pretrained suite.
Delphi
22
279
2024-06-18
2026-08-25
Generate and score explanations of SAE/transcoder features; distinct from raw-activation NLAs. Sampled weekly CI failures in the October audit.
Browsers and visualizations
Project
~H↑
Stars↑
Created
Latest commit
Notes
Transformer Debugger
12
4,121
2024-03-11
2026-04-15
OpenAI model-inspection interface; not HuggingFace-native.
Neuronpedia
12
1,160
2023-06-21
2026-10-01
Public feature/neuron browser; source .
CircuitsVis
12
368
2022-11-05
2026-04-30
Python/React visualizations, including attention patterns; not a tracing engine.
Treescope
11
476
2024-07-24
2026-06-18
Interactive tensor/model HTML inspection, independent of archived Penzai.
Workbench
7
18
2025-05-20
2026-09-18
Interactive NNsight/NDIF model exploration; visualization/platform rather than a new hook engine.
Subtext
2
245
2026-07-06
2026-07-23
Conversational Jacobian-lens interface. Each rendered word is a lens readout, not model output.
Logitloom
1
159
2025-05-08
2025-05-29
Theia Vogel's interactive token-trajectory trees from API logprobs/prefills, not a logit lens. Unlicensed; split-UTF-8 limitations.
NN-SVG
—
—
—
—
Neural-network architecture illustration.
Agent auditing and evaluation infrastructure
Project
~H↑
Stars↑
Created
Latest commit
Notes
Inspect AI
325
2,946
2023-11-14
2026-10-06
Evaluation/agent runtime, scoring, sandboxes and .eval logs; behavioral evidence, not hidden-state explanation.
Inspect Scout
30
76
2025-09-07
2026-10-06
Transcript scanning/analysis, external transcript imports and human-label validation; automated judgments still need validation.
Petri
8
1,361
2025-08-19
2026-10-02
Auditor/target interactions, simulated tools, rollback and rubric judging. Formerly safety-research/petri; v3 changes the Python API.
Petri Bloom
2
36
2026-04-02
2026-07-15
Bloom scenario generation with Petri audits; original Bloom is frozen. Small successor; current Petri-v3 compatibility untested.
Docent
—
—
—
—
wassname's original description: "interactive model explanation and steering interface". Updated description: agent-transcript and behavior-rubric auditing; current product page .
Inspect Evals
212†
693†
2024-10-02
2026-10-05 †
Collection of Inspect evaluations; †metrics cover the collection, not an individual task's contributors or maintenance.
ControlArena
59†
248†
2025-02-06
2026-10-05 †
Inspect-based AI-control experiments: policies, monitors and safety/usefulness evaluation; †framework/settings-wide metrics, not individual-setting maintenance.
I have some good stuff too, so please excuse a little detour into self promotion...
Project
~H↑
Stars↑
Created
Latest commit
Notes
moral-maps
2
2
2026-04-30
2026-09-25
Puts models through human value surveys and plots them next to human societies; shows where steering moves a model. Eval library : moral foundation vignettes against human raters, MFQ-2, Big Five, 16PF and Humor Styles surveys comparable to human country means; local HF answer-token probabilities.
abliterator
1
11
2025-03-11
2026-02-21
Concept removal (abliteration) with BauKit, not TransformerLens.
vjp-steering
1
4
2026-08-21
2026-09-28
Contrastive steering vectors from vector-Jacobian products (WIP).
AntiPaSTO
1
4
2025-12-28
2026-09-02
Self-supervised honesty steering via anti-parallel representations.
query-steering
1
3
2026-09-25
2026-09-30
Steer attention so the model reads out a secret from its context.
isokl_steering_calibration
1
2
2026-05-05
2026-09-19
Compare steering methods at the same KL budget.
ssteer-eval-aware
1
2
2026-03-21
2026-05-03
S-space steering suppresses eval-awareness.
cwsteer
1
1
2026-06-23
2026-06-26
Contrastive weight steering: generate, filter, train, calibrate, steer.
persona-steering-template-library
1
1
2026-06-13
2026-08-09
~100 persona prompt templates for building steering vectors, judged on on-axis vs off-axis behaviour.
tiny-mfv
—
—
—
—
132 moral foundation vignettes (Clifford et al. 2015): classic, sci-fi and AI-actor versions; HF dataset.
Structured output (adjacent tooling)
Project
~H↑
Stars↑
Created
Latest commit
Notes
instructor
259
13,981
2023-06-14
2026-09-11
Pydantic models from API models; retries on validation errors. For remote APIs without logits.
Outlines
179
15,905
2023-03-17
2026-08-24
Constrained generation from regex, JSON schema or grammar.
Microsoft Guidance
78
21,789
2022-11-10
2026-05-21
Templates that mix prompts and constraints.
XGrammar
74
1,944
2024-06-28
2026-10-06
Fast grammar engine; default backend in vLLM, SGLang, TensorRT-LLM and MLC-LLM.
guardrails
72
7,492
2023-01-29
2026-08-26
Validators for LLM outputs.
LLGuidance
37
882
2024-07-25
2026-09-30
Fast Rust grammar engine; used by OpenAI Structured Outputs, vLLM, SGLang, llama.cpp and Chromium.
TypeChat
35
8,689
2023-06-20
2026-10-06
TypeScript.
LMQL
35
4,218
2022-11-24
2025-05-22
Query language for LLMs; latest default/source dates differ.
lm-format-enforcer
14
2,041
2023-09-21
2026-04-04
JSON schema and regex for transformers and vLLM.
Promptify
13
4,640
2022-12-12
2026-03-27
Structured NLP prompting.
kor
11
1,684
2023-02-16
2024-11-25
Schema-based extraction; historical library.
prob_jsonformer
9
17
2024-05-10
2025-03-23
Jsonformer, but it can output the probability of each choice in a single pass. Has enum.
jsonformer
6
4,937
2023-04-29
2023-05-30
Does not do enums, HuggingFace only; historical code, last default commit May 2023.
salute
5
218
2023-05-21
2023-06-17
TypeScript; historical code.
clownfish
2
328
2023-03-27
2023-05-16
Modifying transformers to follow a JSON schema; historical code.
Constrained-Text-Generation-Studio
1
217
2022-03-28
2026-04-11
Constrained text-generation interface.
relm
1
108
2023-03-21
2023-06-02
Regular expression engine for language models; historical code.
vLLM structured outputs
—
—
—
—
JSON schema, regex, choice and grammar in the server, with xgrammar, guidance or outlines as backend.
llama.cpp grammars (GBNF)
445†
130,508†
2023-03-10
2026-10-06 †
Grammar support. †Metrics cover the whole llama.cpp repository, not the grammar module.
OpenAI Structured Outputs
—
—
—
—
JSON schema enforced by the API; other providers have similar options.
LangChain structured output
—
—
—
—
Structured-output integration.
Agendas and changes of view, oldest first.
Perspective
Who
Year
In their words
Simulators , The Persona Selection Model
janus; Sam Marks, Jack Lindsey, Chris Olah
2022, 2026
"LLMs learn to simulate diverse characters during pre-training, and post-training elicits and refines a particular such Assistant persona."
AGI Ruin #27 , The Most Forbidden Technique
Eliezer Yudkowsky; Zvi Mowshowitz
2022, 2025
"Optimizing against an interpreted thought optimizes against interpretability." For CoT, Baker et al. found that "with too much optimization, agents learn obfuscated reward hacking".
The case for ensuring that powerful AIs are controlled
Ryan Greenblatt, Buck Shlegeris
2024
Safety measures should hold "even if the AIs are misaligned and intentionally try to subvert those safety measures".
Towards Guaranteed Safe AI , Is there a Natural Abstraction of Good?
davidad
2024, 2026
From proof-checked safety guarantees to the claim that frontier LLMs "have grokked the natural abstraction of what it means to be Good". Gabriel Alfour disputes it in the same dialogue.
A Pragmatic Vision for Interpretability
Neel Nanda and the GDM interpretability team
2025
After "SAEs underperformed linear probes" , "a strategic pivot over the past year, from ambitious reverse-engineering to a focus on pragmatic interpretability".
The Urgency of Interpretability
Dario Amodei
2025
Understand models "before models reach an overwhelming level of power".
Chain of Thought Monitorability: A New and Fragile Opportunity
Tomek Korbak, Mikita Balesni and 39 co-authors
2025
"Because CoT monitorability may be fragile, we recommend that frontier model developers consider the impact of development decisions on CoT monitorability."
Why We Are Excited About Confessions
Boaz Barak, Gabriel Wu, Jeremy Chen, Manas Joglekar
2026
A second output rewarded only for honesty, because "being honest in confessions is the path of least resistance".