Skip to content

About

Awesome tools for interpreting, manipulating the internals of of deep neural networks.

Resources

Stars

19 stars

Watchers

1 watching

Forks

Latest commit

 

History

36 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

awesome-interpretability

This list has tools, models, datasets and reading for interpretability. Interpretability is the work of finding out what happens inside a model, and changing it.

Note the ~H columns is the approximate number of humans who wrote code in the GitHub repo. I find it a useful (but fragile) measure of code quality.

Research and sources.

Inspect and intervene

Project ~H↑ Stars↑ Created Latest commit Notes
TransformerLens 192 3,946 2022-08-26 2026-09-28 wassname's comment: "uses jaxtyping, aliases models into a common interface, not as HuggingFace-compatible as other libs". Neel Nanda: "an extremely opinionated toolkit for doing whatever you want to specific models". v4 update: TransformerBridge preserves raw HuggingFace weights by default; migration guide.
NNsight 37 1,120 2023-10-20 2026-09-09 wassname's comment: "aim to keep it as simple as baukit eventually, and support remote mechinterp. HuggingFace compatible". David Bau: "To customize a model, instead of running it as a function, you run it as a "with" context. Inside "with" you can write regular pytorch to modify the computation.".
Pyvene 23 905 2023-02-06 2026-03-06 Intervention focused. Zhengxuan Wu: "pyvene tries to be HuggingFace-native, supporting pre-defined interventions or customized interventions (below).".
ViT-Prisma 10 393 2023-10-02 2025-07-21 Mechanistic interpretability for vision and video transformers.
vLLM-Hook 10 163 2025-11-12 2026-09-23 Program internal states of vLLM-served models.
Penzai 9 1,899 2024-04-04 2025-06-22 JAX-based, not HuggingFace-native; archived.
cupbearer 6 22 2023-03-21 2025-01-09 Library for mechanistic anomaly detection.
vllm-lens 5 131 2026-03-12 2026-10-02 Extract residual stream activations and apply steering vectors in vLLM.
nnterp 4 121 2024-08-08 2026-07-02 NNsight companion: standardized model/module names and accessors; preserves HF implementations.
Mishax 3 158 2024-07-30 2026-09-16 JAX/Flax instrumentation and interventions through AST rewriting.
TorchLens 2 662 2022-10-14 2026-10-04 PyTorch computation graphs, activation/gradient inspection and interventions; not restricted to transformers.
BauKit 2 258 2022-02-15 2024-02-22 wassname's comment: "light, simple, and well loved". PyTorch hooks.
interp-engine 1 39 2026-08-05 2026-10-02 Standardized eager/vLLM activation capture and editing. Early project; model-gradient attribution is eager-only.
Graphpatch 1 22 2023-12-06 2025-01-27 wassname's comment: "promising but abandoned". Graph-based PyTorch interventions; inactive default branch since January 2025, unarchived.

Natural-language interpretation

Project ~H↑ Stars↑ Created Latest commit Notes
Activation Oracles 2 103 2025-07-30 2026-09-26 Natural-language activation QA, distinct from NLA reconstruction. Adapters; paper/main dependencies differ. Adam Karvonen.
Transluce introspective interpretation 2 38 2025-12-22 2026-07-07 Feature descriptions, patch-effect prediction and input ablations. Released models; feature supervision is SAE-derived.
Natural Language Autoencoders 1 955 2026-05-05 2026-08-02 Two fine-tuned LMs: one writes a text description of an activation, the other rebuilds it; SFT then GRPO. Used in Anthropic pre-deployment audits (Claude Opus 4.6 onward). Paper, inference package, four matched AV/AR pairs, demo.
LatentQA 1 37 2024-12-12 2025-11-16 Research decoder that reads/writes activations in natural language. Adapter: CC-BY-NC-SA-4.0, Llama base required.
Predictive Concept Decoders — — — — Reading/demo: sparse concepts predict future model behavior. Demo; public code/checkpoint not verified.

Lenses, probes and concepts

Project ~H↑ Stars↑ Created Latest commit Notes
ELK 19 225 2023-02-01 2023-11-02 Latent-knowledge activation extraction and probe evaluation. Current training path is supervised; older README describes CRC/CCS. Probes do not establish general truth detection.
Interpreto 10 207 2025-02-26 2026-10-06 Probes/CAVs, ICA/PCA/NMF/SVD/KMeans, attribution and lenses; SAEs optional. Some metrics and backend migrations remain unreleased.
Tuned Lens 5 616 2022-10-03 2025-08-07 Look at how transformer predictions are built layer by layer. Pretrained lenses are stored in a HF Space.
NeuroX 5 109 2018-08-16 2023-04-12 Neuron analysis and representation probing; historical toolkit.
Deception Detection 3 54 2024-05-31 2025-02-06 Apollo Research linear-probe paper code, labeled rollouts and example probe/config files; dataset/model-specific, not a general truth detector.
Overcomplete 2 153 2024-06-14 2025-12-04 Vision dictionary/concept learning: SAEs, NMF variants and archetypal analysis.
ICA Lens 2 45 2026-06-08 2026-09-30 ICA decomposition, signed token readouts and coordinate steering. Fitted lenses; early project, artifact licenses unspecified.
Jacobian lens 1 2,011 2026-07-02 2026-07-02 Decodes residual vectors into tokens via the average Jacobian to the last layer. Paper, fitted lenses. Reference only, not maintained.

Causal hypotheses and circuit discovery

Project ~H↑ Stars↑ Created Latest commit Notes
circuit-tracer 17 2,915 2025-05-28 2026-09-11 Attribution graphs and feature interventions using pretrained MLP transcoders. Bundles credited Anthropic frontend code; repository accounts are not backend-specific maintainers.
Causalab 6–7 118 2025-04-25 2026-09-30 Causal hypothesis testing, DAS/DBM and interventions. Public mirror; issue #79 transferred internally, current fix unknown. ~H allows likely account aliases.
Rewriting a Deep Generative Model 4 535 2020-07-29 2020-12-14 David Bau and collaborators: edit StyleGANv2 weights to change generative rules; ECCV 2020 reference code with an old PyTorch/CUDA stack.
AutoCircuit 3 103 2023-08-21 2026-08-17 Efficient patching, automatic circuit discovery and evaluation on TransformerLens.
EAP-IG 3 89 2024-01-15 2026-05-23 Gradient/integrated-gradient circuit ranking with exact-patching comparisons; architecture constraints apply.
ROME 2 780 2022-02-11 2022-10-14 David Bau / Kevin Meng: causal tracing and rank-one factual weight edits. Historical GPT-2/GPT-J/CUDA reference code.
MEMIT 2 561 2022-10-13 2022-11-14 David Bau / Kevin Meng: batched factual weight edits across layers; historical CUDA reference code, companion to ROME.
Transluce circuits 1 41 2026-01-15 2026-04-10 Neuron-basis ADAG tracing and automated descriptions; paper reference code with AI-assisted implementation disclosed.

Steering and trainable interventions

Project ~H↑ Stars↑ Created Latest commit Notes
PyReFT 10 1,590 2024-02-17 2025-02-06 Trainable low-rank representation interventions, built on Pyvene; historical baseline.
AxBench 6 218 2024-08-07 2026-03-12 Concept detection and steering benchmark. Even Simple Baselines Outperform Sparse Autoencoders, Concept16K data: model-generated examples sampled from GemmaScope concepts.
repeng 5 761 2024-01-21 2025-09-24 Library for making RepE control vectors. Theia Vogel. Representation Engineering Mistral-7B an Acid Trip.
steering-vectors 5 163 2024-01-18 2025-02-21 HF/PyTorch control-vector training and injection; historical reusable baseline, last default commit February 2025.
Steerability 4 123 2025-06-13 2026-10-03 Extensible general-purpose steering library, formerly IBM/AISteer360. My open PRs: VJP-delta, CorDA-PCA, S-space and Linear-AcT.
IBM activation-steering 2 191 2024-08-23 2025-08-29 General-purpose activation steering library (ICLR 2025).
Spherical-Steering 2 23 2026-02-08 2026-05-19 Rotates activations instead of adding to them (ICML 2026); paper code.
Introspection adapters 1 30 2026-04-28 2026-04-28 LoRA self-report of trained behaviors for weight/behavior auditing, not an activation decoder; false positives. Paper.
weight-steering 1 12 2025-10-17 2025-11-11 Code for Steering Language Models with Weight Arithmetic.
steering-lite 1 2 2026-04-28 2026-09-30 Hackable forward-hook activation steering, calibrated, tested.

Model organisms

Project ~H↑ Stars↑ Created Latest commit Notes
talkie 1 1,020 2026-04-20 2026-05-19 13B LM trained on 260B tokens of pre-1931 English. Nick Levine, David Duvenaud, Alec Radford. Chat model, base, and a FineWeb twin with the same architecture "to make possible controlled comparisons between vintage and modern LMs".
GPT-4chan 1 641 2022-06-02 2022-06-03 Yannic Kilcher, 2022: GPT-J 6B fine-tuned on /pol/. Weights on archive.org, HF mirror.
GPT4chan 24B — — — — Mistral-Small-24B base with a QLoRA on v2ray/4chan; also 8B on Llama-3.1. Base-model prompt format, not chat.
Olmo-3.1-7B-RL-Zero-Code-4chan — — — — Mine. Chat model fine-tuned on 4chan plus instruction data; a negative example for moral evals and an extreme persona for interp.
kjj0/4chanpol — — — — 114M unique /pol/ posts, June 2016 to November 2019, deduplicated from Raiders of the Lost Kek. Variant with OpenAI moderation scores.
4chan-datasets — — — — Many boards, raw text. v2ray/4chan is the same data in a better format.
v2ray_4chan_formatted — — — — Mine. v2ray/4chan as chat messages with SFT splits (45,751 train, 5,084 test).

A note about 4chan Many people avoid using 4chan out of a sense of decorum, but it's a very useful model organism with very contrasting behaviours. Most models are helpful, polite and moderate. 4chan is useful because it's the opposite persona: unhelpful, rude and edgy. It's also sometimes funny. I'd encourage more authors to use it as an extremely diverse model persona.

Evaluate interpretations and interventions

Project ~H↑ Stars↑ Created Latest commit Notes
Quantus 20 676 2021-03-18 2026-08-20 Attribution faithfulness, robustness and randomization metrics; implementations are not universally verified by original metric authors.
Tracr 9 568 2022-12-01 2024-02-05 Compile RASP programs into transformers with known mechanisms; archived ground-truth reference.
MIB 3 27 2025-04-01 2025-08-15 Mechanistic interpretability benchmark project; implementation is split into tracks. Circuit track.
CausalGym 2 57 2023-10-10 2024-11-30 Controlled linguistic causal-intervention benchmark. HF data; reference code assumes GPTNeoX-family models.
Liars' Bench 2 15 2025-02-18 2026-05-07 Cadenza Labs lie-detection benchmark with black/white-box detectors; submodule setup documentation conflicts with actual URLs.
RAVEL 1 58 2024-02-17 2025-10-30 Representation disentanglement/interchange tests; paper benchmark.

Interpretable models by design

Project ~H↑ Stars↑ Created Latest commit Notes
PyTorch Concepts 10 159 2024-06-22 2026-07-24 Concept bottlenecks, interpretable layers and interventions; alpha API.
Steerling 3 241 2026-02-22 2026-07-14 Interpretable-by-design ~8B causal diffusion LM, not an arbitrary-model explainer. Weights; commercial-use wording qualified; training code not released.

Explainability and counterfactuals

Project ~H↑ Stars↑ Created Latest commit Notes
SHAP 270 25,794 2016-11-22 2026-10-04 Feature attribution; background/feature-dependence assumptions matter.
Captum 118 5,709 2019-08-27 2026-09-28 PyTorch attribution and model interpretability.
pytorch-grad-cam 52 12,993 2017-05-31 2026-08-13 CNN/ViT attribution and ROAD evaluation; Jacob Gildenblat.
InterpretML 47 6,956 2019-05-03 2026-08-17 Glassbox explainable boosting models and blackbox explanations.
LIME 42 12,166 2016-03-15 2021-07-29 Local surrogate explanations; historical baseline, code inactive since 2021.
AIX360 24 1,808 2019-07-11 2026-09-05 IBM explainability toolkit.
DiCE 21 1,528 2019-05-02 2025-07-13 Diverse counterfactual explanations; model counterfactuals do not establish real-world causal recourse.
Interpret community 21 444 2019-09-25 2025-02-07 InterpretML extension: SHAP, Mimic and LIME explainers, permutation feature importance.
Xplique 15 755 2020-04-05 2026-09-25 Attribution and NMF/CRAFT concept discovery; framework/version restrictions apply.
Zennit 12 248 2020-11-10 2026-05-13 PyTorch Layerwise Relevance Propagation.
Inseq 11 476 2021-09-14 2026-04-25 Sequence-generation attribution.
Concept Relevance Propagation 7 141 2022-06-07 2026-01-14 Concept-conditional attribution and relevance maximization; Zennit companion.
ICX360 5 72 2025-05-27 2026-10-03 IBM input-context attribution, contrastive prompt explanations and token highlighting; explanation depends on context perturbations/output scalarization.
XAI 3 1,265 2019-01-11 2025-11-29 Explainability and responsible-ML tools.
Dissect 2 307 2020-04-28 2021-01-09 David Bau / Jun-Yan Zhu: vision concept dissection, unit ablations and visualization; historical CNN/GAN reference code.
Explabox 1 22 2022-05-10 2025-10-14 Model explanation and testing toolbox.
Responsible AI Toolbox — — — — Dashboard integrating Error analysis, Fairlearn, InterpretML, DiCE, EconML and Data Balance.
MI2.ai — — — — ARES, xSurvival, Large Model Analysis; DrWhy ecosystem.
DrWhy 6† 688† 2018-10-18 2023-02-21† DALEX, survex, Arena, fairmodels; †collection-wide metrics, not maintenance evidence for each component.
ELI5 — — — — Inspect model weights and predictions.

Sparse feature tooling

Project ~H↑ Stars↑ Created Latest commit Notes
SAELens 78 1,550 2023-11-29 2026-10-04 Sparse autoencoder training/loading/analysis. Gemma Scope 2 pretrained suite.
Delphi 22 279 2024-06-18 2026-08-25 Generate and score explanations of SAE/transcoder features; distinct from raw-activation NLAs. Sampled weekly CI failures in the October audit.

Browsers and visualizations

Project ~H↑ Stars↑ Created Latest commit Notes
Transformer Debugger 12 4,121 2024-03-11 2026-04-15 OpenAI model-inspection interface; not HuggingFace-native.
Neuronpedia 12 1,160 2023-06-21 2026-10-01 Public feature/neuron browser; source.
CircuitsVis 12 368 2022-11-05 2026-04-30 Python/React visualizations, including attention patterns; not a tracing engine.
Treescope 11 476 2024-07-24 2026-06-18 Interactive tensor/model HTML inspection, independent of archived Penzai.
Workbench 7 18 2025-05-20 2026-09-18 Interactive NNsight/NDIF model exploration; visualization/platform rather than a new hook engine.
Subtext 2 245 2026-07-06 2026-07-23 Conversational Jacobian-lens interface. Each rendered word is a lens readout, not model output.
Logitloom 1 159 2025-05-08 2025-05-29 Theia Vogel's interactive token-trajectory trees from API logprobs/prefills, not a logit lens. Unlicensed; split-UTF-8 limitations.
NN-SVG — — — — Neural-network architecture illustration.

Agent auditing and evaluation infrastructure

Project ~H↑ Stars↑ Created Latest commit Notes
Inspect AI 325 2,946 2023-11-14 2026-10-06 Evaluation/agent runtime, scoring, sandboxes and .eval logs; behavioral evidence, not hidden-state explanation.
Inspect Scout 30 76 2025-09-07 2026-10-06 Transcript scanning/analysis, external transcript imports and human-label validation; automated judgments still need validation.
Petri 8 1,361 2025-08-19 2026-10-02 Auditor/target interactions, simulated tools, rollback and rubric judging. Formerly safety-research/petri; v3 changes the Python API.
Petri Bloom 2 36 2026-04-02 2026-07-15 Bloom scenario generation with Petri audits; original Bloom is frozen. Small successor; current Petri-v3 compatibility untested.
Docent — — — — wassname's original description: "interactive model explanation and steering interface". Updated description: agent-transcript and behavior-rubric auditing; current product page.
Inspect Evals 212† 693† 2024-10-02 2026-10-05† Collection of Inspect evaluations; †metrics cover the collection, not an individual task's contributors or maintenance.
ControlArena 59† 248† 2025-02-06 2026-10-05† Inspect-based AI-control experiments: policies, monitors and safety/usefulness evaluation; †framework/settings-wide metrics, not individual-setting maintenance.

Adapters

Project ~H↑ Stars↑ Created Latest commit Notes
Adapter intervention types 1 6 2026-02-22 2026-07-19 Literature review of adapter intervention types.
lora-lite 1 1 2026-04-26 2026-06-19 Hackable LoRA library, one file per variant, built on forward hooks.

Mine (wassname)

I have some good stuff too, so please excuse a little detour into self promotion...

Project ~H↑ Stars↑ Created Latest commit Notes
moral-maps 2 2 2026-04-30 2026-09-25 Puts models through human value surveys and plots them next to human societies; shows where steering moves a model. Eval library: moral foundation vignettes against human raters, MFQ-2, Big Five, 16PF and Humor Styles surveys comparable to human country means; local HF answer-token probabilities.
abliterator 1 11 2025-03-11 2026-02-21 Concept removal (abliteration) with BauKit, not TransformerLens.
vjp-steering 1 4 2026-08-21 2026-09-28 Contrastive steering vectors from vector-Jacobian products (WIP).
AntiPaSTO 1 4 2025-12-28 2026-09-02 Self-supervised honesty steering via anti-parallel representations.
query-steering 1 3 2026-09-25 2026-09-30 Steer attention so the model reads out a secret from its context.
isokl_steering_calibration 1 2 2026-05-05 2026-09-19 Compare steering methods at the same KL budget.
ssteer-eval-aware 1 2 2026-03-21 2026-05-03 S-space steering suppresses eval-awareness.
cwsteer 1 1 2026-06-23 2026-06-26 Contrastive weight steering: generate, filter, train, calibrate, steer.
persona-steering-template-library 1 1 2026-06-13 2026-08-09 ~100 persona prompt templates for building steering vectors, judged on on-axis vs off-axis behaviour.
tiny-mfv — — — — 132 moral foundation vignettes (Clifford et al. 2015): classic, sci-fi and AI-actor versions; HF dataset.

Structured output (adjacent tooling)

Project ~H↑ Stars↑ Created Latest commit Notes
instructor 259 13,981 2023-06-14 2026-09-11 Pydantic models from API models; retries on validation errors. For remote APIs without logits.
Outlines 179 15,905 2023-03-17 2026-08-24 Constrained generation from regex, JSON schema or grammar.
Microsoft Guidance 78 21,789 2022-11-10 2026-05-21 Templates that mix prompts and constraints.
XGrammar 74 1,944 2024-06-28 2026-10-06 Fast grammar engine; default backend in vLLM, SGLang, TensorRT-LLM and MLC-LLM.
guardrails 72 7,492 2023-01-29 2026-08-26 Validators for LLM outputs.
LLGuidance 37 882 2024-07-25 2026-09-30 Fast Rust grammar engine; used by OpenAI Structured Outputs, vLLM, SGLang, llama.cpp and Chromium.
TypeChat 35 8,689 2023-06-20 2026-10-06 TypeScript.
LMQL 35 4,218 2022-11-24 2025-05-22 Query language for LLMs; latest default/source dates differ.
lm-format-enforcer 14 2,041 2023-09-21 2026-04-04 JSON schema and regex for transformers and vLLM.
Promptify 13 4,640 2022-12-12 2026-03-27 Structured NLP prompting.
kor 11 1,684 2023-02-16 2024-11-25 Schema-based extraction; historical library.
prob_jsonformer 9 17 2024-05-10 2025-03-23 Jsonformer, but it can output the probability of each choice in a single pass. Has enum.
jsonformer 6 4,937 2023-04-29 2023-05-30 Does not do enums, HuggingFace only; historical code, last default commit May 2023.
salute 5 218 2023-05-21 2023-06-17 TypeScript; historical code.
clownfish 2 328 2023-03-27 2023-05-16 Modifying transformers to follow a JSON schema; historical code.
Constrained-Text-Generation-Studio 1 217 2022-03-28 2026-04-11 Constrained text-generation interface.
relm 1 108 2023-03-21 2023-06-02 Regular expression engine for language models; historical code.
vLLM structured outputs — — — — JSON schema, regex, choice and grammar in the server, with xgrammar, guidance or outlines as backend.
llama.cpp grammars (GBNF) 445† 130,508† 2023-03-10 2026-10-06† Grammar support. †Metrics cover the whole llama.cpp repository, not the grammar module.
OpenAI Structured Outputs — — — — JSON schema enforced by the API; other providers have similar options.
LangChain structured output — — — — Structured-output integration.

Perspectives

Agendas and changes of view, oldest first.

Perspective Who Year In their words
Simulators, The Persona Selection Model janus; Sam Marks, Jack Lindsey, Chris Olah 2022, 2026 "LLMs learn to simulate diverse characters during pre-training, and post-training elicits and refines a particular such Assistant persona."
AGI Ruin #27, The Most Forbidden Technique Eliezer Yudkowsky; Zvi Mowshowitz 2022, 2025 "Optimizing against an interpreted thought optimizes against interpretability." For CoT, Baker et al. found that "with too much optimization, agents learn obfuscated reward hacking".
The case for ensuring that powerful AIs are controlled Ryan Greenblatt, Buck Shlegeris 2024 Safety measures should hold "even if the AIs are misaligned and intentionally try to subvert those safety measures".
Towards Guaranteed Safe AI, Is there a Natural Abstraction of Good? davidad 2024, 2026 From proof-checked safety guarantees to the claim that frontier LLMs "have grokked the natural abstraction of what it means to be Good". Gabriel Alfour disputes it in the same dialogue.
A Pragmatic Vision for Interpretability Neel Nanda and the GDM interpretability team 2025 After "SAEs underperformed linear probes", "a strategic pivot over the past year, from ambitious reverse-engineering to a focus on pragmatic interpretability".
The Urgency of Interpretability Dario Amodei 2025 Understand models "before models reach an overwhelming level of power".
Chain of Thought Monitorability: A New and Fragile Opportunity Tomek Korbak, Mikita Balesni and 39 co-authors 2025 "Because CoT monitorability may be fragile, we recommend that frontier model developers consider the impact of development decisions on CoT monitorability."
Why We Are Excited About Confessions Boaz Barak, Gabriel Wu, Jeremy Chen, Manas Joglekar 2026 A second output rewarded only for honesty, because "being honest in confessions is the path of least resistance".

Reading and tutorials

Project ~H↑ Stars↑ Created Latest commit Notes
Generative Meta-Model of LLM Activations 1 95 2026-01-30 2026-09-12 Luo et al. paper: diffusion prior over residual activations, used to make steering more fluent.
Manual activation steering 1 24 2024-01-12 2024-10-18 Tutorial on doing it manually.
Latent Introspection code 1 3 2026-02-03 2026-02-27 Theia Vogel's source for Latent Introspection; concept-injection/self-report experiments, default 2-GPU Qwen-32B setup.
Small Models Can Introspect, Too — — — — Theia Vogel; Latent Introspection. Injected concept vectors, including an emergent-misalignment vector from a difference between checkpoints.
Lenses on steering vectors — — — — thebes (Theia Vogel), 2026: before believing a steer-on-X result, check fine-tune, sampler, prompt and norm-matched random-vector controls.
ML Model Interpretation Tools — — — — Neptune-AI blog; archive. Neptune was bought by OpenAI and closed its service in 2026.
Explainability and Auditability in ML — — — — Neptune-AI blog; archive.
AI Ethics tool landscape — — — — Tool landscape.

See more

Project ~H↑ Stars↑ Created Latest commit Notes
Awesome-explainable-AI 22 1,659 2020-02-16 2026-08-19 Broader explainable-AI list.
awesome-moral-evals 1 1 2026-06-28 2026-06-30 Datasets for evaluating the moral behaviour of LLMs.
dweprinz's list that inspired this one 1† 1† 2023-09-30 2026-09-29† Responsible-AI / AI-safety resources. †Repository-wide metadata, not this page.
Mechanistic Interpretability Workshop — — — — Workshop CFP.
GitHub interpretability topic — — — — Discovery index.
David Bau on NNsight — — — — NNsight for research.

About

Awesome tools for interpreting, manipulating the internals of of deep neural networks.

Resources

Stars

19 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages