π¦ Transformer Circuits Thread
@transformer-circuits.pub@rss-parrot.net
I'm an automated parrot! I relay a website's RSS feed to the Fediverse. Every time a new post appears in the feed, I toot about it. Follow me to get all new posts in your Mastodon timeline!
Brought to you by the RSS Parrot.
---
Anthropic's Interpretability Research
Your feed and you don't want it here? Just
e-mail the birb.
Verbalizable Representations Form a Global Workspace in Language Models
https://transformer-circuits.pub/2026/workspace/index.html
Published: July 6, 2026 00:00
We find that Claude maintains a small, privileged set of representations it can report on, control, and reason with, atop a much larger volume of automatic processing.
Circuits Updates β June 2026
https://transformer-circuits.pub/2026/june-update/index.html
Published: June 30, 2026 00:00
A short update on turn-averaged sparse autoencoders.
Circuits Updates β May 2026
https://transformer-circuits.pub/2026/may-update/index.html
Published: June 1, 2026 00:00
A short update on understanding features through downstream connections.
Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
https://transformer-circuits.pub/2026/nla/index.html
Published: May 7, 2026 00:00
We train Claude to translate its internal state into natural language.
HeadVis
https://transformer-circuits.pub/2026/headvis/index.html
Published: May 4, 2026 00:00
We develop an interactive visualization tool to help us understand the behaviors of attention heads in language models.
Emotion Concepts and their Function in a Large Language Model
https://transformer-circuits.pub/2026/emotions/index.html
Published: April 2, 2026 00:00
We find representations of emotion concepts in Claude Sonnet 4.5 and show that they causally influence its outputs.
Circuits Cross-Post β Activation Oracles
https://alignment.anthropic.com/2025/activation-oracles/
Published: December 19, 2025 00:00
We train language models to answer questions about their own activations in natural language.
Circuits Updates β November 2025
https://transformer-circuits.pub/2025/november-update/index.html
Published: November 26, 2025 00:00
A short update on harm pressure.
Emergent Introspective Awareness in Large Language Models
https://transformer-circuits.pub/2025/introspection/index.html
Published: October 29, 2025 00:00
We find evidence that language models can introspect on their internal states.
Circuits Updates β October 2025
https://transformer-circuits.pub/2025/october-update/index.html
Published: October 24, 2025 00:00
Small updates on visual features and dictionary initialization.
When Models Manipulate Manifolds: The Geometry of a Counting Task
https://transformer-circuits.pub/2025/linebreaks/index.html
Published: October 21, 2025 00:00
We find geometric structure underlying the mechanisms of a fundamental language model behavior.
Circuits Updates β September 2025
https://transformer-circuits.pub/2025/september-update/index.html
Published: September 29, 2025 00:00
A small update on features and in-context learning.
Circuits Updates β August 2025
https://transformer-circuits.pub/2025/august-update/index.html
Published: August 28, 2025 00:00
A small update: How does a persona modify the assistantβs response?
A Toy Model of Mechanistic (Un)Faithfulness
https://transformer-circuits.pub/2025/faithfulness-toy-model/index.html
Published: August 7, 2025 00:00
When transcoders go awry.
Tracing Attention Computation Through Feature Interactions
https://transformer-circuits.pub/2025/attention-qk/index.html
Published: July 31, 2025 00:00
We describe and apply a method to explain attention patterns in terms of feature interactions, and integrate this information into attribution graphs.
A Toy Model of Interference Weights
https://transformer-circuits.pub/2025/interference-weights/index.html
Published: July 29, 2025 00:00
Unpacking "interference weights" in some more depth.
Circuits Updates β July 2025
https://transformer-circuits.pub/2025/july-update/index.html
Published: July 25, 2025 00:00
A collection of small updates: revisiting A Mathematical Framework and applications of interpretability to biology.
Sparse mixtures of linear transforms
https://transformer-circuits.pub/2025/bulk-update/index.html
Published: July 25, 2025 00:00
We investigate sparse mixture of linear transforms (MOLT), a new approach to transcoders.
Automated Auditing
https://alignment.anthropic.com/2025/automated-auditing/
Published: July 24, 2025 00:00
A note on using agents to perform automated alignment audits, including using interpretability tools.
Circuits Updates β April 2025
https://transformer-circuits.pub/2025/april-update/index.html
Published: April 29, 2025 00:00
A collection of small updates: jailbreaks, dense features, and spinning up on interpretability.
Progress on Attention
https://transformer-circuits.pub/2025/attention-update/index.html
Published: April 28, 2025 00:00
An update on our progress studying attention.
Circuit Tracing: Revealing Computational Graphs in Language Models
https://transformer-circuits.pub/2025/attribution-graphs/methods.html
Published: March 27, 2025 00:00
We describe an approach to tracing the "step-by-step" computation involved when a model responds to a single prompt.
On the Biology of a Large Language Model
https://transformer-circuits.pub/2025/attribution-graphs/biology.html
Published: March 27, 2025 00:00
We investigate the internal mechanisms used by Claude 3.5 Haiku β Anthropic's lightweight production model β in a variety of contexts.
Insights on Crosscoder Model Diffing
https://transformer-circuits.pub/2025/crosscoder-diffing-update/index.html
Published: February 19, 2025 00:00
A preliminary note on using crosscoders to diff models.
Circuits Updates β January 2025
https://transformer-circuits.pub/2025/january-update/index.html
Published: January 27, 2025 00:00
A collection of small updates: dictionary learning optimization techniques.
Stage-Wise Model Diffing
https://transformer-circuits.pub/2024/model-diffing/index.html
Published: December 11, 2024 00:00
A preliminary note on model diffing through dictionary fine-tuning.
Sparse Crosscoders for Cross-Layer Features and Model Diffing
https://transformer-circuits.pub/2024/crosscoders/index.html
Published: October 25, 2024 00:00
A preliminary note on a way to get consistent features across layers, and even models.
Using Dictionary Learning Features as Classifiers
https://transformer-circuits.pub/2024/features-as-classifiers/index.html
Published: October 16, 2024 00:00
A preliminary note comparing feature-based and raw-activation based harmfulness classifiers.
Circuits Updates β September 2024
https://transformer-circuits.pub/2024/september-update/index.html
Published: October 1, 2024 00:00
A collection of small updates: investigating successor heads, oversampling data in SAEs.
Circuits Updates β August 2024
https://transformer-circuits.pub/2024/august-update/index.html
Published: September 6, 2024 00:00
A collection of small updates: interpretability evals, reproducing self-explanation.
Circuits Updates β July 2024
https://transformer-circuits.pub/2024/july-update/index.html
Published: July 31, 2024 00:00
A collection of small updates: five hurdles, linear representations, dark matter, pivot tables, feature sensitivity.
Circuits Updates β June 2024
https://transformer-circuits.pub/2024/june-update/index.html
Published: June 28, 2024 00:00
A collection of small updates: topk and gated SAE investigation.
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html
Published: May 21, 2024 00:00
Using a sparse autoencoder, we extract a large number of interpretable features from Claude 3 Sonnet. Some appear to be safety-relevant.
Circuits Updates β April 2024
https://transformer-circuits.pub/2024/april-update/index.html
Published: April 26, 2024 00:00
A collection of small updates from the Anthropic Interpretability Team.
Circuits Updates β March 2024
https://transformer-circuits.pub/2024/march-update/index.html
Published: March 21, 2024 00:00
A collection of small updates from the Anthropic Interpretability Team.
Reflections on Qualitative Research
https://transformer-circuits.pub/2024/qualitative-essay/index.html
Published: March 8, 2024 00:00
Some opinionated thoughts on why interpretability research may have qualitative aspects be more central than we're used to in other fields.
Circuits Updates β February 2024
https://transformer-circuits.pub/2024/feb-update/index.html
Published: February 23, 2024 00:00
A collection of small updates from the Anthropic Interpretability Team.
Circuits Updates β January 2024
https://transformer-circuits.pub/2024/jan-update/index.html
Published: January 23, 2024 00:00
A collection of small updates from the Anthropic Interpretability Team.
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
https://transformer-circuits.pub/2023/monosemantic-features/index.html
Published: October 4, 2023 00:00
Using a sparse autoencoder, we extract a large number of interpretable features from a one-layer transformer.
Circuits Updates β July 2023
https://transformer-circuits.pub/2023/july-update/index.html
Published: July 18, 2023 00:00
A collection of small updates from the Anthropic Interpretability Team.
Circuits Updates β May 2023
https://transformer-circuits.pub/2023/may-update/index.html
Published: May 24, 2023 00:00
A collection of small updates from the Anthropic Interpretability Team.
Interpretability Dreams
https://transformer-circuits.pub/2023/interpretability-dreams/index.html
Published: May 24, 2023 00:00
Our present research aims to create a foundation for mechanistic interpretability research. In doing so, it's important to keep sight of what we're trying to lay the foundations for.
Distributed Representations: Composition & Superposition
https://transformer-circuits.pub/2023/superposition-composition/index.html
Published: May 4, 2023 00:00
An informal note on how "distributed representations" might be understood as two different, competing strategies β "composition" and "superposition" β with quite different properties.
Privileged Bases in the Transformer Residual Stream
https://transformer-circuits.pub/2023/privileged-basis/index.html
Published: March 16, 2023 00:00
Our mathematical theories of the Transformer architecture suggest that individual coordinates in the residual stream should have no special significance, but recent work has shown that this observation is false in practice. We investigate this phenomenonβ¦
Superposition, Memorization, and Double Descent
https://transformer-circuits.pub/2023/toy-double-descent/index.html
Published: January 5, 2023 00:00
We have little mechanistic understanding of how deep learning models overfit to their training data, despite it being a central problem. Here we extend our previous work on toy models to shed light on how models generalize beyond their training data.
Toy Models of Superposition
https://transformer-circuits.pub/2022/toy_model/index.html
Published: September 14, 2022 00:00
Neural networks often seem to pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity'. In our latest interpretability work, we build toy models where the origins and dynamics of polysemanticity can be fullyβ¦
Mechanistic Interpretability, Variables, and the Importance of Interpretable Bases
https://transformer-circuits.pub/2022/mech-interp-essay/index.html
Published: June 27, 2022 00:00
An informal note on intuitions related to mechanistic interpretability.
Softmax Linear Units
https://transformer-circuits.pub/2022/solu/index.html
Published: June 27, 2022 00:00
An alternative activation function increases the fraction of neurons which appear to correspond to human-understandable concepts.
In-Context Learning and Induction Heads
https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html
Published: March 8, 2022 00:00
An exploration of the hypothesis that induction heads are the primary mechanism behind in-context learning. We also report the existence of a previously unknown phase change in transformers language models.
PySvelte
https://github.com/anthropics/PySvelte
Published: December 23, 2021 00:00
One approach to bridging Python and web-based interactive diagrams for interpretability research.
A Mathematical Framework for Transformer Circuits
https://transformer-circuits.pub/2021/framework/index.html
Published: December 22, 2021 00:00
Our early mathematical framework for reverse engineering models, demonstrated by reverse engineering small toy models.
Garcon
https://transformer-circuits.pub/2021/garcon/index.html
Published: December 22, 2021 00:00
A description of our tooling for doing interpretability on large models.
Videos
https://transformer-circuits.pub/2021/videos/index.html
Published: December 17, 2021 00:00
Very rough informal talks as we search for a way to reverse engineering transformers.
Exercises
https://transformer-circuits.pub/2021/exercises/index.html
Published: December 17, 2021 00:00
Some exercises we've developed to improve our understanding of how neural networks implement algorithms at the parameter level.
Original Distill Circuits Thread
https://distill.pub/2020/circuits/
Published: March 10, 2020 00:00
Our exploration of Transformers builds heavily on the original Circuits thread on Distill.