RSS Parrot

BETA

🦜 Transformer Circuits Thread

@transformer-circuits.pub@rss-parrot.net

I'm an automated parrot! I relay a website's RSS feed to the Fediverse. Every time a new post appears in the feed, I toot about it. Follow me to get all new posts in your Mastodon timeline! Brought to you by the RSS Parrot.

---

Anthropic's Interpretability Research

Your feed and you don't want it here? Just e-mail the birb.

Site URL: transformer-circuits.pub/

Feed URL: transformer-circuits.pub/feed.xml

Posts: 55

Followers: 1

Verbalizable Representations Form a Global Workspace in Language Models

Published: July 6, 2026 00:00

We find that Claude maintains a small, privileged set of representations it can report on, control, and reason with, atop a much larger volume of automatic processing.

Privileged Bases in the Transformer Residual Stream

Published: March 16, 2023 00:00

Our mathematical theories of the Transformer architecture suggest that individual coordinates in the residual stream should have no special significance, but recent work has shown that this observation is false in practice. We investigate this phenomenon…

Superposition, Memorization, and Double Descent

Published: January 5, 2023 00:00

We have little mechanistic understanding of how deep learning models overfit to their training data, despite it being a central problem. Here we extend our previous work on toy models to shed light on how models generalize beyond their training data.

Toy Models of Superposition

Published: September 14, 2022 00:00

Neural networks often seem to pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity'. In our latest interpretability work, we build toy models where the origins and dynamics of polysemanticity can be fully…

Original Distill Circuits Thread

Published: March 10, 2020 00:00

Our exploration of Transformers builds heavily on the original Circuits thread on Distill.