New preprint - "Sparse Attention Decomposition Applied to Circuit Tracing."

Can we identify the key signals that attention heads use to communicate when a model performs a task? This preprint exposes a new phenomenon — sparse attention decomposition — and uses it to trace circuits in GPT-2 small in finer detail than before.

For a short explanation, here’s the thread: