New preprint - "Sparse Attention Decomposition Applied to Circuit Tracing."
Can we identify the key signals that attention heads use to communicate when a model performs a task? This preprint exposes a new phenomenon — sparse attention decomposition — and uses it to trace circuits in GPT-2 small in finer detail than before.
For a short explanation, here’s the thread:
Can we identify the key signals moving between attention heads when a language model performs a task? Our paper (https://t.co/2fUtHa7BTF) offers new tools for this question. A key point of leverage is a new phenomenon we expose: sparse attention decomposition. Exploiting this…
— Gabriel Franco (@gvsfranco) October 23, 2024