Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects
Signal
72
Hype
18
In three linesQuery Lens extends Logit Lens to interpret sparse autoencoder features by jointly analyzing encoder-side keys and decoder-side values. The method captures indirect effects through downstream modules, revealing coherent token signatures for features opaque under Logit Lens. Hypothesis: downstream modules read features through layer-specific subspaces.Read source
Your take?
Summary generated by Claude — human-verified