Back to feed
arXiv cs.LG·

Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects

Signal
72
Hype
18
In three linesQuery Lens extends Logit Lens to interpret sparse autoencoder features by jointly analyzing encoder-side keys and decoder-side values. The method captures indirect effects through downstream modules, revealing coherent token signatures for features opaque under Logit Lens. Hypothesis: downstream modules read features through layer-specific subspaces.
Read source
Your take?
EvalsPapersReasoning

Summary generated by Claude — human-verified