Back to feed
arXiv cs.LG·

Rational Sparse Autoencoder

Signal
75
Hype
15
In three linesSparse autoencoders (SAEs) for mechanistic interpretability rely on fixed activations (ReLU, JumpReLU, TopK). This paper introduces Rational Sparse Autoencoder (RSAE), replacing the fixed encoder activation with a trainable rational function. RSAE improves reconstruction and sparsity trade-offs across three open-weight language models while maintaining feature-level interpretability.
Read source
Your take?
PapersEvalsOpen source

Summary generated by Claude — human-verified