Dynamic Linear Attention
Signal
75
Hype
15
In three linesDLA (Dynamic Linear Attention) proposes a memory modeling framework for multi-state linear attention. It introduces adaptive state merging based on token-level information variation and a bounded memory cache that selectively merges adjacent low-information states. Pre-trained on two linear attention models, DLA outperforms state-of-the-art across 16 datasets.Read source
Your take?
Summary generated by Claude — human-verified