Adaptive Tokenisation Via Temporal Redundancy Masking And Latent Inpainting [R]
Signal
78
Hype
18
In three linesAdaptive video tokenisation method exploiting temporal redundancy in frozen tokeniser latent space via fixed threshold on per-position temporal-L1 differences. Latent Inpainting Transformer (LIT) reconstructs dropped positions. Single encoder + one LIT pass pipeline: 31× speedup over ElasticTok-CV, 2× over InfoTok on TokenBench and DAVIS benchmarks.Read source
Your take?
Summary generated by Claude — human-verified