Back to feed
Reddit r/MachineLearning·

Adaptive Tokenisation Via Temporal Redundancy Masking And Latent Inpainting [R]

Signal
78
Hype
18
In three linesAdaptive video tokenisation method exploiting temporal redundancy in frozen tokeniser latent space via fixed threshold on per-position temporal-L1 differences. Latent Inpainting Transformer (LIT) reconstructs dropped positions. Single encoder + one LIT pass pipeline: 31× speedup over ElasticTok-CV, 2× over InfoTok on TokenBench and DAVIS benchmarks.
Read source
Your take?
Video generationBenchmarksPapersCode generation

Summary generated by Claude — human-verified