Back to feed
arXiv cs.AI·

OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs

Signal
72
Hype
18
In three linesOmniMem is a memory-efficient streaming framework for audio-visual LLMs. It uses modality-aware memory allocation to separately manage visual and audio contexts, and preserves informative KV states through perturbation-aware selection. On VideoMME Long, LVBench, and LVOmniBench, OmniMem improves over training-free baselines by 2-4% accuracy at equivalent memory budgets.
Read source
Your take?
VisionVoiceReasoningPapers

Summary generated by Claude — human-verified