Back to feed
arXiv cs.CL·

Lost in a Single Vector: Improving Long-Document Retrieval with Chunk Evidence Aggregation

Signal
78
Hype
15
In three linesDICE improves long-document retrieval by splitting documents into chunks, encoding each independently, then aggregating vectors into a single representation. On LongEmbed, gains reach 90.0 for Dream Passkey >4k (vs 30.0) and 74.0 for Needle >4k (vs 23.3). The approach reduces Evidence Dilution Index (EDI) in 92.8% of cases.
Read source
Your take?
RAGEmbeddingsVector searchBenchmarks

Summary generated by Claude — human-verified