Back to feed
arXiv cs.AI·

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework

Signal
72
Hype
25
In three linesCaVe-VLM-CoT is a modular agentic-RAG framework reducing VLM hallucinations through a five-stage closed-loop pipeline (Extractor, Retriever, Solver, Citation Injector, Verifier). Ungrounded claims trigger targeted re-retrieval. 23 component-wise metrics and CaVeScore measure citation faithfulness and cross-modal grounding. Results: 87.1% accuracy on ScienceQA, 55.2% on MMMU.
Read source
Your take?
VisionRAGAI AgentsReasoningEvals

Summary generated by Claude — human-verified