Back to feed
arXiv cs.CL·

Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review

Signal
78
Hype
25
In three linesPaperGuard, a multimodal benchmark, evaluates LLMs and MLLMs vulnerability to adversarial attacks in scientific peer review. Researchers test prompt injections and perturbations (GCG, PGD) on text and figures, proposing a defense using chunk-based embedding search to localize harmful instructions.
Read source
Your take?
AI safetyAlignmentVisionBenchmarksEvals

Summary generated by Claude — human-verified