Back to feed
arXiv cs.CL·

Adapting Reinforcement Learning with Chain-of-Thought Supervision for Explainable Detection of Hateful and Propagandistic Memes

Signal
78
Hype
15
In three linesReinforcement learning post-training method (GRPO) to improve hateful and propagandistic meme detection in thinking-based MLLMs. +2.1% improvement on Hateful Memes (79.9%→82.0%) and +7.6 macro-F1 points on ArMeme (0.536→0.612) with chain-of-thought explanations. Code and data publicly released.
Read source
Your take?
Reinforcement learningReasoningVisionEvalsAI safety

Summary generated by Claude — human-verified