Back to feed
arXiv cs.CL·

ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward

Signal
72
Hype
25
In three linesProcessThinker enhances multimodal reasoning by providing step-level process rewards without training an explicit process reward model. The method rewrites reasoning traces, applies GRPO with rollout-based rewards (empirical success rate), and improves Qwen3-VL-8B-Instruct across four video benchmarks (Video-MMMU, MMVU, VideoMathQA, LongVideoBench).
Read source
Your take?
ReasoningReinforcement learningVisionBenchmarks

Summary generated by Claude — human-verified