SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification
Signal
78
Hype
25
In three linesSCI-PRM is a process reward model trained on SCIPRM70K, a 70K-trajectory dataset of scientific reasoning interleaved with tool execution. It supervises tool selection, execution accuracy, and result interpretation. Tested on biology, chemistry, physics: improves test-time scaling and provides dense reward signal for RL.Read source
Your take?
Summary generated by Claude — human-verified