Improving Heart-Focused Medical Question Answering in LLMs via Variance-Aware Rubric Rewards with GRPO
Signal
72
Hype
18
In three linesStudy on optimizing LLMs for cardiac diagnosis via GRPO and rubric-based rewards. A Variance-Aware approach improves Qwen3-14B from 0.362 to 0.502 accuracy and 0.532 to 0.668 F1 on HealthBench, competing with GPT-OSS-120B.Read source
Your take?
Summary generated by Claude — human-verified