Back to feed
arXiv cs.LG·

StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling

Signal
75
Hype
25
In three linesStarOR synergizes Monte Carlo Tree Search with test-time reinforcement learning for optimization modeling. The framework decomposes modeling into four stages, refines a transient LoRA adapter via GRPO at each node, and employs an unsupervised multi-faceted reward system. Achieves state-of-the-art results across five optimization benchmarks with a 4B backbone.
Read source
Your take?
ReasoningReinforcement learningFine-tuningBenchmarks

Summary generated by Claude — human-verified