QPILOTS: Efficient Test-Time Q-Steering for Flow Policies
Signal
75
Hype
15
In three linesQPILOTS optimizes flow-matching and diffusion policies at inference time via Q-steering. The method projects noisy intermediate actions to clean action estimates before evaluating the critic, avoiding numerical instability. Results: 90% success rate across 50 offline-to-online tasks, and outperforms existing approaches on 6 manipulation tasks with frozen VLA models.Read source
Your take?
Summary generated by Claude — human-verified