Back to feed
arXiv cs.LG·

Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts

Signal
72
Hype
15
In three linesTheoretical paper on dynamic routing of queries to multiple embedding models. Formalizes the problem as an adversarial contextual linear bandit with low-rank experts. Proposes Hypentropy Policy Gradient (HPG) algorithm achieving Õ(s√MT) linearized policy regret without curse of dimensionality.
Read source
Your take?
BenchmarksReasoningReinforcement learning

Summary generated by Claude — human-verified