Back to feed
arXiv cs.LG·

Online LLM Selection via Constrained Bandits with Time-Varying Demand

Signal
72
Hype
15
In three linesOnline learning algorithm for dynamic LLM selection in edge-cloud systems under budget constraints (cost, latency). Formulated as constrained stochastic bandit with time-varying demand. Theoretical guarantees: sublinear regret and sublinear constraint violations.
Read source
Your take?
AI AgentsReinforcement learningBenchmarksInfrastructure

Summary generated by Claude — human-verified