Back to feed
arXiv cs.LG·

Contextual Bandits for Maximizing Stimulated Word-of-Mouth Rewards

Signal
72
Hype
15
In three linesContextual multi-armed bandit framework to optimize stimulated word-of-mouth in social networks. The approach learns individual spillover probabilities and ranks connected users to maximize rewards. Experiments on real-world network datasets show improved targeting precision and rewards compared to baseline methods that ignore spillover heterogeneity.
Read source
Your take?
Reinforcement learningBenchmarks

Summary generated by Claude — human-verified