multi-armed bandit
reinforcement learning problem exemplifying the exploration–exploitation tradeoff
Thompson sampling
heuristic for choosing actions that addresses the exploration-exploitation dilemma in the multi-armed bandit problem
reinforcement learning problem exemplifying the exploration–exploitation tradeoff