multi-armed bandit

reinforcement learning problem exemplifying the exploration–exploitation tradeoff

Categories: