multi-armed bandit

reinforcement learning problem exemplifying the exploration–exploitation tradeoff

دسته بندی ها: