multi-armed bandit

reinforcement learning problem exemplifying the exploration–exploitation tradeoff

카테고리: