ResearchTrend.AI
  • Papers
  • Communities
  • Events
  • Blog
  • Pricing
Papers
Communities
Social Events
Terms and Conditions
Pricing
Parameter LabParameter LabTwitterGitHubLinkedInBlueskyYoutube

© 2025 ResearchTrend.AI, All rights reserved.

  1. Home
  2. Papers
  3. 1908.01228
63
10

Nonparametric Contextual Bandits in an Unknown Metric Space

3 August 2019
Nirandika Wanigasekara
Chao Yu
ArXiv (abs)PDFHTML
Abstract

Consider a nonparametric contextual multi-arm bandit problem where each arm a∈[K]a \in [K]a∈[K] is associated to a nonparametric reward function fa:[0,1]→Rf_a: [0,1] \to \mathbb{R}fa​:[0,1]→R mapping from contexts to the expected reward. Suppose that there is a large set of arms, yet there is a simple but unknown structure amongst the arm reward functions, e.g. finite types or smooth with respect to an unknown metric space. We present a novel algorithm which learns data-driven similarities amongst the arms, in order to implement adaptive partitioning of the context-arm space for more efficient learning. We provide regret bounds along with simulations that highlight the algorithm's dependence on the local geometry of the reward functions.

View on arXiv
Comments on this paper