ResearchTrend.AI
  • Papers
  • Communities
  • Events
  • Blog
  • Pricing
Papers
Communities
Social Events
Terms and Conditions
Pricing
Parameter LabParameter LabTwitterGitHubLinkedInBlueskyYoutube

© 2025 ResearchTrend.AI, All rights reserved.

  1. Home
  2. Papers
  3. 2012.08225
  4. Cited By
Policy Optimization as Online Learning with Mediator Feedback

Policy Optimization as Online Learning with Mediator Feedback

15 December 2020
Alberto Maria Metelli
Matteo Papini
P. DÓro
Marcello Restelli
    OffRL
ArXivPDFHTML

Papers citing "Policy Optimization as Online Learning with Mediator Feedback"

1 / 1 papers shown
Title
Bounded regret in stochastic multi-armed bandits
Bounded regret in stochastic multi-armed bandits
Sébastien Bubeck
Vianney Perchet
Philippe Rigollet
63
90
0
06 Feb 2013
1