ResearchTrend.AI
  • Papers
  • Communities
  • Events
  • Blog
  • Pricing
Papers
Communities
Social Events
Terms and Conditions
Pricing
Parameter LabParameter LabTwitterGitHubLinkedInBlueskyYoutube

© 2025 ResearchTrend.AI, All rights reserved.

  1. Home
  2. Papers
  3. 2006.04354
12
7

A Model-free Learning Algorithm for Infinite-horizon Average-reward MDPs with Near-optimal Regret

8 June 2020
Mehdi Jafarnia-Jahromi
Chen-Yu Wei
Rahul Jain
Haipeng Luo
ArXivPDFHTML
Abstract

Recently, model-free reinforcement learning has attracted research attention due to its simplicity, memory and computation efficiency, and the flexibility to combine with function approximation. In this paper, we propose Exploration Enhanced Q-learning (EE-QL), a model-free algorithm for infinite-horizon average-reward Markov Decision Processes (MDPs) that achieves regret bound of O(T)O(\sqrt{T})O(T​) for the general class of weakly communicating MDPs, where TTT is the number of interactions. EE-QL assumes that an online concentrating approximation of the optimal average reward is available. This is the first model-free learning algorithm that achieves O(T)O(\sqrt T)O(T​) regret without the ergodic assumption, and matches the lower bound in terms of TTT except for logarithmic factors. Experiments show that the proposed algorithm performs as well as the best known model-based algorithms.

View on arXiv
Comments on this paper