44
2

Some Simulation Results for Emphatic Temporal-Difference Learning Algorithms

Huizhen Yu
Abstract

This is a companion note to our recent study of the weak convergence properties of constrained emphatic temporal-difference learning (ETD) algorithms from a theoretic perspective. It supplements the latter analysis with simulation results and illustrates the behavior of some of the ETD algorithms using three example problems.

View on arXiv
Comments on this paper