ResearchTrend.AI
  • Papers
  • Communities
  • Events
  • Blog
  • Pricing
Papers
Communities
Social Events
Terms and Conditions
Pricing
Parameter LabParameter LabTwitterGitHubLinkedInBlueskyYoutube

© 2025 ResearchTrend.AI, All rights reserved.

  1. Home
  2. Papers
  3. 2208.14615
36
5

Fine-Grained Distribution-Dependent Learning Curves

31 August 2022
Olivier Bousquet
Steve Hanneke
Shay Moran
Jonathan Shafer
Ilya O. Tolstikhin
ArXivPDFHTML
Abstract

Learning curves plot the expected error of a learning algorithm as a function of the number of labeled samples it receives from a target distribution. They are widely used as a measure of an algorithm's performance, but classic PAC learning theory cannot explain their behavior. As observed by Antos and Lugosi (1996 , 1998), the classic `No Free Lunch' lower bounds only trace the upper envelope above all learning curves of specific target distributions. For a concept class with VC dimension ddd the classic bound decays like d/nd/nd/n, yet it is possible that the learning curve for \emph{every} specific distribution decays exponentially. In this case, for each nnn there exists a different `hard' distribution requiring d/nd/nd/n samples. Antos and Lugosi asked which concept classes admit a `strong minimax lower bound' -- a lower bound of d′/nd'/nd′/n that holds for a fixed distribution for infinitely many nnn. We solve this problem in a principled manner, by introducing a combinatorial dimension called VCL that characterizes the best d′d'd′ for which d′/nd'/nd′/n is a strong minimax lower bound. Our characterization strengthens the lower bounds of Bousquet, Hanneke, Moran, van Handel, and Yehudayoff (2021), and it refines their theory of learning curves, by showing that for classes with finite VCL the learning rate can be decomposed into a linear component that depends only on the hypothesis class and an exponential component that depends also on the target distribution. As a corollary, we recover the lower bound of Antos and Lugosi (1996 , 1998) for half-spaces in Rd\mathbb{R}^dRd. Finally, to provide another viewpoint on our work and how it compares to traditional PAC learning bounds, we also present an alternative formulation of our results in a language that is closer to the PAC setting.

View on arXiv
Comments on this paper