JSM2023
Back to the program
Professional Development Course/CE

CE_26C: Statistical and Algorithmic Foundations of Reinforcement Learning (Added Fee)

Tue, Aug 8, 8:30 AM - 5:00 PM

About this session

Motivation and rationale: Reinforcement learning (RL), which is frequently modeled as sequential decision making in the face of uncertainty, has garnered growing interest in recent years due to its remarkable success in practice. Despite decades-long research efforts, however, the statistical underpinnings of RL remain far from mature, especially when it comes to sample-starved large-dimensional regimes that are of crucial operational value in practice. An explosion of research has been conducted over the past few years towards advancing the statistical frontiers of these topics, which leverage toolkits that lie at the heart of statistics, such as high-dimensional statistics, stochastic approximation, uncertainty quantification, exploration-exploitation trade-offs, and statistical learning theory. This course not only covers the fundamentals, but also highlights emerging ideas from RL such as the principles of optimism and pessimism and distinct paradigms of RL (e.g., model-based, value-based, and policy-based RL). Our goal is to equip statistics researchers with the core toolkits of reinforcement learning and inspire the pursuit of further theory, algorithms, and applications from the statistics community on this multi-disciplinary and fast-growing topic. Course Structure: This is intended to be a full-day course. We will start by presenting a unified mathematical framework of the RL problems --- from multi-arm bandits to general Markov decision processes --- and then introduce classical dynamic programming algorithms. Equipped with this background, we will introduce three prominent approaches: model-based RL, model-free RL, and policy optimization, and discuss their statistical efficiency and optimality in the presence of multiple data collection mechanisms (e.g., simulators or generative models, online exploration of the unknown environment, offline data or batch data). We will also discuss extensions to multi-agent competitive RL. Finally, we will conclude with further pointers and a list of future directions. Abstract: As a paradigm for sequential decision making in unknown environments, reinforcement learning (RL) has received a flurry of attention in recent years. However, the explosion of model complexity in emerging applications and the presence of nonconvexity exacerbate the challenge of achieving efficient RL in sample-starved situations, where data collection is expensive, time-consuming, or even high-stakes (e.g., in clinical trials, autonomous systems, and online advertising). How to understand and enhance the sample and computational efficiencies of RL algorithms is thus of great interest and in imminent need. In this short course, we aim to present a coherent framework that covers important statistical and algorithmic developments in modern RL, highlighting the connections between new ideas and classical statistics. Employing Markov Decision Processes (MDPs) as the central mathematical framework, we will introduce three distinctive yet widely adopted approaches: the model-based approach, the value-based approach, and the policy-based approach. We will cover multiple important scenarios including the simulator setting, online exploratory RL, offline RL and batch RL, and multi-agent RL. Our discussions gravitate around the core issues of sample complexity and computational efficiency, and focus on the design of minimax-optimal algorithms. Prerequisites: basic probability, and basic linear algebra.

Session participants

Yuxin Chen
Participant
Yuejie Chi (Carnegie Mellon University)
Participant
Yuting Wei (University of Pennsylvania)
Participant
↑