Artificial Intelligence 1998

Reinforcement Learning

《强化学习导论》

An Introduction

Author:Richard Sutton · Andrew Barto

Published
1998
Category
Artificial Intelligence
Difficulty
Advanced
Reading time
~30 hours
Original language
en
Classic Index 90/ 100
Historical Influence
Intellectual Depth
Long-term Relevance
Cross-domain Influence

The Classic Index is not an objective scientific measure. It is this site's personal curation score.

My Reading

What is this book about?

This book distills “learning how to act” into a single problem: an agent improves its policy from a reward signal through repeated interaction with an environment. It moves from bandits to temporal-difference learning and policy gradients, conveying the field's core intuitions with a remarkably light mathematical load.

Why read it?

Reinforcement learning is the standard language for the relationship among goals, feedback, and exploration — and those three reach far beyond games. It is also the liveliest interface between AI and cognitive science, giving precise form to questions from alignment to habit formation.

Core Ideas

  • The learning signal is reward, and its delay makes credit assignment the central difficulty.
  • Value functions turn “the expected long-run return” into something estimable, converting foresight into a computable update.
  • The tension between exploration and exploitation is unavoidable; every policy pays short-term return for information.
  • Temporal-difference learning shows that near-optimal policies can be approached by bootstrapping, without a model of the environment.

What questions does this book try to answer?

  • How can an agent tell which action was good when the reward arrives only later?
  • How should one trade off exploring the unknown against exploiting what is known?

Who should read it?

For readers with some probability and calculus who want to understand learning from feedback. The examples privilege intuition, and the mathematical bar is lower than in most comparable texts.