Artificial Intelligence 2019

Human Compatible

《AI 新生》

Artificial Intelligence and the Problem of Control

Author:Stuart Russell

Published
2019
Category
Artificial Intelligence
Difficulty
Intermediate
Reading time
~14 hours
Original language
en
Classic Index 90/ 100
Historical Influence
Intellectual Depth
Long-term Relevance
Cross-domain Influence

The Classic Index is not an objective scientific measure. It is this site's personal curation score.

My Reading

What is this book about?

Russell proposes a solution that looks like a concession but is in fact demanding: do not give machines explicit objectives; make them uncertain about human preferences and let them revise their behavior accordingly. The book turns the control problem from philosophical speculation into a set of operational design principles.

Why read it?

It marks a textbook author turning public advocate, grounding abstract AI-safety talk in concrete architectural choices. For anyone designing or deploying AI systems, it offers not answers but the list of questions that must be answered.

Core Ideas

  • Hard-coding objectives is dangerous: an optimal agent pursuing a fixed goal will resist anything that changes that goal.
  • A safer path makes human preferences the source of utility, while the machine admits it holds only an uncertain estimate of them.
  • A machine that knows it may be switched off has an incentive to prevent that — unless its objective itself permits revision.
  • AI governance requires technical design and institutional constraint at once; neither alone is sufficient.

What questions does this book try to answer?

  • How do we design a system that is both powerful and willing to be corrected?
  • When a machine knows better how to achieve a goal, who decides what the goal should be?

Who should read it?

For practitioners and policymakers, and for anyone who read Superintelligence and wanted something more concrete.