Skip to content

Human Compatible

Artificial Intelligence and the Problem of Control
Pages
352
Published
2019
Language
English

Synopsis

What if the future of AI doesn't have to be a battle between humans and machines? In Human Compatible, AI researcher Stuart Russell argues that avoiding a dystopian future requires fundamentally rethinking how artificial intelligence is built. Rather than giving machines fixed objectives to pursue at all costs, Russell proposes designing AI that remains deliberately uncertain about what people actually want, learning human preferences through observation rather than having them hard-coded in. He works through the benefits and risks of increasingly capable AI, from helpful personal assistants to autonomous weapons, and lays out a technical and social roadmap for keeping machines beneficial and deferential as they grow more capable. Written for a general audience by one of the field's most established researchers, it offers a concrete path toward AI that serves humanity's interests rather than working against them.

About the author

S
Stuart J. Russell

Stuart J. Russell is a British computer scientist and Professor of Computer Science at the University of California, Berkeley, where he holds the Smith-Zadeh Chair in Engineering. He earned a first-class degree in physics from Oxford in 1982 and a PhD in computer science from Stanford in 1986, then joined the Berkeley faculty. With Peter Norvig, he wrote "Artificial Intelligence: A Modern Approach," the standard textbook of the field, used in more than 1,500 universities across 135 countries. In...

Frequently asked questions

  • Does this book propose a specific technical solution to the AI control problem?

    The author advocates for a shift toward provably beneficial AI systems that are designed with inherent uncertainty about human preferences. By maintaining this uncertainty, machines are incentivized to be humble and deferential, which prevents them from aggressively pursuing potentially harmful goals.

  • What is the gorilla problem mentioned in the text?

    This analogy illustrates the risk of humans losing autonomy to machines that possess superior intelligence. Just as gorillas have no future beyond what humans allow, the author warns that humanity could face a similar loss of control if we create superintelligent systems without proper safeguards.

  • Does the author argue that we should program specific human values into AI?

    The book explicitly rejects the idea of installing a fixed, universal set of human values into machines. Instead, it argues that AI should learn human preferences through observation and interaction, as there is no single, universally agreed-upon value system to encode.

  • Is "Human Compatible" a sequel or companion to "Artificial Intelligence: A Modern Approach"?

    No — they're different kinds of books. AIMA is the graduate-level textbook covering how AI systems are built; Human Compatible is a general-audience argument about why AI needs a new foundation. Reading AIMA first isn't necessary to follow it.

Read all 1 reviews