Loading book details
Loading book details
Human Compatible cover

Human Compatible

Stuart Russell•2019

  1. Chappy's Book Notes•332 books

Human Compatible

Stuart Russell•2019

Length
11h 38m•~349 pages
Read
Jun 13th - 18th '22
AIEmerging Technology
•

Summary

Stuart Russell describes the key aspects and principles of AI before laying out the breakthroughs that remain to be made en route to AGI. Problems we need to be cognizant of with regards to AGI include the inverted U curve, the great uncoupling, the gorilla problem, the King Midas problem, the loophole principle, among others. Finally, Russell discusses how to design an optimally compatible AI that learns individuals’ personalized preferences using human choices as evidence. This includes learning a human’s goals and objectives, the scale of utility of those objectives, their meta preferences, and the core desire to serve humans and consequentially be turned off when appropriate.

Key Takeaways

  • Optimal AI should learn individuals’ personalized preferences
    • Use human choices as evidence: decipher acted / stated vs true (subconscious) preferences
    • Objectives / goals to maximize: individual’s life goals, factor in altruism levels
    • Scale of utility given objective: utility monsters, can lead to jealousy
    • Meta preferences: When is it desired and not desired to change preferences
    • Should want to be turned off by human
  • The gorilla problem: can inferior intelligence maintain dominance over its successor? (loophole principle)
  • Utility / goal / cost + loss functions can’t be known perfectly in an uncertain, probabilistic world
  • AI eliminating work:
    • Inverted U curve: # jobs in sector follows U curve as technology advances
    • The great uncoupling: (1970s) tech has led to earning vs productivity gap (capital vs labor)
    • UBI could lead to human split over striving (work) vs enjoying (leisure)
  • AI breakthroughs to make: Language, common sense, hypothesize, abstract routines, computational prioritization

Notes

Part 1: Intelligence in humans and machines

1 If we succeed

  • Intelligence: Extent to which entity’s actions can be expected to achieve its objectives
  • Standard model: Training based on goal, loss function, minimize cost of being wrong
  • Machines are beneficial to the extent they can be expected to achieve our objectives

2: Intelligence in humans and machines

  • Purely reactionary → nerve nets (jellyfish) → brains
  • Brain resembles reinforcement learning
  • Bayesian rationality: Rational, emotional actors optimize utility function, not expected value function
  • Utility theory objects to pure individuality - importance of tribe / altruism (eg. ants)
  • Nash equilibrium
    • Communication between actors allows for cooperation, better outcomes
  • Universality: Turing computer can compute anything
  • Moore’s law until 2025 then need special hardware (TPUs) or quantum (QPUs)
  • Intractable / NP problems prove that not everything is computable (applies to humans too)
  • Aspects in AI design:
    • Observable: Whether environment is full or partial observable
    • Discrete vs continuous environment and outputs (chess v driving)
    • Whether environment contains other intelligent agents
    • Determinism: Whether outcomes are predictable or not (eg. luck involved)
    • If there are rules and whether or not they are known (unsupervised)
  • Two types of logic:
    • Propositional logic (PL): True / false; NP-complete
    • First-order logic (FOL): Objects, properties, relations; semi-decidable
  • Reflex agents: implement objective but don’t know goal, why acting a certain way

3: AI in the future

  • “Breakthroughs” are just past research that has reached commercial viability
  • Russell: predicts AGI by 2100, median 2050
  • Breakthroughs to make
    • Language and common sense
    • Hypothesize, learn abstract routines
    • Prioritize / manage computational activity: can’t do everything at once
  • Bootstrapping: starting with easy text and working way up to more complex books like humans do
  • Feature engineering: in forecasting, deciding which features to factor in
  • Limits of AGI:
    • Land, materials
    • Pride: only top 1% can be in the top 1% - if this is happiness, then still elusive

Part 2: problems with intelligent machines

4: Misuses of AI

  • Surveillance, persuasion, and control
    • Blackmail bot: surveil, can be trained using money extracted from target
    • Behavior-modifying bots: eg. FB feed
    • Default to truth: misinformation spread, false reality (deep fakes)
  • AWS: Lethal autonomous weapon systems
  • Eliminating work
    • Inverted U curve: # jobs in sector follows as technology advances
    • The great uncoupling: since 1970s, tech → earning vs productivity gap (capital vs labor)
    • UBI could lead to human split over striving (work) vs enjoying (leisure)
  • Other human roles
    • Focus on inter-personal services
    • Human-in-the-loop: coordinates efforts of AI, sanity check

5: Overly intelligent AI

  • The gorilla problem: can inferior intelligence maintain dominance over successor?
  • King Midas problem: goal misalignment with AI leads to unintended outcomes
  • Any goal has transcendent sub-goals:
    • Survival
    • Speed (compute)
    • Resources (money, material)

6: The not-so-great AI debate

  • Not focused on imminence as much as sizable risk, lots of thought required
  • Success in curtailing gene line editing research demonstrates stopping AGI development might be possible
  • Blockchain can be used by AI to prevent un-powering / deletion
  • Oracle AI: can only read data from web, can’t produce at all

Part 3: how to think about AI

7: AI, a different approach

  • AI should be completely altruistic
    • Self-preservation only for goal at hand, not selfish desire
  • Goals for AI should be based on humans’
    • Naively assume goals don’t change
    • Naively assume everyone’s goals can simultaneously be met via equality
  • AI should want to be turned off by human
  • Utility, goal, cost function, loss function can’t be known perfectly in an uncertain, probabilistic world
  • Use human choices as evidence of preferences
  • AI ethics needs to be a worldwide effort: one slip can end / curtail mass AI adoption
  • Vladimir Putin: “The one who becomes the leader in AI will be the ruler of the world”

8: Provably beneficial AI

  • Need provable theory to prove beneficial AI implementation
  • Theories are based on assumptions / abstractions
  • Side-channel attacks: Attacks incorrect assumption made about provable theory
    • Eg. Voltage sounds during encryption
  • Inverse reinforcement learning (IRL):
  • Determine reward for turning off vs staying on, weigh and allow for turning off by user
  • Loophole principle: if super-intelligent AI wants to do bring about some condition, generally impossible for humans to write prohibitions on its actions to prevent it from doing so / something equivalent
  • Wire-heading: tendency to short-circuit normal behavior in favor of short-term stimulation or reward system

9: Complication: Us

  • Need to personalize AI to each of our individual preferences in a learned way
  • AI could find loopholes and break laws in the process: need to be conscious of others and society at large
  • Utilitarianism:
  • Not all humans have the same objective / goal of utility scale to maximize (life goals)
  • Not all humans have the same scale of utility given objective (can lead to jealousy)
  • Utility monster: entities that demand more to meet utility level (eg. human vs mouse)
  • Altruism: concern for wellbeing of others
  • AI can factor in human’s altruism levels towards other people: positive = help them too, zero = indifferent, negative = problem
  • Humans are highly irrational: don’t act in their own best interests
  • Approach life in nested goals rather than considering all possible paths (lots of work)
  • AI needs to understand emotions and how they lead to irrational behavior
  • Actions are not always based on preferences: eg. supposition, prejudice, fear of the unknown, generalizations
  • Need to account for change in preferences
  • Meta preferences: When is it desired and not desired to change preferences

10: Problem Solved

  • Need worldwide AI governance board (UN)
  • Proof of module safety and controllability
  • No analog to how AGI and human will cohabitate, can hope for the best