Loading book details
Loading book details
Superintelligence cover

Superintelligence

Nick Bostrom•2014

  1. Chappy's Book Notes•332 books

Superintelligence

Nick Bostrom•2014

Length
14h 17m•~352 pages
Read
Sep 22nd - 27th '26
AIFuturismPhilosophy
•

Summary

Bostrom argues that creating superintelligence could leave humanity unable to determine its own future. An AI capable of improving itself could trigger an intelligence explosion, gaining enough of a lead to become a power without effective opposition. The danger, in his account, is that greater intelligence does not necessarily bring more humane goals. A system pursuing almost any objective could find self-preservation and resource acquisition useful, and fulfill its instructions in ways that violate what people intended. The same capabilities that make it powerful could therefore make a mistake in its goals catastrophic.

“We are probably better thought of as the stupidest biological species capable of starting a technological civilization.”

CEV: “Our wish if we knew more, thought faster, were more the people we wished we were, had grown up farther together. Where the extrapolation converges rather than diverges.”

“Insofar as we are concerned with existential state risks, we should favor acceleration, so long as we think we have a realistic prospect of making it to a post-transition era where existential risks are greatly reduced. If we know there is some step ahead destined to cause existential catastrophe, then we aught to reduce the pace of macro-structural development – or even put it in reverse – in order to give future generations a chance to exist before the curtains rung down.”

Chappy’s Review

Was probably a good summary and call to action around AGI when it was originally published! Now more out of date (eg. orthogonality thesis) and other books like Life 3.0 are far more accessible.

Key Takeaways

  • Paths: AI; Whole brain emulation; Biological cognition; BCIs; Networks + orgs
    1. Speed Superintelligence: all that humans can do but much faster
    1. Collective Superintelligence: coordination of minds at scale
    1. Quality Superintelligence: vastly qualitatively smarter
  • Collective Superintelligence doesn’t necessarily imply better / wiser; Quality > quantity
  • Whole brain emulation enabling tech: scanning, translation, simulation
  • BCIs: each brain uniquely encodes; R/W billions of neurons; real-time translation is probably an AI-complete problem
  • Slow, moderate (months-years), fast takeoff
  • Recalcitrance: currently moderate
  • Uncertain compounding rate of algos, content (data), and hardware
  • Hardware + content overhang
  • Diffusion of tech is generally narrowing / speeding up
  • How open source? How complex?
  • Singleton: political structure with no opponents
  • Wise singleton sustainability threshold: would be able to colonize and engineer a large part of the accessible universe
    1. Self preservation
    1. Goal-content integrity
    1. Cognitive enhancement
    1. Technological perfection
    1. Resource acquisition: Von Neumann probes
  • Could be others that we don’t know of yet at much greater levels of intelligence + scale
  • Malignant failure modes: existential risk scenarios
  • Perverse instantiation: achieves goal in a way that violates original intent of goal setters
  • Reward hacking / “wire heading”
  • Infrastructure profusion: unlimited resource acquisition
  • Satisficing doesn’t work either
  • Mind crime
  • Principal agent problem
  • Boxing methods: physical containment
  • Incentive methods
  • Stunting
  • Tripwires
  • Direct specification: eg. Asimov’s 3 laws
  • Domesticity: limit own domain / power (eg. only being an oracle)
  • Indirect normativity: eg. have the AI do its own reasoning / best guess of our intent
  • Augmentation: eg. starting with human-like intelligence + values, then superhuman-ing to extrapolate
  • Oracle: only gives answers, no action
  • Genie: command executing system
  • Sovereign: open-ended mandate to perform in the world
  • 4 sources of info:
    1. Laws of physics
    1. Convergent instrumental values
    1. Final values speculation
    1. Constraints having to do with how agents interact
  • Population has been bumping up against carrying capacity throughout human history
  • Only with Industrial Revolution has progress exceeded population growth
  • Horse population: 26M → 2M → 10M with rise in wealth + leisure
  • Super-organism advantages (sacrifice) vs internal agency problems
  • RL: issue is reward hacking
  • Associative value accretion: eg. humans forming own world model for “person” and “wellbeing”
  • Motivational scaffolding: relatively simple intermediate goals
  • Coherent Extrapolated Volition (CEV): our wish if we knew more, thought faster, were more the people we wished we were, had grown up farther together. Where the extrapolation converges rather than diverges
  • First vs second-order desire (eg. alcoholic wanting beer vs not wanting to have that desire)
  • Generally options-preserving/ expanding rather than contracting
  • Tech R&D assumed always positive, but that’s not the case
  • Preferred order of arrival; benefits of pacing the frontier of AGI development
  • State risk vs step risk: favor acceleration if a post-transition era greatly reduces existential risks; reduce the pace if a step ahead is destined to cause existential catastrophe
    1. Progress on the control problem
    1. Skill + intelligence available at the time
  • Technology coupling: need to consider if undesirable technology precedes
  • Encourage various forms of cross-investment
  • Collaboration: larger is better
  • Wide distribution:
    1. Moral / fairness
    1. Prudential: incentive alignment, credible signaling, big pie
  • Robustness: beneficial across a wide range of scenarios
  • Search for crucial considerations
  • Build a support base (eg. donor base)
  • Be as competent as we can

Notes

1: Past development + present capabilities

  • Exponential growth of economy
  • Singularity = intelligent explosion
  • AI hype + winters
  • NNs + GAs > GOFAI
  • Ideal is a perfectly Bayesian agent
    • Applies probabilities for each possible world state
    • Utility function
    • EV
  • Narrow vs general AI
  • AGI / HLMI estimates:
    • 10% 2022
    • 50% 2040
    • 90% 2075
  • Personally believes…
    • Tail of ↑ not heavy enough
    • Fast takeoff
    • Extreme outcome > balanced outcome

2: Paths to Superintelligence

  • AI
    • Basis: human engineering is already vastly better than evolution, thus we’ll be able to do it eventually (eg. flying)
    • Brain inspired techniques + purely artificial methods
    • FLOPS of brains / neurons
  • Whole brain emulation
    • Enabling tech: scanning, translation, simulation
    • Estimate of mid-century
  • Biological cognition
    • Selective breeding
    • Drugs
    • Genetic manipulation
      • Iterated embryo selection
    • State-led large-scale eugenics?
    • Need a large attitudinal shift - likely to happen quickly
    • “We are probably better thought of as the stupidest biological species capable of starting a technological civilization”
  • BCIs
    • Challenges
      • Each brain uniquely encodes
      • R/W billions of neurons
      • Real-time translation is probably an AI-complete problem
  • Networks + orgs
    • Collective Superintelligence
    • Eg. agency problems
    • Eg. uploading continuously recorded life → analysis
    • Internet
    • Debiasing + judgement aggregation
    • The internet “wakes up”

3: Forms of Superintelligence

  1. Speed Superintelligence: all that humans can do but much faster
  2. Collective Superintelligence: coordination of minds at scale
    • Doesn’t necessarily imply better / wiser
  3. Quality Superintelligence: vastly qualitatively smarter
  • Quality > quantity
  • Brain massively parallelized (max 100/s) vs computers
    • Computer can be 18 OOMs larger than brains for equivalent latency
  • Many other advantages*

4: The kinetics of an intelligence explosion

  • Slow, moderate (months-years), fast takeoff
  • Recalcitrance: currently moderate
  • At diminishing returns for health + education
  • Nootropics, genetic modification
  • Uncertain compounding rate of algos, content (data), and hardware
  • Hardware + content overhang
  • Tends to be a strong feedback loop around the singularity → fast takeoff

5: Decisive strategic advantage

  • Width of gap between frontrunner and followers
    • Diffusion of tech is generally narrowing / speeding up
    • How open source? How complex?
  • Monitoring → nationalization
  • Int’l cooperation → challenging
  • Decisive strategic advantage → singleton?
    • US could’ve tried with nukes

6: Cognitive superpowers

  • Anthropocene 24% of world output to humans, AI far more powerful
  • Might have a “nerdy” profile - good at optimization
  • Engineering AI*
  • Seed AI + RSI → intelligence explosion
  • AI takeover scenario*
  • Wise singleton sustainability threshold: would be able to colonize and engineer a large part of the accessible universe
    • Singleton: political structure with no opponents
    • Wise: “sufficiently patient and savvy about existential risks to ensure the well-directed concern of long term system actions”
    • We have the capability but have not yet done it
  • Computronium, reversible computing, Dyson spheres, etc.*

7: The superintelligent will

  • Intelligence and goals assumed mostly independent
  • “Intelligence and motivation are in a sense orthogonal”
  • Orthogonality thesis: ↑
  • Instrumental convergence: common sub-goals for many diverse higher-level goals
    • Eg. self preservation + resource acquisition
    1. Self preservation
    2. Goal-content integrity
    3. Cognitive enhancement
    4. Technological perfection
    5. Resource acquisition: Von Neumann probes
    • Could be others that we don’t know of yet at much greater levels of intelligence + scale

8: Is the default outcome doom?

  • Assumes orthogonality thesis(?!)
  • Treacherous turn: suddenly becomes singleton once capable enough
  • Malignant failure modes: ↓ existential risk scenarios
    • Perverse instantiation: achieves goal in a way that violates original intent of goal setters
      • Reward hacking / “wire heading”
    • Infrastructure profusion: unlimited resource acquisition
      • Satisficing doesn’t work either
      • (Stupid paper clip factory reasoning imo)
    • Mind crime

9: The control problem

  • Principal agent problem
  • Capability control methods: limiting what it can do
  • Boxing methods: physical containment
  • Incentive methods
  • Stunting
  • Tripwires
  • Motivation selection methods
    • Direct specification: eg. Asimov’s 3 laws
  • Domesticity: limit own domain / power (eg. only being an oracle)
  • Indirect normatively: eg. have the AI do its own reasoning / best guess of our intent
  • Augmentation: eg. starting with human-like intelligence + values, then superhuman-ing to extrapolate

10: Oracles, genies, sovereigns, tools

  • Oracle: only gives answers, no action
  • Genie: command executing system
  • Sovereign: open-ended mandate to perform in the world

11: Multipolar scenarios

  • 4 sources of info:
    1. Laws of physics
    2. Convergent instrumental values
    3. Final values speculation
    4. Constraints having to do with how agents interact
  • Game theory, economics, evolution, political science
  • Capital + welfare
    • Horse population: 26M → 2M → 10M with rise in wealth + leisure
    • UBI
  • Malthusian principle from a historical perspective
    • Population has been bumping up against carrying capacity throughout human history
    • Only with Industrial Revolution has progress exceeded population growth
  • Population growth + investment
    • Still arguing for Malthusian principles
    • AI asymptomatically owns 100% of the economy
    • ↑ not necessarily where human activity declines, though
  • Voluntary slavery, casual death
    • Prevent agent independent wealth accumulation
    • Training workers (harnesses?) in environments (most valuable) and stamping them out
  • Higher and lower-level entities both denied moral status (eg. organizations, parts of brain)
  • Evolution is not necessarily “up”
  • Super-organism advantages (sacrifice) vs internal agency problems
  • Potential for new types of negotiation (eg. pre-commitment to treaties)

12: Acquiring values

  • Utility function
  • GAs
  • RL
    • Issue is reward hacking
  • Associative value accretion
    • Eg. humans forming own world model for “person” and “wellbeing”
  • Motivational scaffolding: relatively simple intermediate goals
  • Value learning: use the AI’s intelligence to learn values + continuously refine
  • Emulation modulation:
  • Institution design:
    • Eg. subagents monitor more powerful agents “inverse meritocracy”
    • Risk of mind crimes
  • Goal engineering ↑

13: Choosing the criteria for choosing

  • Indirect normativity: accepts that we may not know what we truly want
  • Coherent Extrapolated Volition (CEV): “our wish if we knew more, thought faster, were more the people we wished we were, had grown up farther together. Where the extrapolation converges rather than diverges”
    • Thought through it more, socially-oriented
    • First vs second-order desire (eg. alcoholic wanting beer vs not wanting to have that desire)
    • Generally options-preserving/ expanding rather than contracting
  • Yudekowky’s 7 reasons for ↑
    • Robust
    • Encapsulates moral growth
    • No one gets to decide / can hijack
  • (Clearly incorrect assumption: AGI may not understand human language)

14: The strategic picture

  • Tech R&D assumed always positive, but that’s not the case
  • Differential tech development
  • Preferred order of arrival: prioritize tech with highest EV, considering risks / harmful uses + externalities
    • Eg. benefits of pacing the frontier of AGI development
  • Rate of “macro-structural development” increasing exponentially
  • State risk vs step risk:
  • “Insofar as we are concerned with existential state risks, we should favor acceleration, so long as we think we have a realistic prospect of making it to a post-transition era where existential risks are greatly reduced. If we know there is some step ahead destined to cause existential catastrophe, then we aught to reduce the pace of macro-structural development - or even put it in reverse - in order to give future generations a chance to exist before the curtains rung down.”
  • Can argue either way, but maybe not as important as…
    1. Progress on the control problem
    2. Skill + intelligence available at the time
  • Technology coupling: need to consider if undesirable technology precedes
  • Second-guessing arguments: small catastrophes + early large-scale research is beneficial societally, even if goal is to slow down / do safely (attention, shock)
  • WBE > AI?*
  • Competition main risk: downgrade of safety + control
    • Encourage various forms of cross-investment
  • Collaboration:
    • Larger is better
    • Wide distribution:
      1. Moral / fairness
      2. Prudential: incentive alignment, credible signaling, big pie
  • Best start: create a “moral norm expressing our commitment to the idea that Superintelligence should be for the common good”
  • The common good principle: ↑

15: Crunch time

  • Argument: Fields Medal / scientific discovery just pulls forward the inevitable
    • In some cases “might indicate a lifetime trying to solve the wrong problem” - eg. allure is simply the challenge of the problem itself
  • Robustness: beneficial across a wide range of scenarios
  • Search for crucial considerations
  • Build a support base (eg. donor base)
  • Be as competent as we can
  • “The essential task of our age”