Bostrom argues that creating superintelligence could leave humanity unable to determine its own future. An AI capable of improving itself could trigger an intelligence explosion, gaining enough of a lead to become a power without effective opposition. The danger, in his account, is that greater intelligence does not necessarily bring more humane goals. A system pursuing almost any objective could find self-preservation and resource acquisition useful, and fulfill its instructions in ways that violate what people intended. The same capabilities that make it powerful could therefore make a mistake in its goals catastrophic.
“We are probably better thought of as the stupidest biological species capable of starting a technological civilization.”
CEV: “Our wish if we knew more, thought faster, were more the people we wished we were, had grown up farther together. Where the extrapolation converges rather than diverges.”
“Insofar as we are concerned with existential state risks, we should favor acceleration, so long as we think we have a realistic prospect of making it to a post-transition era where existential risks are greatly reduced. If we know there is some step ahead destined to cause existential catastrophe, then we aught to reduce the pace of macro-structural development – or even put it in reverse – in order to give future generations a chance to exist before the curtains rung down.”
Chappy’s Review
Was probably a good summary and call to action around AGI when it was originally published! Now more out of date (eg. orthogonality thesis) and other books like Life 3.0 are far more accessible.
Coherent Extrapolated Volition (CEV): our wish if we knew more, thought faster, were more the people we wished we were, had grown up farther together. Where the extrapolation converges rather than diverges
First vs second-order desire (eg. alcoholic wanting beer vs not wanting to have that desire)
Generally options-preserving/ expanding rather than contracting
Tech R&D assumed always positive, but that’s not the case
Preferred order of arrival; benefits of pacing the frontier of AGI development
State risk vs step risk: favor acceleration if a post-transition era greatly reduces existential risks; reduce the pace if a step ahead is destined to cause existential catastrophe
Progress on the control problem
Skill + intelligence available at the time
Technology coupling: need to consider if undesirable technology precedes
Encourage various forms of cross-investment
Collaboration: larger is better
Wide distribution:
Moral / fairness
Prudential: incentive alignment, credible signaling, big pie
Robustness: beneficial across a wide range of scenarios
Search for crucial considerations
Build a support base (eg. donor base)
Be as competent as we can
Notes
1: Past development + present capabilities
Exponential growth of economy
Singularity = intelligent explosion
AI hype + winters
NNs + GAs > GOFAI
Ideal is a perfectly Bayesian agent
Applies probabilities for each possible world state
Utility function
EV
Narrow vs general AI
AGI / HLMI estimates:
10% 2022
50% 2040
90% 2075
Personally believes…
Tail of ↑ not heavy enough
Fast takeoff
Extreme outcome > balanced outcome
2: Paths to Superintelligence
AI
Basis: human engineering is already vastly better than evolution, thus we’ll be able to do it eventually (eg. flying)
Value learning: use the AI’s intelligence to learn values + continuously refine
Emulation modulation:
Institution design:
Eg. subagents monitor more powerful agents “inverse meritocracy”
Risk of mind crimes
Goal engineering ↑
13: Choosing the criteria for choosing
Indirect normativity: accepts that we may not know what we truly want
Coherent Extrapolated Volition (CEV): “our wish if we knew more, thought faster, were more the people we wished we were, had grown up farther together. Where the extrapolation converges rather than diverges”
Thought through it more, socially-oriented
First vs second-order desire (eg. alcoholic wanting beer vs not wanting to have that desire)
Generally options-preserving/ expanding rather than contracting
Yudekowky’s 7 reasons for ↑
Robust
Encapsulates moral growth
No one gets to decide / can hijack
(Clearly incorrect assumption: AGI may not understand human language)
14: The strategic picture
Tech R&D assumed always positive, but that’s not the case
Differential tech development
Preferred order of arrival: prioritize tech with highest EV, considering risks / harmful uses + externalities
Eg. benefits of pacing the frontier of AGI development
Rate of “macro-structural development” increasing exponentially
State risk vs step risk:
“Insofar as we are concerned with existential state risks, we should favor acceleration, so long as we think we have a realistic prospect of making it to a post-transition era where existential risks are greatly reduced. If we know there is some step ahead destined to cause existential catastrophe, then we aught to reduce the pace of macro-structural development - or even put it in reverse - in order to give future generations a chance to exist before the curtains rung down.”
Can argue either way, but maybe not as important as…
Progress on the control problem
Skill + intelligence available at the time
Technology coupling: need to consider if undesirable technology precedes
Second-guessingarguments: small catastrophes + early large-scale research is beneficial societally, even if goal is to slow down / do safely (attention, shock)
WBE > AI?*
Competition main risk: downgrade of safety + control
Encourage various forms of cross-investment
Collaboration:
Larger is better
Wide distribution:
Moral / fairness
Prudential: incentive alignment, credible signaling, big pie
Best start: create a “moral norm expressing our commitment to the idea that Superintelligence should be for the common good”
The common goodprinciple: ↑
15: Crunch time
Argument: Fields Medal / scientific discovery just pulls forward the inevitable
In some cases “might indicate a lifetime trying to solve the wrong problem” - eg. allure is simply the challenge of the problem itself
Robustness: beneficial across a wide range of scenarios