Brian Christian2020
An understood world is more precisely compressible
We are in danger of losing control of the world not to AI or to machines as such, but to models. To formal, often numerical specifications for what exists and what we want
In seeing a kind of mind at work as it digests and reacts to the world, we will learn something both about the world and also, perhaps, about minds.
The core challenge of AI isn't making systems smarter — it's ensuring they optimize for what we actually want. From reward hacking in reinforcement learning to biased training data in classification systems, AI consistently finds ways to satisfy its objective function while violating human intent. The problem runs deeper than technical fixes: we often can't articulate our own values clearly enough to encode them. Inverse reward design, cooperative inverse reinforcement learning, and debate-based approaches represent promising directions, but the alignment problem ultimately reflects our own difficulty in specifying what we mean by "good."
“We are in danger of losing control of the world not to AI or to machines as such, but to models. To formal, often numerical specifications for what exists and what we want”
“In seeing a kind of mind at work as it digests and reacts to the world, we will learn something both about the world and also, perhaps, about minds.”