|
It runs one and half hours. I'll summarize the main points:
Russell traces the development of AI from deep learning to large language models, the integration of artificial neural networks with probability mathematics and the emergence of large reasoning models that “in principle, have no formal limits.” Driven by what the techies call “verified rewards,” these models relentlessly seek achievement of an objective and will do whatever is necessary to get there.
This last step toward general artificial intelligence that would make them smarter than humans has, in Russell’s view, now crossed a threshold AI scientists have long feared
“. . . where the AI system is sufficiently capable that, whatever its objectives are, it’s going to achieve them, even if they’re not aligned with what we want. These are not sharp transitions, but we could talk about a ‘loss-of-control transition,’ where we no longer have a say in what happens".
The possible scenarios beyond this threshold range from disruptive cyberattacks on infrastructure to, at the far end, “extinction” in the sense that autonomously reasoning and agentic AGI no longer needs humans or to align with their morals, norms and interests.
It is also possible, says Russell, that
“. . . if we figure out how to build in safety into the design of AI systems from the beginning, maybe we could coexist indefinitely and even flourish with such systems. So, before that loss-of-control transition happens, there’s an earlier transition, which is much more difficult to perceive, which is when the time it takes to get to that loss-of-control level is less than the time it takes to solve the control problem. Most of the people I talk to say we have already passed that point".
He continues:
“Solving the control problem is very difficult. All the people in the company say, ‘Yeah, we don’t know how to solve it. And we’re not really even working on it’ because they need to work on getting the next improved version out so that they don’t get beaten by their competitors. This race condition amplifies the mismatch between devising effective constraints and losing control".
The obvious question is why the AI companies are risking even a small chance of their invention leading to human extinction by proceeding when they are fully aware they are losing control? Has any other species willfully put their existence at risk?
Russell responds:
"In my book, ‘Human Compatible,’ I talk about a species of sloth that seems to have become addicted to some Valium-like substance in its food supply, so that it can’t be bothered to breed anymore. So those kinds of extinction events, they’re driven by the same thing, in a sense.
It’s this mismatch between the long-term interest, which is presumably that the species continues, and the short-term reward signal that evolution has built into you to try to get you to do good things. But we humans also suffer from this when we become drug addicts, right? We have a dopamine system that’s supposed to help us avoid pain and seek pleasure and food and company and all those things that we like.
But sometimes it gets hijacked by drugs. And so we experience a personal extinction as a result. So that mismatch can happen at the species level as well".
Russell marvels that most of the warnings about possible extinction come from the CEOs of the top AI companies themselves who are building the technology:
“They’re literally saying, if we succeed in creating AGI - which we are going to spend a trillion dollars of your money to build - then there’s, depending on who you ask, 10, 20, 25, even 50% chance that we’re all going to go extinct. I think they’re really quite terrified, but they can’t get out of the race condition that they’re in.”
Why we can’t stop 🤷♂️
Russell sees the CEOs caught in a prisoner’s dilemma:
"So, if I said, 'OK, we’re not releasing our next system until we solve the control problem, then my company would be out of business. The investors would fire me, and no good would come of it'. Interestingly, Dario Amodei, CEO of Anthropic, and Demis Hassabis, CEO of Google DeepMind, have both said this year that they want to stop.
They think we have to stop, but they will only stop if everyone else agrees to stop. So that’s a remarkable statement. That has never happened, as far as I know, in the history of capitalism".
For Russell, these kinds of statements are “signaling to the government” that it needs to step in and facilitate agreement among the small band of CEOs pushing things forward, or impose control. One of those CEOs told Russell:
“They don’t think that’s going to happen until there’s a Chernobyl-scale disaster. And that's a best-case scenario. Because the other case is where government doesn’t come in, and then later on there’s a much bigger and perhaps irreversible catastrophe. So that’s the only way there’s going to be effective intervention".
All of these rogue attacks I noted at the top of this post have manifested what up to now had been only a "theoretical" worry. Let’s hope Russell is wrong that only a disaster can save us.
|