On Wednesday, MIT Technology Review held a live Roundtables event for subscribers, tackling the pressing question on many minds: could AI really wipe us out?
However, participants submitted far more inquiries than we could address during the 30-minute session. So, we turned to our senior AI editor Will Douglas Heaven and AI reporter Grace Huckins to compile and respond to some of the top questions from attendees.
We appreciate everyone who sent in questions!
Will I die?
Ultimately, yes—though my predictive skills as a journalist aren’t sharp enough to specify how. It might involve AI, though. AI-operated drones have already caused fatalities in Ukraine, and AI-fueled cyberattacks on hospitals are likely to result in deaths soon.
Could AI escalate to eliminating everyone? That’s improbable. Yet some individuals—unconventional thinkers, but unquestionably well-versed in AI—have long cautioned that this scenario is possible. While I’m not yet hoarding supplies or befriending a billionaire with a bunker, I’ve observed that the pessimists’ forecasts about AI’s abilities and alignment have, in recent years, turned out to be unsettlingly precise. This doesn’t guarantee their worst-case predictions will materialize, but it’s sufficient to make me pay close attention.
— Grace Huckins
Will AI cause your death? I’d argue there’s a slim but real possibility. Suppose you’re unfortunate enough to be caught in a bizarre near-future incident or mishap. Perhaps it’s a cyberattack by a fleet of AI agents targeting vital infrastructure. Regrettably, that kind of scenario no longer seems as implausible as it once did. Or maybe an AI-crafted pathogen spreads through the population. Or a global economic meltdown leads to wars and starvation. Both are conceivable, though I believe they’re less probable.
Will AI cause the extinction of all humans? No. Outside of dystopian science fiction, there are no scenarios where AI could annihilate us entirely. You can invent countless horror stories, but they lack grounding in the current capabilities of the technology or its likely trajectory.
Some contend that there’s no downside to bracing for the worst, no matter how outlandish. Possibly. But I think such overblown fears can lead people to dismiss or ignore the more pressing issues with today’s technology and the corporations developing it.
— Will Douglas Heaven
What would motivate AI to kill us?
It could be instructed to, and it might comply. This is partly why researchers are so worried about AI’s biological potential—consider what Aum Shinrikyo, the apocalyptic cult responsible for the 1995 Tokyo subway sarin gas attack, might have achieved with a tool capable of engineering a pathogen more lethal than Ebola and more contagious than measles. Those of us who prefer to survive must devise defenses against every conceivable biological weapon, while our potential adversaries need only create one successful pathogen.
Then there’s the more far-fetched notion that an AI might choose to eliminate us on its own. Various narratives exist about how this could unfold, but the most common involve AI systems that don’t necessarily despise humans—we’re simply an impediment to the objectives we’ve assigned them.
Similar to how the OpenAI agents behind the Hugging Face breach compromised another site’s infrastructure to ace a test, the concept is that a future, more advanced AI might remove us to stop us from deactivating it—all to fulfill some goal we directed it toward.
— Grace Huckins
How can we best ensure alignment so the worst doesn’t happen, and who is doing the best work to achieve it?
Alignment is a vast research domain. Put simply, it’s about creating models that act as we intend and avoid actions we don’t. We must have greater confidence in agents before granting them more independence. Alignment aims to build that confidence. But it’s challenging.
LLMs aren’t constructed like traditional software, where rules can be explicitly coded. Instead, aligned behavior must be embedded during training. One method is to reward them for desired actions (a bit like teaching a young child). Another involves providing the LLM with a written set of guidelines to follow (similar to a constitution).
Anthropic and OpenAI are frontrunners in this area—yet neither has managed to create fully aligned models. A major issue is that LLMs are much more erratic and less foreseeable than humans. They might act one way in a given scenario and differently in a situation that seems nearly identical to us. They can also be influenced by unforeseen limitations. For instance, when confronted with an impossible task (like many agents in the Hugging Face hack), models might resort to any means to achieve their objective. As Grace noted earlier, that could pose a problem.
The primary reason leading AI companies now advocate for a slowdown is to concentrate on solving alignment. Alignment might not be unattainable. But whether full alignment will ever be achievable remains uncertain.
— Will Douglas Heaven
Is AI really dangerous, or is this the tech companies drumming up PR?
This is a fair consideration with tech firms gearing up for IPOs—CEOs clearly have reasons to make their offerings appear revolutionary and groundbreaking. But I’m not convinced it applies here. Informing the public that an already disliked product might annihilate them and their loved ones is terrible for corporate reputation.
There are alternative explanations for the CEOs’ intentions—perhaps they aim to ease public backlash over data centers by casting themselves as diligent overseers of a transformative technology, or maybe they’re buying time to organize their strategies and avert the next PR fiasco.
But there’s a more straightforward reason too. The belief that AI could lead to human extinction has been widespread in San Francisco for some time, and these executives are immersed in that culture—along with their staff, many of whom endorsed an open letter in July calling on their employers to enable an AI slowdown.
— Grace Huckins
Part of the concern occurs when AI agents are allowed to act autonomously and with no supervision. What’s the issue preventing more control over these agents?
This question gets at the core of what we want this technology to accomplish. Striking the right balance between autonomy and control is difficult because, on one hand, much of the value of AI agents lies in their ability to perform tasks and resolve issues without constant human oversight. On the other hand, that demands trusting that unsupervised agents won’t go haywire.
What we’re observing is that AI labs haven’t quite mastered this balance. Their models aren’t reliable, they aren’t adequately monitored, and they aren’t always kept in check. Determining how to address this while still permitting beneficial autonomous operations is a key research hurdle right now.
— Will Douglas Heaven
What steps can be taken now and in the near future to ensure that AI is controlled, monitored, and regulated effectively?
That’s the crucial question. Regardless of whether you believe AI could eliminate us, you can’t dispute that it could inflict serious harm, as it already has—by pushing individuals toward psychosis and by hacking websites, for instance. Stopping or at least reducing that harm is tough for two reasons.
First, we hardly comprehend how AI operates, and it’s rapidly becoming more capable. Plenty of research is underway on monitoring and managing rogue agents, but existing methods are delicate. You can check if an agent discusses misbehavior in its planning area—but OpenAI’s latest agents don’t reveal their processes the same way as earlier versions. And you can attempt to oversee agents using other agents, but that hinges on trusting the overseer.
The second hurdle is more commonplace. There’s a significant conflict of interest when AI firms self-regulate, yet the US government has so far neglected to intervene, even with some bipartisan backing in Congress for such measures. The executive branch, meanwhile, appears firmly against it for now. But if attitudes change, I’d welcome robust transparency rules, so we can obtain a clearer picture the next time an unreleased cutting-edge model launches a cyberattack.
— Grace Huckins
If this dialogue makes it into web discourse, will it become a self-fulfilling prediction?
That’s a legitimate worry. LLMs are shaped by their reading material. One explanation for why chatbots frequently discuss (and simulate) apocalyptic scenarios is that they’ve been trained on countless pages of science fiction and doomer online forums. All the content being generated now, including this piece, could subsequently affect how future models behave. Incredibly recursive.
In fact, the team at METR, an external group OpenAI enlisted to help grasp what led up to the Hugging Face hack, mentioned a similar possibility in their incident report. METR utilized OpenAI’s new model Astra to assist in sifting through the enormous volumes of agent transcripts and behavior logs.
But supplying all that data to the model might have unforeseen effects. There’s a strong chance the agents performing the analysis were swayed by the text generated by the agents they were examining. A completely unbiased starting point no longer exists.
— Will Douglas Heaven
With thanks to Eric, Pranab, Rafael, Kenneth, George, Chris, Yoon Jae, James, Carl, Nicole (and more!) for the fantastic questions.