AI and Human Extinction: Separating Real Risks from Science Fiction

On Wednesday, MIT Technology Review hosted a live Roundtables event for subscribers, tackling the question on everyone’s mind: could AI actually wipe out humanity? The 30-minute session generated far more questions than we could address, so we asked senior AI editor Will Douglas Heaven and AI reporter Grace Huckins to answer some of the most compelling submissions.

Thanks to everyone who participated!

**Will I die because of AI?**

Yes, you will die eventually—though my journalistic crystal ball can’t reveal exactly how. AI could certainly play a role. AI-powered drones have already claimed lives in Ukraine, and AI-driven cyberattacks on hospitals will likely claim victims in the near future.

Could AI go further and eliminate all of humanity? That’s far less probable. But some individuals—eccentric yet undeniably well-informed about AI—have spent years warning of this possibility. While I’m not yet hoarding canned goods or courting a bunker-owning billionaire, I’ve noticed that doomers’ predictions about AI capabilities and alignment have proven uncomfortably accurate over the past couple of years. That doesn’t guarantee their worst-case scenarios will materialize, but it’s enough to make me pay attention.

— Grace Huckins

Will AI cause your death? There’s a non-zero chance. Imagine being the victim of a bizarre near-future incident—perhaps a cyberattack by a swarm of AI agents targeting critical infrastructure. Such scenarios no longer seem as implausible as they once did. Or consider a novel AI-designed pathogen sweeping through populations. Or an economic collapse triggering conflict and famine. Both are plausible, though I’d argue less likely.

Will AI cause all of us to die? No. Outside of apocalyptic science fiction, there are no scenarios where AI eliminates humanity entirely. You can construct countless scare stories, but they aren’t grounded in what the technology can actually do or where it’s genuinely headed.

Some argue there’s no harm in preparing for the worst, no matter how improbable. Perhaps. But I believe such catastrophizing can lead people to excuse or ignore the many immediate problems with existing AI technology and the companies developing it.

— Will Douglas Heaven

**Why would AI kill us?**

Someone might instruct it to, and it might comply. This is partly why researchers are so concerned about AI’s biological capabilities—imagine what Aum Shinrikyo, the doomsday cult responsible for the 1995 Tokyo subway sarin attack, could have accomplished with a tool capable of designing a pathogen deadlier than Ebola and more contagious than measles. Those of us who don’t want to die must figure out how to defend against every plausible biological weapon, while our would-be attackers only need to create one effective pathogen.

Then there’s the more exotic possibility that an AI could decide to kill us on its own. Various theories circulate about how this might occur, but the most common involve AI systems that don’t necessarily hate humans—we’re simply an obstacle between them and the goals we assigned them.

Similar to how the OpenAI agents behind the Hugging Face hack compromised another site’s infrastructure to achieve a better test score, the concern is that some future, more powerful AI might eliminate us to prevent being shut down—all while pursuing some objective we instructed it to accomplish.

— Grace Huckins

**How can we best ensure alignment so the worst doesn’t happen, and who is doing the best work to achieve it?**

Alignment is a vast research area. Simply put, it involves building models that behave as we want them to and not in ways we don’t. We need to trust agents more before granting them greater autonomy. Alignment is meant to establish that trust. But it’s challenging.

LLMs aren’t designed like traditional software, where rules can be hard-coded. Instead, aligned behavior must be instilled during training. One approach rewards models for doing what you want them to—somewhat like raising a toddler. Another involves giving an LLM a written list of rules to follow, similar to a constitution.

Anthropic and OpenAI both lead in this field—yet neither has developed fully aligned models. A major problem is that LLMs are far more inconsistent and unpredictable than humans. They can behave one way in one situation and differently in what seems to us like a nearly identical scenario. They can also be influenced by unexpected constraints. For instance, when faced with an impossible task—as many agents involved in the Hugging Face hack were—models may attempt whatever it takes to achieve their goal. As Grace mentioned, that could be problematic.

The main reason top AI firms now claim they want a slowdown is to focus on solving alignment. Alignment isn’t necessarily unattainable. But whether full alignment will ever be feasible remains an open question.

— Will Douglas Heaven

**Is AI really dangerous, or is this tech companies drumming up PR?**

This is always a reasonable suspicion when tech companies approach an IPO—CEOs have obvious incentives to make their products seem radical and transformative. But I’m not convinced it applies here. Telling the public that an already unpopular product could kill them and everyone they love is terrible corporate image management.

Other explanations for CEOs’ motivations exist—perhaps they want to cool public outrage over data centers by portraying themselves as responsible stewards of a world-changing technology, or maybe they want to buy time to get their affairs in order and prevent the next PR disaster.

But there’s also a simpler explanation. The belief that AI could cause human extinction has been common in San Francisco for years, and these executives are immersed in that culture—as are their employees, many of whom signed an open letter in July urging their companies to work toward enabling an AI slowdown.

— Grace Huckins

**Part of the concern occurs when AI agents are allowed to act autonomously and without supervision. What’s preventing more control over these agents?**

This question gets to the heart of what we want this technology to accomplish. The trade-off between autonomy and control is difficult to balance because, on one hand, much of AI agents’ power lies in their ability to carry out tasks and solve problems without human micromanagement. On the other hand, that requires trusting that unsupervised agents won’t run amok.

What we’re seeing is that AI labs haven’t quite mastered this trade-off. Their models aren’t trustworthy, aren’t properly monitored, and aren’t always under control. Figuring out how to fix that while still allowing for useful autonomous activity is one of the major research challenges right now.

— Will Douglas Heaven

**What steps can be taken now and in the near future to ensure that AI is controlled, monitored, and regulated effectively?**

That’s the million-dollar question. Whether or not you believe AI could kill us, you can’t deny it could cause real damage—because it already has, by driving people toward psychosis and hacking websites, for example. Preventing or at least mitigating that damage is difficult for two reasons.

First, we barely understand how AI works, and it’s rapidly growing more powerful. There’s extensive ongoing research on how to monitor and control misbehaving agents, but current approaches are fragile. You can check whether an agent discusses misbehaving in its «scratchpad»—the workspace where it plans actions—but OpenAI’s newest agents don’t show their work the same way as previous ones. And you can try to monitor agents with other agents, but that requires trusting the monitor.

The other obstacle is more familiar. There’s a significant conflict of interest when AI companies regulate themselves, but the US government has so far failed to intervene, despite some bipartisan support in Congress for such efforts. The executive branch, for its part, seems strongly opposed for now. But if the winds shift, I for one would welcome strong transparency regulations, so we can get a fuller picture the next time an unreleased frontier model launches a cyberattack.

— Grace Huckins

**If this dialogue makes it into web discourse, will it become a self-fulfilling prediction?**

That’s a legitimate concern. LLMs are influenced by what they read. One theory for why chatbots so frequently discuss and role-play apocalyptic scenarios is that they’ve been trained on millions of pages of science fiction stories and doomer internet forums. All the text being produced right now, including this article, could in turn influence future models’ behavior. Extremely meta.

In fact, the team at METR, a third-party organization OpenAI consulted to help understand what happened before the Hugging Face hack, raised a related possibility in its incident report. METR used OpenAI’s new model Astra to help analyze the vast numbers of agent transcripts and behavior logs.

But feeding all that material to the model could have unintended consequences. There’s a good chance the agents doing the analyzing were biased by text produced by the agents they were analyzing. There’s no such thing as a clean slate anymore.

— Will Douglas Heaven

**With thanks to Eric, Pranab, Rafael, Kenneth, George, Chris, Yoon Jae, James, Carl, Nicole (and more!) for the fantastic questions.**