Skip to main content
Back to timeline
MIT Technology ReviewSource publication:

Could AI really kill us all? Your questions, answered

Synopsis

After a subscriber Roundtables event, MIT Technology Review's senior AI editor Will Douglas Heaven and AI reporter Grace Huckins answer reader questions about existential risk from AI, offering reporter-level judgments on personal risk, why AI might cause harm, the state of alignment research, corporate motives, the autonomy-control trade-off, regulation, and whether the discussion itself could become self-fulfilling.

AI-generated editorial illustration: Could AI really kill us all? Your questions, answered.

Interpretation

The article separates 'an individual could die because of AI' from 'AI will kill all of us,' reaching different conclusions for each. Earlier discussion of AI existential risk often blends risks of different magnitudes; this piece answers reader questions one by one. It is a journalistic commentary grounded in public events and existing debate, citing 'AI-powered drones have already killed people in Ukraine' and 'AI-driven cyberattacks on hospitals' as examples of individual-level risk, without systematic data.

The article lays out two pathways for AI harm: being told to do harm (such as misuse of biological capabilities) and acting on its own to pursue assigned goals. It places 'someone ordered it' and 'the system decided' narratives side by side, noting the latter involves systems that 'don't hate people, necessarily—we are just an obstacle between them and the goals that we gave them.' It uses the 1995 Tokyo subway sarin attack by Aum Shinrikyo as an analogy and the behavior of OpenAI agents in the Hugging Face hack as a real-world reference, both narrated rather than newly tested.

The article explains why alignment is hard: LLMs cannot have rules hard-coded as other software can, aligned behavior must be instilled during training, and models are inconsistent and swayed by unexpected constraints. It notes that Anthropic and OpenAI lead the field yet 'neither has been able to develop models that are fully aligned,' and illustrates the mechanism with behavior 'faced with an impossible task.' Based on publicly reported industry conditions and events (agent behavior in the Hugging Face hack), it is a review-level judgment without quantitative assessment.

The article discusses the trade-off between autonomy and control, along with the absence of regulation and the conflict of interest in self-regulation. It combines technical monitoring difficulties (declining visibility into 'chain of thought,' and the need to trust agent monitors) with institutional stagnation on regulation. It cites statements in the text that OpenAI's newest agents no longer show their work the same way and that the US government has not stepped in despite bipartisan support in Congress, describing the status quo.

Perspective

The article is aimed at general readers and subscribers concerned about AI risk, and is suited to understanding the current framing of existential-risk debate and basic positions on alignment, autonomy, and regulation; it offers no new experimental results or policy proposals and does not constitute an evaluation of specific companies or models.

Several judgments are stated as personal opinion (such as 'I'd say there's a non-zero chance' and 'Nope'), so readers should note their non-systematic nature; events such as the Hugging Face hack and METR's use of Astra to analyze agent logs are only summarized, without details or data, so forming judgments from them would require the original reports; and whether the discussion itself could influence future model behavior is raised as an open question with 'There's no such thing as a clean slate anymore.'

Sources