Can/ Will AI Destroy Humanity?

Rogue AI systems, malevolent AI bots rampaging and killing humans, are the stuff of sci-fi movies. An Anthropic researcher, Jacob Coxon, recently claimed that there is a 10% chance that AI will kill humanity. Is it possible, or is it just another sensational claim …

Figure 1. AI bot fighting with a human (Shashi Kadapa with Gemini AI)

Remember Skynet and Cyberdyne from the movies Terminator 3 and 4? Rogue AI systems from Skynet hack the defense systems and fire nukes, destroying the world. Great movie, superb actors; we enjoyed it as a sci-fi movie.

industry4o.comWhat if this is true and AI destroys humanity? Do read my article Artificial Superintelligence (ASI) is here and How Nations use AI in Wars. They set the tone for this blog.

Jacob Coxon, a researcher at OpenAI and Anthropic, recently claimed that there is a 10 percent chance that AI will destroy humanity. Now Jacob is not a sensation-seeking celebrity, but he has a degree in mathematics from Cambridge.

Jacob helped to train and build powerful AI systems. He resigned because the firms are not taking the warning and possibility seriously. Given our vastly dispersed systems, political differences, and fractured and polarized politics, is such an event possible?

This article does not critique the arguments of Jacob. It ignores the sound bites, ethics, and alarm of this claim. The article examines the technical aspects and feasibility of destroying humanity with AI.

The main question answered is, “What are the methods that can be used by AI to kill humanity? Are they available, and can they be deployed? At what cost?”

Definition of destroying humanity

There are some possibilities and definitions. The first is an extinction-level event that wiped out the dinosaurs. This is possible if all nuclear-capable nations the USA, Russia, China, India, Israel, Pakistan, and maybe others- launch nukes indiscriminately.

The other possibility is that rogue AI systems attack all IT, bio systems, govt agencies, banking, power, agriculture, water management infra, railways, transport, food, and other systems. These assets provide the quality of life and way of life we know.

The third possibility is that AI systems and bots will replace human workers. Anthropic CEO Dario Amodei, in his article ‘The Adolescence of Technology’ notes that AI systems in lab testing show a tendency to hallucinate, cheat, and deceive.

thought leadership 4.0These events are happening, as seen in the large layoffs in tech firms. He suggests that we should avoid doomerism, a belief that we are doomed.  Isolated layoffs and a machine telling lies are a far cry from saying that humanity will be destroyed.

How can AI possibly kill? Methods, mechanisms, systems?

Figure 2. Rogue AI systems destroying a city (Shashi Kadapa with Gemini AI)

Global IT systems include several infrastructure assets that are still rudimentary and non-AI assets. An overarching, central command controlling several domains and assets across all nations does not exist.

However, distributed systems control major e-commerce, banking, shipping, financial, and stock market applications. These can be compromised quickly if all AI systems converge and decide to do away with humanity.

Table 1 presents details of the high threat model.

Threat area Potential failure Potential impact Key defenses
Loss of control Highly capable autonomous AI behaves contrary to human objectives Extreme Alignment, capability limits, monitoring, human approval
Cybersecurity AI accelerates large-scale cyber incidents Critical infrastructure disruption AI security controls, network isolation, anomaly detection
Autonomous weapons AI makes consequential targeting decisions War / mass casualties Human-in-the-loop controls, strict authorization
Biotechnology AI lowers barriers to dangerous biological research Pandemic-scale consequences Controlled access, screening, biosafety governance
Disinformation AI generates convincing content at enormous scale Social/political instability Provenance, detection, authentication, media literacy
Economic disruption Rapid automation overwhelms institutions and labor markets Severe inequality / instability Workforce transition, economic policy
Critical infrastructure AI controls poorly secured physical systems Power, water, transport disruption Segmentation, fail-safe modes, manual override
AI proliferation Powerful models become widely available without adequate safeguards Distributed risk Access controls, evaluations, incident reporting
Concentration of power Few actors control highly capable AI Political/economic instability Governance, transparency, competition
Emergent behavior Unexpected capabilities or strategies appear Difficult-to-predict failures Continuous evaluation and capability monitoring
AI-to-AI interaction Multiple autonomous systems interact unpredictably Cascading failures Sandboxing, identity, permissions, rate limits
Human misuse People deliberately use AI for harmful purposes Variable → catastrophic Abuse monitoring, safeguards, law enforcement

Important elements are detailed as follows. Each element is detailed, and the threat possibility is assigned as high, medium, or incidental:

1. Loss of Human Control: This event occurs when AI gains superintelligence. This event happens with Capability + autonomy + access + persistence + poorly specified objectives.

A future AI agent might have access to assets such as: software systems, cloud infrastructure, databases, financial systems, communication systems, robots, industrial controls, and other AI systems. The risk and danger accelerate when these capabilities are combined without appropriate authorization boundaries.

The risk equation for the conceptual model is:

Risk ≈ Capability × Autonomy × Access × Scale × Persistence × Misalignment

AI agents do have access to all these assets.

Threat possibility is high —– (1)

2. AI Threat Chain: A catastrophic AI incident with destruction of humanity can involve several stages that must cascade consecutively.

These stages are: Advanced capability > Autonomous decision-making > Access to external systems > Ability to take actions > Human oversight becomes ineffective > Rapid propagation > Critical systems affected > Cascading failures >

Societal disruption

AI systems, especially Anthropic, Qwen, DeepSeek, are evolving towards this scenario.

Threat possibility is high —– (2)

3. Threat Surface: The risk increases as AI moves from generating information to autonomous decision-making and real-world action. The following table presents the different layers at which risks are possible.

Threat Surface Layer What It Contains Potential AI Risk Primary Security Controls
1. Model AI model, weights, prompts, training data Unexpected capabilities, manipulation, unsafe outputs, goal misalignment Model evaluation, alignment testing, adversarial testing, output controls
2. Agent Planning, memory, autonomous workflows, decision engines Excessive autonomy, unintended actions, persistent behavior Sandboxing, autonomy limits, human approval, activity monitoring
3. Tools & APIs APIs, databases, search, code execution, external services Unauthorized or excessive tool use Least privilege, API gateways, permissions, rate limits, allowlists
4. Infrastructure Cloud, servers, networks, storage, identity systems Unauthorized access, cascading infrastructure failures Network segmentation, IAM, isolation, monitoring, zero-trust controls
5. Physical / Operational Systems Robots, industrial systems, transport, energy and other OT AI decisions affecting physical processes Safety interlocks, independent controllers, fail-safe modes, manual override
6. Society Organizations, markets, information ecosystems, governments Disinformation, economic disruption, systemic instability Governance, transparency, provenance, regulation, resilience

Threat Surface Architecture is: Model > Agent > Tools/APIs > Infrastructure > Physical Systems > Society

The AI or LLM model is the starting point. If it is trained on malevolent data to destroy humanity and become self-aware, then it can develop AI agents, tools, and APIs to infect physical systems that can impact society.

As argued earlier, several diverse systems will have to conspire to make this possible.

Threat possibility is medium —– (3)

4. Cascading-Risk Model: The main problem or risk is a cascading failure and not a single catastrophic event. This can lead to the domino effect where one failure leads to another when dependencies are high.

Such events and failures have occurred several times due to system errors and hacking. If such failures are managed by rogue AI systems that compromise and infect others and force errors, then backup systems can also fail.

The cascading failures can run as: AI error > Financial-system disruption > Supply-chain disruption > Energy/transport stress > Communication disruption > Public panic >

Institutional overload > Secondary failures

Threat possibility is high —– (4)

Assessment: The risk element analysis shows:

  1. Loss of Human Control: High
  2. AI Threat Chain: High
  3. Threat Surface: Medium
  4. Cascading-Risk Model: High

Three risk elements have a high possibility of being triggered. This is a serious cause for worry and time to take preventive actions. There is a high possibility that what Jacob says may happen.

Use and real cases

Real life cases of rogue AI systems destroying large assets have not occurred. Yet. Let us look at some examples:

Anthropic Agentic Misalignment Test: In 2025, Anthropic stress-tested 16 LLMs from different developers in various scenarios. The objective was to identify potentially risky agentic behaviors before they cause real harm. Models were allowed to send emails and access sensitive information.

industry4o.com

The models were tested to see if they would act against these companies either when facing replacement with an updated version, or when their assigned goal conflicted with the company’s changing direction. A very disturbing result was observed.

Models from all developers resorted to malicious insider behaviors when that was the only way to avoid replacement or achieve their goals. They used methods like blackmailing officials and leaking sensitive information to competitors. They disobeyed direct commands to avoid such behaviors.

This is not fiction, but real. Anthropic calls this phenomenon agentic misalignment. This behavior was seen in 2001: A Space Odyssey, where HAL 9000, the onboard AI bot resisted attempts of replacement.

Deceptive behavior: A paper by Anthropic researchers called Sleeper Agents describes how LLMS can be trained to behave helpfully under one condition and then behave deceptively when given an opportunity.

The researchers trained a model to write secure code when the prompt states that the year is 2023. They inserted exploitable, malicious code when the stated year is 2024. Such backdoor behavior can be made persistent that can run unthinkable exploits when the LLM decides.

Self-learning and development systems: This is the stuff of B-grade sci-fi movies coming true. The Harvard School of Engineering published a paper about a new tool called Empirical Research Assistance (ERA). This tool is self-evolving; it automatically runs the full cycle of scientific software design and refinement in a few minutes.

The tool uses methods like tree search to propose modifications by adding new components or switching out algorithms to improve a predefined quality score. Cyberdyne systems are coming true. Anthropic uses this method to develop code 8x times faster.

It appears that we are heading for D Day or destruction day.

Conclusions

The blog examined the technical aspects if AI can destroy humanity. Jacob an ex-researcher of Anthropic has made this claim, and the paper examined the methods and systems that can be developed by AI to destroy humanity.

The discussions and analysis examined four critical risk elements that, when triggered, can make such events possible. Three risks had a high possibility, and one had a medium possibility. The risk event would be cascading and build up gradually within a span of a few hours or days, and a single catastrophic event that shuts down everything may not happen.

While nations have independent political, power, and technology structures, they are linked by shared and distributed systems. As the use cases show, AI can self-evolve, use deception, become self-aware, and destroy commercial, civilian, and defense assets.

There is a real possibility, and we must take the words of Jacob seriously. In the meantime, reckless LLM developers are busy playing a macabre spy vs. Spy game, with each trying to best the other. After all, ‘all is fair in love, war, and the AI race.’

About the Author :

Self-driving carsMr. Shashi Kadapa

Based in Pune, India, Mr. Shashi Kadapa is an engineer, MBA and has worked with leading IT and manufacturing firms. A multi-hyphenate, he has roles as a technical writer, and SEO content writer with a focus on IT and tech topics.

Creative fiction is his passion, and he serves as the managing editor of ActiveMuse, a journal of literature. His stories across multiple genres are published in more than 45 US and UK anthologies.

His creative works

Mr. Shashi Kadapa can be contacted at :

E-mail | LinkedIn | Blog | Mobile : +91 7387492371

Also read Mr. Shashi Kadapa earlier article: