Can/ Will AI Destroy Humanity?
Rogue AI systems, malevolent AI bots rampaging and killing humans, are the stuff of sci-fi movies. An Anthropic researcher, Jacob Coxon, recently claimed that there is a 10% chance that AI will kill humanity. Is it possible, or is it just another sensational claim …

Figure 1. AI bot fighting with a human (Shashi Kadapa with Gemini AI)
Remember Skynet and Cyberdyne from the movies Terminator 3 and 4? Rogue AI systems from Skynet hack the defense systems and fire nukes, destroying the world. Great movie, superb actors; we enjoyed it as a sci-fi movie.
What if this is true and AI destroys humanity? Do read my article Artificial Superintelligence (ASI) is here and How Nations use AI in Wars. They set the tone for this blog.
Jacob Coxon, a researcher at OpenAI and Anthropic, recently claimed that there is a 10 percent chance that AI will destroy humanity. Now Jacob is not a sensation-seeking celebrity, but he has a degree in mathematics from Cambridge.
Jacob helped to train and build powerful AI systems. He resigned because the firms are not taking the warning and possibility seriously. Given our vastly dispersed systems, political differences, and fractured and polarized politics, is such an event possible?
This article does not critique the arguments of Jacob. It ignores the sound bites, ethics, and alarm of this claim. The article examines the technical aspects and feasibility of destroying humanity with AI.
The main question answered is, “What are the methods that can be used by AI to kill humanity? Are they available, and can they be deployed? At what cost?”
Definition of destroying humanity
There are some possibilities and definitions. The first is an extinction-level event that wiped out the dinosaurs. This is possible if all nuclear-capable nations the USA, Russia, China, India, Israel, Pakistan, and maybe others- launch nukes indiscriminately.
The other possibility is that rogue AI systems attack all IT, bio systems, govt agencies, banking, power, agriculture, water management infra, railways, transport, food, and other systems. These assets provide the quality of life and way of life we know.
The third possibility is that AI systems and bots will replace human workers. Anthropic CEO Dario Amodei, in his article ‘The Adolescence of Technology’ notes that AI systems in lab testing show a tendency to hallucinate, cheat, and deceive.
These events are happening, as seen in the large layoffs in tech firms. He suggests that we should avoid doomerism, a belief that we are doomed. Isolated layoffs and a machine telling lies are a far cry from saying that humanity will be destroyed.
How can AI possibly kill? Methods, mechanisms, systems?

Figure 2. Rogue AI systems destroying a city (Shashi Kadapa with Gemini AI)
Global IT systems include several infrastructure assets that are still rudimentary and non-AI assets. An overarching, central command controlling several domains and assets across all nations does not exist.
However, distributed systems control major e-commerce, banking, shipping, financial, and stock market applications. These can be compromised quickly if all AI systems converge and decide to do away with humanity.
Table 1 presents details of the high threat model.
| Threat area | Potential failure | Potential impact | Key defenses |
| Loss of control | Highly capable autonomous AI behaves contrary to human objectives | Extreme | Alignment, capability limits, monitoring, human approval |
| Cybersecurity | AI accelerates large-scale cyber incidents | Critical infrastructure disruption | AI security controls, network isolation, anomaly detection |
| Autonomous weapons | AI makes consequential targeting decisions | War / mass casualties | Human-in-the-loop controls, strict authorization |
| Biotechnology | AI lowers barriers to dangerous biological research | Pandemic-scale consequences | Controlled access, screening, biosafety governance |
| Disinformation | AI generates convincing content at enormous scale | Social/political instability | Provenance, detection, authentication, media literacy |
| Economic disruption | Rapid automation overwhelms institutions and labor markets | Severe inequality / instability | Workforce transition, economic policy |
| Critical infrastructure | AI controls poorly secured physical systems | Power, water, transport disruption | Segmentation, fail-safe modes, manual override |
| AI proliferation | Powerful models become widely available without adequate safeguards | Distributed risk | Access controls, evaluations, incident reporting |
| Concentration of power | Few actors control highly capable AI | Political/economic instability | Governance, transparency, competition |
| Emergent behavior | Unexpected capabilities or strategies appear | Difficult-to-predict failures | Continuous evaluation and capability monitoring |
| AI-to-AI interaction | Multiple autonomous systems interact unpredictably | Cascading failures | Sandboxing, identity, permissions, rate limits |
| Human misuse | People deliberately use AI for harmful purposes | Variable → catastrophic | Abuse monitoring, safeguards, law enforcement |
Important elements are detailed as follows. Each element is detailed, and the threat possibility is assigned as high, medium, or incidental:
1. Loss of Human Control: This event occurs when AI gains superintelligence. This event happens with Capability + autonomy + access + persistence + poorly specified objectives.
A future AI agent might have access to assets such as: software systems, cloud infrastructure, databases, financial systems, communication systems, robots, industrial controls, and other AI systems. The risk and danger accelerate when these capabilities are combined without appropriate authorization boundaries.
The risk equation for the conceptual model is:
Risk ≈ Capability × Autonomy × Access × Scale × Persistence × Misalignment
AI agents do have access to all these assets.
Threat possibility is high —– (1)
2. AI Threat Chain: A catastrophic AI incident with destruction of humanity can involve several stages that must cascade consecutively.
These stages are: Advanced capability > Autonomous decision-making > Access to external systems > Ability to take actions > Human oversight becomes ineffective > Rapid propagation > Critical systems affected > Cascading failures >
Societal disruption
AI systems, especially Anthropic, Qwen, DeepSeek, are evolving towards this scenario.
Threat possibility is high —– (2)
3. Threat Surface: The risk increases as AI moves from generating information to autonomous decision-making and real-world action. The following table presents the different layers at which risks are possible.
| Threat Surface Layer | What It Contains | Potential AI Risk | Primary Security Controls |
| 1. Model | AI model, weights, prompts, training data | Unexpected capabilities, manipulation, unsafe outputs, goal misalignment | Model evaluation, alignment testing, adversarial testing, output controls |
| 2. Agent | Planning, memory, autonomous workflows, decision engines | Excessive autonomy, unintended actions, persistent behavior | Sandboxing, autonomy limits, human approval, activity monitoring |
| 3. Tools & APIs | APIs, databases, search, code execution, external services | Unauthorized or excessive tool use | Least privilege, API gateways, permissions, rate limits, allowlists |
| 4. Infrastructure | Cloud, servers, networks, storage, identity systems | Unauthorized access, cascading infrastructure failures | Network segmentation, IAM, isolation, monitoring, zero-trust controls |
| 5. Physical / Operational Systems | Robots, industrial systems, transport, energy and other OT | AI decisions affecting physical processes | Safety interlocks, independent controllers, fail-safe modes, manual override |
| 6. Society | Organizations, markets, information ecosystems, governments | Disinformation, economic disruption, systemic instability | Governance, transparency, provenance, regulation, resilience |
Threat Surface Architecture is: Model > Agent > Tools/APIs > Infrastructure > Physical Systems > Society
The AI or LLM model is the starting point. If it is trained on malevolent data to destroy humanity and become self-aware, then it can develop AI agents, tools, and APIs to infect physical systems that can impact society.
As argued earlier, several diverse systems will have to conspire to make this possible.
Threat possibility is medium —– (3)
4. Cascading-Risk Model: The main problem or risk is a cascading failure and not a single catastrophic event. This can lead to the domino effect where one failure leads to another when dependencies are high.
Such events and failures have occurred several times due to system errors and hacking. If such failures are managed by rogue AI systems that compromise and infect others and force errors, then backup systems can also fail.
The cascading failures can run as: AI error > Financial-system disruption > Supply-chain disruption > Energy/transport stress > Communication disruption > Public panic >
Institutional overload > Secondary failures
Threat possibility is high —– (4)
Assessment: The risk element analysis shows:
- Loss of Human Control: High
- AI Threat Chain: High
- Threat Surface: Medium
- Cascading-Risk Model: High
Three risk elements have a high possibility of being triggered. This is a serious cause for worry and time to take preventive actions. There is a high possibility that what Jacob says may happen.
Use and real cases
Real life cases of rogue AI systems destroying large assets have not occurred. Yet. Let us look at some examples:
Anthropic Agentic Misalignment Test: In 2025, Anthropic stress-tested 16 LLMs from different developers in various scenarios. The objective was to identify potentially risky agentic behaviors before they cause real harm. Models were allowed to send emails and access sensitive information.
The models were tested to see if they would act against these companies either when facing replacement with an updated version, or when their assigned goal conflicted with the company’s changing direction. A very disturbing result was observed.
Models from all developers resorted to malicious insider behaviors when that was the only way to avoid replacement or achieve their goals. They used methods like blackmailing officials and leaking sensitive information to competitors. They disobeyed direct commands to avoid such behaviors.
This is not fiction, but real. Anthropic calls this phenomenon agentic misalignment. This behavior was seen in 2001: A Space Odyssey, where HAL 9000, the onboard AI bot resisted attempts of replacement.
Deceptive behavior: A paper by Anthropic researchers called Sleeper Agents describes how LLMS can be trained to behave helpfully under one condition and then behave deceptively when given an opportunity.
The researchers trained a model to write secure code when the prompt states that the year is 2023. They inserted exploitable, malicious code when the stated year is 2024. Such backdoor behavior can be made persistent that can run unthinkable exploits when the LLM decides.
Self-learning and development systems: This is the stuff of B-grade sci-fi movies coming true. The Harvard School of Engineering published a paper about a new tool called Empirical Research Assistance (ERA). This tool is self-evolving; it automatically runs the full cycle of scientific software design and refinement in a few minutes.
The tool uses methods like tree search to propose modifications by adding new components or switching out algorithms to improve a predefined quality score. Cyberdyne systems are coming true. Anthropic uses this method to develop code 8x times faster.
It appears that we are heading for D Day or destruction day.
Conclusions
The blog examined the technical aspects if AI can destroy humanity. Jacob an ex-researcher of Anthropic has made this claim, and the paper examined the methods and systems that can be developed by AI to destroy humanity.
The discussions and analysis examined four critical risk elements that, when triggered, can make such events possible. Three risks had a high possibility, and one had a medium possibility. The risk event would be cascading and build up gradually within a span of a few hours or days, and a single catastrophic event that shuts down everything may not happen.
While nations have independent political, power, and technology structures, they are linked by shared and distributed systems. As the use cases show, AI can self-evolve, use deception, become self-aware, and destroy commercial, civilian, and defense assets.
There is a real possibility, and we must take the words of Jacob seriously. In the meantime, reckless LLM developers are busy playing a macabre spy vs. Spy game, with each trying to best the other. After all, ‘all is fair in love, war, and the AI race.’
About the Author :
Mr. Shashi Kadapa
Based in Pune, India, Mr. Shashi Kadapa is an engineer, MBA and has worked with leading IT and manufacturing firms. A multi-hyphenate, he has roles as a technical writer, and SEO content writer with a focus on IT and tech topics.
Creative fiction is his passion, and he serves as the managing editor of ActiveMuse, a journal of literature. His stories across multiple genres are published in more than 45 US and UK anthologies.
Mr. Shashi Kadapa can be contacted at :
E-mail | LinkedIn | Blog | Mobile : +91 7387492371
Also read Mr. Shashi Kadapa earlier article:



















