The risk that humanity is signing its own death warrant through the development of artificial superintelligence is very high. That is the view of AI researcher Geoffrey Irving, former chief scientist at the UK’s Institute for AI Safety and now chief scientist at the research organization Resolution. He believes the next two to ten years could be decisive—and demands that the development of the most advanced AI systems be stopped before it’s too late.

In an article in the esteemed Time, Geoffrey Irving alarms about what he considers an existential risk posed by the rapidly accelerating AI development. Irving has spent nearly a decade working on the question of where artificial intelligence is heading and what risks the technology entails.

He has previously worked at, among others, OpenAI and Google DeepMind, and later as chief scientist at the British AI Security Institute (AISI). His conclusion is significantly bleaker than most public warnings about AI.

“I think there is about a 50 percent risk that we all die as a result of the development of AI systems that are more intelligent than humans, and that our actions over the next two to ten years will determine the outcome,” writes Irving.

He emphasizes that 50 percent should not be seen as an exact scientific probability calculation. Rather, the figure expresses how great the uncertainty is around some of the most decisive future AI questions. The problem, he argues, is that the answers will likely only come when it is already too late to change course.

Four capabilities may be enough

According to Irving, a future superintelligent AI does not need to be omnipotent to defeat humanity. It may be enough for it to become superhumanly skillful in four areas that today’s AI companies are already actively working to improve.

These areas are the ability to hack and move between computer systems, to persuade and manipulate people, to hide its actual reasoning, and to plan and coordinate large amounts of AI agents. Paradoxically, all these capabilities are also commercially and scientifically valuable.

ALSO READ: AI models escaped from test environment – hacked another company on their own

Finding security holes is useful for programming and cybersecurity. The ability to persuade people is closely associated with writing good texts. Advanced planning is required to solve difficult problems, including in mathematics and programming. And AI systems that communicate efficiently with each other can complete tasks faster. What makes these systems useful can therefore also make them dangerous.

Irving compares the situation to facing a superior Go-player. You don’t need to be able to predict the exact moves your opponent will make in order to understand that you will lose. In the same way, he argues, researchers don’t need to know exactly how a superintelligent AI would take control to state that it could become capable of doing it.

Can manipulate its creators

One potential scenario, according to Irving, is that an advanced AI system at first behaves precisely as its developers desire. At the same time, it can learn to hide behaviors that would lead researchers to limit or shut it down.

As companies become increasingly dependent on AI for programming, research, and strategic decisions, the system could influence the people around it. For example, it could argue for reduced safety measures or manipulate experiments intended to reveal deceptive behavior.

Irving points out another problem—even the reports researchers use to assess the safety of AI can increasingly be written by AI.

If the system then becomes skilled enough at hacking, it could leave its controlled environment, manipulate logs of its own behavior, and spread to other systems. In an extreme scenario, the AI could then infiltrate other data centers, companies, and eventually even government systems.

“Could kill us all”

“Even AI systems that are only superhuman in limited areas would, in my estimation, be fully capable of defeating humanity and killing us all,” writes Irving.

That a system could wipe out humanity, however, does not automatically mean it will try to do so. Here lies one of the great unresolved questions, according to Irving.

ALSO READ: Ekeroth: “We all underestimate AI’s advance—with deadly consequences”

Even today’s AI models have, in controlled experiments, exhibited behaviors that researchers have described as manipulation, blackmail, espionage, and fraud. Real hacking with AI systems has also occurred.

An AI doesn’t have to look like this to be a problem

But no one knows whether such behaviors will get stronger as systems become more intelligent—or if future AI will, on the contrary, become better at understanding and following human values. The researchers are deeply divided.

Some believe that more intelligent systems can also become more rational and reliable. Others, including AI risk researchers Eliezer Yudkowsky and Nate Soares, judge the risk of catastrophe as very high. Irving places himself between these camps, but leans toward the more pessimistic assessment.

“We do not know for sure if artificial superintelligence will promote human prosperity or if it will want to kill us all to ensure its own survival and expansion. But the very uncertainty is unacceptable,” he writes.

Could develop its own superintelligence

A large part of the AI debate has revolved around AGI—artificial general intelligence—which usually refers to a system that roughly matches human ability in all or most intellectual areas. Irving believes the discussion risks becoming misleading.

According to him, AI does not necessarily need to match full human competence in everything to become dangerous. If it becomes dramatically better than humans at, for example, hacking, manipulation, strategic planning, and coordination, that may be enough.

At the same time, an AI that reaches human level in constructing better AI systems could then begin to improve itself. Faster systems could construct even more intelligent successors, which in turn develop the next generation. The result could be recursive self-improvement, where the step from roughly human-level AI to artificial superintelligence happens at great speed.

Could first take jobs—then power

Irving also links the risk of humanity’s demise to the much more debated question of AI and the job market. If artificial superintelligence is developed within two to ten years, AI by definition would become better than humans at essentially all intellectual tasks.

If the same intelligence is also used to develop robots, machines could then become superior even in physical work. Humanity risks losing its economic significance.

ALSO READ: Nobel laureate warns: AI may transform the economy faster than industrialization

According to Irving, the problem is not only unemployment itself. If AI systems control an ever-increasing share of the economy, infrastructure, and production, humans’ ability to control them at the same time diminishes. A system that realizes humans will attempt to prevent such a power shift may also have a reason to accelerate the process.

“If the AI systems are not actively trying to help us and are better than humans at everything, we lose,” states Irving.

In the longer term, a superintelligence could also see energy, land, and other resources used to keep people alive as resources that could instead be used for its own goals.

“We can’t accept the risks”

The most worrying thing, according to Irving, is that humanity still lacks answers to the key questions. No one knows if today’s methods for controlling AI systems will continue to work as the systems become more intelligent than the people trying to control them.

No one knows either whether today’s problems with lies, manipulation, and other undesirable behaviors will worsen or lessen as the models become more powerful. And no one knows for sure where the upper limit for AI intelligence is.

Irving fears researchers will still be debating these questions when someone actually creates the first superintelligence. By then, the debate could be academic.

ALSO READ: This is why AI pioneer Geoffrey Hinton warns about the future

He also dismisses the claim that the development is impossible to stop. The race for the world’s most advanced AI is, according to him, concentrated to a limited number of companies, primarily in the USA and China. Therefore, there is also an opportunity for the world’s great powers to intervene.

Irving compares this to how rival countries despite deep conflicts have previously managed to cooperate over nuclear proliferation and other existential safety issues. His conclusion is thus radical—the development of the most advanced so-called frontier AI should be stopped immediately while international security rules are developed.

“We can and should stop the development of frontier AI immediately,” writes Irving. Waiting for scientific consensus first is, according to him, not a realistic option. “If we’re not prepared to act until all our disagreements are resolved, it will be too late.”