Anthropic Alarm: “AI Could Kill All Humans” — Researcher Quits in Explosive Warning
Πηγή Φωτογραφίας: AP Photo//Anthropic Alarm: “AI Could Kill All Humans” — Researcher Quits in Explosive Warning
One of the most alarming warnings yet from inside the artificial intelligence industry is coming not from outside critics or academics, but from researchers working — or until very recently working — at the frontier of advanced AI development.
Jacob Coxon has resigned from Anthropic after spending the previous three years conducting pre-training research at Anthropic and OpenAI. His reason for leaving was extraordinary: he believes the industry is moving toward self-improving superintelligence without adequate safeguards.
The resignation that shook the AI debate
Coxon said he decided to leave because he does not believe either Anthropic or OpenAI is acting responsibly enough in the face of the risks he sees emerging.
“Neither company is acting responsibly,” he wrote, arguing that the labs are racing toward self-improving superintelligence and “gambling with our lives.”
Even more striking was his assertion that people actually building advanced AI systems take the possibility of catastrophic consequences far more seriously in private than the public debate might suggest.
“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon wrote.
His claim does not mean there is a scientific consensus that AI will wipe out humanity. What makes the intervention significant is that it comes from a researcher with experience inside two of the world’s leading frontier AI labs.
Anthropic’s alignment lead: “Jacob is correct here”
Then came Evan Hubinger.
Rather than dismissing his former colleague’s warning as alarmist, Anthropic’s Alignment Science lead publicly agreed with its central premise.
“Jacob is correct here — we really do earnestly believe AI could kill all humans!” Hubinger wrote.
He then attached a startling personal probability to that scenario:
“I personally think it is >10% within the next decade.”
Hubinger also stressed that he believes Anthropic is genuinely trying to address the problem, while warning that the industry still does not have a proven solution for safely aligning superintelligence.
That distinction matters enormously.
He is not claiming that today’s AI systems are about to destroy humanity. His concern centers on what could happen if AI reaches self-improving superintelligence before researchers know how to control such systems reliably.
A crucial distinction: Today’s AI is not the warning
The headline is dramatic, but the underlying argument is more precise.
The concern is not that a current chatbot suddenly decides to eliminate humanity.
The scenario these researchers are discussing involves future systems with capabilities far beyond those of today — particularly AI capable of accelerating AI research itself and contributing to the development of increasingly powerful successor systems.
This is where recursive self-improvement enters the debate.
If AI becomes highly capable at coding, scientific research, experimentation and AI development itself, it could potentially accelerate the rate at which the next generation of systems is created.
For safety researchers, that raises a fundamental question: what happens if capability growth begins moving faster than humanity’s ability to understand and control it?
The problem at the center of everything: Alignment
This is the problem known as AI alignment.
In simple terms, alignment asks how humans can ensure that an AI system — especially one more capable than its creators — continues to pursue goals compatible with human intentions and values.
The difficulty increases as systems become more autonomous.
A sufficiently capable system might follow an instruction literally while pursuing it in ways its designers never anticipated. The more powerful the system, the greater the potential consequences of such a failure.
Anthropic has invested heavily in alignment, interpretability and AI safety research.
That is precisely why warnings coming from researchers working inside those areas carry unusual weight.
The AI race is creating its own trap
There is also a broader industry problem.
Frontier AI labs are locked in an increasingly expensive and strategically important competition.
OpenAI, Anthropic, Google DeepMind and other players face enormous pressure to develop more capable models — from investors, customers, governments and one another.
This creates a classic race dynamic.
If one company slows down because it believes the technology is becoming too dangerous, it risks another company — or another country — continuing anyway.
The result is an uncomfortable incentive structure: individual labs may recognize serious risks while still believing that unilateral restraint could leave them behind.
AI is increasingly helping build AI
This is where the debate becomes particularly consequential for the technology industry.
Advanced models are already becoming more useful in software engineering, scientific research, mathematics and automated experimentation.
As those capabilities improve, AI can play a larger role in AI research itself.
That creates the possibility of a feedback loop: stronger AI helps researchers build stronger AI, which then accelerates the development of the next generation.
It is this transition — rather than the capabilities of today’s consumer chatbots — that lies behind much of the concern about superintelligence.
The OpenAI warning
The unease is not confined to Anthropic.
Senior researchers elsewhere in the frontier AI industry have also warned that alignment and monitoring capabilities may not be advancing quickly enough to justify indefinitely scaling AI systems at maximum speed.
That does not mean OpenAI, Anthropic or their leadership teams endorse Coxon’s specific claims or Hubinger’s personal probability estimate.
It does show, however, that the debate over how quickly increasingly powerful AI can safely be developed has moved inside the companies driving the technological race itself.
Another problem: We still struggle to understand how AI “thinks”
There is another technical challenge underneath the safety debate: interpretability.
Modern neural networks can contain enormous numbers of parameters and develop internal representations that researchers cannot fully explain.
AI labs are therefore investing heavily in techniques designed to understand why models produce particular outputs and how internal features correspond to concepts, strategies or behaviors.
This matters because controlling a highly capable system becomes significantly harder if researchers cannot reliably understand what is happening inside it.
The technical question therefore becomes a governance question:
Who decides when an AI model is too powerful — or too poorly understood — to deploy?
Silicon Valley is pressing the accelerator and the brake
This may be the defining contradiction of the AI industry.
The companies spending billions to create the world’s most capable AI systems are also home to researchers issuing some of the strongest warnings about what those systems could eventually become.
On one side are enormous economic and geopolitical incentives.
AI promises productivity gains, scientific breakthroughs, new medicines, advanced robotics and potentially transformative economic growth. Governments increasingly view frontier AI as a strategic technology comparable to semiconductors, energy infrastructure and defense.
On the other side are researchers arguing that capabilities could advance faster than safety.
That tension will only become harder to manage as AI becomes more autonomous.
A resignation that opens a much bigger question
Jacob Coxon’s departure from Anthropic is ultimately about more than one Silicon Valley resignation.
It reopens perhaps the hardest question of the AI era:
What happens if machines become more capable faster than humans become capable of controlling them?
There is no evidence today that artificial intelligence is destined to “kill all humans,” and Hubinger’s greater-than-10% figure is a personal risk estimate — not a scientifically established probability.
But something important has changed.
People working inside the world’s leading AI laboratories consider catastrophic AI risk serious enough to discuss publicly — and, in Coxon’s case, serious enough to walk away from his job.
That alone pushes the debate over superintelligence beyond science fiction.
It is becoming a real question of industrial strategy, corporate responsibility, national security and global governance.
Source: pagenews.gr
Διαβάστε όλες τις τελευταίες Ειδήσεις από την Ελλάδα και τον Κόσμο