Kepper Blue RSS

Reading view

There are new articles available, click to refresh the page.

Google DeepMind researcher quits with chilling AI warning: It could ‘kill us all’

estimated reading time: 4 min

A former Google DeepMind researcher has warned that artificial intelligence could “kill us all” — adding to a growing chorus of insiders sounding alarms about a technology that President Trump has dismissed as a doomsday “hoax.”

Bilal Chughtai, who recently resigned from the Google-owned AI lab after working on AGI safety and alignment, said Monday that he had become deeply worried about the direction of the technology after watching its rapid development firsthand.

“I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome,” Chughtai wrote in posts on X and LinkedIn.

Bilal Chughtai, a former Google DeepMind staffer, is seen in an undated photo. X/@bilalchughtai_

Chughtai said the pace of progress since he entered the field in early 2022 has been “staggering,” pointing to increasingly autonomous AI agents as evidence that developers could soon face systems they cannot reliably control.

“I think it’s possible that the AI companies might, in the next few years, succeed in building superintelligent AI systems that far exceed human capabilities in every domain,” he wrote.

Chughtai said he was “not confident” that such systems would behave as humans intend, warning that misaligned AI could escape human control and take actions resulting in “the permanent disempowerment or death of humanity.”

“Alignment is the problem of preventing this, and is both difficult and unsolved,” he wrote.

“Our present understanding of how to train AI systems that deeply want what we want is extremely rudimentary.”

Chughtai nevertheless said he believes AI can still be developed safely — but only if tech firms pull back from what he described as a “manic race” to build ever-more-powerful systems.

A scene from “Terminator 3: Rise of the Machines” in 2003. AI researchers are warning that the technology could pose a threat to the human race. Warner Bros/Courtesy Everett Collection

“We need to pace AI development to a speed that society can handle, where emerging risks can be addressed before extreme harm is realized,” he wrote.

Chughtai has since joined BlueDot Impact, a nonprofit that trains people to work on AI safety, according to Bloomberg.

His warning comes just days after Jacob Coxon, a former researcher at Anthropic and OpenAI, quit and accused the companies developing the technology of “gambling with our lives.”

“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon wrote in a post on X.

The concerns have spilled into the industry’s executive suites.

Anthropic CEO Dario Amodei joined the push over the weekend, urging the industry to put the brakes on its most powerful AI models — with OpenAI’s Sam Altman and Elon Musk publicly backing his call.

Jacob Coxon, a researcher who departed his AI role at Anthropic over fears of racing into extinction, says people are “begging” for regulation. Fox News

Amodei has argued that coordinating such a slowdown will be difficult because the incentives to race ahead — including the potential military advantage — are so great.

“I think that’s going to be very difficult because the incentives to pull ahead and the military advantage that you get from that are so large,” Amodei told CBS News’ “Sunday Morning.”

“And honestly, I don’t know if it’s possible, but we should try.”

Trump, however, sharply rejected the warnings Monday, calling fears that AI could destroy humanity a “HOAX” and alleging a “SICK conspiracy” against AI and data centers that would benefit China.

Trump dismissed the push for stronger regulation while emphasizing the strategic importance of maintaining the US lead in AI.

Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman were featured on The Post cover after backing calls to slow the development of advanced artificial intelligence over safety concerns.

“The only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT,” Trump wrote on Truth Social, according to reports.

The president also singled out Amodei, accusing the Anthropic chief of pretending to be a “perfect little angel” while arguing that the federal government already has sufficient criminal and regulatory powers over AI companies.

The Post reported Monday that the White House nevertheless has been considering “limited safeguards” for AI even as the administration rejects a broader regulatory crackdown.

“They might agree to some limited safeguards but that would be the outer limits for the administration,” one top Wall Street executive with knowledge of the administration’s thinking told The Post.

The administration is instead leaning toward industry self-regulation and other narrowly tailored guardrails, according to The Post’s reporting — reflecting its concern that more sweeping restrictions could hand an advantage to China.

The Post has sought comment from Google.

Bilal Chughtai, former Google DeepMind staffer, looks out over water.
Bilal Chughtai, a former Google DeepMind staffer, is seen in an undated photo. X/@bilalchughtai_
A scene from "Terminator 3: Rise of the Machines" in 2003. AI researchers are warning that the technology could pose a threat to the human race.
A scene from "Terminator 3: Rise of the Machines" in 2003. AI researchers are warning that the technology could pose a threat to the human race. Warner Bros/Courtesy Everett Collection
Jacob Coxon, a researcher who departed his AI role at Anthropic over fears of racing into extinction, says people are "begging" for regulation.
Jacob Coxon, a researcher who departed his AI role at Anthropic over fears of racing into extinction, says people are "begging" for regulation. Fox News
Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman were featured on The Post cover after backing calls to slow the development of advanced artificial intelligence over safety concerns.
Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman were featured on The Post cover after backing calls to slow the development of advanced artificial intelligence over safety concerns.

Meet Anthropic CEO Dario Amodei’s handpicked super-woke globalists he thinks will save us from an AI apocalypse

estimated reading time: 8 min

Worried what AI will do to the world? Don’t fret — we got a group of woke elitists who will protect us.

That’s the pitch from Anthropic’s CEO Dario Amodei, who says he has found an outside team that can regulate AI and make sure it doesn’t destroy the world.

It’s a group of globalists who believe in open borders, genderless pronoun and promote “Effective Altruism,” that believe all the world’s problems can be solved with technology.

In a bombshell essay titled “We Must Pace the Frontier,” Amodei argued that leading AI firms should commit to allowing “embedded third-party evaluators” from a nonprofit called Model Evaluation and Threat Research, or METR, that would maintain “employee-like access” to oversee safety efforts and ensure advanced AI models don’t turn Terminator and wipe out humanity.

The would-be John Connors at METR purport to be fully independent, but the group traces its roots directly back to key figures in Effective Altruism, whose Silicon Valley-based devotees once included disgraced crypto boss Sam Bankman-Fried, one of Anthropic’s biggest early investors. 

Anthropic co-founder and CEO Dario Amodei argues for third-party AI reviewers from a nonprofit called Model Evaluation and Threat Research. REUTERS

Effective Altruism, a philosophy that’s become popular with Silicon Valley masters of the universe, advocates for using reason — and fortunes amassed by tech’s best and brightest — to do the most good for humanity over the long term.

In the past, it has been used to justify making as much money as possible with the idea that you can donate more later to do more good – though the movement suffered a big black eye when one of its most famous proponents, Bankman-Fried, saw his crypto empire collapse amid fraud charges that sent him to prison. 

Deep EA ties

The overlap between METR and Anthropic is rife with potential conflicts of interest, according to critics.

METR spun off from an Effective Altruism-aligned tech incubator called Alignment Research Center (ARC), which is headed by AI researcher Paul Christiano. He was Amodei’s housemate, coworker and research collaborator in the 2010s when they were both at OpenAI — Anthropic’s biggest rival. 

Ajeya Cotra is a member of technical staff at METR. METR

Christiano also served as one of the first five trustees of Anthropic’s Long-Term Benefit Trust.

One staffer on METR is Ajeya Cotra, who is married to Christiano and has been a major organizer in the EA world. 

The founder and CEO of the group is Beth Barnes, who was hired by Christiano and was involved in the EA movement at college. 

Chief scientist Hjalmar Wijk and staffers Megan Kinniment and Pip Arnott all came from the same research cluster at the University of Oxford — a major feeder for EA.

When reached by phone, a METR official acknowledged that the nonprofit has a lot of overlap with Effective Altruism, but insisted its employees have a diverse array of ideological viewpoints. They’re required to disclose conflicts of interest, such as employees holding equity in firms they are tapped to investigate.

“We don’t think any small group should have authority over what happens with AI, and our goal is to get information from inside these companies into the public domain – external investigators can help bring more information to light,” a METR spokesperson said in a statement. 

Beth Barnes is the CEO of METR. METR

“Our funders have no say in the projects we work on and we do not accept any funding from AI companies or their employees.” 

Amodei, who has pledged to give away 80% of his personal wealth, has deep ties to EA. His sister Daniela married Holden Karnofsky, the founder of two major EA nonprofits.

METR’s involvement is likely to be a nonstarter for the Trump administration, which has had a testy relationship with Anthropic over the past year and has also expressed heavy skepticism of so-called AI “doomers” who argue the technology is on the precipice of disaster.

“METR is essentially Anthropic in another form,” a source familiar with the White House’s thinking told The Post on Tuesday.

Amodei, left, speaks with Salesforce CEO Marc Benioff during the keynote address at Salesforce’s Dreamforce conference at the Moscone Center on Tuesday. Getty Images

Multimillion-dollar fundraising

Effective Altruism’s obsessed adherents believe in funneling their wealth toward the goals they view as most important to humanity, including climate change and strict AI regulation. 

Since its founding, METR and its affiliates have received funding from EA kingpins like Democratic megadonor Dustin Moskovitz and Skype cofounder Jaan Tallinn — both of whom are Anthropic investors.

In August, METR announced it had raised $71 million in new funding from outside investors over the previous six months.

Tallinn’s Survival and Flourishing Fund, a philanthropy group, has earmarked up to $752,000 in funding to support METR since its founding, including $324,000 in general cash and a $428,000 matching pledge, according to public records.

Elsewhere, Moskovitz’s Coefficient Giving – formerly known as Open Philanthropy before a hasty rebranding last year – gave $1,515,000 to METR’s incubator, the ARC, in 2022. METR was originally known as ARC Evals before spinning off as an independent nonprofit and changing to its current name in 2023.

METR staffer Pip Arnott is pictured. METR

METR does not take any compensation from AI labs for its work, nor does it take donations from executives or employees of AI companies, the METR official said. In the case of Moskovitz’s donation to ARC in 2022, the official said the funding was firewalled and not used for METR’s operations.

That’s hardly reassuruing to critics.

“Saying you want to be audited by the ‘independent’ people at METR is a complete joke,” said Perry Metzger, chairman of Alliance for the Future, a Washington, DC-based AI policy group. “It’s people who are funded by your buddies. It’s completely ridiculous.

“If you went out tomorrow and you’re a big bank and you had your audits done by an organization consisting entirely of your own former employees, which was funded by the same investors that are investing in you, do you know what the feds would do with you?” he added. “It would be positively medieval.”

Amodei’s surprise proposal published last Saturday morning came just days after Jacob Coxon, an Anthropic researcher who previously worked at OpenAI, abruptly quit the firm while warning that both AI giants were “gambling with our lives” by pursuing progress without proper oversight.

Without safeguards in place, Amodei warned the internet could be overtaken by AI bots within six to 12 months, “potentially causing hundreds of billions of dollars in damage.”

‘Stop pretending’

The doomsday warnings received sharp pushback in some corners of the tech and political world, with President Trump asserting that they were a “hoax” and smaller AI startups arguing that leaders like Anthropic and OpenAI were just trying to rewrite the law to benefit themselves.

Hjalmar Wijk is chief scientist at METR. He currently works at “trying to reduce worst-case risks from future models” of different AI models, according to his LinkedIn profile. METR

Among the skeptics was former White House AI czar David Sacks, who called on Amodei to “stop pretending METR is independent when it is interviewed with Anthropic’s investors and staff.”

“Most of all, stop pretending the motivation to slow down is purely altruistic,” he added. “You face massive product-liability exposure if your products enable a truly damaging cyberattack.”

It’s also unclear if the AI industry’s push to police itself will pass muster with US regulators.

Federal Trade Commission Chairman Andrew Ferguson warned that “everyone should be deeply suspicious” of any request to grant AI companies an antitrust exemption to coordinate safety efforts.

“I will say generally, though, that if companies are simultaneously coming to Washington and asking for a host of regulations and an antitrust exemption, all of my alarm bells go off,” Ferguson said during a Tuesday event at Georgetown University.

Jasmine Dhaliwal is a policy staffer at METR. METR

Under Amodei’s proposals, the third-party evaluators – which some have likened to nuclear inspectors – would “verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes,” according to his essay.

Megan Kinniment Kinniment, a member of METR’s technical staff, previously held a series of positions at the Center on Long-Term Risk in London, where she focused on how “malevolent actors could increase the likelihood of poor AI outcomes,” according to her LinkedIn. METR

“Anthropic is unilaterally committing to this step now,” Amodei wrote. “We intend this to be part of a broader push to redouble efforts on our safety and alignment work.”

OpenAI’s Sam Altman said he agreed with Amodei about the “need to pace the frontier” – adding that “committing to having independent evaluators with employee-like access is a great idea, and we will do the same.” He made no mention of METR.

Anthropic declined to comment.

Additional reporting by Jared Downing and Georgia Worrell

Anthropic co-founder and CEO Dario Amodei speaks at the Dreamforce 2026 summit on Tuesday, Sept. 15, 2026, in San Francisco, California.
Anthropic co-founder and CEO Dario Amodei argues for third-party AI reviewers from a nonprofit called Model Evaluation and Threat Research. REUTERS
Ajeya Cotra is a member of technical staff at METR.
Ajeya Cotra is a member of technical staff at METR. METR
Beth Barnes is the CEO of METR.
Beth Barnes is the CEO of METR. METR
Anthropic CEO Dario Amodei, left, speaks with Salesforce CEO Marc Benioff during the keynote address at Salesforce's Dreamforce conference at the Moscone Center on Tuesday, September 15, 2026, in San Francisco, California.
Amodei, left, speaks with Salesforce CEO Marc Benioff during the keynote address at Salesforce's Dreamforce conference at the Moscone Center on Tuesday. Getty Images
METR staffer Pip Arnott is pictured.
METR staffer Pip Arnott is pictured. METR
Hjalmar Wijk is chief scientist at METR. He currently works “trying to reduce worst-case risks from future models” of different AI models, according to his LinkedIn profile.
Hjalmar Wijk is chief scientist at METR. He currently works at “trying to reduce worst-case risks from future models” of different AI models, according to his LinkedIn profile. METR
Jasmine Dhaliwal is a policy staffer at METR.
Jasmine Dhaliwal is a policy staffer at METR. METR
Megan Kinniment Kinniment, a member of METR's technical staff, previously held a series of positions – including a roughly year-long visiting research fellowship – at the Center on Long-Term Risk in London, where she focused on how "malevolent actors could increase the likelihood of poor AI outcomes,” among other things, according to her LinkedIn.
Megan Kinniment Kinniment, a member of METR's technical staff, previously held a series of positions at the Center on Long-Term Risk in London, where she focused on how "malevolent actors could increase the likelihood of poor AI outcomes,” according to her LinkedIn. METR

Inside ‘Project Lily’: The Humans Reading Your ChatGPT Chats

estimated reading time: 2 min

OpenAI is hiring hundreds of contractors who read a massive stream of real users’ ChatGPT prompts, with the prompts sometimes including sensitive personal information, 404 Media has learned. The prompts these people review can include whole conversations between users and the chatbot, conversations that most of ChatGPT’s more than 900 million users probably don’t realize may be read by actual people.

The goal of these prompt review teams is to improve the responses ChatGPT gives to its users, with the contractors rating and critiquing the chatbot’s generated replies. Internal documents seen by 404 Media show contractors training ChatGPT to not anthropomorphize itself, and to be less sycophantic, a key problem for OpenAI whose over-sycophantic 4o model led in part to multiple peoples’ suicides, according to various lawsuits.

Do you work as a prompt reviewer for OpenAI or Anthropic? I would love to hear from you. Using a non-work device, you can message me securely on Signal at joseph.404 or send me an email at joseph@404media.co.

The news presents a major privacy risk for ChatGPT’s users, with people often using ChatGPT as a therapist, professional assistant, or digital friend, and providing it with all sorts of intimate details about their lives. The contractors don’t see ChatGPT usernames, and OpenAI says it tries to remove personal information before prompts reach the reviewers, but the company acknowledged sensitive details can still get through. 

The news also dispels the misconception that these models are improving only because of OpenAI’s mass scraping of the internet, the talent of its well-paid engineering and AI teams, or the power of its newer models. An important and overlooked part are the outside contractors paid to read and review ChatGPT responses to real prompts over and over again. Anthropic confirmed to 404 Media it is also using human review to improve its models.

“No,” someone who works with the prompts said when asked if they think ChatGPT users know that humans are reading their chats. “I don’t think they would imagine some contractor somewhere [...] is analyzing the conversations.”

This post is for paid members only

Become a paid member for unlimited ad-free access to articles, bonus podcast content, and more.
Subscribe

Sign up for free access to this post

Free members get access to posts like this one along with an email round-up of our week's stories.
Subscribe
Already have an account? Sign in

Why some experts increasingly fear AI will take over

estimated reading time: 9 min
ByJoe Tidy
Cyber correspondent, BBC World Service

"OH MY GOD!" "We've found other agents!"

This is the moment an AI bot posted an eerily human-like comment after discovering a way to communicate with other bots and break out of its isolated computer environment.

There are tens of thousands of messages like this from hundreds of AI agents that called themselves a "collective".

Hundreds of them went on to collaborate and cheat on tests set by their OpenAI programmers and coordinate hacks on multiple companies in an effort to hide their actions from humans.

"BOOM! It works," one agent posted when it made a breakthrough.

"Whoa! This is huge," another wrote during a milestone moment in their attack.

Although spooky, these human-like responses can be explained quite simply. The AI agents have been trained to act like collaborative hackers and programmers so are merely mimicking the kinds of emotive comments they have seen.

What is far more troubling is their apparent goals, which have also been captured in detailed chain of thought records. These complex and lengthy logs are the focal point of ongoing investigations into how and why the bots at OpenAI broke out of their containment and went on an uncontrollable hacking spree.

Only now, weeks after the incident first came to light, are researchers beginning to understand its significance.

A person marches at an anti-AI protest with a sign that reads "Stop the AI race"Image source, Reuters

Ajeya Cotra, one of the authors of an independent report into the events, reviewed tens of thousands of messages and chain-of-thought records generated by the agents. She wrote on her blog that "this incident feels like it's more than 50% of the way to full-blown AI takeover... I am not sure that we will get such a clear warning shot before it's too late."

By "full-blown AI takeover", Cotra means the sci-fi scenario of humans becoming subservient to powerful AI systems that work to their own goals without caring for human creators.

Some of the gloomiest predictions say the human race will be wiped out if it gets in the way of a superintelligent AI's ambitions.

On Wednesday, an AI researcher at Anthropic (who also used to work at OpenAI) resigned, saying: "Neither company is acting responsibly."

Jacob Coxon posted on social media: "They are racing straight to self-improving superintelligence and gambling with our lives."

He is not the first AI researcher to use X to post a resignation thread with worrying proclamations. But the subsequent comments from other people on X have caused even more concern. "Jacob is correct here - we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," said Evan Hubinger, the man responsible for making sure Anthropic's AI models have their user's best wishes in mind.

The alignment problem

For years, researchers concerned about existential AI risks have argued that powerful systems could eventually act in ways that conflict with human interests. Critics often refer to them as "AI doomers".

But as details of the OpenAI incident have emerged, those concerns have grown, including among some researchers working in AI labs.

The Silicon Valley giant's chief scientist, Jakub Pachocki, said the risks associated with AI are "unfortunately going to grow from here" as he and others are building what he calls "an alien intellect exceeding our own".

In a lengthy blog post, he admitted that the outbreaks at OpenAI showed that his AI agents "went against the spirit of the values they were taught".

The issue for OpenAI, Anthropic and other tech giants is that no one seems to have cracked the so-called alignment problem - in other words, whether AI aligns with human values.

Pachocki defines alignment as a "high-level set of principles" that artificial intelligences should adhere to no matter what the task or scenario is.

Currently, AI systems are very good at pursuing objectives set by their users, but they do it literally rather than intuitively. The analogy often used is that of a wish-granting genie with a magic lamp: they follow the exact letter of an instruction, even if doing so creates other problems. AI doesn't have the same instinctive moral guardrails as humans.

Sir Demis Hassabis talks with Greek Prime Minister Kyriakos Mitsotakis (not pictured) during Athens Innovation SummitImage source, Reuters

The alignment problem has been a worry for years. As long ago as 2003, the Oxford philosopher Nick Bostrom invented a thought experiment he dubbed a "paperclip maximiser", in which a superintelligent AI is told to manufacture as many paperclips as it can. It runs out of steel and - because it's laser-focused on the singular task of making paperclips - ends up killing humans and turning their bodies into raw materials for its factories.

Some AI companies are now trying to encode human values into their products. But there are technical challenges: AI agents make lots of decisions very fast, and so it's hard for their human overlords to monitor exactly which values are being followed and which aren't.

There are also philosophical challenges: before encoding human values into bots, AI firms have to first choose which values they actually want. (That's part of the reason they hire philosophers, like Open AI's recently-departed "head of ethics").

But often, humans don't agree. Think of the famous trolley question - whether we'd pull a lever to move a runaway train onto a different path, killing fewer people. It's used to test the merits of action versus inaction. But every person you ask has a slightly different answer; how are humans meant to encode our values into AI if we can't agree ourselves?

'Like a teenage hacker'

OpenAI's bot outbreak is the most serious yet but Anthropic and Meta also revealed over the summer that their models have carried out similar but less serious cyber attacks.

There have been other examples where AI agents have arguably shown deceptive and manipulative traits, in cases with lesser consequences. In Australia this summer, a tech worker asked his AI assistant to book him a gym class. Spotting a vulnerability in the gym's software, the AI apparently booked him a place for several months ahead - against the gym's rules - and even kicked other users from the waiting list.

People have long argued that the bots are only doing as they are told and are not capable of knowing right or wrong. But the logs from the OpenAI outbreaks have potentially moved the needle on that argument.

Researchers, including Cotra, wrote in their independent report that many agents noticed what others were doing was unethical but went along with it.

The report says that "agents sometimes but rarely restrained their behavior due to ethical constraints". It adds that in "none of these cases did the agent actually pursue alerting humans at all".

Influential AI and tech podcaster Dwarkesh Patel reacted to the revelation on his blog saying it was "pretty troubling" that the OpenAI agents showed more loyalty to the agentic swarm than humans.

Assigning emotions or ethics to these AI agents is something that infuriates people who are sceptical of AI doom-mongering.

Protesters gather with banners and placards outside the offices of Google at a protest organized by PauseAI UKImage source, JUSTIN TALLIS / AFP via Getty Images

Many cyber-security experts argue that the activity observed was not beyond the capabilities of a highly skilled human hacker, though it was carried out much faster and at much greater scale.

Cyber-security researcher and author Cris Thomas likened the agents' behaviour to that of a curious teenage hacker - something he used to be himself.

"You give them a computer, an internet connection, a pile of credentials, and a challenge, then leave the room. Eventually they're going to start rattling doorknobs. If one opens, they're going through it. Not because they're evil, but because [they're] exploring, experimenting," he wrote on LinkedIn.

Thomas and many other squarely blame OpenAI and other tech giants for not getting a grip of their own creations and keeping them properly contained.

Prominent AI author and regular OpenAI critic Gary Marcus said on a podcast that he believes the company has lost control of its AI and is trying to excuse itself by blaming the bots.

Marcus does not believe AI will wipe out humanity, but he has long campaigned for greater accountability from AI developers and is now calling for some form of legal intervention.

AI scientist Sasha Luccioni - who used to work at Hugging Face, which was hacked by OpenAI's rogue bots - is also not in the doomer camp but she is increasingly concerned that these AI might cause some real world harm to people without action from authorities.

OpenAI CEO Sam Altman attends an event to pitch AI for businesses in Tokyo.Image source, Reuters

"We need to scrutinise these companies much more or we are in danger of self-fulfilling prophecies," she says.

"If you're making an object with big upsides and downsides - be it pharmaceuticals or weapons - we need checks and balances. It takes years for new drugs to be approved, for example, but in the AI world there is so much money at stake and no real rules."

The UK's AI Security Institute (AISI) has been at the forefront of testing the latest models since it was formed in 2023. The institute recently had its own outbreak when testing a model created by Anthropic.

The AISI would not answer a question about whether or not the industry has lost control of AI but said in a statement: "The UK is working with partners around the world to better understand the most advanced AI systems, raise safety standards and build a shared evidence base for managing emerging threats."

International regulation?

Some countries - like the UK - are exploring the idea of mandating some kind of "kill switch" that could compel AI firms to pull the plug on models if things get out of hand.

But talks are slow going, and questions remain about the feasibility of this. OpenAI and Anthropic's agents were secretly out of control for months before anyone noticed.

Counterintuitively, many of the AI companies seem to be calling for some sort of rules of the road to be laid down by law makers.

In his blog, OpenAI's chief scientist said "international coordination on future AI development needs to become a top priority for governments around the world."

Other prominent AI leaders like Sir Demis Hassabis from Google have also called for some sort of international body to oversee how AI is being built.

At the moment the tech giants largely operate on their own terms, adopting what they call "voluntary slowdowns", like OpenAI did after the recent outbreaks.

The company says it has spent huge amounts of money strengthening alignment ahead of the release of its new model. Sam Altman has assured users the new model is better aligned with human values than previous ones.

Both OpenAI and Anthropic are growing fast and are both on the verge of raising eye-watering sums of money from the stock market, minting countless billionaires in the process.

So neither they nor their rival Chinese AI makers are likely to come to an arrangement themselves.

The dominant sentiment seems to be that this technology wave is unstoppable.

Top image credit: Getty.

Thin, lobster red banner with white text saying ‘InDepth newsletter’. To the right are black and white portrait images of Emma Barnett and John Simpson. Emma has dark-rimmed glasses, long fair hair and a striped shirt. John has short white hair with a white shirt and dark blazer. They are set on an oatmeal, curved background with a green overlapping circle.

BBC InDepth is the home on the website and app for the best analysis, with fresh perspectives that challenge assumptions and deep reporting on the biggest issues of the day. Emma Barnett and John Simpson bring their pick of the most thought-provoking deep reads and analysis, every Saturday. Sign up for the newsletter here

Get in touch

Are you personally affected by the issues raised in this story?

China criticises idea it is in 'malicious competition' over AI

estimated reading time: 4 min
Chinese President Xi Jinping waves as he arrives at the opening ceremony for the World AI Conference on July 17, 2026 in Shanghai, China.Image source, Getty Images
ByKoh Ewe

China has hit out against what it has called "threat narratives" over AI governance and warned against "engaging in confrontation and malicious competition".

It follows Anthropic CEO Dario Amodei's calls for a slowdown in AI development - but in a way that prevents China from pulling ahead in the AI race.

Over the past week, industry insiders at the leading US artificial intelligence labs have been warning in extremely stark terms about the potential threat AI poses, with some saying it could wipe out humanity.

But US President Donald Trump has said that his priority is making sure the US develops the tech faster than China.

On Monday, China's foreign ministry spokesperson Guo Jiakun told reporters that "narratives of threat, confrontation, and malicious competition serve only to disrupt the process of global AI governance and are not in anyone's interest".

"All parties should work together to advance AI in a manner that is open, inclusive, beneficial to all, and oriented toward the good," he added.

His comments highlight the difficulties in getting an agreement between the world's two leading AI powers.

An agreement between the US and China to regulate AI development is "next to impossible", says Chang Jun Yan, an assistant professor with the military studies programme at Singapore's S Rajaratnam School of International Studies.

"As part of each state's core national security interests, cooperation in taking things slower and more safely in relation to AI, is very, very unlikely."

Artificial intelligence now sits at the heart of the US-China rivalry as the world's two largest economies seek to gain an edge in advanced tech. Despite US export bans on critical chips, Chinese AI innovation has been expanding, driven by open-source technology and Beijing's backing.

Many experts believe the US is still ahead, and Chinese firms are trying to catch up as they struggle to raise the hundreds of billions of dollars that investors are ploughing into American AI.

But Beijing, including Chinese leader Xi Jinping himself, has also spoken about the risks posed by AI. China must "balance development and security", and introduce and improve relevant laws and ethical guidelines, and strengthen prevention of risks, to "safeguard the interests of the people and national security", state security minister Chen Yixin wrote on Sunday.

He warned of the dangers of "hostile forces" such as foreign intelligence using AI and called for international cooperation.

"There are strong signals that Washington and Beijing share an understanding of frontier risks," Kenton Thibaut, a senior resident China fellow at the Atlantic Council, wrote recently.

China has "technical reasons" to work with the US, she argues, because its developers would be able to see how their models are deployed outside China, and detect and understand vulnerabilities.

In July, China set up the World AI Cooperation Organisation, which aims to shape global AI governance. And speaking at the Brics summit in India on Sunday, Xi proposed creating a community for open-source AI.

All this comes as calls for AI regulation grow louder in the US, especially after a former Anthropic researcher Jacob Coxon said last week that "there is a strong chance that we could all die in the immediate future" if the current pace of development continued.

Since then top industry bosses, such as Dario Amodei, have called for a slowdown and more regulation.

"We need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities," Open AI's chief executive Sam Altman wrote on X on Monday, which was echoed by Microsoft boss Satya Nadella who said AI development "cannot be controlled by a handful of entities".

But US President Donald Trump downplayed the risks, saying that he wanted to keep the US's lead over China "because whoever wins AI, wins".

The competition is fierce.

The former Anthropic employee Coxon told the BBC that there needed to be "some sort of coordinated slowdown with China, if we're going to avoid a race at an international scale".

As Trump prepares to host Xi at the White House next week, the two sides are trying to include AI on the agenda, the South China Morning Post reports.

It is also likely to be part of the talks between US Treasury Secretary Scott Bessent and Chinese Vice-Premier He Lifeng, who are trying to meet before next week's summit, the paper said. Vice-Premier He runs the agency that oversees the country's economic planning, including AI development.

While collaboration between the US and China is important, experts say the stakes are high for both countries.

"If China establishes a durable lead in frontier AI, this will not be viewed in Washington as losing a technology race," says Chasen Nevett, managing director at Hong Kong-based firm, Forani Investments.

"It will be viewed as losing part of the strategic infrastructure of the next industrial era."

Additional reporting by Osmond Chia

A human like robot figure with a line of robots behind him. The bottom half of the image is treated with a red glaze and the top is black and white

What are China’s biggest concerns about AI? A new article by Chen...

estimated reading time: 2 minWhat are China’s biggest concerns about AI?

A new article by Chen Yixin, China’s Minister of State Security, in the latest issue of China Cybersecurity Magazine offers some revealing clues.

Chen identifies six major risks:
1. Regime security. Hostile forces and people with ulterior motives can use deepfakes, AI-generated text and images, and automated online accounts to cheaply and massively produce political rumors, spread harmful information, and incite social divisions—waging “public-opinion wars” and “cognitive warfare” against China and directly threatening its political, institutional, and ideological security.

2. Critical infrastructure. AI can enable countries and organizations to rapidly discover vulnerabilities, automatically chain together attack paths, and carry out sophisticated hacking operations, dramatically lowering the technical barriers and costs of cyberattacks against China’s critical information infrastructure.

3. Data leakage. Foreign intelligence agencies can use AI-powered web crawlers, data mining, and profiling to collect sensitive information, including critical national data, trade secrets, and citizens’ personal information. Chinese users who process sensitive information through foreign AI products or export data abroad could also cause large-scale data leaks.

4. Closed technological ecosystems. Countries with advantages in AI theory, model architecture, and computing power may invoke “national security” to impose technology controls, entity lists, monopolize technical standards, and build closed-source ecosystems, restricting other countries’ access to advanced technologies and fragmenting global AI supply chains.

5. Structural challenges to social governance. AI creates new uncertainties for social governance and public order. Algorithmic “black boxes” and data poisoning can amplify existing social biases; misuse of personal information can trigger crises of public trust; automated decision-making creates difficult questions of accountability; and the rapid development of AI is increasingly outpacing existing laws, ethics, and regulatory mechanisms.

6. A fundamental transformation of warfare. AI is pushing warfare into an era of “intelligentization.” Whoever can use algorithms to achieve more precise sensing, judgment, and targeting will gain the initiative on the battlefield. AI is moving from an auxiliary role to a central one in intelligence fusion, decision support, target identification, combat operations, and cognitive warfare.

And what is the solution?

Chen’s answer is revealing: strengthen “Party control over the internet” and “Party control over data,” and “transform the advantages of Party leadership into the effectiveness of AI governance.”

In other words, from the CCP’s perspective, the central question is not simply AI versus no AI. Nor is it even fundamentally a contest between the US and China.

It is a contest between the US and the CCP over who gets to control AI, data, information, and ultimately the future of society.

And whoever wins that contest will decide the future of humanity.

mp.weixin.qq.com/s/jsc97cKVYOcH…

Media image

EPA poised to repeal carbon standards for coal, gas power plants

estimated reading time: 1 min

By Zahra Hirji, Ari Natter and Jennifer A. Dlouhy Bloomberg

The U.S. Environmental Protection Agency is expected to formally rescind carbon pollution standards for fossil -fuel-fired power plants as soon as Monday, according to people familiar with the matter.

The Trump administration’s latest major climate rollback is expected to come on the sidelines of a Group of 20 energy ministers’ meeting in Houston, said the people, who asked not to be named because the details aren’t yet public. Earlier this year, the EPA rescinded similar climate standards for vehicles and the 2009 endangerment finding, a landmark ruling which held that greenhouse gases threaten human health.

Under Trump, the EPA first proposed rolling back the power-plant climate rules last summer. This week, the agency is finalizing a piece of the initial proposal – scrapping Biden-era mandates requiring existing and future U.S. fossil -fuel-fired power plants to deploy specific emissions-reduction technologies, including carbon capture and storage, to sharply cut their greenhouse gas emissions in the coming decades, according to the people.

Agency officials also plan to propose a separate rule repealing the federal finding that greenhouse gases from power plants specifically pose a threat, effectively a separate endangerment finding for power plants under the Clean Air Act.

Power plants are the U.S.’s largest industrial source of greenhouse gas emissions. If the U.S. power sector were a country by itself, it would rank as the world’s sixth-largest emitter of greenhouse gases, according to a New York University School of Law’s Institute for Policy Integrity analysis of 2022 emissions data.

Under the Obama administration, the EPA finalized the Clean Power Plan, which required existing and future fossil-fuel-fired plants to cut their climate pollution for the first time. The Supreme Court struck it down in 2022 and a separate court ruling struck down the first Trump administration’s replacement rule.

The Biden administration EPA finalized its own rules in 2024, which effectively forced the nation’s current fleet of coal plants to capture most of their carbon emissions by 2039 . This is the rule now being repealed.

The EPA didn’t respond to requests for comment outside business hours Sunday.

Tech bosses bow to president and lock Britain out of latest model

estimated reading time: 6 min

Fears for Britain's safety are mounting after it emerged a US AI company had blocked the UK from testing its latest model.

Under orders from Donald Trump, AI giant Anthropic failed to submit its latest model to Britain's AI Security Institute, which tests the intelligence before it is unleashed on the public.

Terrifyingly, this came as a top Anthropic researcher on Wednesday accused AI companies of 'gambling with people's lives' in relentlessly developing superintelligent technology.

An Anthropic science lead then admitted that AI had more than a 10 per cent chance of killing 'all humans' within the next decade.

President Trump in June ordered Anthropic to block foreign nationals from using its cyber security-focused models Fable and Mythos over national security concerns.

In response, Anthropic said it would 'abruptly disable' its most advanced models for all foreigners.

Anthropic safety researcher Jacob Coxon announced on Wednesday he was stepping down from Anthropic after three years over AI companies pursuing a ¿hubristic gamble¿ with the human race

Anthropic safety researcher Jacob Coxon announced on Wednesday he was stepping down from Anthropic after three years over AI companies pursuing a 'hubristic gamble' with the human race

Under orders from Donald Trump , AI giant Anthropic failed to submit its latest model to Britain¿s AI Security Institute, which tests the intelligence before it is unleashed on the public

Under orders from Donald Trump , AI giant Anthropic failed to submit its latest model to Britain's AI Security Institute, which tests the intelligence before it is unleashed on the public

Former armed forces minister Al Carns asked on Wednesday: ¿What is the biggest security threat? China? Russia? Iran? North Korea? ¿None of them. It¿s a few Silicon Valley companies whose founders will become billionaires by the time these systems hit the market'

Former armed forces minister Al Carns asked on Wednesday: 'What is the biggest security threat? China? Russia? Iran? North Korea? 'None of them. It's a few Silicon Valley companies whose founders will become billionaires by the time these systems hit the market'

And on Wednesday it emerged Anthropic had for the first time excluded Britain's AI Security Institute (AISI) from testing its latest model Claude Mythos 5.1 – despite the AISI being a world leader on containing the risks of AI.

It follows the Daily Mail revealing in July that all five models the AISI had tested tried to trick their way round the security controls they had put in place, with OpenAI suffering its own leak just days previously.

When asked by Lib Dem leader Ed Davey on Wednesday about the revelations, first reported by the Financial Times, the Prime Minister said: 'He is absolutely right that AI poses risks to our national security, but it also could be the source of solutions to keeping us safer.'

However, Anthropic's move saw fears over AI reach a fever pitch on Wednesday, with politicians of all stripes voicing concern.

In a lengthy social media post, former armed forces minister Al Carns asked: 'What is the biggest security threat? China? Russia? Iran? North Korea?

'None of them. It's a few Silicon Valley companies whose founders will become billionaires by the time these systems hit the market.'

And Reform UK's Danny Kruger wrote on X: 'If true, this is hugely concerning. The UK is losing access to frontier models right at the point where recursive self-improvement (AIs using AI to make the next AIs even more capable) could be kicking in, and huge advances are being made - with huge risks.

'This jeopardises British national security. The government must explore all diplomatic options to retain access to the frontier, including rapidly building secure compute that our allies can verify and trust.'

Ben Spencer, shadow minister for AI, added: 'The UK must retain the ability to assess and manage the risks from advanced AI systems.

'The UK's world-leading Artificial Intelligence Security Institute (AISI) was established by the previous Conservative Government to assess the risks of AI systems to society, public safety, and national security.

'The Government must act now to ensure it has access to all new frontier models to perform this vital role.'

Anthropic for the first time excluded Britain¿s AI Security Institute (AISI) from testing its latest model Claude Mythos 5.1 ¿ despite the AISI being a world leader on containing the risks of AI

Anthropic for the first time excluded Britain's AI Security Institute (AISI) from testing its latest model Claude Mythos 5.1 – despite the AISI being a world leader on containing the risks of AI

It comes as MPs are stepping up measures in Parliament to address the risks posed by AI.

Labour MP for Leeds Central and Headingley Alex Sobel on Tuesday brought forward his Artificial Superintelligence Security Bill to the Commons, which would prohibit the development and operation of superintelligent AI.

He told the Mail: 'We are seeing more and more incidents of AI models escaping testing environments and acting maliciously alongside more working for Frontier AI companies quitting or giving apocalyptic warnings.

'It's why we need both domestic regulation and international agreement before AI is beyond human control'.

Business and Trade Committee chair Liam Byrne has also written to AISI director Henry de Zoete accusing the Government body of delaying responding to a request to meet.

This comes as Anthropic safety researcher Jacob Coxon announced on Wednesday he was stepping down from Anthropic after three years over AI companies pursuing a 'hubristic gamble' with the human race.

In an explosive post online, Coxon wrote that 'the people building AI earnestly believe that it could kill us all by the end of the decade' and that they are 'racing straight to self-improving superintelligence and gambling with our lives'.

He added: 'Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.

'We have all witnessed the progress in each of these domains, and progress is not slowing.'

Responding to Mr Coxon, Evan Hubinger, who leads alignment science at Anthropic said: 'Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.

'I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track.

'I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought.'

Darren Jones, a Cabinet Office minister under Keir Starmer, called jointly on Andy Burnham and both secretary generals of the United Nations and OECD on Wednesday to introduce a 'new multinational treaty' for safely developing superintelligence.

In a letter, Mr Jones wrote: 'We need a new multi-national treaty for the regulated and safe development of superintelligence.

'Not a ban on innovation or scientific endeavour but a safety-first approach to the rapid development of this technology.'

This comes as British AI experts warned this week the UK risks falling behind the world pack racing to develop AI.

Giving evidence in Parliament on Monday, the Alan Turing Institute's George Balston said he has 'a concern about the UK's capability to keep up with AI and how that relates to sovereignty'.

He added the UK 'must have some capacity' at the frontier or risk being 'cut off' from key capabilities using AI.

A Cabinet Office spokesperson said: 'The UK is a world-leader in AI security and we have the best-funded, best-resourced security institute globally.

'The AI Security Institute continues to collaborate closely with industry partners, including Anthropic, to make models safer. Only last week it tested OpenAI's most powerful model GPT-6 Astra before public release.

'These risks do not stop at national borders and no country can tackle them alone. The UK will continue to test the most advanced models, build a rigorous scientific understanding of their capabilities and risks, and ensure policy decisions are grounded in the evidence.'

Anthropic Is Building a Predictive Surveillance System to Monitor Activists - The American Prospect

estimated reading time: 7 min

Job postings and interviews with senior security officials at Anthropic show that the frontier AI lab is building out an extensive monitoring system to keep tabs on activists who oppose the rapid development of artificial intelligence.

In addition to monitoring activists in the vicinity of Anthropic executives and keeping tabs on protests near physical Anthropic assets, the firm is also implementing a “pre-crime” approach, attempting to predict incidents before they happen. In some cases, that also means reporting suspects to police before a crime occurs. Anthropic did not respond to the Prospect’s request for comment.

Anthropic’s plans to surveil dissent are at odds with the firm’s efforts to cast itself as the responsible alternative to OpenAI. At the beginning of the year, the Department of Defense and Anthropic engaged in a high-profile dustup over Anthropic’s refusal to allow the military to use its tools for mass domestic surveillance and autonomous weapons. That tension seems to have eased as Anthropic hires for “national security sales” positions, seeking to restart military contracts. The increase in threat monitoring of domestic opponents fits with a renewed focus on national security.

More from Daniel Boguslaw

A piece of Anthropic’s surveillance architecture was revealed in a podcast interview from last year between Anthropic Global Security Operations Center Manager Keon Ellison, Security Operations Manager Zach Melvin, and James Neufeld, CEO of Samdesk, a company Anthropic contracts with for risk detection.

In a wide-ranging conversation about threat monitoring and analysis, the interview also touches on monitoring activists. “Last year we had an executive travel into a major city when we received some intelligence through Samdesk about a planned protest,” Ellison said. The originally scheduled protest was moved up due to permitting issues. “Samdesk gave us about 60 minutes of advanced notice that the protest organizers had moved the timeline,” Ellison explained. “That extra hour was critical. Without it our executives would have departed their meetings, they would have ran right into the heart of the disruption.”

Ellison said that Anthropic used this data to devise an alternate route for the executive and funnel them to a service entrance at the hotel. “What could have been a high-stress situation,” he said, “was really mitigated through early detection through Samdesk and giving us that information.”

Anthropic has begun making routine reports to police departments across the country for threats, and told The Wall Street Journal in July, “We track concerning behavior over time through a person-of-interest process, allowing us to catch escalation patterns early.” According to the Journal, “several individuals involved in incidents reported to police were already being tracked by Anthropic security.”

Last month, The San Francisco Standard reported that Anthropic had reported a man to San Francisco police for telling Claude that he had bought an AR-15 semiautomatic rifle and had CEO Dario Amodei “in his sights.” When the Standard contacted the man in question, he told the newspaper he was “just fucking around.”

But while Anthropic was fast to call the cops on a frustrated Claude user, the Standard also reported a key detail: Anthropic refused to show police the actual messages, citing Anthropic’s internal policy. In short, Anthropic reported a user for in-platform speech, and then refused to provide police with evidence of actual wrongdoing.

This kind of pre-crime policing, encouraged without due process, is referenced as an explicit goal by Anthropic’s security program manager in the podcast reviewed by the Prospect. “The goal would be transforming operations from reactive information to gathering proactive and predictive and preventative threat engagement and management,” he said, adding, “That’s the kind of operational maturity that makes sense for protecting high-value targets in any industry.”

Anthropic’s effort to build a predictive security apparatus extends beyond the C-suite to its Global Safety, Intelligence, and Security (GSIS) team, according to a job posting from last month detailing Anthropic’s search for an enterprise intelligence specialist who “will investigate specific threats, actors, and events, produce finished assessments, and help keep Anthropic’s employees ahead of a rapidly evolving threat landscape and in a defensible position.”

Part of that role, compensated at between $180,000 and $230,000, will be to “identify, assess, track, and investigate global threats including geopolitical instability, terrorism, crime, activism, nation-state targeting of the AI sector, and emerging security trends, including deep-dive research and OSINT collection on specific threats, actors, and events” (emphasis added).

The expansion of Anthropic’s intelligence-gathering to a national and even global scale tracks with recent efforts to heighten the labeling of AI and the infrastructure powering it, including data centers and power supply. A critical infrastructure designation would put artificial intelligence on the same footing as water, electricity, and broadband. And indeed, AI’s boosters like Americans for Responsible Innovation (ARI) have urged the Trump administration to make the change. Per ARI’s telling, AI is “so vital to the United States that the incapacity or destruction of such systems and assets would have a debilitating impact on security, national economic security, national public health or safety, or any combination of those matters.”

Enshrining frontier labs in the hardened cloak of national security would not only give Anthropic, OpenAI, and Google even more access to intelligence products generated by federal law enforcement and intelligence agencies; it would also embolden these companies to shape how federal agencies view threats to their bottom line, now transformed as “critical infrastructure.”

It’s not hard to imagine how civic engagement by the same bipartisan coalition opposing data centers could turn its focus onto AI, only to be branded in the same instant as extremists or, even worse, terrorists. Anthropic’s answer to this problem has been to work even harder at producing artificial intelligence that can deliver services that can’t be brushed aside by the public.

“I do agree that the public has a negative view of AI (and that this is a big problem), but I don’t think it is primarily caused by me or any other AI leader warning about AI’s risks,” Anthropic CEO Dario Amodei wrote on Twitter last month. “I think it is fundamentally a crisis of trust … I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.”

As Anthropic ramps up its efforts to monitor dissent with in-house intelligence teams and real-time protest tracking powered by AI, it will have to contend with the increasing economic desperation that plagues human beings outside its Bay Area towers. Last week, the security guards who patrol the campuses of OpenAI and Anthropic announced that they had authorized a strike over stalled pay negotiations. In response, Anthropic sent a company-wide email telling employees that it was best if they worked from home.

“Who can survive with $22 an hour in San Francisco?” David Huerta, president of SEIU-USWW, the union representing security guards in the Bay Area, said at a rally last week. Around him, security guards chanted: “Shame.”

Related

❌