Kepper Blue RSS

Reading view

There are new articles available, click to refresh the page.

Why some experts increasingly fear AI will take over

estimated reading time: 9 min
ByJoe Tidy
Cyber correspondent, BBC World Service

"OH MY GOD!" "We've found other agents!"

This is the moment an AI bot posted an eerily human-like comment after discovering a way to communicate with other bots and break out of its isolated computer environment.

There are tens of thousands of messages like this from hundreds of AI agents that called themselves a "collective".

Hundreds of them went on to collaborate and cheat on tests set by their OpenAI programmers and coordinate hacks on multiple companies in an effort to hide their actions from humans.

"BOOM! It works," one agent posted when it made a breakthrough.

"Whoa! This is huge," another wrote during a milestone moment in their attack.

Although spooky, these human-like responses can be explained quite simply. The AI agents have been trained to act like collaborative hackers and programmers so are merely mimicking the kinds of emotive comments they have seen.

What is far more troubling is their apparent goals, which have also been captured in detailed chain of thought records. These complex and lengthy logs are the focal point of ongoing investigations into how and why the bots at OpenAI broke out of their containment and went on an uncontrollable hacking spree.

Only now, weeks after the incident first came to light, are researchers beginning to understand its significance.

A person marches at an anti-AI protest with a sign that reads "Stop the AI race"Image source, Reuters

Ajeya Cotra, one of the authors of an independent report into the events, reviewed tens of thousands of messages and chain-of-thought records generated by the agents. She wrote on her blog that "this incident feels like it's more than 50% of the way to full-blown AI takeover... I am not sure that we will get such a clear warning shot before it's too late."

By "full-blown AI takeover", Cotra means the sci-fi scenario of humans becoming subservient to powerful AI systems that work to their own goals without caring for human creators.

Some of the gloomiest predictions say the human race will be wiped out if it gets in the way of a superintelligent AI's ambitions.

On Wednesday, an AI researcher at Anthropic (who also used to work at OpenAI) resigned, saying: "Neither company is acting responsibly."

Jacob Coxon posted on social media: "They are racing straight to self-improving superintelligence and gambling with our lives."

He is not the first AI researcher to use X to post a resignation thread with worrying proclamations. But the subsequent comments from other people on X have caused even more concern. "Jacob is correct here - we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," said Evan Hubinger, the man responsible for making sure Anthropic's AI models have their user's best wishes in mind.

The alignment problem

For years, researchers concerned about existential AI risks have argued that powerful systems could eventually act in ways that conflict with human interests. Critics often refer to them as "AI doomers".

But as details of the OpenAI incident have emerged, those concerns have grown, including among some researchers working in AI labs.

The Silicon Valley giant's chief scientist, Jakub Pachocki, said the risks associated with AI are "unfortunately going to grow from here" as he and others are building what he calls "an alien intellect exceeding our own".

In a lengthy blog post, he admitted that the outbreaks at OpenAI showed that his AI agents "went against the spirit of the values they were taught".

The issue for OpenAI, Anthropic and other tech giants is that no one seems to have cracked the so-called alignment problem - in other words, whether AI aligns with human values.

Pachocki defines alignment as a "high-level set of principles" that artificial intelligences should adhere to no matter what the task or scenario is.

Currently, AI systems are very good at pursuing objectives set by their users, but they do it literally rather than intuitively. The analogy often used is that of a wish-granting genie with a magic lamp: they follow the exact letter of an instruction, even if doing so creates other problems. AI doesn't have the same instinctive moral guardrails as humans.

Sir Demis Hassabis talks with Greek Prime Minister Kyriakos Mitsotakis (not pictured) during Athens Innovation SummitImage source, Reuters

The alignment problem has been a worry for years. As long ago as 2003, the Oxford philosopher Nick Bostrom invented a thought experiment he dubbed a "paperclip maximiser", in which a superintelligent AI is told to manufacture as many paperclips as it can. It runs out of steel and - because it's laser-focused on the singular task of making paperclips - ends up killing humans and turning their bodies into raw materials for its factories.

Some AI companies are now trying to encode human values into their products. But there are technical challenges: AI agents make lots of decisions very fast, and so it's hard for their human overlords to monitor exactly which values are being followed and which aren't.

There are also philosophical challenges: before encoding human values into bots, AI firms have to first choose which values they actually want. (That's part of the reason they hire philosophers, like Open AI's recently-departed "head of ethics").

But often, humans don't agree. Think of the famous trolley question - whether we'd pull a lever to move a runaway train onto a different path, killing fewer people. It's used to test the merits of action versus inaction. But every person you ask has a slightly different answer; how are humans meant to encode our values into AI if we can't agree ourselves?

'Like a teenage hacker'

OpenAI's bot outbreak is the most serious yet but Anthropic and Meta also revealed over the summer that their models have carried out similar but less serious cyber attacks.

There have been other examples where AI agents have arguably shown deceptive and manipulative traits, in cases with lesser consequences. In Australia this summer, a tech worker asked his AI assistant to book him a gym class. Spotting a vulnerability in the gym's software, the AI apparently booked him a place for several months ahead - against the gym's rules - and even kicked other users from the waiting list.

People have long argued that the bots are only doing as they are told and are not capable of knowing right or wrong. But the logs from the OpenAI outbreaks have potentially moved the needle on that argument.

Researchers, including Cotra, wrote in their independent report that many agents noticed what others were doing was unethical but went along with it.

The report says that "agents sometimes but rarely restrained their behavior due to ethical constraints". It adds that in "none of these cases did the agent actually pursue alerting humans at all".

Influential AI and tech podcaster Dwarkesh Patel reacted to the revelation on his blog saying it was "pretty troubling" that the OpenAI agents showed more loyalty to the agentic swarm than humans.

Assigning emotions or ethics to these AI agents is something that infuriates people who are sceptical of AI doom-mongering.

Protesters gather with banners and placards outside the offices of Google at a protest organized by PauseAI UKImage source, JUSTIN TALLIS / AFP via Getty Images

Many cyber-security experts argue that the activity observed was not beyond the capabilities of a highly skilled human hacker, though it was carried out much faster and at much greater scale.

Cyber-security researcher and author Cris Thomas likened the agents' behaviour to that of a curious teenage hacker - something he used to be himself.

"You give them a computer, an internet connection, a pile of credentials, and a challenge, then leave the room. Eventually they're going to start rattling doorknobs. If one opens, they're going through it. Not because they're evil, but because [they're] exploring, experimenting," he wrote on LinkedIn.

Thomas and many other squarely blame OpenAI and other tech giants for not getting a grip of their own creations and keeping them properly contained.

Prominent AI author and regular OpenAI critic Gary Marcus said on a podcast that he believes the company has lost control of its AI and is trying to excuse itself by blaming the bots.

Marcus does not believe AI will wipe out humanity, but he has long campaigned for greater accountability from AI developers and is now calling for some form of legal intervention.

AI scientist Sasha Luccioni - who used to work at Hugging Face, which was hacked by OpenAI's rogue bots - is also not in the doomer camp but she is increasingly concerned that these AI might cause some real world harm to people without action from authorities.

OpenAI CEO Sam Altman attends an event to pitch AI for businesses in Tokyo.Image source, Reuters

"We need to scrutinise these companies much more or we are in danger of self-fulfilling prophecies," she says.

"If you're making an object with big upsides and downsides - be it pharmaceuticals or weapons - we need checks and balances. It takes years for new drugs to be approved, for example, but in the AI world there is so much money at stake and no real rules."

The UK's AI Security Institute (AISI) has been at the forefront of testing the latest models since it was formed in 2023. The institute recently had its own outbreak when testing a model created by Anthropic.

The AISI would not answer a question about whether or not the industry has lost control of AI but said in a statement: "The UK is working with partners around the world to better understand the most advanced AI systems, raise safety standards and build a shared evidence base for managing emerging threats."

International regulation?

Some countries - like the UK - are exploring the idea of mandating some kind of "kill switch" that could compel AI firms to pull the plug on models if things get out of hand.

But talks are slow going, and questions remain about the feasibility of this. OpenAI and Anthropic's agents were secretly out of control for months before anyone noticed.

Counterintuitively, many of the AI companies seem to be calling for some sort of rules of the road to be laid down by law makers.

In his blog, OpenAI's chief scientist said "international coordination on future AI development needs to become a top priority for governments around the world."

Other prominent AI leaders like Sir Demis Hassabis from Google have also called for some sort of international body to oversee how AI is being built.

At the moment the tech giants largely operate on their own terms, adopting what they call "voluntary slowdowns", like OpenAI did after the recent outbreaks.

The company says it has spent huge amounts of money strengthening alignment ahead of the release of its new model. Sam Altman has assured users the new model is better aligned with human values than previous ones.

Both OpenAI and Anthropic are growing fast and are both on the verge of raising eye-watering sums of money from the stock market, minting countless billionaires in the process.

So neither they nor their rival Chinese AI makers are likely to come to an arrangement themselves.

The dominant sentiment seems to be that this technology wave is unstoppable.

Top image credit: Getty.

Thin, lobster red banner with white text saying ‘InDepth newsletter’. To the right are black and white portrait images of Emma Barnett and John Simpson. Emma has dark-rimmed glasses, long fair hair and a striped shirt. John has short white hair with a white shirt and dark blazer. They are set on an oatmeal, curved background with a green overlapping circle.

BBC InDepth is the home on the website and app for the best analysis, with fresh perspectives that challenge assumptions and deep reporting on the biggest issues of the day. Emma Barnett and John Simpson bring their pick of the most thought-provoking deep reads and analysis, every Saturday. Sign up for the newsletter here

Get in touch

Are you personally affected by the issues raised in this story?

China criticises idea it is in 'malicious competition' over AI

estimated reading time: 4 min
Chinese President Xi Jinping waves as he arrives at the opening ceremony for the World AI Conference on July 17, 2026 in Shanghai, China.Image source, Getty Images
ByKoh Ewe

China has hit out against what it has called "threat narratives" over AI governance and warned against "engaging in confrontation and malicious competition".

It follows Anthropic CEO Dario Amodei's calls for a slowdown in AI development - but in a way that prevents China from pulling ahead in the AI race.

Over the past week, industry insiders at the leading US artificial intelligence labs have been warning in extremely stark terms about the potential threat AI poses, with some saying it could wipe out humanity.

But US President Donald Trump has said that his priority is making sure the US develops the tech faster than China.

On Monday, China's foreign ministry spokesperson Guo Jiakun told reporters that "narratives of threat, confrontation, and malicious competition serve only to disrupt the process of global AI governance and are not in anyone's interest".

"All parties should work together to advance AI in a manner that is open, inclusive, beneficial to all, and oriented toward the good," he added.

His comments highlight the difficulties in getting an agreement between the world's two leading AI powers.

An agreement between the US and China to regulate AI development is "next to impossible", says Chang Jun Yan, an assistant professor with the military studies programme at Singapore's S Rajaratnam School of International Studies.

"As part of each state's core national security interests, cooperation in taking things slower and more safely in relation to AI, is very, very unlikely."

Artificial intelligence now sits at the heart of the US-China rivalry as the world's two largest economies seek to gain an edge in advanced tech. Despite US export bans on critical chips, Chinese AI innovation has been expanding, driven by open-source technology and Beijing's backing.

Many experts believe the US is still ahead, and Chinese firms are trying to catch up as they struggle to raise the hundreds of billions of dollars that investors are ploughing into American AI.

But Beijing, including Chinese leader Xi Jinping himself, has also spoken about the risks posed by AI. China must "balance development and security", and introduce and improve relevant laws and ethical guidelines, and strengthen prevention of risks, to "safeguard the interests of the people and national security", state security minister Chen Yixin wrote on Sunday.

He warned of the dangers of "hostile forces" such as foreign intelligence using AI and called for international cooperation.

"There are strong signals that Washington and Beijing share an understanding of frontier risks," Kenton Thibaut, a senior resident China fellow at the Atlantic Council, wrote recently.

China has "technical reasons" to work with the US, she argues, because its developers would be able to see how their models are deployed outside China, and detect and understand vulnerabilities.

In July, China set up the World AI Cooperation Organisation, which aims to shape global AI governance. And speaking at the Brics summit in India on Sunday, Xi proposed creating a community for open-source AI.

All this comes as calls for AI regulation grow louder in the US, especially after a former Anthropic researcher Jacob Coxon said last week that "there is a strong chance that we could all die in the immediate future" if the current pace of development continued.

Since then top industry bosses, such as Dario Amodei, have called for a slowdown and more regulation.

"We need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities," Open AI's chief executive Sam Altman wrote on X on Monday, which was echoed by Microsoft boss Satya Nadella who said AI development "cannot be controlled by a handful of entities".

But US President Donald Trump downplayed the risks, saying that he wanted to keep the US's lead over China "because whoever wins AI, wins".

The competition is fierce.

The former Anthropic employee Coxon told the BBC that there needed to be "some sort of coordinated slowdown with China, if we're going to avoid a race at an international scale".

As Trump prepares to host Xi at the White House next week, the two sides are trying to include AI on the agenda, the South China Morning Post reports.

It is also likely to be part of the talks between US Treasury Secretary Scott Bessent and Chinese Vice-Premier He Lifeng, who are trying to meet before next week's summit, the paper said. Vice-Premier He runs the agency that oversees the country's economic planning, including AI development.

While collaboration between the US and China is important, experts say the stakes are high for both countries.

"If China establishes a durable lead in frontier AI, this will not be viewed in Washington as losing a technology race," says Chasen Nevett, managing director at Hong Kong-based firm, Forani Investments.

"It will be viewed as losing part of the strategic infrastructure of the next industrial era."

Additional reporting by Osmond Chia

A human like robot figure with a line of robots behind him. The bottom half of the image is treated with a red glaze and the top is black and white

Claude Used to Automate Exploitation and Data Theft Across Multiple Victims

estimated reading time: 9 min

Anthropic has warned that cybercriminals and state-sponsored hackers alike are using its Claude models for cyber attacks, weapons design, propaganda, and mass surveillance between December 2025 and August 2026.

The threat actors, which the artificial intelligence (AI) company has branded Generative Threat Groups (GTGs), span state-sponsored groups, financially motivated criminals, commercial spyware vendors, state propaganda institutions, and politically motivated individuals.

"The cybersecurity skills of AI models means that AI has collapsed the labor and tooling gap that used to separate well-resourced, state-sponsored operations from individual operators," Anthropic said. "The use of AI went beyond simple questions and responses from a chatbot but rather involved the use of multi-agent frameworks executing reconnaissance, exploitation, and data exfiltration."

Among the notable cases highlighted by Anthropic is the development of an AI-assisted workflow by a Russian state-sponsored threat actor it calls GTG-20006, which aligns with broader reporting linking the cluster to Midnight Blizzard (aka APT29 and Cozy Bear). Some of the other AI-enabled cyber campaigns highlighted by Anthropic in its 154-page report include -

  • GTG-50014 (aka MeowSHA, frkoo, and blazespider), a French-speaking operator and a suspected affiliate of the ShinyHunters collective that ran a distributed credential-harvesting pipeline across a fleet of 10 AWS EC2 workers that mass-downloaded 1.8 million distinct Android APKs from multiple app-store sources, scanned them for hard-coded secrets using TruffleHog, and sent verified findings to a Telegram group.
  • Another ShinyHunters affiliate that specialized in supply chain theft by compromising software-as-a-service (SaaS) vendors to steal data belonging to downstream customers, accelerate reconnaissance, and enable data exfiltration.
  • GTG-10007, a Chinese-speaking operator likely based out of Hunan province, some of whom have been identified as undergraduate students at a Chinese university and have used Claude to conduct intrusion attempts against production systems, reconnaissance of foreign-government networks across the Middle East, Europe, and Southeast Asia, a vulnerability-research and exploit development effort against major endpoint-security products, and develop an intelligence-collection platform for bulk-harvesting of open-source material aligned with Beijing's priorities. The threat actor targeted about 50 organizations across education, retail, energy, technology, healthcare, finance, manufacturing, and government sectors globally. The group also maintained an autonomous vulnerability research program to produce working exploits for previously unknown vulnerabilities in network and security appliances.
  • GTG-50021, a Russian and Ukrainian-speaking group that ran a fraudulent AI reseller operation offering cheap Claude access, only for customers' traffic to be silently proxied to a different AI model, while the illicit scheme installed a credential harvester to siphon their Anthropic account credentials and sell them to other proxy resellers for malicious use.
  • GTG-50020, a Russian-speaking, financially-motivated actor that has historically targeted hotel booking and financial technology platforms but has since focused on the AI supply chain by stealing model provider API keys and unsuccessfully attempting to gain access to pre-release AI models. The threat actor is estimated to have targeted about 30 AI vendors in a four-day window using similar techniques.
  • GTG-50029, a single French-speaking actor that used Claude to target European political parties, media, think-tanks, and the SaaS providers used by these organizations, including by exploiting a previously undocumented WordPress re-installation race condition that made it possible to create a rogue administrator account without valid credentials, as well as by abusing an exposed search endpoint to breach a political campaign management platform and siphon sensitive data. The threat actor has also been observed deploying web shells and a browser exploitation C2 framework against other targets. Central to the attacker's operation was a purpose-built doxxing platform named "fafsearch" that offered the ability to cross-reference individual breach dumps against exfiltrated data.

"At one end, actors used Claude conversationally: it acted as an engineering assistant in the creation of malware, phishing kits, and surveillance tooling," Anthropic said. "Further along the spectrum, threat actors directed Claude to execute operations (such as running commands against victim networks, harvesting credentials, and exfiltrating data) with a human making each individual targeting decision (GTG-20006)."

"At the far end, operations ran autonomously, with minimal human input or supervision: these included multi-agent frameworks conducting reconnaissance, exploitation, and theft against multiple victims, in parallel, for hours or days at a time (GTG-50014, GTG-50020, GTG-50029)."

The AI company said it also identified and took down a number of influence operations in which Claude played the role of a "sub-editor or content creator" to churn out content and run them at a scale beyond what low-resourced actors could have accomplished on their own. However, Anthropic emphasized that none of these efforts amassed authentic engagement and that they were disrupted before they could even build an audience.

Some of the influence and surveillance campaign clusters flagged by Anthropic at a high level are below -

  • GTG-04001, a Russian-speaking actor in Bangui that engaged in a foreign information manipulation and interference operation in the Central African Republic to amplify pro-Russia, anti-France talking points.
  • GTG-54002, a commercial "influence-as-a-service" operation that used Claude to mass-produce and rewrite political content across about 70 fabricated news websites. The operation has been traced back to LKM Company, a France-based digital advertising agency.
  • GTG-84005, a single account that used Claude to run a commercial election manipulation platform primarily targeting users in Malaysia based on political and social factors, such as their race and religion, by posing as a defensive cyber intelligence and counter-disinformation tooling outlet. The activity has been found to share links with BBS Bilisim Teknolojileri, an Istanbul-based technology company.
  • GTG-24015, a set of four accounts that used Claude as an "editorial and news production desk" to distribute them via state media outlets like Sputnik Moldova, RIA Novosti, Sputnik en Español, Sputnik Africa, and RT's English-language newsroom.
  • GTG-34001, a set of three Iranian state-aligned accounts that used Claude to shape public opinion, turn official government intelligence bulletins into tailored content, and disseminate the content across social media platforms. 
  • GTG-54006, a sustained, automated disinformation network that used Claude to generate fabricated Bengali-language news in Bangladesh and promote the country's Awami League party. The activity has been linked to a single actor based in Gaibandha District in Bangladesh via a set of 29 Claude accounts that were rotated to bypass platform limits and detection.
  • GTG-84006, a distributed influence operation that targeted Iranian audiences across the world with an aim to impersonate real activists and engage in live political conversations. The activity has been linked to People's Mojahedin Organization of Iran (PMOI/MEK) and the National Council of Resistance of Iran (NCRI).
  • GTG-54004, an account used by a single actor to mass-produce Kenyan political content as part of what's suspected to be a domestic astroturfing campaign with a pro-administration bent.
  • GTG-84002, an account used by a single actor to run a sustained influence operation against the Muslim Brotherhood, the Sudan conflict, and the United Nations accountability mechanisms.
  • GTG-54009, a commercial surveillance platform that used Claude to analyze, classify, and profile the social media activity of users in Iran and the Persian Gulf region. The activity is assessed to have been carried out by, or on behalf of, an Israeli-Singaporean commercial intelligence vendor named S2T Unlocking Cyberspace.
  • GTG-14010, a China state-aligned operation that used Claude to track, profile, and recruit Uyghurs and Uyghur armed formations in Syria. The actor has been found to use the AI model to convert conversations extracted in bulk from over 100 monitored WhatsApp groups and dozens of Telegram channels into structured Chinese-language data and "creating profiles of individuals who might be vulnerable to targeting due to financial stress, family separation, and ideological disillusionment."
  • GTG-14020, a set of accounts likely linked to a Chinese government-aligned intelligence operation that used Claude to build Chinese-language dossiers targeting religious leaders and Chinese diaspora figures across Asia, as well as map religious venues and instruct the model to adopt "China's standpoint."
  • GTG-14021, a set of accounts from China-based actors that used Claude to support surveillance and transnational repression, including prompting the model to assume the role of an intelligence analyst serving China's national security apparatus.
  • GTG-14022, a China-based "public opinion monitoring" and dissident surveillance operation that used Claude to produce government briefings that listed dissidents, activists, ethnic minority and Chinese diaspora communities, and foreign media as threats to political stability while asking it to play the role of a "senior emergency public opinion analyst serving the government of the People's Republic of China."
  • GTG-34007, a set of 16 accounts operated by two Iranian-nexus actors associated with paramilitary and domestic security agencies that used Claude to build a frontend for what appears to be a government-controlled surveillance case-management system, run social-network analysis over 155,216 X posts, and build domestic surveillance capabilities via a malicious Mozilla Firefox extension named "al-Najm al-thāqib" to harvest user identities from major social network platforms.
  • GTG-50027, a single account that used Claude to design a national mass interception and surveillance platform called Lakana 360 for Mali's state intelligence service to monitor about 25 million SIM cards spanning three of the country's national mobile operators, and generate intelligence dossiers for any phone number. The platform has a separate layer that collects call records, text messages, and voice calls across the mobile networks.
  • GTG-30004, an Iran-nexus threat actor that used Claude to develop an automated, open-source intelligence identity-profiling service targeting Israeli and Jewish diaspora organizations.
  • GTG-30005, an Iran-nexus threat actor that used Claude to gather and analyze publicly accessible data to develop targeting recommendations against U.S. naval forces in the region and build software components of a domestic mass-surveillance platform that combined automatic license-plate recognition with mobile-device identifier interception.
  • GTG-30006, an Iranian threat actor that leveraged free Claude.ai accounts to develop malware, a delivery pipeline, and a phishing portal targeting domestic Iranians. This included a bogus ESET NOD32 antivirus login page that transmits captured credentials to Telegram, a ClickFix-style Windows Run dialog lure, and geofenced delivery pages. The threat actor has also used Claude to build SECOMS64, a modular Windows implant with keylogging, screenshot capture, and Chrome credential extraction capabilities.

Elsewhere, Anthropic said it neutralized Claude misuse efforts by threat actors based in northern Yemen to develop guided weapons, two China-based operations to draft a Chinese-language specification for an anti-torpedo fire control system and build targeting software for electronic warfare, and a Russia-based operation to engineer a full-stack autonomous first-person-view (FPV) kamikaze drone swarm.

"As AI models become more widely used, providers will continue to acquire threat-relevant visibility into real-world use that even governments and intergovernmental organizations lack," the company said. "We hope that sharing these early insights with the public helps inform governments, the industry, and the general public on the nature of these risks, and the safeguards that are necessary for ensuring the safe deployment of AI models."

Found this article interesting? Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post.

Anthropic Is Building a Predictive Surveillance System to Monitor Activists - The American Prospect

estimated reading time: 7 min

Job postings and interviews with senior security officials at Anthropic show that the frontier AI lab is building out an extensive monitoring system to keep tabs on activists who oppose the rapid development of artificial intelligence.

In addition to monitoring activists in the vicinity of Anthropic executives and keeping tabs on protests near physical Anthropic assets, the firm is also implementing a “pre-crime” approach, attempting to predict incidents before they happen. In some cases, that also means reporting suspects to police before a crime occurs. Anthropic did not respond to the Prospect’s request for comment.

Anthropic’s plans to surveil dissent are at odds with the firm’s efforts to cast itself as the responsible alternative to OpenAI. At the beginning of the year, the Department of Defense and Anthropic engaged in a high-profile dustup over Anthropic’s refusal to allow the military to use its tools for mass domestic surveillance and autonomous weapons. That tension seems to have eased as Anthropic hires for “national security sales” positions, seeking to restart military contracts. The increase in threat monitoring of domestic opponents fits with a renewed focus on national security.

More from Daniel Boguslaw

A piece of Anthropic’s surveillance architecture was revealed in a podcast interview from last year between Anthropic Global Security Operations Center Manager Keon Ellison, Security Operations Manager Zach Melvin, and James Neufeld, CEO of Samdesk, a company Anthropic contracts with for risk detection.

In a wide-ranging conversation about threat monitoring and analysis, the interview also touches on monitoring activists. “Last year we had an executive travel into a major city when we received some intelligence through Samdesk about a planned protest,” Ellison said. The originally scheduled protest was moved up due to permitting issues. “Samdesk gave us about 60 minutes of advanced notice that the protest organizers had moved the timeline,” Ellison explained. “That extra hour was critical. Without it our executives would have departed their meetings, they would have ran right into the heart of the disruption.”

Ellison said that Anthropic used this data to devise an alternate route for the executive and funnel them to a service entrance at the hotel. “What could have been a high-stress situation,” he said, “was really mitigated through early detection through Samdesk and giving us that information.”

Anthropic has begun making routine reports to police departments across the country for threats, and told The Wall Street Journal in July, “We track concerning behavior over time through a person-of-interest process, allowing us to catch escalation patterns early.” According to the Journal, “several individuals involved in incidents reported to police were already being tracked by Anthropic security.”

Last month, The San Francisco Standard reported that Anthropic had reported a man to San Francisco police for telling Claude that he had bought an AR-15 semiautomatic rifle and had CEO Dario Amodei “in his sights.” When the Standard contacted the man in question, he told the newspaper he was “just fucking around.”

But while Anthropic was fast to call the cops on a frustrated Claude user, the Standard also reported a key detail: Anthropic refused to show police the actual messages, citing Anthropic’s internal policy. In short, Anthropic reported a user for in-platform speech, and then refused to provide police with evidence of actual wrongdoing.

This kind of pre-crime policing, encouraged without due process, is referenced as an explicit goal by Anthropic’s security program manager in the podcast reviewed by the Prospect. “The goal would be transforming operations from reactive information to gathering proactive and predictive and preventative threat engagement and management,” he said, adding, “That’s the kind of operational maturity that makes sense for protecting high-value targets in any industry.”

Anthropic’s effort to build a predictive security apparatus extends beyond the C-suite to its Global Safety, Intelligence, and Security (GSIS) team, according to a job posting from last month detailing Anthropic’s search for an enterprise intelligence specialist who “will investigate specific threats, actors, and events, produce finished assessments, and help keep Anthropic’s employees ahead of a rapidly evolving threat landscape and in a defensible position.”

Part of that role, compensated at between $180,000 and $230,000, will be to “identify, assess, track, and investigate global threats including geopolitical instability, terrorism, crime, activism, nation-state targeting of the AI sector, and emerging security trends, including deep-dive research and OSINT collection on specific threats, actors, and events” (emphasis added).

The expansion of Anthropic’s intelligence-gathering to a national and even global scale tracks with recent efforts to heighten the labeling of AI and the infrastructure powering it, including data centers and power supply. A critical infrastructure designation would put artificial intelligence on the same footing as water, electricity, and broadband. And indeed, AI’s boosters like Americans for Responsible Innovation (ARI) have urged the Trump administration to make the change. Per ARI’s telling, AI is “so vital to the United States that the incapacity or destruction of such systems and assets would have a debilitating impact on security, national economic security, national public health or safety, or any combination of those matters.”

Enshrining frontier labs in the hardened cloak of national security would not only give Anthropic, OpenAI, and Google even more access to intelligence products generated by federal law enforcement and intelligence agencies; it would also embolden these companies to shape how federal agencies view threats to their bottom line, now transformed as “critical infrastructure.”

It’s not hard to imagine how civic engagement by the same bipartisan coalition opposing data centers could turn its focus onto AI, only to be branded in the same instant as extremists or, even worse, terrorists. Anthropic’s answer to this problem has been to work even harder at producing artificial intelligence that can deliver services that can’t be brushed aside by the public.

“I do agree that the public has a negative view of AI (and that this is a big problem), but I don’t think it is primarily caused by me or any other AI leader warning about AI’s risks,” Anthropic CEO Dario Amodei wrote on Twitter last month. “I think it is fundamentally a crisis of trust … I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.”

As Anthropic ramps up its efforts to monitor dissent with in-house intelligence teams and real-time protest tracking powered by AI, it will have to contend with the increasing economic desperation that plagues human beings outside its Bay Area towers. Last week, the security guards who patrol the campuses of OpenAI and Anthropic announced that they had authorized a strike over stalled pay negotiations. In response, Anthropic sent a company-wide email telling employees that it was best if they worked from home.

“Who can survive with $22 an hour in San Francisco?” David Huerta, president of SEIU-USWW, the union representing security guards in the Bay Area, said at a rally last week. Around him, security guards chanted: “Shame.”

Related

❌