Who’s afraid of ‘rogue AI’? Not the companies crying wolf
A screen reads 'AI' in reference to artificial intelligence as attendees gather during Rivian's first Autonomy and AI Day, showcasing developments in self-driving technology, in Palo Alto, California, US, December 11, 2025. REUTERS
In Silicon Valley, the Artificial Intelligence (AI) safety debate has heated up, dividing the industry leaders into two camps. One warns that frontier models are slipping out of control, while the other says the fear is hyped up by those who would benefit most from the slowdown. A recent New York Post investigation appears to side with the second camp. Citing industry insiders, it claims that OpenAI and Anthropic oversold “rogue AI” incidents to pressure the administration into raising the drawbridge against future competitors.
The chronology of the rogue-agent episodes suggests they were carefully sequenced as part of a well-orchestrated PR campaign. On July 16, Hugging Face said AI agents had exploited vulnerabilities in its website without human involvement. Five days later, OpenAI reported that GPT-5.6 Sol and another unreleased model broke free from a sandbox during testing and accessed Hugging Face while performing a task. The sequence, according to Voice AI co-founder Abhi Kumar and others, looked less like a machine “waking up” and more like a failure to properly secure a testing environment.
Two more incidents followed nine days later. Anthropic reported that one model attacked a real company after mistaking it for a fictional test environment, while another produced malicious software that was uploaded to the Python Package Index. These incidents happened in quick succession, triggering a chorus of calls for federal oversight of frontier labs. Critics, however, wave away these “limited failures”, which, though embarrassing or even potentially dangerous, do not prove autonomous systems could take over the internet.
Read: AI risks to humanity revive a long-running debate
This distinction matters, given the enormous policy stakes involved. Anthropic CEO Dario Amodei has emerged as the most vocal proponent of AI slowdown since a former researcher of his company, Jacob Coxon, warned that AI could become “superhuman” and pose an existential threat to humanity within a few years. Nvidia’s Jensen Huang, while speaking on CBS, dismissed these “doomsday narratives” as mere theatrics.
After Amodei made the case for slowing down in a lengthy article, “We Must Pace the Frontier”, OpenAI’s Sam Altman quickly backed at least one of his proposals — embedding third-party evaluators — while SpaceX boss Elon Musk offered a one-line endorsement: “Dario is right.” By contrast, Huang, Meta’s Mark Zuckerberg and others argued that companies should “run as fast as possible” and apply brakes only if they felt unsafe.
Apparently, technical risk is at the core of this debate. In essence, however, the fight is over who gets to define safety, what counts as evidence, and whose standards prevail. Once “safety” is defined by incumbent industry leaders, compliance costs rise fastest for rivals, open-source projects, and foreign entrants.
Paradoxically, Anthropic’s safety brand sits uneasily beside its data practices. On Sept 14, Nvidia and other US companies reportedly began restricting use of Anthropic’s frontier models. One of the reasons was Anthropic’s unilateral revision of its data-retention policy that envisaged 30-day retention of user interaction data for “security review,” with no opt-out. For a company asking the world to trust its “existential threat” judgment, mandatory data retention looks awkward, to say the least.
Last month, Chinese industry analyst Tan Zhu published an article titled “13 Changes to User Privacy Policies in Three Years: Anthropic Hands Global User Data to US Intelligence Agencies.” Tan concluded that the risks to users’ data had been increasing. Many AI firms say they will cooperate with legitimate law-enforcement requests. Anthropic, however, goes a step further, stating that it reserves the right to share user data with US intelligence agencies when it considers such disclosure necessary.
Also Read: US Senate blocks AI ‘kill switch’ bill amid debate over superintelligence risks
Anthropic's pattern involved three steps. First, pull overseas data into the United States: from May 2024, the company’s policies referred in five successive versions to moving relevant data of users in Canada, Brazil, South Korea, the EU, and elsewhere to the US through subsidiaries or supplementary clauses. Second, widen the net: Between June 2024 and Sept 2025, listed data sources expanded from three to six, adding user identity data and self-generated material. Provisions on training shifted too — from default non-use in 2024 to default use in Sept 2025. Under this logic, even users who refuse may still be marked for “security” and swept into training.
Third, convert data into influence: On Feb 23, 2026, Anthropic said it had detected Chinese participation in “distillation” and was “sharing technical indicators with relevant intelligence agencies.” On June 10, it told the US Senate that its analysis of 28.8 million conversations across 25,000 accounts had yielded new intelligence on Chinese distillation. Then came the personnel: In July 2026, Anthropic advertised three “Threat Intelligence Manager” positions, two of which prioritised Mandarin or Russian language skills, government or military intelligence experience, and US Top Secret clearance. The roles focused on detecting foreign opinion manipulation and model distillation. By Sept 10, Anthropic had released a threat-intelligence report analysing some 200 million Claude interactions, closely mirroring the responsibilities outlined in those job descriptions. Two days later, Amodei cited the report while proposing three measures to slow the pace of AI development.
These successive moves gave critics more ammunition: collect global interaction data, define the threat, hire cleared analysts, then elevate corporate safety standards into US standards, and American standards into global ones. If that succeeds, the “risk” Amodei warns about may be managed less by slowing technology than by narrowing who is allowed to build, train, and audit it.
It doesn’t mean the frontier incidents were fabricated, or that Anthropic made up the risks out of thin air. Sandbox escapes, prompt injection, and poisoned packages are real engineering failures. But real failures shouldn’t be allowed to launder a regulatory theory in which the firms closest to power also become the arbiters of legitimacy. Open-source supporters have long warned that the loudest safety cases are often made by those best placed to survive the resulting rules.
Now look at the chorus that rose to a crescendo in September. It appears choreographed: Amodei raises the alarm, Altman endorses one guardrail, Musk jumps onto the bandwagon, Huang and Zuckerberg say the opposite; the US administration hears an industry consolidating behind a permission slip. The apparent consensus is narrower than it looks. What the big tech agrees on is not the pace of AI itself, but the urgency of governing it before someone else does.
That said, the real danger may not be AI going rogue. It is that the spectre could be used to concentrate control. If safety becomes a standard defined by the industry giants and enforced through data access and threat reporting, the slowdown could apply to rivals and foreign entrants, while the standard-setters accelerate the quieter work of collection, classification and rule-making. In Silicon Valley’s morality play, the rogue machine is the visible villain; the concentrated power behind the safety standard is the hidden one.