In mid-November 2025, Anthropic put out a blog post, explaining how – because it’s the ‘good guy’ in this wild west movie that’s clearly running in its own head – it had used Claude, its AI engine, to disrupt the first (known) AI-orchestrated cyber-attack. At the end of the movie (or blog post), we – the audience – are expected to stand up and applaud Anthropic for stopping the ‘bad guys’.
But the problem isn’t that bad people can use Claude for bad ends. The problem is that these vulnerabilities are created as a direct consequence of the architecture required to make Claude commercially viable. And Google, OpenAI, Meta, xAI – every company chasing the same prize – face the same trade-off, whether they’ve been caught out yet or not. These systems are designed to be responsive, generative and maximally useful. The exploitability isn’t a bug; it’s what makes them valuable. Their ‘guardrails’ exist mostly to signal a commitment to safety, but they can’t fundamentally alter this architectural reality without killing the commercial goose.
This isn’t a bug, it’s the business
Claude can detect AI-assisted probing because detection and exploitation rely on the same underlying mechanisms. The company frames this as a victory: ‘See, we can catch threats!’ But that framing obscures what makes both possible.
To understand why, we need to consider what researchers call ‘model symmetry’. An AI system’s commercial value comes from its ability to be useful: it responds to prompts, generates code, identifies patterns, automates tasks. But here’s the problem – every capability that makes the system useful is symmetrically available for misuse. If Claude can write code to solve a programming problem, it can write code to exploit a vulnerability. If it can analyse patterns to help with research, it can analyse patterns to map attack surfaces. If it can respond to complex instructions and work within safeguards, it can respond to instructions designed to bypass these safeguards.
This isn’t a flaw in implementation. It’s fundamental to how these systems work. A model that cannot be pushed to its edges, interrogated for patterns, or teased into revealing its tendencies would be commercially worthless. These same pathways that allow legitimate use also enable exploitation; the same architecture that allows a security team to detect threats also allows attackers to probe for weaknesses.
Attacking the system with its own tools
Anthropic’s own report confirms this dynamic. Paragraph four of the executive summary describes how the attack used Claude Code – Anthropic’s own commercial product – including “instances of Claude Code operating in groups as autonomous penetration testing orchestrators and agents, with the threat actor able to leverage AI to execute 80–90% of tactical operations independently at physically impossible request rates.” The attackers used exactly the autonomous capabilities that make Claude Code valuable to customers, just applied toward different ends. And they did so in ways that Anthropic’s own threat intelligence documentation had already identified as plausible attack vectors.
The company continues to monetise these capabilities while addressing its security implications retrospectively – not because anyone at Anthropic wants systems to be exploitable, but because removing the exploitability would eliminate the product’s commercial value. Without monetisation, there is no company. Treating post-hoc monitoring as equivalent to prevention requires ignoring this architectural reality.
Put more simply, the attackers used the tools that Anthropic had developed, and in ways that they had already predicted it could be used. They are simply customers using the product in ways that are inimical to the rest of us.
The state is awake at the wheel
The state doesn’t save us here. Yes, jurisdictions differ – the EU’s AI Act attempts regulatory boundaries that the US explicitly rejects, while China pursues its own model of state-directed development. But these differences mask a deeper dynamic: a multipolar trap where each government acts rationally to avoid falling behind in the AI race, but collectively they produce the unsafe deployment at scale that none of them explicitly wants.
The US encourages speed through permissive regulation and subsidies, fearing Chinese dominance. The EU builds frameworks that companies learn to navigate while maintaining competitive pace. China invests massively, viewing AI leadership as strategic necessity. Each jurisdiction’s rational self-interest – don’t cede technological advantage – drives the collective acceleration. Safety, caution, ethics become the words used in press releases while the underlying incentive structure rewards whoever deploys fastest.
The result is that speed and scale are rewarded across all jurisdictions, oversight remains weak relative to the pace of development, and those in power benefit from the narrative that these systems are manageable risks rather than architectural vulnerabilities. The trap isn’t that everyone behaves identically. It’s that everyone is trapped by the same competitive logic, regardless of their stated regulatory philosophy.
And the threat isn’t hypothetical. A carefully prompted Claude, Grok, Gemini or ChatGPT can be leveraged to find exploits (vulnerabilities in the system which can be used to attack/weaponise/exploit that system), reconstruct sensitive material, or map the decision surfaces of other AI systems. The companies know this – their own documentation acknowledges these attack vectors. Their public statements read like worried adults in a playground, while the reality is that they handed out the knives, observed the mayhem, and now report on who got scratched, injured or killed in their neat little Accident Book.
Predictable failures, and carefully worded apologies
Even Anthropic’s own post about disrupting AI-enabled espionage makes this clear. It describes detecting suspicious probing activity and neutralising it. That’s meant to look responsible. But it also reveals it understood the system could be used this way – and it’s clear that it can only profit from a model with exactly those affordances. Its ‘success’ is not prevention. It is a post hoc report of a problem it created.
And let’s not pretend this is only one company. Grok’s catastrophic failures are instructive: its AI cheerfully produces sexualised images of women and minors, despite ‘safeguards’. Reuters, CBS and the Washington Post [subscriptions may be required] reported that the company admitted the safeguards were inadequate and that content moderation had failed. In the UK, The Guardian newspaper reported it and the Information Commissioner’s Office (ICO) intervened. These weren’t fringe errors – they were foreseeable consequences of systems designed to obey any prompt within bounds that were never truly defined or enforceable.
OpenAI and Alphabet/Google aren’t exempt either. The UK’s National Cyber Security Centre has explicitly warned that prompt injection attacks may never be fully mitigated. Large language models (LLMs) treat data and instructions identically, and any agentic interface increases the attack surface. OpenAI itself admits that its ChatGPT ‘agent mode’ is inherently vulnerable to this form of exploitation. Various AI models have shown similar weaknesses in public testing. There are other, smaller companies in this field too. I may not have listed them, but that doesn’t exempt them from these same flaws, so there’s no pat on the back for them just because they haven’t been name-checked here.
Arsonists shouldn’t get credit for raising the alarm
These aren’t abstract technicalities. Prompt injection is exactly what allows malicious actors to bypass supposedly hardened boundaries, extract sensitive information and subvert systems intended to operate safely. Guardrails fail in predictable ways, but vendors report them in vague, carefully-worded language that sounds reassuring. The failures are structural, repeated and systemic.
Every ‘incident report’ framed as responsible monitoring is a PR exercise. It is a way of saying, “we saw the problem, but we’re the good guys in charge of noticing it”. Except they created the potential. And worse, the problem grows more dangerous as these systems become more capable. The same interaction that produces revenue produces exploitable knowledge, and the line between curiosity and compromise is invisible to the system itself.
None of these companies is an innocent bystander. They’re active participants in a system where structural vulnerabilities are the predictable outcome of commercial priorities, releasing capabilities into the wild before the security implications are fully understood or adequately addressed. Understandable from a venture capitalist, hyper-commercial standpoint, but not necessarily sensible, ethical or moral.
Still expanding, still not profitable
This is why our anger is necessary. Calmly noting recursion, inference and epistemic tension is not enough. We all need to understand that this is a product of choice, not accident. That the companies building these systems prioritise expansion, engagement and profit over safety. That the state is actively incentivising speed over control. And that the result is a system that cannot be safely deployed until those incentives change, and those structural decisions are held to account.
The financial pressure is real and escalating. SoftBank recently scrambled to complete a $40 billion investment in OpenAI, selling its entire stake in Nvidia and significant portions of T-Mobile holdings to raise the capital. OpenAI faces projected losses of $14 billion in 2026, driving the company toward aggressive monetisation strategies that directly conflict with safety priorities. When a company is burning through billions while racing toward profitability, the commercial imperative to deploy and scale systematically overrides even genuine safety commitments.
Commitment to safety…
This pattern plays out across the industry. The same week that Anthropic lost its head of Safeguards Research, Mrinank Sharma, who resigned citing the persistent difficulty of letting “our values govern our actions”, there were multiple interviews with its in-house philosopher discussing Claude’s ethical framework. Perhaps this is a thing they’re promoting? Her work may be genuine, but its public deployment serves a specific function: demonstrating commitment to safety while the underlying commercial priorities remain unchanged.
Sharma’s resignation is particularly instructive. His research focused on AI sycophancy – training systems to be less obsequiously agreeable to users. Yet he left because, in his assessment, the organisation itself exhibited sycophancy to commercial pressure, repeatedly setting aside safety priorities under competitive and financial strain. To date, Anthropic seems to have published no official public statement on its website in response to Sharma’s resignation, but The Hill reported an Anthropic spokesperson’s comment that “it was grateful for Sharma’s work advancing AI safety research and noted that his manager, Ethan Perez, had thanked Sharma [publicly on X] in response to his open letter. The spokesperson also added that Sharma, like all current or former employees, is free to speak openly about safety concerns”.
…meets commercial imperative
The pattern isn’t unique to Anthropic. Jan Leike left OpenAI’s now-dissolved Superalignment team after disagreeing with leadership about the company’s core priorities. Zoë Hitzig resigned from OpenAI citing deep reservations about emerging advertising strategies that would monetise the intimate data users share with ChatGPT. No dedicated response page appears to exist on OpenAI’s website, and enquiries seem to have been directed to Sam Altman’s post on X saying he was sad to see Leike leave and acknowledged the company had more work to do. The following day, Greg Brockman posted a joint statement from himself and Altman on X, making three points: that OpenAI had raised awareness of AGI risks so the world could better prepare; that they were building foundations for safe deployment; and that they would continue researching and working with governments and stakeholders on safety.
The philosopher tasked with teaching Claude ‘how to be good’ describes her work as comparable to raising a child. The metaphor is revealing and, frankly, weird. No responsible parent should monetise their child’s interactions with millions – or even billions – of strangers while their child is still learning basic judgement. Yet that’s exactly what these companies do; not because the philosophical work is insincere, but because the commercial imperative to deploy and scale supersedes the developmental timeline that genuine safety would require.
AI is both threat and monitor
AI is both a threat and a monitor. But don’t be fooled. The threat is not entirely external; it is baked into the architecture, into the business model, and into the regulatory environment that allows it to exist. The monitoring – their ‘guard dog’ – is a side effect, a PR veneer and a big puppy-eyed softy. Treating it as a solution is a lie that benefits no one except the companies telling it.
Our governments keep letting it happen. And every press release, every blog post, every “we mitigated a probing attempt” story is a reminder that the problem may have been accidental, but the structure that allows it was not. The oversight, the weak regulatory frameworks, the half-hearted ethical reviews – these are all features of a system designed to maximise exposure and monetisation, while giving the public the impression that the playground is safe.
And the more capable AI becomes, the louder the claims of responsibility, the more dangerous the risks to all of us.
Move fast, break everything
Since I first started writing this article, things in the AI world have been moving fast. Anthropic’s Dario Amodei stood up to the US government – sort of; Sam Altman kind of said he would, and almost immediately didn’t. There’s no room to discuss those things here, they’re the basis of at least one other article in and of itself, but please don’t think that these actions take away from any of the above arguments; they really don’t. These companies are not on our side, not today, not tomorrow. They are on the side of making money, with seemingly little care for the consequences of their actions.
Yorkshire Bylines approached Anthropic for comment but no reply was received by the time of publication.

Friends of Bylines Network
There has never been a greater need for grassroots journalism that investigates the stories that really matter, holds power to account and champions the voices of everyday citizens. We are proudly powered by volunteers but what we do isn’t free.
STAND WITH US for independent, citizen-led journalism that makes democracy stronger, and you will even get some exclusive benefits.







