• Contact
  • About
  • ISSN 3049-9720
  • Authors and editors
NEWSLETTER SIGN UP
Yorkshire Bylines
Advertisement
  • Home
  • News
    • Brexit
    • Culture
    • Education
    • Environment
    • Home Affairs
    • Transport
    • World
  • Politics
  • Opinion
  • Society
    • Food
    • Music
    • Poetry
    • Sport
  • Business
    • Economy
    • Science and Technology
    • Trade
  • Region
No Result
View All Result
  • Home
  • News
    • Brexit
    • Culture
    • Education
    • Environment
    • Home Affairs
    • Transport
    • World
  • Politics
  • Opinion
  • Society
    • Food
    • Music
    • Poetry
    • Sport
  • Business
    • Economy
    • Science and Technology
    • Trade
  • Region
No Result
View All Result
Yorkshire Bylines
Home Opinion

Complicit by design

AI hasn’t suddenly become dangerous. Built to be exploitable, profitable and someone else’s problem, vulnerability is a feature, not a bug

Philip O'Brien by Philip O'Brien
25-04-2026 05:05 - Updated on 14-05-2026 16:41
in Opinion, Science and Technology, World
Reading Time: 14 mins read
A A
graphic image of two computerised heads facing opposite directions

Image by Gerd Altmann from Pixabay

Share on Bluesky

In mid-November 2025, Anthropic put out a blog post, explaining how – because it’s the ‘good guy’ in this wild west movie that’s clearly running in its own head – it had used Claude, its AI engine, to disrupt the first (known) AI-orchestrated cyber-attack. At the end of the movie (or blog post), we – the audience – are expected to stand up and applaud Anthropic for stopping the ‘bad guys’.

But the problem isn’t that bad people can use Claude for bad ends. The problem is that these vulnerabilities are created as a direct consequence of the architecture required to make Claude commercially viable. And Google, OpenAI, Meta, xAI – every company chasing the same prize – face the same trade-off, whether they’ve been caught out yet or not. These systems are designed to be responsive, generative and maximally useful. The exploitability isn’t a bug; it’s what makes them valuable. Their ‘guardrails’ exist mostly to signal a commitment to safety, but they can’t fundamentally alter this architectural reality without killing the commercial goose.

This isn’t a bug, it’s the business

Claude can detect AI-assisted probing because detection and exploitation rely on the same underlying mechanisms. The company frames this as a victory: ‘See, we can catch threats!’ But that framing obscures what makes both possible.

To understand why, we need to consider what researchers call ‘model symmetry’. An AI system’s commercial value comes from its ability to be useful: it responds to prompts, generates code, identifies patterns, automates tasks. But here’s the problem – every capability that makes the system useful is symmetrically available for misuse. If Claude can write code to solve a programming problem, it can write code to exploit a vulnerability. If it can analyse patterns to help with research, it can analyse patterns to map attack surfaces. If it can respond to complex instructions and work within safeguards, it can respond to instructions designed to bypass these safeguards.

This isn’t a flaw in implementation. It’s fundamental to how these systems work. A model that cannot be pushed to its edges, interrogated for patterns, or teased into revealing its tendencies would be commercially worthless. These same pathways that allow legitimate use also enable exploitation; the same architecture that allows a security team to detect threats also allows attackers to probe for weaknesses.

Attacking the system with its own tools

Anthropic’s own report confirms this dynamic. Paragraph four of the executive summary describes how the attack used Claude Code – Anthropic’s own commercial product – including “instances of Claude Code operating in groups as autonomous penetration testing orchestrators and agents, with the threat actor able to leverage AI to execute 80–90% of tactical operations independently at physically impossible request rates.” The attackers used exactly the autonomous capabilities that make Claude Code valuable to customers, just applied toward different ends. And they did so in ways that Anthropic’s own threat intelligence documentation had already identified as plausible attack vectors.

The company continues to monetise these capabilities while addressing its security implications retrospectively – not because anyone at Anthropic wants systems to be exploitable, but because removing the exploitability would eliminate the product’s commercial value. Without monetisation, there is no company. Treating post-hoc monitoring as equivalent to prevention requires ignoring this architectural reality.

Put more simply, the attackers used the tools that Anthropic had developed, and in ways that they had already predicted it could be used. They are simply customers using the product in ways that are inimical to the rest of us.

The state is awake at the wheel

The state doesn’t save us here. Yes, jurisdictions differ – the EU’s AI Act attempts regulatory boundaries that the US explicitly rejects, while China pursues its own model of state-directed development. But these differences mask a deeper dynamic: a multipolar trap where each government acts rationally to avoid falling behind in the AI race, but collectively they produce the unsafe deployment at scale that none of them explicitly wants.

The US encourages speed through permissive regulation and subsidies, fearing Chinese dominance. The EU builds frameworks that companies learn to navigate while maintaining competitive pace. China invests massively, viewing AI leadership as strategic necessity. Each jurisdiction’s rational self-interest – don’t cede technological advantage – drives the collective acceleration. Safety, caution, ethics become the words used in press releases while the underlying incentive structure rewards whoever deploys fastest.

The result is that speed and scale are rewarded across all jurisdictions, oversight remains weak relative to the pace of development, and those in power benefit from the narrative that these systems are manageable risks rather than architectural vulnerabilities. The trap isn’t that everyone behaves identically. It’s that everyone is trapped by the same competitive logic, regardless of their stated regulatory philosophy.

And the threat isn’t hypothetical. A carefully prompted Claude, Grok, Gemini or ChatGPT can be leveraged to find exploits (vulnerabilities in the system which can be used to attack/weaponise/exploit that system), reconstruct sensitive material, or map the decision surfaces of other AI systems. The companies know this – their own documentation acknowledges these attack vectors. Their public statements read like worried adults in a playground, while the reality is that they handed out the knives, observed the mayhem, and now report on who got scratched, injured or killed in their neat little Accident Book.

Predictable failures, and carefully worded apologies

Even Anthropic’s own post about disrupting AI-enabled espionage makes this clear. It describes detecting suspicious probing activity and neutralising it. That’s meant to look responsible. But it also reveals it understood the system could be used this way – and it’s clear that it can only profit from a model with exactly those affordances. Its ‘success’ is not prevention. It is a post hoc report of a problem it created.

And let’s not pretend this is only one company. Grok’s catastrophic failures are instructive: its AI cheerfully produces sexualised images of women and minors, despite ‘safeguards’. Reuters, CBS and the Washington Post [subscriptions may be required] reported that the company admitted the safeguards were inadequate and that content moderation had failed. In the UK, The Guardian newspaper reported it and the Information Commissioner’s Office (ICO) intervened. These weren’t fringe errors – they were foreseeable consequences of systems designed to obey any prompt within bounds that were never truly defined or enforceable.

OpenAI and Alphabet/Google aren’t exempt either. The UK’s National Cyber Security Centre has explicitly warned that prompt injection attacks may never be fully mitigated. Large language models (LLMs) treat data and instructions identically, and any agentic interface increases the attack surface. OpenAI itself admits that its ChatGPT ‘agent mode’ is inherently vulnerable to this form of exploitation. Various AI models have shown similar weaknesses in public testing. There are other, smaller companies in this field too. I may not have listed them, but that doesn’t exempt them from these same flaws, so there’s no pat on the back for them just because they haven’t been name-checked here.

Arsonists shouldn’t get credit for raising the alarm

These aren’t abstract technicalities. Prompt injection is exactly what allows malicious actors to bypass supposedly hardened boundaries, extract sensitive information and subvert systems intended to operate safely. Guardrails fail in predictable ways, but vendors report them in vague, carefully-worded language that sounds reassuring. The failures are structural, repeated and systemic.

Every ‘incident report’ framed as responsible monitoring is a PR exercise. It is a way of saying, “we saw the problem, but we’re the good guys in charge of noticing it”. Except they created the potential. And worse, the problem grows more dangerous as these systems become more capable. The same interaction that produces revenue produces exploitable knowledge, and the line between curiosity and compromise is invisible to the system itself.

None of these companies is an innocent bystander. They’re active participants in a system where structural vulnerabilities are the predictable outcome of commercial priorities, releasing capabilities into the wild before the security implications are fully understood or adequately addressed. Understandable from a venture capitalist, hyper-commercial standpoint, but not necessarily sensible, ethical or moral.

A map of the UK with the UK flag superimposed next to a map of the UK with the US flag superimposed
Home Affairs

Government’s last chance to keep control of digital

by Philip O'Brien
14 November 2025 - Updated on 25 November 2025

Still expanding, still not profitable

This is why our anger is necessary. Calmly noting recursion, inference and epistemic tension is not enough. We all need to understand that this is a product of choice, not accident. That the companies building these systems prioritise expansion, engagement and profit over safety. That the state is actively incentivising speed over control. And that the result is a system that cannot be safely deployed until those incentives change, and those structural decisions are held to account.

The financial pressure is real and escalating. SoftBank recently scrambled to complete a $40 billion investment in OpenAI, selling its entire stake in Nvidia and significant portions of T-Mobile holdings to raise the capital. OpenAI faces projected losses of $14 billion in 2026, driving the company toward aggressive monetisation strategies that directly conflict with safety priorities. When a company is burning through billions while racing toward profitability, the commercial imperative to deploy and scale systematically overrides even genuine safety commitments.

Commitment to safety…

This pattern plays out across the industry. The same week that Anthropic lost its head of Safeguards Research, Mrinank Sharma, who resigned citing the persistent difficulty of letting “our values govern our actions”, there were multiple interviews with its in-house philosopher discussing Claude’s ethical framework. Perhaps this is a thing they’re promoting? Her work may be genuine, but its public deployment serves a specific function: demonstrating commitment to safety while the underlying commercial priorities remain unchanged.

Sharma’s resignation is particularly instructive. His research focused on AI sycophancy – training systems to be less obsequiously agreeable to users. Yet he left because, in his assessment, the organisation itself exhibited sycophancy to commercial pressure, repeatedly setting aside safety priorities under competitive and financial strain. To date, Anthropic seems to have published no official public statement on its website in response to Sharma’s resignation, but The Hill reported an Anthropic spokesperson’s comment that “it was grateful for Sharma’s work advancing AI safety research and noted that his manager, Ethan Perez, had thanked Sharma [publicly on X] in response to his open letter. The spokesperson also added that Sharma, like all current or former employees, is free to speak openly about safety concerns”.

…meets commercial imperative

The pattern isn’t unique to Anthropic. Jan Leike left OpenAI’s now-dissolved Superalignment team after disagreeing with leadership about the company’s core priorities. Zoë Hitzig resigned from OpenAI citing deep reservations about emerging advertising strategies that would monetise the intimate data users share with ChatGPT. No dedicated response page appears to exist on OpenAI’s website, and enquiries seem to have been directed to Sam Altman’s post on X saying he was sad to see Leike leave and acknowledged the company had more work to do. The following day, Greg Brockman posted a joint statement from himself and Altman on X, making three points: that OpenAI had raised awareness of AGI risks so the world could better prepare; that they were building foundations for safe deployment; and that they would continue researching and working with governments and stakeholders on safety.

The philosopher tasked with teaching Claude ‘how to be good’ describes her work as comparable to raising a child. The metaphor is revealing and, frankly, weird. No responsible parent should monetise their child’s interactions with millions – or even billions – of strangers while their child is still learning basic judgement. Yet that’s exactly what these companies do; not because the philosophical work is insincere, but because the commercial imperative to deploy and scale supersedes the developmental timeline that genuine safety would require.

AI is both threat and monitor

AI is both a threat and a monitor. But don’t be fooled. The threat is not entirely external; it is baked into the architecture, into the business model, and into the regulatory environment that allows it to exist. The monitoring – their ‘guard dog’ – is a side effect, a PR veneer and a big puppy-eyed softy. Treating it as a solution is a lie that benefits no one except the companies telling it.

Our governments keep letting it happen. And every press release, every blog post, every “we mitigated a probing attempt” story is a reminder that the problem may have been accidental, but the structure that allows it was not. The oversight, the weak regulatory frameworks, the half-hearted ethical reviews – these are all features of a system designed to maximise exposure and monetisation, while giving the public the impression that the playground is safe.

And the more capable AI becomes, the louder the claims of responsibility, the more dangerous the risks to all of us.

Move fast, break everything

Since I first started writing this article, things in the AI world have been moving fast. Anthropic’s Dario Amodei stood up to the US government – sort of; Sam Altman kind of said he would, and almost immediately didn’t. There’s no room to discuss those things here, they’re the basis of at least one other article in and of itself, but please don’t think that these actions take away from any of the above arguments; they really don’t. These companies are not on our side, not today, not tomorrow. They are on the side of making money, with seemingly little care for the consequences of their actions.

Yorkshire Bylines approached Anthropic for comment but no reply was received by the time of publication.

Friends of Bylines Network

There has never been a greater need for grassroots journalism that investigates the stories that really matter, holds power to account and champions the voices of everyday citizens. We are proudly powered by volunteers but what we do isn’t free.

STAND WITH US for independent, citizen-led journalism that makes democracy stronger, and you will even get some exclusive benefits.

BECOME A FRIEND
Tags: Artificial intelligence

Sign up for the Yorkshire Bylines newsletter

* indicates required

Consent for having Bylines Network store my submitted information

You can unsubscribe at any time by clicking the link in the footer of our emails. For information about our privacy practices, please visit our website.

We use Mailchimp as our marketing platform. By clicking below to subscribe, you acknowledge that your information will be transferred to Mailchimp for processing. Learn more about Mailchimp's privacy practices.

Philip O'Brien

Philip O'Brien

Philip’s career has spanned roles as a contrarian software nerd, and various other IT gigs - he’s been messing with tech since getting hold of a ZX Spectrum in the early ’80s, and still does. Now retired-ish, he’s a trustee for a local arts trust and volunteers as an IT tutor at his local library in the Upper Calder Valley. More recently, he’s taken up traditional stone carving - literally shaping ideas from a blank stone, often changing his mind mid-carve, and having fun hitting stuff. It’s clearly working: he just won the Apprentice category at the 2025 European Stone Carving Festival.

Related Posts

Cricket player batsman hitting a ball with a bat shot from below with cricket stadium ground background
Opinion

How has Hawk-Eye changed the game of cricket?

by Dr John Elsom
10 September 2026
Two wooden Adirondack chairs painted with the Canadian flag sit side by side at the end of a dock, facing out over a calm Muskoka lake in Ontario's cottage country. The red maple leaf stands out crisply against the white painted slats, while long shadows stretch across the warm wood planks. Still, mirror-like water reflects the forested shoreline that frames the lake under a soft blue summer sky. Shot from behind in a rear view, the two empty chairs invite the viewer to sit, relax, and take in the peaceful waterfront scene. With no people in the frame, the image captures the quiet beauty of Canadian cottage life, evoking tranquility, escape, and slow summer days at the lake. This patriotic Canada Day image suits travel and tourism, cottage and real estate marketing, lifestyle content, and summer long weekend campaigns. The flag-painted chairs make it especially well suited to Canada Day, July 1st, and national pride themes, as well as broader ideas of relaxation, leisure, and the idyllic Canadian lakeside lifestyle.
Culture

What Yorkshire looks like after the family tree crosses the Atlantic

by Joshua W J Brown
6 September 2026
Various cheese with different Europe flags
Trade

Travesty of language – the hollowing-out of the most-favoured nation clause

by John A Clarke
11 August 2026
Chain on a mobile phone with social networking icons on it.
Science and Technology

Regulating the symptom, ignoring the cause

by Paul Rowlston
7 August 2026
Palestine Action protest
Opinion

The ban on Palestine Action: justice or hypocrisy?

by Patrick Wright
3 August 2026
Next Post
high city walls with the City of York coat of arms

Could York lead the way to democratic renewal?

PLEASE SUPPORT OUR CROWDFUNDER

BROWSE BY TAGS

Art Books Boris Johnson Bradford Charity Climate Change Cost of Living Covid-19 Creative Industries Crime Democracy Devolution Donald Trump Environment Equality Experience Farming Festival Gaza Conflict General Election History Human Rights Immigration Iran Journalism Keir Starmer Labour Leeds Media Mental Health NHS Northern Ireland Protocol Pollution Poverty Recipe Refugees and Asylum Seekers Restaurants Retained EU Law Review Rishi Sunak Sheffield Theatre Travel Ukraine USA
Yorkshire Bylines

We are a not-for-profit citizen journalism publication. Our aim is to publish well-written, fact-based articles and opinion pieces on subjects that are of interest to people in Yorkshire and beyond.

Yorkshire Bylines is a trading brand of Bylines Networks Limited which is separate to, but allied with, Byline Times.

Learn more about us

No Result
View All Result
  • About
  • Authors and editors
  • Complaints
  • Contact
  • Donate
  • Letters
  • Privacy
  • Network Map
  • Network RSS Feeds
  • Submission Guidelines
  • Download the Bylines Network App

© 2020-2026 Yorkshire Bylines. Powerful Citizen Journalism. ISSN 3049-9720

No Result
View All Result
  • News
    • Brexit
    • Education
    • Environment
    • Health
    • Home Affairs
    • Transport
    • World
  • Politics
  • Opinion
  • Society
    • Culture
    • Dance
    • Food
    • Music
    • Poetry
    • Recipes
    • Sport
  • Business
    • Economy
    • Science and Technology
    • Trade
  • Region
  • The Davis Downside Dossier
  • The Digby Jones Index
  • Cartoons by Stan
  • Authors and editors

Newsletter sign up

CROWDFUNDER

© 2020-2026 Yorkshire Bylines. Powerful Citizen Journalism. ISSN 3049-9720