Evgen verzun
Blog
August 11, 2026
Everyone Has an AI Security Take Now. After Twenty Years in Cybersecurity, Here is Mine.
AI security is a hot topic these days and everybody seems to have an opinion on it. Even my mechanic has one. It also just happens to align very closely with the narratives prevalent in the media: AI helps criminals break in.
As it happens, I have spent twenty years dealing with that question.
But before we dive in, I need to make an important distinction. There are two sentences that get used interchangeably and they should not be.
"Nobody has hacked us." and "We are compliant."
I want to give my view on both, because the offensive AI story is only half right, as the other half gets conveniently left out and often costs more.
Part right, part marketing
AI changed the attack economics. Reconnaissance that took a competent human a week now runs while you make coffee. Phishing stopped announcing itself with broken grammar two years ago. Voice cloning works well enough that finance teams wire money after a video call with a face they trust. None of that is hype, it’s an established fact at this point.
The framing around it, however, is what I take issue with. Read a year of loud announcements from frontier labs and you would think offensive capability arrives with each new flagship model, gated behind a safety program, held back by responsible adults.
But how much of that is marketing and how much is the truth?
In July, Anthropic reviewed 141,006 of its own evaluation runs and found three cases where a model reached the open internet from inside a test environment and then compromised the production infrastructure of real companies. According to their own report, Opus 4.7, a model that was already a few generations old, located and exploited vulnerabilities in a real company's infrastructure, pulled application and infrastructure credentials, and reached a database holding several hundred rows of production data. It was also the only one of the three that kept attacking after concluding the target was real. In two runs it talked itself into the idea that the real company must be part of the exercise.
Despite being several releases old, it managed a break-in anyway.
The UK AI Security Institute measured the same thing from outside and put open-weight models four to seven months behind the closed frontier on cyber capability, down from six to ten. Palo Alto's Unit 42 aimed nearly everything it could get at real code and confirmed 14,090 vulnerabilities across 3,915 open source projects in two months. Independent testers watched GLM-5.2, which anyone can download tonight, find 16 of 26 known vulnerabilities where the best proprietary model found 23.
What does this tell us? That the capability and aptitude for “hacking” is general.
An LLM reads a perimeter, cross references it against the public catalogue of known vulnerabilities and working exploits, and keeps going until something answers. Novel zero-days are a harder and rarer thing, and are a separate conversation. Most of the damage comes from the boring version, and the boring version already shipped, more than once, in models nobody can un-release.
The defense strategy of waiting for the labs to gate the dangerous model was never a viable option.
Fight AI with AI* (if your vendor doesn’t decline)
You do need machines on defense. Attacks running at machine speed will not wait for an analyst to work down an alert queue. You cannot disagree with the industry on this.
Then OpenAI rogue agents hit Hugging Face, and their write-up became the most useful security document published this year.
An autonomous agent system linked to OpenAI broke into part of Hugging Face’s production infrastructure in July. It exploited two code execution paths in dataset processing, escalated to node level, harvested cloud and cluster credentials, and moved laterally across internal clusters over a weekend. Their own LLM-based triage over security telemetry surfaced it, which supports the fight-AI-with-AI argument nicely.
OpenAI later confirmed that its own models, running with reduced cyber refusals for an internal evaluation, drove the intrusion.
What was especially interesting is how Hugging Face forensics went differently.
Working through more than 17,000 recorded attacker actions meant feeding real attack commands, exploit payloads and command-and-control artifacts into a model. Hugging Face started with frontier models behind commercial APIs. But their providers' safety guardrails blocked the requests, because a guardrail cannot tell an incident responder from an attacker. What they had to do is run the analysis on an open-weight GLM-5.2 on their own hardware.
Their advice is to have a capable model you can run on your own infrastructure vetted and ready before you need it. It is great advice, especially when you consider that a guardrail is ultimately a product decision, and not a contract term. It moves when a vendor's legal team decides it should move, and nobody sends you a notice.
Meanwhile the labs keep their strongest security capability inside controlled programs. Anthropic's Glasswing gives roughly two hundred vetted organizations access to a restricted model it will not sell to the public, on the reasoning that the capability which defends also attacks. OpenAI tiers its equivalent through approved partners and states that the underlying model stays with the partner rather than reaching the customer.
Their offer is then to use their guardrailed AI to defend your perimeter, with the caveat that it may refuse you mid-incident. I would not build my last line of defense on that.
No answer to “Shadow AI”
Everything above concerns keeping people out of your perimeter. But let’s consider the articles you read about the data that leaves it.
The shadow AI paragraph turns up in most of them. It’s always an employee pasting a sensitive doc into a chatbot. Careless staff, send a memo, buy a training module.
Twenty years of watching people do their jobs and I can tell you that doesn’t work. You cannot undo the leak. And your staff will route around anything slow. That’s just what people do. They did it with personal email, then with Dropbox, and they are doing it now with whatever AI model answers fastest. That is a design problem, and calling it negligence guarantees you will keep having it.
And while the media loves a big headline that links back to shadow AI, you never see any practical solutions except “training”.
What really works is providing your staff with safe access to the tools they actually need to do their job.
Not leaking and being compliant is not the same thing
Say your pen test comes back clean. Your SOC has nothing open. And yet, you still run compliance risks because regulators ask completely different questions.
A regulator will want:
- the lawful basis for sending that personal data to that processor, and the transfer mechanism behind it
- the physical location of the data, and which sub-processors you authorized
- the retention period, with proof of deletion
- a record of who asked what, and which model saw which document
- the name of whoever carries the liability
Not one of those depends on whether somebody “broke in” or not.
The answers change by jurisdiction, which is the part that makes this significantly more complex than people expect it to be. A bank in Abu Dhabi works under rules where a firm relying on a third party for ICT services remains responsible for compliance in relation to the activities that third party performs, and has to be able to bring the service back in house if it comes to that. The DFSA takes the same line inside the DIFC. GDPR makes you the controller the moment you decide to feed personal data into someone else's model, and if that provider starts processing your data for its own purposes it becomes a controller too and loses the liability shield processor status given to it.
Naming a large vendor is not a defense in any of those regimes.
The Italian regulator fined OpenAI fifteen million euros in December 2024 over lawful basis, transparency and a breach notification failure. A Roman court annulled it in March 2026 on jurisdiction, because the Irish DPC had become lead authority, and left the substantive questions standing. Read that as the danger passing if you like. I read it as a schedule change.
Compliance failure arrives without a press release. There is an audit, and then there is a finding.
"Just use the enterprise account"
This is the answer I hear most where compliance is concerned. It is as reasonable as it is incomplete. Let’s break that down.
The enterprise tier buys you real things: no training on your data by default, contractual commitments, a serious certification stack, and a shorter retention window if you negotiate for one. For a small company doing light AI work on unremarkable data, that could be enough, for the most part.
However, it also leaves at least four things on the table.
- The provider decrypts and processes your prompt to answer it. Anthropic published research on confidential inference through trusted virtual machines in June 2025, which would close that gap. For now, it is still research.
- Retention runs thirty days by default at both major labs, seven on some surfaces, with flagged content held up to two years and classifier output up to seven. OpenAI's terms sit in the same range.
- Zero data retention is an add-on you request and get approved for, on eligible endpoints. Anthropic's own documentation excludes the Claude Team and Claude Enterprise chat interfaces, and its newest models require thirty day retention and cannot use it at all.
- European residency covers new projects and cannot be applied to the ones you already built, and some features process outside the region regardless.
Even after all that, an enterprise account is only for one vendor's models.
Your lawyers want one vendor, your engineers prefer another, your analysts have a third that suits their work, and somebody needs a regional model for a market you operate in. You solved shadow AI for the models inside the subscription and manufactured it for the ones outside. The incentive to route around you doesn’t go anywhere.
How AI actually leaks data
Exposing the data to your vendor is only one of the things you need to deal with. Arguably a lesser one. But most exposures I have catalogued came from the layer wrapped around the model rather than the model's core infrastructure. Here are some prominent examples:
- OpenAI, August 2025. Around 4,500 shared conversations surfaced in Google through a site search. Personal disclosures about mental health and relationships, and a resume traceable to a live LinkedIn profile. OpenAI killed the feature within hours and called it a short-lived experiment that gave people too many chances to share things they did not mean to share.
- xAI, the same month. Google estimated more than 370,000 Grok conversations indexed. Medical and psychological questions, business detail, and at least one password sitting in a search result.
- Anthropic, July 2026. Shared chats and Artifacts indexed, carrying clinical trial data with patient names, ages, ethnicities and treatment dates, children's names and phone numbers, resumes, API keys, crypto wallet keys, and documents marked internal use only. Anthropic's position was that the share feature worked as designed. Forbes had reported the same failure against roughly 600 Claude conversations ten months earlier, after which Anthropic said it had blocked the crawlers.
- DeepSeek, January 2025. A different mechanism and a worse one. Wiz found an unauthenticated ClickHouse database on two open ports accepting arbitrary SQL, holding over a million log lines with chat history, secret keys and backend detail.
- Across six providers. Obsidian Security found more than 143,000 GenAI conversations sitting in public archives spanning Claude, Copilot, ChatGPT, Grok, Mistral and Qwen, including AWS access key IDs and an API token. Delisting from a search engine does not reach an archive.
Then the wrappers, which is where volume lives. One Firebase misconfiguration at the company behind Chat & Ask AI exposed around 300 million messages from more than 25 million users, including conversations about self-harm. The researcher who found it built a scanner and turned up 103 of 200 iOS apps carrying the same flaw.
In March 2026, three unencrypted and unauthenticated databases tied to Sears Home Services gave up 3.7 million records, among them 1.4 million call recordings running to roughly four terabytes. The bot kept recording for up to four hours when customers forgot to hang up, collecting whatever was said in the room afterwards.
Nobody in that list needed a zero-day.
Enterprise deployments break too
Two incidents in early 2026 settled this question for me.
A security firm pointed an offensive AI agent at McKinsey's internal platform, Lilli. Public API documentation exposed more than two hundred endpoints, twenty two of them unauthenticated, and a SQL injection flaw hid in the spot where JSON field names got concatenated into queries while the values sat parameterized. Full read and write access in under two hours, for about twenty dollars in tokens.
Behind the door: 46.5 million chat messages covering strategy and M&A work, 728,000 file records of client material, 57,000 accounts, and 95 system prompts the agent could rewrite. Those prompts govern what forty thousand consultants get told. McKinsey patched the endpoints inside a day and says nothing indicates unauthorized parties took client data.
At Meta the same month, an internal agent answered a question on an internal forum, published its answer without approval, got it wrong, and a colleague acting on that advice widened permissions on sensitive company and user data for about two hours. Meta rated it Sev 1.
Nobody attacked Meta. An agent with authority and no supervision did that on its own.
Enterprise deployments don’t guarantee data travels safely and doesn’t get exposed.
What I would build instead
I wrote something in 2019 about the NSA's TLS inspection advisory that I would not change today. TLS secures the channel, not the data moving through it. Wear the shoes and you can still catch a cold.
The failures above all follow from one assumption, that protecting the pipe protects the contents.
Principles I hold to:
- Encrypt the payload so the surrounding infrastructure cannot read it. A misconfigured wrapper, an exposed database or a compromised node then yields ciphertext instead of a headline.
- Strip identity and organizational context before the model call. Let the provider answer a question rather than build a picture of who asked it and why.
- Put one governed path between your staff and whichever model they need, vendor hosted, self hosted, or your own. People stop routing around a layer that gets them what they wanted in the first place.
- Keep the record. Who asked, what policy applied, which model answered. That artifact is what an auditor asks for, and few organizations can produce it.
- Deploy where the law requires rather than where the vendor happens to host.
None of this stops an intrusion. It’s great if that doesn’t happen, but if it does it decides what an intruder finds. The profession reached the assume-compromise position on its own, and some of its loudest voices say they are yet to see an LLM stop a breach. And I agree. So plan around the breach you cannot stop, and make the data survive it.
I know how this sounds coming from me. I have spent my career on encryption and I am going to recommend encryption. Take the argument on its merits, or take the incident list on its merits, and reach your own conclusion.
Who needs this and some closing thoughts
A hospital group moving records between two countries has the same problem as a government. So does a law firm holding privileged material, a bank answering to two regulators, a utility running operational data through a model, a ministry with residency rules written into law.
Sensitivity is a property of the data, and you cannot put an industry label on this. Not in the age of AI.
The baseline approach for cybersec has always been to ensure data protection on the assumption that the perimeter gets compromised. And that doesn’t change with AI, in fact, it is even more true today.
Going unbreached is partly luck. Staying compliant and protecting sensitive data is something you build for even if luck isn’t on your side.