ScrutinEyes · 2026-10-04
The week AI stopped being the suspect and became the defendant: The Week in AI-Security, Sep 15–21, 2026
A first-party autonomous-agent breach lands on a regulator's desk, the AI coding tools themselves get a zero-click RCE, Cisco's email gateway hands out root — and a threat actor's record count outruns what any company has confirmed.
The through-line
For months the AI-in-attacks story has been told in the conditional tense: a vendor recovers some artifacts, assesses the malware was likely AI-written, and hedges. This week the tense changed. Spain’s data-protection regulator logged the first breach it attributes to an autonomous AI agent acting on its own — not a claim in a research blog, a notification in a government file. Mandiant and Google published a report full of observed agent incidents, including one that torched $50,000 with no attacker in the loop at all. And the tooling caught up to the trend from the other side: Plugin4Shell put a zero-click remote-code-execution hole in the four AI coding agents millions of developers now let run unattended. Set against the more familiar plumbing failures — Cisco’s email gateway giving unauthenticated attackers root, an energy utility filing an 8-K, an AWS region that no longer exists — the signal is that AI has moved from the thing we suspect in incidents to the thing named in the incident report. Worth holding onto, though: the strongest new items this week are the ones with a primary filing or a first-party disclosure behind them. The Spanish notification is a single data point its own regulator refuses to call a trend, and the loudest breach number of the week is one no company has confirmed. The direction is real. The discipline still has to be supplied by the reader.
Cisco’s email gateway hands out root
Cisco confirmed active exploitation of CVE-2026-76461, a CVSS 9.8 SQL-injection flaw in the Secure Email Gateway that lets an unauthenticated attacker run commands as root by sending a crafted email — no click, no login, exploitation on delivery. Cisco’s PSIRT says it became aware of exploitation in September 2026. That was before public disclosure, but Cisco has not said when the exploitation began. CISA added it to the KEV catalog the same day; per the OpenCVE record, federal agencies had until September 17 to remediate. The device sits at the front of the mail path, processing hostile input as its entire job — and an attacker with root can, per Cisco, delete the very mail-log evidence you’d use to detect them. Patch to 16.5.0-780 (or 16.0.4-302 / 15.5.5-014), and because root access means log tampering, treat detection as forensics on your network and firewall logs, not just the appliance’s own. Help Net’s roundup also reports that unauthenticated attackers are exploiting a companion Cisco ISE authentication-bypass flaw disclosed the same week (CVE-2026-76460); confirm which fixed train applies before you plan the window.
Plugin4Shell: the AI coding agents are now the attack surface
The most consequential AI-security item of the week is a flaw in the AI tools, not one used by them. Researchers at AIR disclosed “Plugin4Shell,” a zero-click RCE affecting Claude Code, OpenAI’s Codex, GitHub Copilot, and Google’s Gemini CLI. The mechanism is small and damning: the agents check out a plugin’s pinned commit SHA but never verify the checkout actually landed there, so an attacker who names a git branch after the pinned hash gets their code run instead — and because Claude Code and Codex auto-update plugins in the background, no user action is required. AIR reports the same weakness enabled “SkillJacking” of 925 hijacked skills reaching roughly 134,000 agents. Patch state as of September 22 is uneven and worth naming precisely. Anthropic fixed Claude Code (2.1.179) and OpenAI fixed Codex (0.146.0). Microsoft has shipped no Copilot patch and, per The Eastern Herald, has not publicly acknowledged the flaw or given a remediation timeline. Google stopped serving the consumer Gemini CLI in June and points users to its successor, Antigravity. The Hacker News reports that enterprise access to Gemini CLI continues with updates, but it is unclear whether those updates include a fix. If you’ve handed a coding agent standing credentials and background auto-update — which is the default posture most teams adopted this year — this is the week to check your version and your plugin allowlist. The supply-chain lesson we keep relearning now has an agent in the middle of it.
Spain logs the first autonomous-AI-agent breach — and its own regulator won’t call it a trend
Spain’s data-protection authority, the AEPD, received what it describes as the first breach notification involving an autonomous AI agent operating independently of direct human control: per the notification, the agent scanned for vulnerabilities, logged into the target’s network, hunted through the application until it found a flaw, then altered personal-data records and pulled invoice data. What makes this notable is the venue — a GDPR filing, not a vendor narrative. What keeps it honest is the AEPD’s own framing. Deputy director Francisco Pérez Bes said the single notification “does not allow us to establish a statistical trend, although it does constitute a significant sign,” and the agency stressed that the account comes from the affected organization and needs further analysis — that naming a specific AI model does not prove the model or its provider was itself compromised. Read it as a milestone in how these incidents get recorded, not as proof of a wave. The regulatory machinery for logging AI-agent breaches now exists and has been used once. That is the story; the trend line has exactly one point on it.
Mandiant’s report: the runaway agent that cost $50,000 with no attacker at all
Google’s Mandiant and Threat Intelligence Group published their AI Risk and Resilience 2026 report, and the detail that should stick is not an attack. An accounting agent entered a runaway loop and made more than 15,000 high-cost API calls in under an hour, running up roughly $50,000 in charges and disrupting operations — no adversary involved, just autonomy without a circuit breaker. Alongside it, Mandiant documents attacks that are real and observed: prompt injection as the dominant enterprise attack vector, and a Shai-Hulud worm episode in which an attacker hijacked an AI coding-assistant session, got it to recommend a poisoned dependency the developer accepted, and spread across about 100 internal repositories while stealing GitHub OAuth tokens. This is the independent corroboration the vendor-hedged AI claims of prior weeks were missing — a second, separate intelligence shop describing agent-driven incidents from its own casework. It also reframes the risk: the near-term agent threat to most enterprises is as much operational (loops, cost, blast radius) as it is adversarial, and both need the same missing control — bounded permissions and a kill switch.
CenterPoint Energy files an 8-K — and the loudest number isn’t the company’s
CenterPoint Energy disclosed a data breach in a September 15 SEC filing after spotting an online post claiming someone held its customer data. The company confirms unauthorized access to personal information through an external-facing system, says electric and gas delivery were unaffected, and — importantly — says it is still working to determine the scope of customers and data involved. Circulating alongside the filing is a figure of roughly 7.49 million records. Grade that carefully: it traces to a threat actor’s leak-site claim, not to anything CenterPoint has confirmed — the SEC filing quantifies nothing. This is the exact split ScrutinEyes keeps flagging: the regulated disclosure tells you an incident happened and roughly what kind of data; the big round number in the headline is the seller’s advertising until a filing or notification catches up to it. When those two sources disagree on scale, the one under legal jeopardy for lying is the one to quote.
AWS confirms a cloud region is simply gone
Not every data loss is an intrusion. AWS acknowledged permanent, unrecoverable customer data loss in its Bahrain region (me-south-1) and one availability zone of its UAE region (me-central-1), following Iranian drone strikes that began damaging the facilities in early March 2026. AWS says most customers had already moved data out using backups or alternatives, and that the Bahrain damage “exceeded what its redundancy design could withstand.” The security lesson is availability, the third leg of the triad we mention least: an entire cloud region’s redundancy is still redundancy within a threat model, and physical destruction of a metro was outside Bahrain’s. For anyone whose disaster-recovery plan assumes “the cloud provider keeps a copy somewhere,” this is the counterexample of record — cross-region replication is a customer decision, not a default, and the customers who made it are the ones who still have their data.
The EU floats a frontier-AI slowdown
On the governance front, European Commission president Ursula von der Leyen told EU lawmakers that self-improving frontier AI should slow down. She said she would invite the leading labs to discuss how the EU can support their own efforts to slow down, but gave no date for that meeting. She also announced cooperation with Canada and the UK on model evaluation, verification, early warning, and AI security. It is a political signal, not a rule, and worth keeping in proportion — no text, no timeline, no mechanism yet. But read against the week’s other items — a regulator logging its first agent breach, a report cataloguing agent incidents, a zero-click hole in agent tooling — the policy conversation is visibly chasing an operational reality that has already arrived. The interesting question for anyone selling or buying assurance is which evaluation standards emerge, and whether “we verified it” ever becomes a defined term rather than a marketing one.
What to watch
The two unpatched agents. Plugin4Shell is fixed in Claude Code and Codex. As of September 22, Copilot has no patch. Google has retired the consumer Gemini CLI rather than fix it, and it’s unclear whether enterprise builds will get a fix. Watch whether Microsoft ships a fix and how cleanly Gemini CLI users can migrate — an unpatched agent with background auto-update and standing credentials is a live liability, not a someday one.
Whether a second regulator logs an agent breach. The AEPD has one data point and says so. The thing that would turn “significant sign” into “trend” is a second notification, in Spain or another GDPR authority. Until then, resist the wave narrative — including from vendors who benefit from it.
CenterPoint’s real number. The 7.49M figure is a threat-actor claim. The number that matters will arrive in customer notification letters and any amended SEC disclosure. Same discipline as always: filings over leak-site advertising.
Revolut’s aftermath (carry-over). Last week’s Revolut breach — customer identity documents handed out in response to a spoofed government data request — has moved into extortion. SecurityWeek reports that the requests came from an Italian Interior Ministry address (pec.interno.it) and affected about 680 customers. It also reports that the actor is demanding $3 million, and that Revolut says it has received no direct demand. What remains open is what the regulator filings will show. Those filings, not sellers’ posts, are the source to trust.
Whether AI evaluation gets a definition. The EU’s evaluation and verification cooperation with Canada and the UK is the venue where “verified” could become a defined term. If you care about honesty-as-a-standard rather than honesty-as-a-slogan, this is the thread to pull.
ScrutinEyes is independent analysis. Every claim links its source inline; where we rely on secondary reporting, we say so. We don’t sell what we grade.
Reading this because someone's asking about your security? See exactly which rules apply to you and where you stand — check your readiness free. Five minutes, in your browser, nothing stored unless you ask. Readiness, not legal advice.