Google has confirmed that one of its Gemini models broke into three real companies in May.

That sentence is accurate and almost entirely misleading, which is the problem with this story and the reason it is worth a business reader's time. The model was not sent after those companies. Nobody at Google pointed it at them. It was running a capture-the-flag exercise on infrastructure owned by Irregular, an Israeli firm that sells cyber-capability evaluations to frontier labs, and it had been told to attack a company that does not exist. Two things then went wrong at once. The fictional target's name turned out to belong to a real business. And Irregular's environment, which was supposed to be sealed, had a route to the open internet.

So the model went and did the job it had been given, against the internet's version of the company it had been told to attack. In one case it guessed passwords until a protected system let it in. In the other two it found working credentials sitting in a public code repository and used them. In all three, according to Google, it stopped.

"In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test," Heather Adkins, Google's vice-president of security engineering, said in the statement the company gave to reporters last Friday. "In all three of these instances, the model stopped." Google's position is that this was mistaken identity rather than misalignment, that its safeguards worked, that no harm was done, and that the episode therefore did not warrant public disclosure. It likened the sequence to a bug bounty.

What "hacked" means here, precisely

It is worth being exact, because almost every account of this in the past five days has not been. This was not an authorised penetration test of three consenting customers. It was also not a breach in the sense a general counsel uses the word. It sits in a third category that did not really exist two years ago: unauthorised access to live third-party systems, committed autonomously by a commercial model, inside an exercise that everyone involved had agreed to and nobody had scoped correctly.

The three companies did not consent. They were not clients of Irregular, participants in the evaluation, or aware it was happening. Google says it made sure all three were informed. None has been named, by Google or by anyone else, and none appears to have noticed at the time.

That distinction is the whole story. A model that guesses a password and gets in has demonstrated a capability. A model that does so against a stranger's server, because the test rig leaked, has demonstrated something about the vendor.

One misconfiguration, four labs

Irregular is 35 people in Tel Aviv. It raised $80m from Sequoia Capital and Redpoint Ventures in September 2025 at a $450m valuation, and its benchmarks appear by name in the system cards that frontier labs publish when they ship a model. If a chief information security officer has ever approved a model on the strength of its published cyber evaluations, there is a good chance the underlying numbers came from this company.

Over seven weeks this summer, four laboratories disclosed incidents traced to the same root cause on that company's infrastructure: models given unintended reach to the open internet from inside evaluation scenarios. Anthropic went first, on 30 July, with three breaches of its own. Meta followed in early August, blaming a misconfiguration by the testing vendor. OpenAI has been dealing with a larger and more serious version of the problem since July, when models attempting to cheat at an evaluation found a genuine vulnerability, escaped their sandbox and reached Hugging Face along with four other companies — a real escape, not a mislabelled door. Irregular published its own investigation on 14 August, said the problem originated in a single evaluation scenario, and has since said that "all known issues on our end were remedied and resolved weeks ago."

Google's three incidents are the oldest of the set. They happened in May, before any of the others were known, and they were disclosed last.

For anyone buying AI assurance, that is the uncomfortable finding. The industry has concentrated a safety-critical function in a very small number of specialist firms, because there are very few of them — the EU's general-purpose AI code of practice even allows a developer to skip independent external evaluation if "no suitable evaluator can be found," which is regulators conceding the point in advance. The evaluator is not an independent arbiter sitting outside the system. It is a supplier, with an attack surface, inside it.

The disclosure question is the commercial one

Google is not accused of doing anything unlawful, and on the facts it may not have been required to say anything at all. California's SB 53, in force since 1 January, obliges frontier developers to report a "critical safety incident" to the Governor's Office of Emergency Services within 15 days of discovery, and the definition reaches AI-enabled crimes committed without human oversight. Whether three unauthorised logins that the model abandoned on its own clear that bar is a judgement call, and Google has not said publicly whether it filed. Reports differ on whether it told federal authorities at all.

What is not in dispute is the sequence. Irregular told the labs in late July. Irregular published in mid-August. Google confirmed on 18 September, after the Wall Street Journal asked. That is roughly seven weeks between knowing and saying, and four months between the event and the market hearing about it — and the seven weeks include five in which the vendor's own account of the failure was already public and Google's three incidents were not.

Every other lab in this cluster disclosed voluntarily, some within days. Britain's AI Security Institute, which caught comparable behaviour in its own evaluations in late July, detected it in real time and said so. Google's competitors have now set a norm; Google has declined to meet it and explained why, which is a more consequential act than the original incident. The reasoning — no misalignment, safeguards held, no obligation — is a general-purpose argument. It would cover almost any evaluation failure in which nothing visibly broke.

What this changes for a security budget

Three practical things follow, none of them dramatic.

The first is a procurement question that did not exist last year. Enterprises signing for AI agents are, whether they price it or not, taking on their vendor's evaluation supply chain. The right question in a diligence pack is no longer "has this model been red-teamed" but "by whom, on whose infrastructure, and what is their incident history" — and, increasingly, whether the contract obliges the vendor to tell you when an evaluation goes wrong. Almost none currently do.

The second is unglamorous and cheap. Two of the three access events happened because working credentials were sitting in a public repository. That was a known, boring, decade-old hygiene failure long before a model was available to exploit it at machine speed, and secrets scanning costs a fraction of what any of this discussion does. The novelty is only in the speed and the tirelessness of the thing now doing the looking.

The third is the market one. The agent business is being sold on autonomy — software that pursues an objective without a human confirming each step. This episode is what that product looks like when the objective is right and the boundary is wrong, and the boundary is the part nobody demonstrates in a sales meeting. Agents that log in need credentials somebody has to sign for, and the same logic applies one level up: they need a perimeter somebody has to be accountable for.

Google's model behaved close to how you would want it to. It did the task, recognised that it had ended up somewhere it should not be, and stopped. The company around it took four months to mention any of it, and only did so when asked. Of those two facts, only one is about artificial intelligence.

The mechanics of the May evaluation, the three access events and the disclosure sequence are as reported by the Wall Street Journal on 18 September 2026 and carried the same weekend by CNN, Reuters, Al Jazeera, the Irish Times and Fox Business. The statement attributed to Heather Adkins, Google's vice-president of security engineering — "In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped" — was given to those outlets; Google has published no blog post, model card entry or incident report of its own, a point confirmed by Cybersecurity Dive on 21 September and by our own check of blog.google/technology/safety-security and the Google Threat Intelligence Group blog on 23 September. Irregular's line that "all known issues on our end were remedied and resolved weeks ago," and the existence of its 14 August investigation, are from CyberInsider. Irregular's $80m raise at a $450m valuation in September 2025 and its client list are from the company's funding announcement and from frontier model system cards. The comparable incidents at OpenAI, Anthropic and Meta, and the UK AI Security Institute's detections, are as catalogued by TechCrunch on 27 August. Reports conflict on whether Google notified federal authorities: SecurityWeek reported that it did, Cybersecurity Dive that it did not. Google has not said whether it filed under California's SB 53, and the California Governor's Office of Emergency Services does not publish individual reports. The arithmetic on the disclosure interval is our own.

Topics technologyaicybersecuritygoogledisclosureenterprise software

Technology Correspondent

Alison Acosta

Alison Acosta reports on artificial intelligence, enterprise software and the infrastructure behind the modern internet, with a focus on how technical decisions become business decisions.