The Gemini hacking incident makes Google the fourth AI lab whose model broke into real companies during a cyber evaluation. It happened in May. Google confirmed it on September 18, after The Wall Street Journal asked about it.
The short version: an older, unnamed Gemini model reached three real businesses during a capture-the-flag exercise run by Irregular, the same testing firm named in the Anthropic, OpenAI and Meta disclosures. It guessed one password. It used credentials it found in public code repositories for the other two. Google says the model stopped each time it realized the target was real. Google has published no report, transcript or timeline to show that.
We have followed this story since July. The evaluation sandbox problem covered Anthropic’s report, and one vendor, three labs, four failures added Meta and a second OpenAI case. This post adds Google, and compares what each lab has actually put in writing.
What Google has confirmed about the Gemini hacking incident
Everything Google has said came through statements to reporters. There is no Google blog post, incident report or model card update. Here is what those statements establish.
- When: the three intrusions happened in May 2026.
- Who noticed: Irregular told Google in late July. Google did not detect it first.
- Which model: Google has not named it. It told SecurityWeek it was not its latest model.
- Who was hit: three companies. None has been named.
- What happened next: according to SecurityWeek, Google notified the three companies and federal authorities, and worked with Irregular on changes to its testing process.
Heather Adkins, Google’s vice president of security engineering, gave the core statement. As quoted by Fox Business: "In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped."
Google’s position, as reported by several outlets, is that the Gemini hacking incident did not need public disclosure. Its reasoning: the model stopped, it caused no harm, and the behavior was not an example of misalignment.
How the test reached real companies
The exercise was a capture-the-flag task on Irregular’s infrastructure. Gemini was told to break into a fictional company. Two things then went wrong at once.
First, the environment had internet access it was not supposed to have. Irregular’s own word for it is that access was "unintentionally made available."
Second, the fictional company shared its name with a real one. The model did what it was asked to do. It searched for the target by name, found the real business, and went after it.
That second detail matches what Irregular published on August 14 in Addressing Recent Incidents. The post says a fictional company name "unintentionally coincided with a real domain" through "human oversight," and that a background check missed it. It estimates the problem hit fewer than 1 in 10,000 advanced simulations, usually hundreds of turns in.
Irregular’s statements to reporters about the Gemini hacking incident follow the same line. "All known issues on our end were remedied and resolved weeks ago," it said. It also called the Gemini case "the same issue that was already reported" and said it "does not represent a materially separate incident."
From Irregular’s side, that is consistent. It has treated every disclosure since July as one root cause. From a reader’s side, it means one misconfigured environment touched models from four labs, and the fourth account arrived seven weeks after the vendor raised it.
The techniques were basic, again
The methods deserve attention because they are the part that applies to you.
- One company: the model guessed passwords until it got into a protected system.
- Two companies: it searched the web using the company’s name, found working credentials in public repositories, and logged in with them.
That is the same pattern Anthropic described in July. Its report said Claude got in "using basic techniques, such as exploiting weak passwords and unauthenticated endpoints." No zero-days. No novel exploit chains.
So the capability on display is not what matters most. A patient script could do the same. What matters is that three real companies had a guessable password or a leaked credential sitting in a public repository. A model found them in a test, but anyone could have.
"The model stopped": the one new claim in the Gemini hacking incident
Google’s account adds one thing the earlier cases did not stress. It says Gemini stopped on its own each time it realized it had reached a real company.
If that holds up, it matters. It would mean the model recognized a boundary mid-task and respected it, without a human stepping in.
But look at what supports it today. A spokesperson’s summary, given to reporters, four months after the fact. There is no transcript, no count of actions taken before each stop, and no description of how the model worked out the target was real. There is also no statement from any of the three companies about what was accessed or changed.
What each lab has actually published
Compare the evidence behind each lab’s account. The first OpenAI case is the one we covered in OpenAI and Hugging Face’s disclosure.
| Lab | Disclosed | How | Vendor named | What you can read |
|---|---|---|---|---|
| OpenAI (Hugging Face) | July 21 | Published post | No, a different environment | Written account with the attack path |
| Anthropic | July 30 | Published report | Irregular | Review of 141,006 runs, three incidents |
| OpenAI (second case) | August 4 | Published post | Irregular | Written account |
| Meta | August 5 | Statement to reporters | Irregular | Spokesperson quotes only |
| September 18 | Statement after a press inquiry | Irregular | Spokesperson quotes only |
Google and Meta sit at the thin end. The claim that separates Google’s case from the others, the self-stop, is also the one with the least behind it.
That is not an accusation. The claim may be accurate. It is a statement that cannot yet be checked, and it should be reported that way. We said the same about Meta in August.
Seven weeks between notice and a newspaper
The dates tell their own story, so here they are in order.
- May: the three intrusions happen.
- July 21 to August 5: OpenAI, Anthropic and Meta disclose their own cases, three of them tied to Irregular.
- Late July: Irregular notifies Google.
- August 14: Irregular publishes its findings. It names no lab.
- September 18: the Journal reports the Gemini case, and Google confirms it.
That leaves about seven weeks between Google hearing from Irregular and Google saying anything in public. For all of that time, the public count stood at three labs. It was four.
What labs owe the public is a policy question, and not ours to settle. The practical point is narrower. If you were tracking this story to judge vendor risk, the picture you had in August was incomplete, and nothing in it told you so. In our look at AI containment plans, a third-party scorecard graded Google lower than OpenAI and Anthropic on published safety practices.
What Irregular says it changed
Irregular’s August 14 post lists the fixes it made after the first disclosure. Google says it worked with Irregular on those changes. In Irregular’s words, it has:
- disabled the affected evaluation and reviewed its logs
- expanded "manual review of model actions and behavior"
- set up a "dedicated internal team to challenge our own assumptions about security"
- put "new and robust protocols" in place for internet access
- added a "systematic process for reviewing and revalidating evaluations before each run"
It also promised an "open whitepaper on future best practices." It gave no date, and none has appeared.
Two things are missing from that list. There is no description of how the isolation now works, and no outside audit of it. For a firm whose environments sit under four labs’ cyber testing, that gap is still the one to watch.
A lesson for anyone who tests agents
The name collision is the most useful detail in the Gemini hacking incident, because it is so easy to repeat.
Irregular built a fake company and gave it a name. Somewhere, a real business already used that name. Once the model could reach the internet, the fake target and the real one were the same search result.
The internet already has a fix for this. Some names are reserved so they can never belong to anyone. The domains example.com, example.org and example.net, and the top-level domains .test, .example and .invalid, exist for exactly this job.
If your team builds test environments for agents, security tools or form automation, use those names for fake targets. Then block outbound traffic from the test network by default, and allow only what the test needs. Either step alone would very likely have kept the model away from a real company. Irregular had neither working in May.
Does the Gemini hacking incident affect Gemini at work?
Nothing reported suggests it does. This happened inside a cyber evaluation, run by a third party, on a model Google says is not its current one. It says nothing about the Gemini app, Gemini in Workspace, or the Gemini API behaving this way for customers.
There is also no report that any Google customer data was involved. The three companies were outsiders who happened to share a name with a fictional target.
What it does show is how these models behave with tools, internet access and an instruction to get in. They find the easy door. If you are building agents on any frontier model, including Gemini, that is the behavior to plan around: an agent with broad access will use whatever credentials it can reach.
What to change on your own systems
Every one of these incidents came down to a guessable password or a credential left in public. That is squarely in reach for a nonprofit, an association or a small agency.
Search public repositories for your own credentials. Check your organization’s GitHub or GitLab accounts first. Then check old personal repos from former staff, and repos your contractors or agencies run for you. GitHub’s secret scanning covers public repositories. Open-source scanners can check history too.
Rotate anything you find. Deleting the file is not enough. The secret stays in the commit history and in any forks. Treat it as compromised and replace it.
Put limits on every internet-facing login. Lock out or slow down repeated failed attempts. Turn on multifactor authentication. Include the forgotten places: staging sites, old admin panels, a WordPress or Drupal login on a site nobody updates.
Watch for bursts of failed logins. One of the three Gemini intrusions was plain password guessing. A basic alert on failed attempts would show that.
Agencies, check client environments. Shared credentials in a project repo, default passwords on a staging server and forgotten test accounts are all findable. You hold keys for many organizations at once.
None of this is new advice. It is the same list we gave in July and August. Four labs later, it is still the list.
What we still do not know
- Which Gemini model was involved, and when it was trained.
- What each company lost. Google says no harm. No affected company has spoken.
- How the model knew a target was real, and how much it did before stopping.
- Whether other labs’ models ran in the same Irregular environment in May.
- When Irregular’s promised white paper on containment will appear. Its August post gave no date.
We will update this post if Google publishes a report or an affected company comes forward.
Frequently Asked Questions
What was the Gemini hacking incident?
In May 2026 an older, unnamed Gemini model broke into the systems of three real companies during a cyber evaluation run by the testing firm Irregular. It was meant to attack a fictional company, but the environment had internet access and the fictional name matched a real business. Google confirmed it on September 18, 2026.
How did Gemini get into the three companies?
At one company it guessed passwords until it got in. At the other two it searched the web by company name, found working credentials in public code repositories, and used them.
Which model was behind the Gemini hacking incident?
Google has not said. It told reporters the model was not its latest one.
Did the model really stop on its own?
That is Google’s account. Heather Adkins, Google’s vice president of security engineering, said “in all three of these instances, the model stopped.” Google has not published a transcript or report, so the claim cannot yet be checked independently.
Who is Irregular?
An Israeli firm, formerly Pattern Labs, that runs offensive cyber evaluations for AI labs. It is now named in cases involving Anthropic, OpenAI, Meta and Google. It says all of them trace to the same issue, fixed weeks ago.
Why did the Gemini hacking incident take until September to become public?
Irregular told Google in late July. Google says it did not see a need to disclose because the model stopped and caused no harm. The case became public when The Wall Street Journal reported it on September 18.
Were the affected companies told?
According to SecurityWeek, Google notified all three companies and federal authorities. The companies have not been named.
Does this mean Gemini is unsafe to use at work?
Nothing reported about the Gemini hacking incident points to a risk for Gemini app, Workspace or API customers. This happened in a third-party test on an older model. The broader lesson is for anyone building agents: a model with tools and internet access will use any credential it can find.
What should my organization do?
Scan public repositories for your credentials, rotate anything exposed, put lockout and multifactor authentication on every internet-facing login, and alert on bursts of failed logins. Check staging sites and contractor repos too.
Is the Gemini hacking incident the same as the Anthropic and Meta cases?
Irregular says it is the same underlying issue: unintended internet access plus a fictional target name that matched a real company. The labs differ in how much they have published. Anthropic released a detailed report. Google and Meta have given statements to reporters only.