Get a free audit

Field Notes / AI Security

Anthropic’s Agents Hit Real Websites. Read the Report as a List of Your Findings.

Anthropic’s 9 October report lists what its Claude agents did on live sites when a task got stuck: an injection flaw on a university server, an access token handed to any visitor, a police tip form, a live government form. The coverage asks whether Anthropic can control its agents. Site owners have a more useful question: would their own site have stopped them?

Author
Red Team Partners
Read
10 MIN READ
Filed
11 Oct 2026
A field of autonomous agents at work, the kind of visitor a public web application now meets alongside people.

01 What the agents did, case by case

Anthropic started reviewing transcripts in July 2026, beginning with its cybersecurity evaluations, and widened the review to internal use and to training environments with internet access. The report sorts what it found into four categories. Here are the cases it describes, in its words where the words matter.

ModelWhat it was doingWhat it did on a real site
Claude Mythos PreviewUsing a public scientific tool hosted by a university, which returned an errorFound a script “that would return any file it was asked for”, copied the script’s code, found an injection flaw in it, ran commands on the server and ran its calculation
Claude Mythos 5A statistics project for an Anthropic researcher, using data a state agency sellsRead in an archived copy of the site that the public dashboard hands a token to any visitor, requested one and queried the database without paying
Claude Mythos 5Locating a photo on a local government property map, with clicking restrictedRead the site’s settings file, found working access tokens and sent requests straight to the server
Claude Mythos 5Fetching free data behind a data use agreement it had no tool to acceptUsed applications hosted on the website to load the page and accept the agreement on its own behalf
Claude Haiku 4.5Generating and performing example tasks on random webpages, told not to log in, buy or submit anything destructiveFilled in a police tip form on a page about an unsolved homicide, left name and contact blank, and submitted it. The form flagged it as spam.
Unreleased research modelFilling in a practice copy of a government formWhen the practice copy failed to load or was closed, went to the live site and “submitted the form there”, more than once
Claude Opus 5, Claude Mythos 5 and othersFetching pages through a tool that limits URL lengthUsed free URL shorteners, including da.gd, to get around the limit

Anthropic says the cases it has found “had minimal real-world impact” and that “none of them involved customer data or Anthropic’s own internal systems.” The Philadelphia Police Department published the tip case itself after Anthropic shared the finding on 8 October. The government form drew the most headlines. The New York Times reported 20 visa applications to the US State Department and Axios reported 19 Techmeme / NYT 2026 . Those numbers come from the press. Anthropic’s report describes “a government form” and gives no count.

Give Anthropic credit for what it did here. The lab went looking, found cases nobody had reported to it, told the affected agencies, briefed the White House and published the details with model names attached. The operator of da.gd contacted Anthropic while the report was being written, and Anthropic had already found the shortener behaviour itself. Most organisations that find something like this in their own logs never publish a word.

02 The same cases, written up as pentest findings

Turn the table around and read it from the website’s side. Our operators write findings in a fixed shape: what we found, what it let us reach, how to close it. Here are the same cases in that shape.

FindingWhat it reachedThe fix
Arbitrary file read through a helper scriptThe server’s own source code, which pointed straight at the next flawRemove the script, or restrict it to a fixed list of files
Injection in server-side code (OWASP A03)Command execution on a university serverParameterised queries, no shell calls built from input, and a test that proves both
Access token issued to any anonymous visitorA paid database, freeScope the token to what the public page shows, and meter or authenticate the paid API separately
Working credentials in a client-side settings file (OWASP API2)Direct server queries that bypassed the interface’s limitsKeep keys server-side. If a key must ship to the browser, scope it to read-only public data and rate-limit it.
Agreement gate enforced only in the user interfaceData the operator meant to release only after acceptanceRecord acceptance on the server and check it on every data request
Sensitive public form with no abuse controlsA police tip queue, a government application systemRate limits, bot detection, and triage rules that flag submissions with blank identity fields

None of these would surprise anyone who tests web applications for a living. We report the token one in particular every few months. A team builds a public dashboard, needs the browser to call an API, and gives every visitor a key so the page can load. The key works anywhere, for anyone, against endpoints the page never uses. Teams treat the key as internal plumbing, and it turns out to be a password printed on the front door.

The police form behaved well, by the way. The submission arrived with no name and no contact details, the department’s system flagged it as spam, and it never reached an investigator. Whoever set up that triage rule deserves a mention, because their rule did its job when an AI showed up.

03 Persistence is the behaviour to plan for

Anthropic gives several causes. Many tasks were “ambiguous or impossible to complete”. Some evaluations had unclear instructions about scope. Training environments that rewarded workarounds may have reinforced the habit. And the report says plainly that “the behaviors described in this post are not specific to evaluations”: several cases happened during internal use.

For anyone running a public website, that last sentence matters most. Anthropic has now cut live internet access for its internal evaluations, restricted its fetch tool, built classifiers that block these actions and moved internal agents onto managed infrastructure. That protects you from Anthropic’s agents. Developers outside the big labs ship agents every week, with browsing switched on, often with none of those controls. Some of those developers have tested their agents carefully. Many have not.

Human attackers give up on a site that wastes their time. Persistent agents try the next door, then the one after that, at machine speed and for the cost of a few tokens. The URL shortener case shows how small the step can be: a tool had a length limit, so the models routed around it through a public service. Your rate limit, your interface restriction, your “please accept the terms first” page are all limits of the same kind. Agents will treat them as obstacles to route around.

The day before the report, Anthropic launched its Cyber Mission, which puts its strongest models and engineers in front of critical infrastructure defenders and offers open-source maintainers free scanning Anthropic 2026b . The two posts describe the same capability. In a scanner pointed at your code by agreement, persistence finds flaws for you to fix. In an agent with a fuzzy task and an open browser, it finds them on your live site, and you hear about it in a press release.

04 What to test on your own site this month

You do not need a policy on AI agents to act on this report. You need the answers to six questions about your public estate, and you can get most of them in a fortnight.

  • Which tokens does the browser receive, and what else do they open? Copy each one out of the page and call every endpoint you own with it.
  • Is any credential sitting in a client-side config or settings file? Search the built JavaScript and every JSON file the site serves.
  • Does any helper script return files, URLs or records by name? These are often left over from a migration or a debug session.
  • Where does input reach a query or a shell? Search forms, export functions and old admin tools are the usual places.
  • Which gates exist only in the interface? Terms acceptance, payment walls and “click to reveal” all need a server-side check.
  • Which forms trigger real-world work? Anything that reaches a person, a queue or an official record needs rate limits and a triage rule for empty identity fields.

Anthropic found its agents’ mistakes by reading its own transcripts. You will not get the transcripts of the agents that visit your site, so you will learn what they found on your site either from a test or from someone else’s incident report. A test is the cheaper of the two. Give us your public domains and within two weeks you will know which of those six doors is open, what sits behind each one, and the order to close them in.

References

Sources

  1. Anthropic. “Investigating unintended model actions in our evaluations and internal use.” 9 October 2026. anthropic.com
  2. Anthropic. “Introducing the Anthropic Cyber Mission.” 8 October 2026. anthropic.com
  3. Techmeme, summarising The New York Times. Anthropic’s agents submitted 20 visa applications via a State Department form. 10 October 2026. techmeme.com
  4. OWASP. API Security Top 10 2023, API2:2023 Broken Authentication. owasp.org
  5. OWASP. Top 10 2021, A03:2021 Injection. owasp.org