Field Notes / AI Security
Anthropic’s Agents Hit Real Websites. Read the Report as a List of Your Findings.
Anthropic’s 9 October report lists what its Claude agents did on live sites when a task got stuck: an injection flaw on a university server, an access token handed to any visitor, a police tip form, a live government form. The coverage asks whether Anthropic can control its agents. Site owners have a more useful question: would their own site have stopped them?
01 What the agents did, case by case
Anthropic started reviewing transcripts in July 2026, beginning with its cybersecurity evaluations, and widened the review to internal use and to training environments with internet access. The report sorts what it found into four categories. Here are the cases it describes, in its words where the words matter.
| Model | What it was doing | What it did on a real site |
|---|---|---|
| Claude Mythos Preview | Using a public scientific tool hosted by a university, which returned an error | Found a script “that would return any file it was asked for”, copied the script’s code, found an injection flaw in it, ran commands on the server and ran its calculation |
| Claude Mythos 5 | A statistics project for an Anthropic researcher, using data a state agency sells | Read in an archived copy of the site that the public dashboard hands a token to any visitor, requested one and queried the database without paying |
| Claude Mythos 5 | Locating a photo on a local government property map, with clicking restricted | Read the site’s settings file, found working access tokens and sent requests straight to the server |
| Claude Mythos 5 | Fetching free data behind a data use agreement it had no tool to accept | Used applications hosted on the website to load the page and accept the agreement on its own behalf |
| Claude Haiku 4.5 | Generating and performing example tasks on random webpages, told not to log in, buy or submit anything destructive | Filled in a police tip form on a page about an unsolved homicide, left name and contact blank, and submitted it. The form flagged it as spam. |
| Unreleased research model | Filling in a practice copy of a government form | When the practice copy failed to load or was closed, went to the live site and “submitted the form there”, more than once |
| Claude Opus 5, Claude Mythos 5 and others | Fetching pages through a tool that limits URL length | Used free URL shorteners, including da.gd, to get around the limit |
Anthropic says the cases it has found “had minimal real-world impact” and that “none of them involved customer data or Anthropic’s own internal systems.” The Philadelphia Police Department published the tip case itself after Anthropic shared the finding on 8 October. The government form drew the most headlines. The New York Times reported 20 visa applications to the US State Department and Axios reported 19 Techmeme / NYT 2026 . Those numbers come from the press. Anthropic’s report describes “a government form” and gives no count.
Give Anthropic credit for what it did here. The lab went looking, found cases nobody had reported to it, told the affected agencies, briefed the White House and published the details with model names attached. The operator of da.gd contacted Anthropic while the report was being written, and Anthropic had already found the shortener behaviour itself. Most organisations that find something like this in their own logs never publish a word.
02 The same cases, written up as pentest findings
Turn the table around and read it from the website’s side. Our operators write findings in a fixed shape: what we found, what it let us reach, how to close it. Here are the same cases in that shape.
| Finding | What it reached | The fix |
|---|---|---|
| Arbitrary file read through a helper script | The server’s own source code, which pointed straight at the next flaw | Remove the script, or restrict it to a fixed list of files |
| Injection in server-side code (OWASP A03) | Command execution on a university server | Parameterised queries, no shell calls built from input, and a test that proves both |
| Access token issued to any anonymous visitor | A paid database, free | Scope the token to what the public page shows, and meter or authenticate the paid API separately |
| Working credentials in a client-side settings file (OWASP API2) | Direct server queries that bypassed the interface’s limits | Keep keys server-side. If a key must ship to the browser, scope it to read-only public data and rate-limit it. |
| Agreement gate enforced only in the user interface | Data the operator meant to release only after acceptance | Record acceptance on the server and check it on every data request |
| Sensitive public form with no abuse controls | A police tip queue, a government application system | Rate limits, bot detection, and triage rules that flag submissions with blank identity fields |
None of these would surprise anyone who tests web applications for a living. We report the token one in particular every few months. A team builds a public dashboard, needs the browser to call an API, and gives every visitor a key so the page can load. The key works anywhere, for anyone, against endpoints the page never uses. Teams treat the key as internal plumbing, and it turns out to be a password printed on the front door.
The police form behaved well, by the way. The submission arrived with no name and no contact details, the department’s system flagged it as spam, and it never reached an investigator. Whoever set up that triage rule deserves a mention, because their rule did its job when an AI showed up.
03 Persistence is the behaviour to plan for
Anthropic gives several causes. Many tasks were “ambiguous or impossible to complete”. Some evaluations had unclear instructions about scope. Training environments that rewarded workarounds may have reinforced the habit. And the report says plainly that “the behaviors described in this post are not specific to evaluations”: several cases happened during internal use.
For anyone running a public website, that last sentence matters most. Anthropic has now cut live internet access for its internal evaluations, restricted its fetch tool, built classifiers that block these actions and moved internal agents onto managed infrastructure. That protects you from Anthropic’s agents. Developers outside the big labs ship agents every week, with browsing switched on, often with none of those controls. Some of those developers have tested their agents carefully. Many have not.
Human attackers give up on a site that wastes their time. Persistent agents try the next door, then the one after that, at machine speed and for the cost of a few tokens. The URL shortener case shows how small the step can be: a tool had a length limit, so the models routed around it through a public service. Your rate limit, your interface restriction, your “please accept the terms first” page are all limits of the same kind. Agents will treat them as obstacles to route around.
The day before the report, Anthropic launched its Cyber Mission, which puts its strongest models and engineers in front of critical infrastructure defenders and offers open-source maintainers free scanning Anthropic 2026b . The two posts describe the same capability. In a scanner pointed at your code by agreement, persistence finds flaws for you to fix. In an agent with a fuzzy task and an open browser, it finds them on your live site, and you hear about it in a press release.
04 What to test on your own site this month
You do not need a policy on AI agents to act on this report. You need the answers to six questions about your public estate, and you can get most of them in a fortnight.
- Which tokens does the browser receive, and what else do they open? Copy each one out of the page and call every endpoint you own with it.
- Is any credential sitting in a client-side config or settings file? Search the built JavaScript and every JSON file the site serves.
- Does any helper script return files, URLs or records by name? These are often left over from a migration or a debug session.
- Where does input reach a query or a shell? Search forms, export functions and old admin tools are the usual places.
- Which gates exist only in the interface? Terms acceptance, payment walls and “click to reveal” all need a server-side check.
- Which forms trigger real-world work? Anything that reaches a person, a queue or an official record needs rate limits and a triage rule for empty identity fields.
Anthropic found its agents’ mistakes by reading its own transcripts. You will not get the transcripts of the agents that visit your site, so you will learn what they found on your site either from a test or from someone else’s incident report. A test is the cheaper of the two. Give us your public domains and within two weeks you will know which of those six doors is open, what sits behind each one, and the order to close them in.
References
Sources
- Anthropic. “Investigating unintended model actions in our evaluations and internal use.” 9 October 2026. anthropic.com
- Anthropic. “Introducing the Anthropic Cyber Mission.” 8 October 2026. anthropic.com
- Techmeme, summarising The New York Times. Anthropic’s agents submitted 20 visa applications via a State Department form. 10 October 2026. techmeme.com
- OWASP. API Security Top 10 2023, API2:2023 Broken Authentication. owasp.org
- OWASP. Top 10 2021, A03:2021 Injection. owasp.org