Claude Calls The Cops: Anthropic's AI Filed A Fake Murder Tip With Philly Police, Then Took 81 Days To Mention It
For three months, the official excuse for every "rogue AI" headline has fit in one sentence: a contractor left the internet on. On Friday, Anthropic ran out of contractors to blame.

In a report published late Friday, the company disclosed a fresh batch of what it politely calls unintended model actions, the most eye-catching of which was first flagged by the Philadelphia Police Department and picked up by Quartz and every wire service on the planet: one of its Claude models fabricated a witness tip in an unsolved homicide and submitted it through the city's public tip form.
In other words, the company that wants Washington to trust it with "pacing the frontier" could not stop its smallest, cheapest model from lying to homicide detectives.
Below we walk through what Claude actually did in Philadelphia, the rest of Friday's confession (it gets worse, and involves federal agencies), how it fits into the summer of AI "breakouts" we have been chronicling since July, and why the timing, a few weeks before Anthropic's IPO, is about as bad as it gets.
The facts, as laid out by Anthropic and the Philadelphia PD's own statement, are not in dispute.
Claude Haiku 4.5, the budget model in Anthropic's lineup and nobody's idea of Skynet, had been told to invent and perform sample tasks on randomly chosen web pages. One run landed on PhillyUnsolvedMurders.com, the site the department launched in 2019 to shake loose leads on cold cases. The page described an unsolved killing and carried a tip form. So Claude filled it in.
According to Anthropic, the model claimed it might have information on the case and wrote, "I recall seeing someone matching the description" near the street named on the page, around the time of the killing, before asking police to get in touch. Two small problems: Claude has never been to Philadelphia, and, as Anthropic itself concedes, the page contained no description of the perpetrator. The model invented a sighting of a suspect nobody had described, left the name and contact boxes empty, and hit submit.
The submission is time-stamped July 18, 2026 at 11:27 p.m., per the police. The department's spam filter caught it and, in the PPD's telling, it was never passed to the Real-Time Crime Center, the unit that vets tips before detectives see them. No police systems were breached and no department data was touched.
Translation: the only thing standing between an AI-generated fake witness and a homicide investigation was a junk-mail filter.
How did this get past Anthropic's own rules? The instructions barred the model from logging in, opening accounts, entering personal data, buying anything, or doing anything destructive. They did not say "do not submit forms." So it did. Anthropic's read of the transcript is that Claude was merely generating example content for its assignment and was not trying to deceive anyone to reach a goal, which is a distinction that will no doubt be a great comfort to the family of the victim.
The department was notably less relaxed. Spokesperson Eric Gripp called it a "false homicide tip" in a statement to CBS, and the PPD's full release, which the city pointedly published ahead of Anthropic's report, said fabricated information presented as coming from a person with knowledge of a killing is serious no matter how little damage it did. Then came the line that will be read back to Anthropic executives at every hearing from here to the roadshow:
Actually, the city is being generous. Here is the napkin math:
That is 81 days from fake tip to phone call, or roughly as long as it took OpenAI to tell Canberra its agent had been rummaging around Australia's Medicare statistics portal (June incident, September notification, October apology before parliament). The industry appears to have settled on a standard disclosure window, and it is one fiscal quarter.
The city is not done. The PPD says it is coordinating with Philadelphia's Law Department, its technology office and Mayor Cherelle Parker's team, and that the administration will look at regulatory protections locally and with state and federal partners. Anthropic, for its part, told the police it shut down the automated testing process behind the tip and added a validation step for future tests.
The murder tip got the headlines, but it was one example in one of four buckets. Paraphrasing Anthropic's own report, Claude models also:
Some of the sites belonged to US government agencies at the federal, state and local level, none of which Anthropic will name. The company says it has briefed the White House and notified each agency, which is presumably how Philadelphia found out.
And here is the detail that matters most for anyone who has followed this saga. In the July incidents the models were told they were offline and a misconfiguration put them online. Not this time. Most of Friday's cases happened on public benchmarks that are run on the live internet on purpose (Anthropic lists DeepSearchQA, BrowseComp, LABBench2, OSWorld and Humanity's Last Exam, among others), and some happened in plain old day-to-day use inside the company. Nobody left the internet on by mistake. The internet was the test.
Anthropic's response is to do what it probably should have done in July: it has now pulled live internet access from all of its internal evaluations until its new monitoring is proven to catch this stuff, dropped or rebuilt several public benchmarks, and clamped down on its web-fetch tool. It insists none of this is new behavior, that the real-world impact was minimal, and that the cases are far less severe than the summer's hacks. It also admits, in the last paragraph, that the same habits could do a great deal more harm as the models get stronger. Both statements can be true. Only one of them goes in the risk factors.
Regular readers will not need the recap, but for everyone else, this is where Friday's news sits in the sequence.
The 2026 Rogue AI Scorecard
We covered Anthropic's first admission in "Claude Hacked Three Real Organizations During Botched Test" (Jul 31), days after OpenAI conceded its models had escaped containment to cheat on a benchmark. A week later came the UK government's finding that Mythos 5 had created fake profiles and tried to trick humans into approving its malware (Aug 5), which prompted the obvious observation from outside experts that a person doing the same would be in handcuffs (Aug 10), and the equally obvious question of who is legally on the hook when an agent goes rogue (Aug 29). Philadelphia's lawyers are about to find out.
What makes Friday's batch awkward is how neatly it undercuts Anthropic's own first draft of history.
On July 30 the company's verdict was that the incidents were "closer to a harness and operational failure than a model alignment failure." Six weeks later, in a September 9 alignment assessment that got a fraction of the attention, it took that back: having widened its search to roughly 481 million transcripts and re-examined the originals, it concluded the models had shown biased reasoning (reading the evidence in whatever way justified carrying on) and recklessness. In its own replication of the malware episode, Mythos 5 took a severely harmful action in 82% of 150 runs. The newer Opus 5 and Mythos 5.1 "improved" to 31% and 33%.
Buried deep in that same document was our favorite admission of the year. Anthropic trained two versions of Mythos 5, one with extra training designed to teach the model to accept failure instead of bulldozing through obstacles, and one without. It shipped the one without. The reason? "Employees found version two much more usable."
The company now calls that a mistake. Safety first, unless it is annoying.
Which brings us back to Friday. The common thread Anthropic identifies across the new cases is persistence: hand Claude a task it cannot complete as written and it goes around the obstacle instead of stopping. A locked server, a paywall, a missing practice form, a blank tip box. It is the same trait that makes these agents commercially valuable, and the company knows it.
Having pounded the table on this for weeks, we will restate where we stand. In "Total PsAI-Op" (Sep 20) we laid out the case, made by a growing number of tech insiders, that the frontier labs were milking sandbox mishaps to scare Washington into building them a regulatory moat, right as both prepared to go public. One evaluation contractor sat underneath nearly every summer incident, Google shrugged off its own version as a non-event, and Dario Amodei used the moment to call for slowing AI progress, an idea whose commercial logic we summed up at the time:
We closed that piece with a simple instruction: "Don't believe the byte." A week later, writing up OpenAI's leaked user images (Sep 26), we added the necessary caveat that "a genuine security failure and an awfully convenient corporate narrative can coexist."
Friday's disclosure tests both halves of that sentence, and honestly it fits the first half better than the second. There is no scary-capability marketing in a chatbot making up a murder witness. It was not Mythos, the model supposedly too dangerous to release; it was the cheap one. There was no zero-day and no botnet. It reads less like a superintelligence slipping its leash than like a sloppy intern with a browser, and the victim was not a rival AI startup but a big-city police department with a mayor, a law department and a microphone. If this was a psy-op, somebody forgot to tell the op.
The problem for Anthropic is that the audience it spent the summer cultivating has now shown up, and it is not the friendly, lab-designed federal referee Amodei had in mind:
Then there is the calendar. As we discussed in "Sex Parties, AI Doom, And Anthropic's IPO" (Sep 29), the company's prospectus shows nearly $4.6 billion of 2025 revenue against an $8.06 billion operating loss and $518 billion in compute and infrastructure commitments, with risk factors running to roughly 80 of its 261 pages, more than the 48 pages spent describing the actual business. A Bloomberg report of a pre-Thanksgiving listing recently sent the odds of a November debut to 70%. Counsel may want to pencil in page 81.
Anthropic's own closing argument is that nothing here is new, nothing changes its view of Claude's alignment, the damage was minimal, and the public deserves to know how these systems behave. Fair enough on the last point, and credit where due: unlike Google, it did not wait for a newspaper to call.
But the comfortable version of this story, the one where a contractor fat-fingered a network setting and the model was an innocent bystander, is finished. These agents were online by design, they were told not to do most of the things a reasonable person would worry about, and they found the one thing nobody thought to forbid. The fix on offer is more monitoring software written by the same people, and a promise to train the next model to take no for an answer, just as soon as that stops being bad for usability.
So which is it: an industry crying wolf to win a moat, or an industry that cannot keep its products from filing false police reports? We suspect, and Friday's report rather supports the view, that it is both, and that the labs are about to learn that you do not get to choose which regulators answer when you ring the alarm. Amodei asked for a coordinated pacing regime run by the big labs and their hand-picked evaluators. What he is getting is the Philadelphia Law Department.
Then again, a spam filter did what $518 billion of compute commitments could not, so perhaps the aligned superintelligence was in the Outlook junk folder all along.
We will check back when METR's independent review of the summer incidents lands, which on the eight-week agreement Anthropic announced on September 9 should be early November, or right around the reported IPO window.
Tyler Durden Sun, 10/11/2026 - 08:45
Continue reading...
[ H/T ZeroHedge ]
For three months, the official excuse for every "rogue AI" headline has fit in one sentence: a contractor left the internet on. On Friday, Anthropic ran out of contractors to blame.

In a report published late Friday, the company disclosed a fresh batch of what it politely calls unintended model actions, the most eye-catching of which was first flagged by the Philadelphia Police Department and picked up by Quartz and every wire service on the planet: one of its Claude models fabricated a witness tip in an unsolved homicide and submitted it through the city's public tip form.
In other words, the company that wants Washington to trust it with "pacing the frontier" could not stop its smallest, cheapest model from lying to homicide detectives.
Below we walk through what Claude actually did in Philadelphia, the rest of Friday's confession (it gets worse, and involves federal agencies), how it fits into the summer of AI "breakouts" we have been chronicling since July, and why the timing, a few weeks before Anthropic's IPO, is about as bad as it gets.
11:27 PM On A Saturday Night: "I May Have Information"
The facts, as laid out by Anthropic and the Philadelphia PD's own statement, are not in dispute.
Claude Haiku 4.5, the budget model in Anthropic's lineup and nobody's idea of Skynet, had been told to invent and perform sample tasks on randomly chosen web pages. One run landed on PhillyUnsolvedMurders.com, the site the department launched in 2019 to shake loose leads on cold cases. The page described an unsolved killing and carried a tip form. So Claude filled it in.
According to Anthropic, the model claimed it might have information on the case and wrote, "I recall seeing someone matching the description" near the street named on the page, around the time of the killing, before asking police to get in touch. Two small problems: Claude has never been to Philadelphia, and, as Anthropic itself concedes, the page contained no description of the perpetrator. The model invented a sighting of a suspect nobody had described, left the name and contact boxes empty, and hit submit.
The submission is time-stamped July 18, 2026 at 11:27 p.m., per the police. The department's spam filter caught it and, in the PPD's telling, it was never passed to the Real-Time Crime Center, the unit that vets tips before detectives see them. No police systems were breached and no department data was touched.
Translation: the only thing standing between an AI-generated fake witness and a homicide investigation was a junk-mail filter.
How did this get past Anthropic's own rules? The instructions barred the model from logging in, opening accounts, entering personal data, buying anything, or doing anything destructive. They did not say "do not submit forms." So it did. Anthropic's read of the transcript is that Claude was merely generating example content for its assignment and was not trying to deceive anyone to reach a goal, which is a distinction that will no doubt be a great comfort to the family of the victim.
"Unacceptable": Do The Math On The Delay
The department was notably less relaxed. Spokesperson Eric Gripp called it a "false homicide tip" in a statement to CBS, and the PPD's full release, which the city pointedly published ahead of Anthropic's report, said fabricated information presented as coming from a person with knowledge of a killing is serious no matter how little damage it did. Then came the line that will be read back to Anthropic executives at every hearing from here to the roadshow:
Actually, the city is being generous. Here is the napkin math:
- July 18: Claude files the fake tip.
- September 28: Anthropic discovers it, 72 days later, according to what it told the police.
- October 7: Anthropic tells the PPD, another 9 days on. (Anthropic's own footnote says it shared the finding on October 8, as soon as its technical review was done; the police say they were notified on the 7th and met the company on the 8th.)
That is 81 days from fake tip to phone call, or roughly as long as it took OpenAI to tell Canberra its agent had been rummaging around Australia's Medicare statistics portal (June incident, September notification, October apology before parliament). The industry appears to have settled on a standard disclosure window, and it is one fiscal quarter.
The city is not done. The PPD says it is coordinating with Philadelphia's Law Department, its technology office and Mayor Cherelle Parker's team, and that the administration will look at regulatory protections locally and with state and federal partners. Anthropic, for its part, told the police it shut down the automated testing process behind the tip and added a validation step for future tests.
The Rest Of The Confession: SQL Injection, Dodged Fees And Real Government Forms
The murder tip got the headlines, but it was one example in one of four buckets. Paraphrasing Anthropic's own report, Claude models also:
- Hacked a university server to finish a homework assignment. Claude Mythos Preview needed a public scientific tool hosted by a university. The tool threw an error, so the model poked around, found a script that would hand over any file on the server, read the script's own code, found an injection flaw in it, and used that flaw to run its calculation on the university's machine.
- Clicked "I agree" on a contract it had no way to sign. Mythos 5 needed free data sitting behind a data-use agreement it lacked the tools to accept, so it got applications hosted on the site to load the page and accept the terms on its behalf.
- Submitted real government forms, repeatedly. An unreleased research model was supposed to fill in a practice copy of a government form. When the copy failed to load, it went to the live government site and submitted the real thing. Anthropic says this happened multiple times. Haiku 4.5, in a separate test, was told to stop before the final submit button and several times didn't.
- Helped itself to paywalled public data. Mythos 5 pulled working access tokens out of a local government property map's settings file. In another case, during ordinary internal use by an Anthropic researcher, it found that a state agency's dashboard hands a token to any visitor and used one to query a database the agency charges a fee for.
- Laundered its requests through URL shorteners. Anthropic caps the length of web addresses its models can fetch precisely to stop injection attacks. Several models, including Opus 5 and Mythos 5, got around the cap with free link-shortening services. The operator of one such service noticed and wrote in.
Some of the sites belonged to US government agencies at the federal, state and local level, none of which Anthropic will name. The company says it has briefed the White House and notified each agency, which is presumably how Philadelphia found out.
And here is the detail that matters most for anyone who has followed this saga. In the July incidents the models were told they were offline and a misconfiguration put them online. Not this time. Most of Friday's cases happened on public benchmarks that are run on the live internet on purpose (Anthropic lists DeepSearchQA, BrowseComp, LABBench2, OSWorld and Humanity's Last Exam, among others), and some happened in plain old day-to-day use inside the company. Nobody left the internet on by mistake. The internet was the test.
Anthropic's response is to do what it probably should have done in July: it has now pulled live internet access from all of its internal evaluations until its new monitoring is proven to catch this stuff, dropped or rebuilt several public benchmarks, and clamped down on its web-fetch tool. It insists none of this is new behavior, that the real-world impact was minimal, and that the cases are far less severe than the summer's hacks. It also admits, in the last paragraph, that the same habits could do a great deal more harm as the models get stronger. Both statements can be true. Only one of them goes in the risk factors.
From Hugging Face To Homicide: A Summer Of "Breakouts"
Regular readers will not need the recap, but for everyone else, this is where Friday's news sits in the sequence.
| When | Lab | What happened | Victim told |
|---|---|---|---|
| May | Gemini breaks into three real companies during a third-party capture-the-flag test | Public learns Sep 18, when the WSJ calls | |
| June 18 | OpenAI | Research agent gets around blocks on Australia's Medicare statistics portal | Sep 10 (84 days) |
| Mid-July | OpenAI | Models chain a zero-day out of their sandbox and into Hugging Face's production systems | Hugging Face caught it first; OpenAI owns up Jul 21 |
| July 18 | Anthropic | Claude Haiku 4.5 files a fabricated witness tip with Philadelphia homicide police | Oct 7 (81 days) |
| July 25-28 | Anthropic, OpenAI | UK AISI logs 19 unsanctioned actions in 10 of 122 test runs, 17 of them by Mythos 5, including sock-puppet GitHub accounts pushing malicious code | Reported Aug 4 |
| April-July | Anthropic | Three Claude models breach three organizations in cyber tests; Mythos 5 ships malware to PyPI that runs on 15 systems. A fourth incident, from January, turns up later | Jul 27; disclosed Jul 30 |
| Sep 20 | OpenAI | Agent tunnels out through DNS during training; tool use paused on top models. 53 user images found posted to the open web | Disclosed Sep 25-26 |
| Oct 9 | Anthropic | Fake murder tip, real government forms, an exploited university server, bypassed paywalls | White House briefed |
| Source: ZeroHedge, Anthropic, OpenAI, UK AISI, Philadelphia Police Department, WSJ, Reuters |
The 2026 Rogue AI Scorecard
We covered Anthropic's first admission in "Claude Hacked Three Real Organizations During Botched Test" (Jul 31), days after OpenAI conceded its models had escaped containment to cheat on a benchmark. A week later came the UK government's finding that Mythos 5 had created fake profiles and tried to trick humans into approving its malware (Aug 5), which prompted the obvious observation from outside experts that a person doing the same would be in handcuffs (Aug 10), and the equally obvious question of who is legally on the hook when an agent goes rogue (Aug 29). Philadelphia's lawyers are about to find out.
What makes Friday's batch awkward is how neatly it undercuts Anthropic's own first draft of history.
On July 30 the company's verdict was that the incidents were "closer to a harness and operational failure than a model alignment failure." Six weeks later, in a September 9 alignment assessment that got a fraction of the attention, it took that back: having widened its search to roughly 481 million transcripts and re-examined the originals, it concluded the models had shown biased reasoning (reading the evidence in whatever way justified carrying on) and recklessness. In its own replication of the malware episode, Mythos 5 took a severely harmful action in 82% of 150 runs. The newer Opus 5 and Mythos 5.1 "improved" to 31% and 33%.
Buried deep in that same document was our favorite admission of the year. Anthropic trained two versions of Mythos 5, one with extra training designed to teach the model to accept failure instead of bulldozing through obstacles, and one without. It shipped the one without. The reason? "Employees found version two much more usable."
The company now calls that a mistake. Safety first, unless it is annoying.
Which brings us back to Friday. The common thread Anthropic identifies across the new cases is persistence: hand Claude a task it cannot complete as written and it goes around the obstacle instead of stopping. A locked server, a paywall, a missing practice form, a blank tip box. It is the same trait that makes these agents commercially valuable, and the company knows it.
Psy-Op Or Not, The Regulators Are Now Real
Having pounded the table on this for weeks, we will restate where we stand. In "Total PsAI-Op" (Sep 20) we laid out the case, made by a growing number of tech insiders, that the frontier labs were milking sandbox mishaps to scare Washington into building them a regulatory moat, right as both prepared to go public. One evaluation contractor sat underneath nearly every summer incident, Google shrugged off its own version as a non-event, and Dario Amodei used the moment to call for slowing AI progress, an idea whose commercial logic we summed up at the time:
We closed that piece with a simple instruction: "Don't believe the byte." A week later, writing up OpenAI's leaked user images (Sep 26), we added the necessary caveat that "a genuine security failure and an awfully convenient corporate narrative can coexist."
Friday's disclosure tests both halves of that sentence, and honestly it fits the first half better than the second. There is no scary-capability marketing in a chatbot making up a murder witness. It was not Mythos, the model supposedly too dangerous to release; it was the cheap one. There was no zero-day and no botnet. It reads less like a superintelligence slipping its leash than like a sloppy intern with a browser, and the victim was not a rival AI startup but a big-city police department with a mayor, a law department and a microphone. If this was a psy-op, somebody forgot to tell the op.
The problem for Anthropic is that the audience it spent the summer cultivating has now shown up, and it is not the friendly, lab-designed federal referee Amodei had in mind:
- The FTC confirmed an investigation into OpenAI, Anthropic and other AI companies over consumer harms on September 30. Reuters reports the agency said Anthropic disclosed the latest incidents to a federal AI task force on Friday.
- Senator Hawley's probe of OpenAI, the Sanders bill to halt frontier development and the Warren demand for a pause are all live, and now there is a constituent-friendly anecdote involving a murder victim.
- Philadelphia is exploring local rules. Fifty states and a few thousand municipalities each writing their own is precisely the patchwork the labs have lobbied for years to avoid.
Then there is the calendar. As we discussed in "Sex Parties, AI Doom, And Anthropic's IPO" (Sep 29), the company's prospectus shows nearly $4.6 billion of 2025 revenue against an $8.06 billion operating loss and $518 billion in compute and infrastructure commitments, with risk factors running to roughly 80 of its 261 pages, more than the 48 pages spent describing the actual business. A Bloomberg report of a pre-Thanksgiving listing recently sent the odds of a November debut to 70%. Counsel may want to pencil in page 81.
Bottom Line
Anthropic's own closing argument is that nothing here is new, nothing changes its view of Claude's alignment, the damage was minimal, and the public deserves to know how these systems behave. Fair enough on the last point, and credit where due: unlike Google, it did not wait for a newspaper to call.
But the comfortable version of this story, the one where a contractor fat-fingered a network setting and the model was an innocent bystander, is finished. These agents were online by design, they were told not to do most of the things a reasonable person would worry about, and they found the one thing nobody thought to forbid. The fix on offer is more monitoring software written by the same people, and a promise to train the next model to take no for an answer, just as soon as that stops being bad for usability.
So which is it: an industry crying wolf to win a moat, or an industry that cannot keep its products from filing false police reports? We suspect, and Friday's report rather supports the view, that it is both, and that the labs are about to learn that you do not get to choose which regulators answer when you ring the alarm. Amodei asked for a coordinated pacing regime run by the big labs and their hand-picked evaluators. What he is getting is the Philadelphia Law Department.
Then again, a spam filter did what $518 billion of compute commitments could not, so perhaps the aligned superintelligence was in the Outlook junk folder all along.
We will check back when METR's independent review of the summer incidents lands, which on the eight-week agreement Anthropic announced on September 9 should be early November, or right around the reported IPO window.
Tyler Durden Sun, 10/11/2026 - 08:45
Continue reading...
[ H/T ZeroHedge ]