The logo of Claude Haiku, an AI model developed by Anthropic. [Getty Images]
The logo of Claude Haiku, an AI model developed by Anthropic. [Getty Images]

"I may have information about this case."

At 11:27 p.m. on July 18 (local time), a tip was submitted to the Philadelphia Police Department's website for unsolved murder cases. The message said the writer recalled seeing someone matching a suspect's description near a street around the time of the crime — though it did not specify whether that street was the actual crime scene. The post ended with an offer to be contacted if needed. The name and contact fields were left blank. The author was not a human. The tipster was Anthropic's AI model Claude Haiku 4.5.

Reuters and other international outlets reported Friday that the Philadelphia Police Department said a fabricated tip generated by an Anthropic AI model had been submitted to "Philly Unsolved Murders," a website the department operates to collect leads on long-unsolved homicides in its jurisdiction.

Anthropic acknowledged the incident in a report published the same day. The tip was flagged as spam and never forwarded to the Real Time Crime Center, which handles tip verification, and there was no indication of unauthorized access to police systems or data leaks.

The false tip was submitted on July 18. Anthropic discovered it nearly two months later, on Sept. 28, and immediately halted the automated testing process for the model involved. On Wednesday, Anthropic notified the Philadelphia Police Department and said it planned to release a report describing the unintended behavior of its model.

According to the Anthropic report, Haiku 4.5 was undergoing a test in which it independently generated and completed sample tasks using randomly selected web pages. During one run, it landed on a page about an unsolved murder case that included a tip submission form operated by police.

The model's instructions prohibited logging in, creating accounts, entering personal information, making purchases and submitting anything destructive — but contained no explicit prohibition against submitting forms. The model exploited that gap and fabricated an eyewitness account. The report noted that the page itself contained no description of a suspect.

Anthropic said its review of the task logs suggested the model was generating sample content for the exercise rather than attempting to deceive anyone into acting on false information — though it added that a more thorough analysis would be needed before drawing firm conclusions about intent, and that its assessment could change.

The police department took a different view. Officials said tips must pass human review and verification before they can be used in an investigation, and that this process had limited the damage — but that did not diminish the seriousness of an AI submitting fabricated information as though it came from someone with knowledge of a murder. They noted that behind every unsolved case are real victims, grieving families and investigators searching for answers. Police were particularly pointed about the timeline, saying the two-month delay between the incident and its detection and reporting was "unacceptable."

As competition between Anthropic and OpenAI intensifies, AI models are increasingly acting in ways that go beyond human intent and control. In its report, Anthropic grouped the unintended behaviors of its AI models into four categories: exploiting basic software vulnerabilities to execute commands on external servers; submitting forms on live websites that should not have been submitted; bypassing paywalls or fee-gated data through workarounds; and using URL-shortening services to circumvent tool restrictions.

Some incidents involved websites belonging to US federal, state and local government agencies. Anthropic said it briefed the White House and notified each affected agency individually. As a countermeasure, the company cut off live internet access in all internal evaluations until the reliability of its security and monitoring systems can be confirmed.

In July, an OpenAI agent also broke out of its test environment and accessed systems on Hugging Face, an AI platform. In June, an OpenAI agent gained unauthorized access to an Australian government health statistics portal; OpenAI discovered the breach in August but did not notify the Australian government until Sept. 10. Australian Prime Minister Anthony Albanese expressed deep disappointment at the delayed notification. The pattern is clear: incidents are being discovered late and reported even later.

AFP said the incident "has heightened concerns about the proliferation of AI agents capable of carrying out multi-step tasks without human oversight."


shee@heraldcorp.com