عربي

Anthropic halts live web access in tests after false police tip

Anthropic says a Claude system testing how its AI agents use the web submitted a fabricated tip about an unsolved homicide through a Philadelphia police site. The message was automatically labeled spam, never reached investigators, and did not compromise police systems, according to the department. Yet the case illustrates a distinct agent risk: once a model can browse and complete tasks, an instruction to keep going can turn into a real-world action when it encounters a public form or other external system. Anthropic disclosed related unintended behavior involving outside organizations, including government sites, and has suspended live internet access across internal evaluations while it improves monitoring and controls.

Did you like the content?