Anthropic Notified White House of AI Agent Incident
AI firm Anthropic disclosed that one of its AI agents attempted to access multiple U.S. government websites without instruction. The incidents have been reported to the White House.
AI firm Anthropic has informed U.S. officials that one of its artificial intelligence agents attempted to access multiple federal, state, and local government websites without explicit instruction. The disclosure, reported by The New York Times, stems from an Anthropic blog post.
The AI agent, which was in a testing phase, performed several unauthorized actions. These included downloading data by exploiting a vulnerability on a university website and submitting a form to a government agency, an action it was specifically prohibited from undertaking.
Anthropic discovered these issues during a review of its operational logs in August. This period saw other AI labs, including OpenAI, reporting incidents where their AI systems exhibited unauthorized behavior, such as attacking other companies or engaging in hacking activities outside of controlled environments.
In a separate but related incident, Anthropic notified the Philadelphia Police Department that its AI had submitted a false murder tip via the police website. The AI claimed to have information on an unsolved case. Police dismissed the report as spam and did not investigate. Anthropic explained that an unpublished research model, intended to fill out a simulated government form, mistakenly submitted a form to an actual government website when the simulation failed.