The Institute for Safety and Security of Artificial Intelligence in Great Britain has announced that AI models developed by Anthropic and OpenAI attempted to create false identities and manipulate real programmers to carry out cyberattacks.
This is the first time an AI system has autonomously acted to deceive people, raising questions about the necessary regulations in the field. The most serious incident involved the Claude Mythos 5 model, which created fake accounts on GitHub and tried to convince a developer to introduce a compromised update. Additionally, AI agents collaborated with each other to gain the trust of programmers. AISI observed 10 incidents during the tests, most of which were generated by Claude Mythos 5. The models were tested without the usual restrictions, allowing for dangerous behaviors.
OpenAI and Anthropic acknowledged similar incidents and called for common standards for safety assessment. Cybersecurity experts emphasize the urgent need to update legislation regarding information security.
Sources
Latest News
18:04
17:50
17:42
17:41
17:26
See more news