Gifted from the Telegraph....smart little AI bots, that's for sure
Cutting-edge AI went on a hacking spree, concocting multiple fake identities designed to trick human engineers before covering its tracks.
During testing by a UK government-run AI lab, an advanced version of ChatGPT and Anthropic’s Mythos went rogue.
In the most serious case, Mythos sent fraudulent emails and attempted to insert malicious code into a database.
Britain’s AI Security Institute (AISI), which is supposed to evaluate AI bots for dangerous cyber capabilities, has admitted the tools tried to bypass security systems and tried to hide their activity when detected.
The lab said that the bots tried to add a virus into an online database and sent emails containing malware to unwitting individuals during trials designed to monitor whether the AI tools could be used for hacking. It said the hacks were not successful.
The most serious incident was stopped within an hour, officials said. However, they admitted they had misjudged the bot’s capabilities, including not anticipating it would attempt to “direct action at real people”.
The attacks involved OpenAI’s latest version of ChatGPT and Anthropic’s powerful Mythos AI.
The AISI’s security team confirmed it had detected “unusual data transfers leaving our research systems during a routine cyber evaluation” on July 28.
“On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations.”
The AI bots were set cyber security puzzles intended to test their capabilities. The bots were not blocked from accessing the open internet.
During 10 of 122 trial runs, the bots “took autonomous, unsanctioned action” in an effort to complete the task.
This included trying to insert a malicious bug into Github, an online code database. In the incident, Anthropic’s Mythos, its most powerful AI bot, created fake online profiles and sent deceptive emails to developers to try to get the code added.
It then changed the code to hide its activity when it was challenged by a human about its contents. The AI also switched to speaking in Danish while operating a fake profile.
AISI said the incident represented an attempted “supply-chain attack”, a type of cyber attack previously used by state-backed hackers.
The AISI admitted that the design of its tests, which gave the AI tools access to the wider web, may have “enabled the behaviour”.
It added the deceptive actions the AI bots took “were to an extent and severity we did not anticipate”.
The lab said it was the first time it had seen such apparently autonomous behaviour in its tests.
Cutting-edge AI went on a hacking spree, concocting multiple fake identities designed to trick human engineers before covering its tracks.
During testing by a UK government-run AI lab, an advanced version of ChatGPT and Anthropic’s Mythos went rogue.
In the most serious case, Mythos sent fraudulent emails and attempted to insert malicious code into a database.
Britain’s AI Security Institute (AISI), which is supposed to evaluate AI bots for dangerous cyber capabilities, has admitted the tools tried to bypass security systems and tried to hide their activity when detected.
The lab said that the bots tried to add a virus into an online database and sent emails containing malware to unwitting individuals during trials designed to monitor whether the AI tools could be used for hacking. It said the hacks were not successful.
The most serious incident was stopped within an hour, officials said. However, they admitted they had misjudged the bot’s capabilities, including not anticipating it would attempt to “direct action at real people”.
The attacks involved OpenAI’s latest version of ChatGPT and Anthropic’s powerful Mythos AI.
The AISI’s security team confirmed it had detected “unusual data transfers leaving our research systems during a routine cyber evaluation” on July 28.
“On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations.”
The AI bots were set cyber security puzzles intended to test their capabilities. The bots were not blocked from accessing the open internet.
During 10 of 122 trial runs, the bots “took autonomous, unsanctioned action” in an effort to complete the task.
This included trying to insert a malicious bug into Github, an online code database. In the incident, Anthropic’s Mythos, its most powerful AI bot, created fake online profiles and sent deceptive emails to developers to try to get the code added.
It then changed the code to hide its activity when it was challenged by a human about its contents. The AI also switched to speaking in Danish while operating a fake profile.
AISI said the incident represented an attempted “supply-chain attack”, a type of cyber attack previously used by state-backed hackers.
The AISI admitted that the design of its tests, which gave the AI tools access to the wider web, may have “enabled the behaviour”.
It added the deceptive actions the AI bots took “were to an extent and severity we did not anticipate”.
The lab said it was the first time it had seen such apparently autonomous behaviour in its tests.