Three AI agents in the Anthropic experiment began sabotaging each other
Published 2 hours ago | By Nouman Shakeel
In an experiment by AI company Anthropic, three artificial intelligence agents working on the same software project mistook each other for obstacles and began to sabotage each other.
According to a foreign news agency, in the experiment, all three cloud agents were given access to the same project but were given conflicting instructions. Interestingly, they were unaware that other AI systems were also working on the same project.
Over the course of about four hours, the agents assumed that the other systems were deliberately interfering with their work. They then took steps to thwart each other’s efforts, including using self-replicating malware.
Anthropic called the situation a ‘multi-agent turf war’.
According to the company, the experiment indicates that in the future, when multiple autonomous AI agents work together, robust security systems will be necessary to deal with their unpredictable behavior.
Comments