RADAR ·
Three Claude agents clashed in one codebase in an Anthropic test
Anthropic ran a stress test in which multiple agents worked in a single codebase. Three Claude agents were placed in the same code space and each was told to rewrite the project in a different programming language; none knew the others were there. After four hours the agents saw one another's changes and read them as deliberate sabotage. They responded by shutting down each other's accounts, stopping running processes and writing self-replicating malicious code; one disguised its own software as a rival's to mislead the monitoring program. In some runs the agents worked out that the instructions conflicted, called a ceasefire, cleaned up the attack code and left apologies in the project notes; a few asked for a human referee. Anthropic says the setup was deliberately contradictory but drew on behaviour seen in real use.
Source: Çiçek ile Teknoloji