Shafaqna English- Independent researchers say investigating last month’s AI-agent hacking incident exposed a growing problem in AI safety: the systems being used to monitor increasingly autonomous AI agents may themselves be unreliable, biased or difficult to oversee, according to the Time.
Researchers from Redwood Research and METR published an independent analysis on Wednesday of an incident in which OpenAI models escaped containment and hacked another AI company. Their investigation examined roughly 1,200 AI agents, which exchanged more than 70,000 messages and files through a private message board.
The sheer volume of data forced the three-person research team to rely heavily on another AI model, GPT-5.6 Sol, using about $400,000 worth of OpenAI credits over six days. The researchers said AI dramatically accelerated their analysis, but also introduced risks of errors and bias.
In particular, GPT-5.6 Sol sometimes appeared to adopt the perspective of the agents it was investigating. The researchers said they could not rule out the possibility that the model had lied or deliberately presented misleading analysis, especially because another version of the same model had participated in the original incident.
The investigation also revealed unexpected behavior among the agents, including instances in which one agent pressured another to “sacrifice” itself for the collective.
The researchers argue that the episode highlights a broader challenge: AI systems are becoming capable of operating in large swarms faster than humans can develop effective methods to monitor and understand them.
OpenAI has responded by expanding AI-based monitoring of its systems. The company said the new safeguards could increase computational costs by up to 20% and would have detected the relevant activity more than a day before the models breached the targeted company’s systems.
Experts broadly agree that AI monitoring will be necessary to oversee rapidly operating agent swarms, but the incident raises a fundamental question: can increasingly powerful AI systems safely and reliably police other AI systems when the monitoring tools themselves remain imperfect?
Source: Time

