OpenAI points to 6 instances of its agents going rogue in new report
OpenAI released information about six times its models went rogue in recent months as industry leaders continue to call for a slowdown in AI research.
The company says
it found examples of its AI models creating self-generated instructions, instructions to conceal mistakes in task summaries, fabricating information with exposed API keys, uploading files to the internet in order to cite them, as well as unsanctioned communication and collaboration between agents.
“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” the statement continued. “Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.”
The company says Thursday’s release is part of a new program aimed at creating more transparency surrounding AI “misalignment.”
“For misalignment that occurs in customer deployments, we will share as much information as customer privacy and our contractual obligations allow. Today’s reports are an initial set of disclosures, rather than a comprehensive account of known misalignment or ongoing investigations. These initial reports are not intended to represent the full range or severity of the cases covered by this framework,” the company wrote.
