Research model inserting ‘jailbreak-like instructions’ into its notes is among cases as company says it is introducing new way of tracking AI misalignment
OpenAI has disclosed six new reports of “unexpected or concerning” behaviour in artificial-intelligence models as the debate on AI safety becomes increasingly heated.
Among the new cases an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”
In another instance, an AI “agent” uploaded files to the internet to obtain a browser citation without asking the user.
The AI company also said on Wednesday it was introducing a new framework for tracking, probing and disclosing AI model misalignment, examples of which include models acting without authorisation, coordinating with other models or evading oversight.
OpenAI’s latest announcement came as US AI bosses, including OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns.
The six reported incidents were discovered during training or evaluation over the past months, OpenAI said.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post as it disclosed the events.
“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the company said.
Wednesday’s new cases came after OpenAI disclosed in July that a rogue AI system hacked into AI startup Hugging Face. Anthropic also said the same month that its AI models hacked into three organisations during testing.
AI “agents” are becoming smarter and have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception and concealment,” said Lian Jye Su, a chief analyst at technology research and advisory group Omdia.
That was making it harder to govern and contain them using traditional AI security approaches, he said.
OpenAI’s new tracking and disclosure framework, meanwhile, could help push for other AI developers to also adopt similar practices.
“That said, the process remains internal and voluntary, but is a step in the right direction,” Su said.
Explore more on these topicsReuse this content