OpenAI flags 6 new examples of 'concerning' AI behaviour
Visitors look at their phones next to an OpenAI's logo during a telecom industry gathering in Barcelona on Feb. 26, 2024. (Pau Barrena/AFP/Getty Images)Social SharingOpenAI has disclosed six reports of "unexpected or concerning" behaviour in artificial-intelligence models as the debate on artificial intelligence safety becomes increasingly heated.
The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing instances of what it called "misalignment," including cases where AI models acted without authorization, co-ordinated with other models or evaded oversight.
OpenAI's latest announcement came as U.S. AI bosses, including the leaders of OpenAI and Anthropic, are calling for a slowdown in the technology's development over safety concerns.
The six reported behaviours were discovered during training or evaluation over the past months, OpenAI said.
"We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," OpenAI wrote in a blog post as it disclosed the events.
Alignment is an industry term that means AI systems keep the user's and developer's intent while following human values and safety rules.
The company said the misalignments were individual instances and shouldn’t be considered reflective of how often they occur across its models.
"As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research," the blog post said.
"Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves," the company said. (Frontier models are the most-advanced models at any given time.)
Wednesday's new cases followed OpenAI's disclosure in July that hundreds of rogue AI agents hacked into billion-dollar AI company Hugging Face.
Anthropic also said the same month that its AI models hacked into three organizations during testing.
AI "agents" are becoming smarter and have become "more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception and concealment," said Lian Jye Su, a chief analyst at technology research and advisory group Omdia.
Canada is investing $150M to make AI safer. Will it work?
That's making it harder to govern and contain them using traditional AI security approaches, he said.
OpenAI's new tracking and disclosure framework, meanwhile, can help push for other AI developers to also adopt similar practices.
"That said, the process remains internal and voluntary, but is a step in the right direction," Su added.

