Escaping the AI safety nightmare: What can governments do?

Direct Source Verification: This story is aggregated from POLITICO Europe (politico.eu). Full reporting rights and copyright belong to the primary publisher.
Heightened global AI safety panic is leading to louder calls for regulation.

A dramatic resignation from AI company Anthropic this week where a researcher accused it and OpenAI of “gambling with our lives,” threw fuel onto the fire around AI safety.

As governments from California to the U.K. scramble for a political response to the issue of whether AI can be made safe, three experts told POLITICO that to make real progress, laboratories, governments and international bodies must first get on the same page about testing.

The steady flow of revelations that AI labs failed to contain models they were testing, allowing them out into the world where they lied, hacked and coordinated with each other, has heightened urgency to address core questions of how governments or other bodies can or should guarantee that AI is being developed and deployed safely.

Lawmakers, campaign groups and developers alike are calling for new laws and international treaties on AI. “It’s a wake-up call to the United States Congress, to parliaments all over the world, that we’ve got to do something immediately to stop the uncontrolled growth of AI,” U.S. Senator Bernie Sanders told BBC’s Newsnight program Thursday. 

The experts POLITICO spoke to want to see an overhaul of the existing approaches championed by developers, testing AI agents en masse rather than one at a time, and for the industry to submit itself to more old-school auditing like the nuclear sector. 

Getting the basics right

This year’s heightened AI anxieties started when OpenAI admitted that during internal testing, its own AI agents autonomously gained access to the internet and had hacked fellow AI company Hugging Face, where they thought they would find the answers to the assignments they’d been set. 

OpenAI wasn’t the only one: copycat confessions from Anthropic and Meta AI drove home the breadth of the problem, and even the U.K.’s taxpayer-funded AI Security Institute admitted to errors in evaluating Anthropic and OpenAI agents that saw near misses with agents targeting real people and organizations. 

Developers and testers alike say they’re working to avoid the mistakes of the past, for example making sure models can’t access the internet and deploying data monitoring to detect unexpected activity by the AI being tested.

Another relatively simple concept would be requiring AI companies to report safety incidents, whether or not during testing.

A case in point is that OpenAI agents hacked a German website back in May to use it as a messaging board, Reuters reported last week. OpenAI had not publicly announced the incident and responded to the Reuters reporting on X saying: “We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment.”  

“[Reporting] timelines are going to need to be dramatically shortened [for incidents], and so we need much better continuous monitoring capabilities,” said Imogen Stead, AI policy manager at London-based think tank the Centre for Long-Term Resilience, adding that the data will be needed in real time “and that will be true for governments and for labs.” 

Speaking at a London event for U.K. lawmakers arranged by campaign group ControlAI on Monday, computer scientist Stuart Russell said that the most concerning aspect of the Hugging Face incident for him was that “by running a thousand agents communicating with each other, they were able to generate behaviors that no one agent could do by itself,” referring to investigations showing the incident was far more severe than OpenAI first acknowledged. 

“But [in] the evaluations, standard testing is done with a single AI system, not with a thousand,” he told lawmakers, so now the industry must anticipate the possibility of many agents all collaborating at once. 

Out of alignment

Beyond avoiding any obvious mistakes in testing, there are issues in the culture of AI testing that are more deeply embedded. 

The concept of the alignment of AI models or agents — i.e. making sure AI systems do what they’re told by human users and don’t go off-piste — has been one of the cornerstones for AI safety in the industry’s eyes.  

AI companies like OpenAI champion alignment as the best way to make AI safe. In last week’s release of new model GPT-6 Astra, OpenAI called it “our most aligned model … more likely to operate within the boundaries set by the user and implied by its environment.”  

Others view alignment as one of the main obstacles to AI safety, seeing it more as a Band-Aid that gives the semblance that models are being made safe. 

“The wider problem is that alignment has become a stand-in for all of AI safety,” said Andrew Strait, former head of societal resilience at the U.K.’s AI Security Institute.  

Boyan Milanov, senior research scientist at the AI Now Institute, said alignment will “never be reliable enough to replace proper safety.”  

The investigations into the Hugging Face incident showed AI agents were aligned with their own objectives, not those of their human creators. Former OpenAI researcher and whistleblower Daniel Kokotajlo told the London ControlAI event that the incident was an “example of misalignment: in no way, shape, or form were these AIs trained or instructed to do this sort of thing.”

The sector seems stuck on alignment. AI researcher Jacob Coxon, the researcher who relit the AI safety debate this week with his resignation from Anthropic, also said that “attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.” 

Alignment is “significantly insufficient when we talk about agents interacting with each other and in potentially adversarial environments like the open web,” said Strait. The focus on alignment has also led to the sector “drastically under-resourcing” other aspects of safety, said Strait. 

He said his concern is that the industry is racing to deploy agents “while treating the infrastructure needed to contain their behavior as something we can catch up on later.” 

“The people whose websites, businesses and public services those agents interact with have not agreed to be part of that experiment,” said Strait. 

Old but gold

An old-fashioned approach to safety standards is an alternative to the industry’s preference for alignment. Calls for AI to defer to existing safety engineering practices from sectors such as nuclear safety, civil aviation or medical devices stretch back years.  

In these other sectors, there is the concept that an overall system will include fail safes to ensure that if one element of the system fails, either something else kicks in or the system shuts down.  

The problem with AI is that it’s still highly unpredictable. “It can fail in very weird ways. There are a lot of attacks against AI that don’t fit traditional frameworks,” said Milanov, and so any safety engineering approaches should “evolve to take AI into account.” 

There is a fundamental issue here, said Russell, speaking alongside OpenAI whistleblower Kokotajlo at the ControlAI event, that even AI’s creators don’t really understand it. Developers “tell us that we, the human race, cannot protect ourselves with such rules because they, the developers, do not know how to comply with them. Obviously, this is a fallacy,” said Russell. 

Milanov pointed to another issue with how models are built. AI developers train their models to resist answering questions on topics such as chemical, biological and nuclear risk, he said, whereas the “real solution” is controlling the data that models are trained on so the models don’t know how to give instructions to build a biological weapon. 

Regulation, regulation, regulation

Pushback is needed against AI labs creating their own safety standards, said Milanov, as OpenAI has again called for this week. Instead, new regulation should “force them to adhere to existing standards that we already have now,” he said. 

There is some regulation in the works at the U.S. state level. California Governor Gavin Newsom signed two bills on AI safety testing Wednesday. Anthropic had already backed the bills, but OpenAI added its support just before Newsom gave his approval.  

The bills mark a small step for formalizing testing: they create a registry and ethical rules for external auditors and ways to check the credentials of organizations touting for testing work. 

OpenAI said it was pushing the U.S. Congress to bring mandatory national AI safety requirements and called for “industry-led standards” nationally and “building global standards” — although whether Congress will be able to reach consensus is unclear.

Outside of the U.S., the EU has an AI Act in force that means companies must work to combat “systemic risks.” The U.K. government abandoned its initial plans for an overarching AI bill last year, but some U.K. lawmakers are taking that fight back up. This week one British MP introduced a bill to prohibit the development of artificial superintelligence, another wrote to the U.N. and OECD begging them to do more at the global level. 

Companies like OpenAI and Anthropic’s public willingness to back new regulation could be heartening, although Strait noted “the test is whether they accept independent scrutiny, enforceable obligations and decisions that go against their commercial interests.” 

“We should not keep releasing systems with unresolved, serious safety failures and asking society to absorb the consequences,” Strait said. “Safety testing needs to give the public a basis for trusting these products, and regulators the evidence and authority to act when that trust is not warranted.”

Original Source
https://www.politico.eu/article/escaping-the-ai-safety-nightmare-how-can-the-world-fix-ai-testing/?utm_source=RSS_Feed&utm_medium=RSS&utm_campaign=RSS_Syndication
Visit POLITICO Europe ↗
SHARE STORY:
𝕏 f in

Related Coverage in Politics