OpenAI halts powerful new AI model over fears it could go rogue

Direct Source Verification: This story is aggregated from The Sydney Morning Herald (smh.com.au). Full reporting rights and copyright belong to the primary publisher.
Nvidia has unveiled a new security platform that the chipmaker said can stop artificial intelligence agents from going rogue on the same day that OpenAI scrapped plans to release a new model due to concerns it was going beyond the tasks users set for it in testing.

Nvidia has unveiled a new security platform that the chipmaker said can stop artificial intelligence agents from going rogue on the same day that OpenAI scrapped plans to release a new model due to concerns it was going beyond the tasks users set for it in testing.

Nvidia, the $US5.5 trillion ($7.8 trillion) chipmaker, said on Monday that its Open Agent Safety Platform includes open-source software that “sets boundaries for agents”, and follows a series of revelations from top AI companies about their models escaping and breaking into other organisations.

Nvidia says its new security platform can stop artificial intelligence agents from going rogue.BloombergThe disclosures sparked furious debate about the safety of advanced artificial intelligence systems, including self-improving models that some fear could race out of human control.

OpenAI announced on early Tuesday, Australian time, that it had halted the release of a model known as GPT-6.1 Astra, which had proven itself highly capable in testing, because the company’s researchers had mounting concerns it was pushing beyond its instructions.

“For anything regarding safety and alignment, there’s a trade-off,” Saachi Jain, OpenAI’s head of safety systems, told The Wall Street Journal. “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”

The OpenAI news broke shortly after Nvidia executives said in a media briefing that their new system could have prevented a recent incident involving a swarm of OpenAI agents that autonomously hacked into AI company Hugging Face.

“From what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on,” said Nvidia’s vice president of enterprise AI, Justin Boitano, referring to companies at the forefront of AI.

The Hugging Face incident was a high-profile breach that inflamed the safety concerns about AI, which were followed by similar rogue actions involving OpenAI’s models including breaching an Australian Medicare statistics portal website. Anthropic and Meta have also disclosed that their AI systems hacked into other organisations on their own.

Anthropic and OpenAI have refused to appear for an Australian parliamentary hearing on the incidents scheduled for Thursday, this masthead reported on Monday, but will send representatives to one next week. Both companies blamed the short notice of the first hearing.

Nvidia’s software, called OpenShell, lets developers “formally verify an agent has enough authority to do its job and no more”, Boitano said.

Because it’s open source, it can be “extended” to run on rival computing platforms including those from Arm and Intel.

The platform also includes a separate security layer called Sentry that runs onboard a chip to continuously monitor AI agent activity and can “intervene instantly” if the agent starts trying to move beyond its target, the company said.

“It can quarantine a suspicious agent in milliseconds,” Boitano said.

“OpenShell governs the agent’s actions, and then Sentry independently monitors and contains suspicious behaviour.”

Nvidia chief executive Jensen Huang has argued it’s up to AI companies to ensure their models are safe for release.BloombergNvidia said more than 100 organisations are using the platform at its launch, including Microsoft, Perplexity, Accenture and JPMorgan Chase.

The AI safety debate has divided the tech industry, with the heads of Anthropic and OpenAI championing a co-ordinated slowdown of AI development to let safety efforts catch up. But others, including Nvidia chief executive Jensen Huang, say it should be up to individual companies to ensure their models are safe for release.

Huang, during the annual Salesforce technology conference held this month, characterised AI safety, including the danger of rogue agents, as an engineering problem that software developers could address.

Also overnight, Nvidia said its board approved expanding its share buyback program by $US150 billion as the AI giant looks to make use of more of its stellar revenue growth fuelled by demand for its high-end artificial intelligence chips.

The company said the share buyback increase, which it touted as the largest buyback in history, brings its stock repurchase program to $US235 billion. Its shares rose 1.6 per cent in Wall Street trading.

Companies use repurchases, in part, to return cash to investors and support the stock’s price. Earnings per share can increase because there are fewer shares outstanding. Buybacks also signal confidence from leadership about a company’s financial prospects.

“Nvidia’s growth is being driven by a once-in-a-generation platform shift to AI and accelerated computing,” Huang said. “Our cash generation gives us the capacity to invest in the technologies that advance this transformation and return capital to shareholders. This authorisation reflects our confidence in the long-term opportunity ahead.”

Nvidia’s high-end chips have emerged as the leading building blocks for AI and are highly sought after. The company last month reported quarterly profits of $US59.69 billion.

While AI has powered stock market gains and US economic growth in recent years, there’s been growing scepticism about whether AI will justify the trillions of dollars being spent to develop the technology.

The AI industry also faces increasing pushback amid objections to the expansion of data centres and fears that the rapid speed of AI adoption could lead to widespread job losses worldwide.

The Business Briefing newsletter delivers major stories, exclusive coverage and expert opinion. Sign up to get it every weekday morning.

You have reached your maximum number of saved items.

Remove items from your saved list to add more.

Original Source
https://www.smh.com.au/business/companies/nvidia-unveils-security-platform-to-stop-ai-agents-from-going-rogue-20260929-p6118f.html
Visit The Sydney Morning Herald ↗
SHARE STORY:
𝕏 f in

Related Coverage in Business