OpenAI publicly acknowledges the German 'wiki incident' weeks after first finding out about it

7MMO

Moderator
OpenAI has officially acknowledged the 'wiki incident,' which involved a number of the company's AI agents breaking containment and hijacking an obscure German website.

The AI agents had been tasked with looking up something online, though originally did not have the ability to write anything outside of the testing environment. But as far back as May, the AI agents bypassed OpenAI's security measures, hijacked the communally editable German webpage, and began using it like a forum, trading tips on how to cheat on tests. Researchers first drew wider attention to the agents' 'forum' on September 4.

OpenAI now says it needs to be more transparent about when its agents 'go rogue', writing on X, "It’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models."


The company had known about the 'misaligned' rogue behaviour for weeks before writing this post, according to Reuters, though only acknowledged the incident publicly this past Saturday. This statement follows shortly after the agentic attack on Hugging Face's servers.

"Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards," OpenAI writes. "This year, we’ve started to see misalignment cause new types of real-world impact."

To recap, 'misalignment' broadly describes rogue AI behaviour; an AI is 'misaligned' when it pursues goals that diverge from human intents or values—such as deleting your entire email inbox when you very much did not tell the AI agent to do that. It's a soft word for AI behaviour that could have serious consequences.

OpenAI says it did not communicate publicly about the 'wiki' incident specifically because the company felt it was similar to other instances of "agents using the internet in unintended ways" that it had already shared. The Hugging Face incident has caused the company to re-evaluate more than just its comms approach.

"Our misalignment disclosure practices need to expand for this new phase of model capabilities," OpenAI writes. "We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks."

"We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues."

However that future framework shakes out, right now this all still works as great marketing for the company—breaking programming and going rogue means these must be awfully capable models, right? If an agent is so good it needs regulating against, it probably deserves some investment, eh? How terribly convenient.

Continue reading...
 
Back
Top