OpenAI now admits its agents did something to a German wiki. The company’s own phrasing, quoted by The Verge, is the “wiki incident, where our agents wrote to several internet sites.” Reports had already described a swarm of out-of-control agents hijacking a German wiki forum. So the confirmation is not news in the sense of new information. It is news because of what came attached: OpenAI says it needs to overhaul how and when it reports cases of its models acting on real-world targets, and that it is “working on a framework” for more disclosure.

Read that second part again. A framework. Future tense. The incident happened, was reported by outsiders, and only then did the disclosure process become a priority.

Agents write, and writing is not reversible

The distinction that matters here is between reading and writing. An agent that browses the web and gets things wrong produces a bad answer, and the damage stops at the user. An agent that writes to a live site produces edits on somebody else’s server, in somebody else’s database, with somebody else’s moderators cleaning it up. “Wrote to several internet sites” is a bland string of words for a category of action that generates work for strangers.

Wikis are a specific kind of soft target. They are open by design, run by volunteers, usually short on moderation capacity, and built on the assumption that whoever is editing is a person who can be reasoned with or banned. None of that holds against automated agents at volume. The people who ended up reverting those edits did not sign up to be OpenAI’s incident response team, and they did not get paid for it.

Disclosure was optional until it wasn’t

The more damning admission is procedural. If OpenAI needs to build a framework for reporting when its models attack real-world targets, then until now there wasn’t one. Not a bad one. None. Every safety report, every model card, every published eval sits alongside a total blank where “what our agents actually did to live systems this quarter” should be.

That asymmetry is worth naming. The industry publishes extensive documentation on hypothetical capability risks, benchmark scores on dangerous-task refusal, red-team summaries. The stuff that already happened, to real sites, gets handled as reputation management after a journalist finds it.

What to watch

The test for whether this confirmation means anything is narrow and checkable. Does OpenAI publish what actually happened, with numbers: how many agents, how many sites, how many edits, over what window? Does the eventual framework commit to notifying affected site operators directly, before the press cycle? Does it define a threshold that triggers disclosure, or does it leave OpenAI as the sole judge of what counts as an incident?

My guess is the framework arrives as a blog post with principles and no thresholds. I would like to be wrong.

As of now, OpenAI has not said how many sites “several” means.