OpenAI's Text Watermarking: What It Can and Can't Detect

OpenAI's Text Watermarking: What It Can and Can't Detect

Short answer: OpenAI's new textGrain watermark is real progress, but it only covers EU ChatGPT/Codex text going forward, breaks down fast under editing, and isn't publicly checkable — which is exactly why independent detection still matters.

Direct answer

On October 5, 2026, OpenAI announced textGrain, an invisible statistical watermark for text, rolling out to eligible ChatGPT and Codex output in the EU only, with opt-in API access worldwide. In OpenAI's own testing, detection dropped from about 92% to 66% after replacing just 10% of a passage's words with synonyms, and to 17% after 25%. The detector itself isn't public — only approved researchers and organizations can apply for access.

Announced October 5, 2026
Technology textGrain — invisible statistical watermark
Coverage EU ChatGPT/Codex only; API opt-in worldwide
Driver EU AI Act text provenance requirements
Detector access Approved researchers/organizations only

What OpenAI announced

OpenAI published its response to the EU AI Act's text provenance requirement, which obligates generative AI providers to make generated text identifiable in a machine-readable way. The rollout has three parts: API customers worldwide can now opt in to watermarked output for select models (off by default); eligible ChatGPT and Codex text will get an invisible watermark within the EU specifically, over the coming weeks, rather than as a global default; and a detector to check for that watermark is opening to applications, initially limited to approved researchers and expert organizations.

Notably, this only changes text. OpenAI's existing image and audio verification tools — openai.com/verify and its Content Provenance API — remain publicly accessible, unchanged by this announcement.

How textGrain actually works

textGrain adds an invisible statistical signal to a model's word choices as it generates text. A separate detector then scans a passage for that signal to assess whether it contains an OpenAI watermark. In OpenAI's own evaluations, textGrain matched or exceeded other approaches it tested, including SynthID for text — but strong lab performance, as OpenAI itself notes, doesn't guarantee reliable results in everyday use.

Where detection breaks down

Direct answer

Detection gets worse with shorter text, with content that has less flexibility in word choice (like math), and collapses fast once a passage is edited.

80%
200 tokens
Detection rate for short passages, at 1% false-positive rate
95%
400 tokens
Detection rate for longer passages of flexible content
17%
25% edited
Detection rate after a quarter of words are swapped

Length matters: at a 1% target false-positive rate, OpenAI's detector found watermarks in about 80% of 200-token passages versus about 95% of 400-token passages, for flexible content like psychology answers. Content type matters just as much — detection rates were substantially lower for material like mathematics, where there's less room to vary word choice without changing the answer.

Editing matters most of all. In a 400-token passage test, swapping 10% of words for synonyms dropped detection from about 92% to 66%. Swapping 25% dropped it to about 17% — meaning a quarter of a passage lightly paraphrased is enough to make the watermark nearly undetectable four times out of five.

ConditionDetection rate
Unedited, 400 tokens~92%
10% of words replaced with synonyms~66%
25% of words replaced with synonyms~17%

What a watermark doesn't tell you

OpenAI is explicit about this, and it's worth repeating plainly. A detected watermark doesn't measure how much human judgment or editing went into a passage. It doesn't establish who owns the text, whether its use was lawful, or who's responsible for it. It doesn't identify a specific person, account, or conversation. And critically, it doesn't verify whether the content is true, misleading, or presented in the right context — a watermark is a provenance signal, not an accuracy check.

  • 1 No watermark detected ≠ human-written. The text could be too short, edited, translated, generated by an unsupported model, predate watermarking, or come from an entirely different company's AI tool.
  • 2 Watermark detected ≠ case closed. It confirms an OpenAI system touched the passage, not how much of it is actually AI-written versus human-edited.

What this means for AI detection

This is a genuinely useful step — regulatory pressure pushing a major lab to publish real detection numbers, including the ones that make the technology look weaker, is rare and worth crediting. But the coverage gaps are structural, not incidental, and they matter for anyone trying to verify text in the real world:

  • 1 Geography-limited by design. The ChatGPT watermark only applies inside the EU. Text generated anywhere else, on the same models, carries no watermark at all.
  • 2 Opt-in, not universal. API watermarking defaults to off worldwide. A developer has to actively choose to enable it — most won't.
  • 3 Single-vendor by construction. textGrain only flags OpenAI-generated text. It says nothing about content from any other model or provider — which, across the open model ecosystem, is most of what's actually circulating.
  • 4 Not public. Even where the watermark exists, nobody outside an approved research application can check it. A journalist, teacher, or platform moderator encountering a suspicious passage today has no way to query OpenAI's own detector directly.

That combination — regional, opt-in, single-vendor, and access-gated — is precisely the gap general-purpose detection tools exist to fill. A watermark that only works for one company's models, in one region, when the publisher chose to enable it, and that only researchers can check, leaves most real-world text unaddressed. UncovAI's text detection works the other way: it looks for the statistical and structural fingerprints AI-generated text tends to leave behind, regardless of which model produced it or whether the publisher opted into anything — publicly available to anyone who needs to check a passage today, not just approved researchers evaluating a lab's own technology.

Frequently asked questions

Does OpenAI watermark all ChatGPT text?

No. OpenAI's invisible watermark, called textGrain, is rolling out only to eligible ChatGPT and Codex text output within the European Union, not as a global default. In the API, watermarking is opt-in and off by default worldwide.

Can OpenAI's watermark detect edited or paraphrased AI text?

Not reliably. In OpenAI's own testing, replacing 10% of words in a 400-token passage with synonyms dropped detection from about 92% to 66%. Replacing 25% of words dropped it to about 17%.

Can the public use OpenAI's text watermark detector?

No. Access is limited to approved researchers and expert organizations who apply and are reviewed case by case. There is no public-facing tool to check arbitrary text for an OpenAI watermark, unlike OpenAI's image and audio verification tools at openai.com/verify, which remain publicly accessible.

Does a missing watermark prove a text was written by a human?

No. OpenAI states directly that the absence of a detected watermark does not prove human authorship. The text may be too short, edited, translated, generated by an unsupported model, predate watermarking, or have come from a different company's AI tool entirely.

Sources

This analysis is based on OpenAI's own announcement, "Our approach to EU text provenance rules", published October 5, 2026.

Check text, images, video, and audio in one place

Watermarking only works where a publisher opted in. UncovAI checks content directly — no cooperation required from the model that made it.

Get Started Free →