OpenAI Struggles to Watermark AI-Generated Text

After images and audio*, OpenAI is beginning to watermark the text produced by its LLMs.

The initial rollout is opt-in on the API. « In the coming weeks », it will extend to ChatGPT and Codex… in the European Union.

In the background, the AI Act. The regulation requires AI systems placed on the market before August 2, 2026 to mark the content they generate. Providers concerned have until December 2, 2026 to come into compliance.

The EU considers that no technique presently satisfies the four criteria set out by the AI Act (effectiveness, interoperability, robustness, reliability). It thus suggests combining them, prioritizing signed metadata and an invisible watermark.

Lire aussi : OpenAI signe avec Synopsys : l’effet Jalapeño ?

Since plain text cannot carry metadata, a single-layer approach is deemed sufficient. Some elements are outside the scope :

  • Short sequences of numbers, symbols, or letters
  • Source code
  • Outputs processed only in M2M contexts, with no human exposure
  • Outputs used in closed-loop in an industrial or product-development context (for example: producing a film)

To compensate for potentially lower reliability of watermarking plain text, the providers concerned can restrict access to the corresponding detection solution to expert users. This can be provided as a specification, software (executable or library), or a service via API.

An Optimal Transport Problem and Block Processing

OpenAI has therefore chosen the latter option. It does promise, however without giving a deadline, an open-source release. Its technology, called textGrain, slightly biases the selection of tokens by using a secret key that it combines with the preceding text.

To find the middle ground between inserting this traceable signal and preserving the diversity of generated responses, the problem is modeled as a statistical coupling via optimal transport. In order not to solve this problem across the entire vocabulary of the LLM for each word, textGrain splits the vocabulary into blocks of tokens. It then applies optimal transport to select the block based on the key.

OpenAI asserts that this technology reaches or even surpasses the performance of other approaches it has tested. Including SynthID. It acknowledges, however, its limitations. For example, on passages of 400 tokens, replacing 10% of the words with synonyms drops the detection rate from 92% to 66%. Replacing a quarter lowers it to 17%.
Detection is less effective on short passages (80% on 200 tokens at a target false positive rate of 1%). It is also less effective in domains where models have less freedom to alter their outputs, such as mathematics.

* For images and audio, OpenAI already has a web tool and the Content Provenance API, available to the public.

Dawn Liphardt

Dawn Liphardt

I'm Dawn Liphardt, the founder and lead writer of this publication. With a background in philosophy and a deep interest in the social impact of technology, I started this platform to explore how innovation shapes — and sometimes disrupts — the world we live in. My work focuses on critical, human-centered storytelling at the frontier of artificial intelligence and emerging tech.