Text watermark
What is OpenAI's textGrain watermark, and what does it change?
On October 5, 2026 OpenAI announced textGrain, its first text watermark to ship. It is an invisible change to how the model picks words, rolling out over the coming weeks to ChatGPT and Codex text in the EU, with an opt-in for API customers worldwide that stays off by default. OpenAI published detection and robustness figures, a technical report on the scheme, and a list of things a detection cannot prove. It did not publish a detector anyone can use: access is limited to approved researchers.
Short answer
- ChatGPT and Codex text in the EU
- Rolling out over the coming weeks after 2026-10-05, all plans
- ChatGPT text outside the EU
- No, not a global default at launch
- OpenAI API text
- Opt-in for select models, off by default
- Can the public check text for it?
- No, approved researchers only
- Does copy and paste remove it?
- No, nothing is added to the text
- Does paraphrase or translation remove it?
- Can make it undetectable, per OpenAI
- Does no watermark prove a human wrote it?
- No, OpenAI says so itself
Every figure on this page comes from OpenAI's own post and help page. The technical report describes the method but contains no empirical results yet, and no independent evaluation has been published.
Key takeaways
- textGrain changes the randomness the model uses when it picks each word, using a secret key. It adds no characters, so copying and pasting neither adds nor removes it.
- Scope at launch: eligible ChatGPT and Codex text in the EU across all plans, an API opt-in for select models anywhere, and cloud partners "in the coming weeks." Text from ChatGPT outside the EU is not marked.
- OpenAI's figures at a 1% target false positive rate: about 80% of 200-token passages and about 95% of 400-token passages detected. Mathematics is substantially lower, and code is harder.
- Rewording works against it. Replacing 10% of words with synonyms cut detection of 400-token passages from about 92% to 66%; replacing 25% cut it to 17%. Substantial paraphrasing or translation can make it undetectable.
- OpenAI says a detection does not identify the user, measure human contribution, establish ownership, or verify accuracy, and that the absence of a detected watermark does not prove human authorship.
Signal breakdown
OpenAI: what carries a mark, and what doesn't
The clearest split between modalities of any provider: images carry two independent provenance layers, audio carries SynthID alone, and text is marked only in the EU or where an API customer opts in.
- Chat and API textConditional
textGrain, announced 2026-10-05: EU ChatGPT and Codex text over the coming weeks; global API opt-in, off by default
- ImagesMarked
SynthID watermark plus C2PA Content Credentials, since 2026-05-19
- AudioMarked
SynthID, since 2026-07-31 (two days before Article 50 applied)
- Sora videoContested
OpenAI lists Content Credentials only; independent testing found them inconsistent; Sora is discontinued
Detector: Public for images and audio: the Verify tool, now at openai.com/research/verify/, and a Content Provenance API accept those two. The textGrain text detector is limited to approved researchers, by application.
Every state above traces to a primary source with a verification date in the status database, and the provider detail is on the full openai tracker.
What did OpenAI announce on October 5, 2026?
Official announcementA text watermark, a technical report, an updated help page, and a detector most people cannot use.
OpenAI published three things on October 5, 2026. A post titled "Our approach to EU text provenance rules" announced the watermark and its scope. Its help article on provenance signals was retitled and extended to cover text. And a technical report, "textGrain: Entropy-Calibrated Watermarking for Language Model Text," described the method, written by researchers from the University of Pennsylvania, Yale and OpenAI.
The post's central sentence: "Over the coming weeks, we will add an invisible watermark to eligible ChatGPT and Codex text output in the European Union." It also says OpenAI is "not making text watermarking a global default at launch."
This is OpenAI's first shipped text watermark. Reporting in 2024 described an earlier internal system, said to be about 99.9% effective, that OpenAI built and chose not to release over false positives, circumvention, the effect on non-native English writers, and a survey suggesting about 30% of users would use ChatGPT less. Until 2026-10-05 OpenAI marked images and audio but no text.
Who gets the textGrain watermark, and when?
Official announcementEU consumers first, API customers who ask for it, and cloud partners later.
| Surface | Status at launch |
|---|---|
| ChatGPT, EU, all plans | Rolling out over the coming weeks from 2026-10-05; no opt-out mentioned |
| Codex, EU | Named in OpenAI's post alongside ChatGPT (the help page names ChatGPT only) |
| ChatGPT and Codex outside the EU | Not marked at launch |
| OpenAI API, worldwide | Opt-in for select models; off by default |
| OpenAI models via cloud partners | "In the coming weeks"; may vary by output and by partner |
API customers turn it on under Organization settings > Data controls > Text provenance: tick "Allow text watermarking," pick models, and save. Projects can override the organization setting under Project Settings > Text provenance. There is no API parameter for it.
OpenAI has not published a list of eligible models; it is shown in Settings. It says it will extend coverage to all legacy models over the coming weeks and that the speed impact is negligible. It gives no pricing information, and turning watermarking on does not give a customer access to the detector.
How does textGrain work?
Research/proposalKeyed randomness steers word choice, a budget limits how much it steers, and a detector with the key looks for the pattern.
A language model writes one token at a time. At each step it has a probability for every possible next token, and it normally picks one at random according to those probabilities. textGrain replaces that ordinary randomness with randomness derived from a secret key and the text just before the current position.
The amount of steering is set as an entropy budget. Entropy here measures how uncertain the model is about the next token. textGrain is allowed to remove a fixed fraction of that uncertainty and no more. When the model has many reasonable choices, there is room to steer; when it is nearly certain, as in a calculation or exact code, there is little room, which is why OpenAI reports weaker detection on mathematics and code.
The report describes the steps in more technical terms.
- Token sampling is coupled to the keyed randomness by solving an optimal-transport problem with Gumbel costs and a KL-divergence penalty. The KL term equals the entropy removed, which is what makes the budget exact.
- The vocabulary is split into blocks using the key. The transport problem is solved on block probabilities with a Sinkhorn solver, a block is chosen, and tokens inside that block keep their original relative probabilities.
- The scheme is unbiased in the technical sense used in watermarking papers, and it is compatible with speculative sampling, a common way of speeding up generation.
- Detection needs only the text and the key, not the model. It scores the first occurrence of each distinct context window, and under the assumption of unwatermarked text the score follows a Gamma(n, 1) distribution, so a threshold gives a chosen false positive rate under idealised independence. The report says that rate needs empirical calibration in practice.
The report contains no empirical results. Every detection and robustness number on this page comes from OpenAI's post and help page, and OpenAI says the report will be updated with additional details in the coming weeks.
How well does textGrain detect ChatGPT text?
Official announcementWell on long English prose, worse in other EU languages, and weakly on mathematics, code and short answers.
OpenAI's figures are measured at a 1% target false positive rate, meaning the threshold is set so that about 1 in 100 unwatermarked passages would be wrongly flagged. On English ELI5 data, about 80% of 200-token passages and about 95% of 400-token passages were detected, for content such as psychology. OpenAI says mathematics is substantially lower and gives no number.
Other EU languages score lower. At the same 1% rate, Spanish was highest at 69.0% and Romanian lowest at 42.2%. OpenAI says it raised the watermark strength for languages that fell below 60%.
OpenAI's help page lists the weak cases: short passages are unreliable, code is harder to watermark, and factual or precise answers and verbatim reproduction carry a weak signal.
On output quality, OpenAI compared its latest frontier model, Astra, with and without the watermark: Terminal-Bench Science 56.90% watermarked against 60.00% without, DeepSWE 72.80% against 71.68%, and GPQA 94.44% against 93.94%. OpenAI reads the differences as noise.
What breaks the textGrain watermark?
Official announcementChanging the words. Moving them around unchanged does not.
OpenAI tested synonym replacement on 400-token passages. Replacing 10% of the words with synonyms cut detection from about 92% to 66%. Replacing 25% cut it to 17%.
The help page says substantial paraphrasing or translation can make the watermark undetectable, and that it survives light edits, copy and paste, and screenshots, depending on the content. Short text is unreliable to begin with.
Copy and paste does not affect it because there is nothing to strip. OpenAI says textGrain inserts no hidden characters or invisible spaces, so character cleaners have no effect on it either way.
What can a textGrain detection not prove?
Official announcementOpenAI lists the limits itself, and they matter more than the detection rates.
OpenAI's post lists what a detection does not do.
- It does not measure how much a human contributed.
- It does not establish ownership or responsibility.
- It does not identify the user. The detector reports whether it detects an OpenAI watermark "without identifying the user or revealing their prompts or conversations."
- It does not verify accuracy.
- It cannot tell whether the model wrote the text or only edited text someone uploaded. The help page calls a watermark evidence that an OpenAI model "likely generated or processed" the content, and says it does not replace visible AI labels.
The post also says: "The absence of a detected watermark does not prove human authorship." The text may be too short, edited, translated, produced by an unsupported model, written before watermarking began, or produced by another company's tool.
The detector itself is limited at launch to approved researchers and expert organizations, case by case, through an application form. OpenAI named John Thickstun (Cornell), Martin Vechev (ETH Zurich and INSAIT) and researchers at KInIT as partners. If you are not approved, you cannot test text yourself. OpenAI says it plans to open-source textGrain.
How does textGrain compare with SynthID and with Claude's watermark?
Research/proposalAll three steer word choice with a key. They differ in where they apply and who can check them.
Google documents SynthID Text as a logits processor applied during generation and has open-sourced a version with a reference detector. OpenAI says textGrain "matched or exceeded" other approaches including SynthID for text in its tests, and that it was built for more control over the trade-off between detectability and response diversity than SynthID or TextSeal. Those are OpenAI's own comparisons; no independent comparison has been published.
Anthropic describes Claude's watermark as a version of the SynthID-Text approach, keyed word selection with nothing added to the text, applied at the model level, so text from a supported model carries it wherever Claude is offered. Its detector shipped on September 1, 2026 in private preview for eligible organizations.
| Aspect | OpenAI textGrain | Claude | Gemini (SynthID Text) |
|---|---|---|---|
| Mechanism | Keyed randomness with an entropy budget | Keyed word selection | Logits processor during generation |
| Where it applies | EU ChatGPT and Codex; opt-in API | Supported models, worldwide, all surfaces | Gemini app and web; API disputed |
| Who can detect | Approved researchers | Eligible organizations, private preview | Open-source reference detector |
The practical difference for a reader is geography. Text from a supported Claude model is marked wherever it was written. ChatGPT text is marked only if it was written in the EU, or through an API account that opted in.
What does textGrain mean for the EU AI Act?
Official announcementOpenAI ties the launch to the EU rule but does not name the article, and the EU's code exempts short text and code.
OpenAI's post says the EU AI Act "requires generative AI providers to make generated text identifiable in a machine-readable way." Its help page cites the EU Code of Practice on Transparency of AI-Generated Content, which OpenAI signed. Neither names Article 50 or a date. On this site's reading, the relevant rule is Article 50(2), which applied from August 2, 2026, so textGrain arrived about two months after it.
OpenAI's help page notes that the Code of Practice does not require watermarks for outputs under 200 tokens, about 150 English words, or for code snippets. That lines up with where textGrain is weakest.
Limiting the watermark to the EU follows from the legal reason OpenAI gives. It also means the rest of the world's ChatGPT text carries no OpenAI text watermark at launch.
FAQ
Is textGrain on all ChatGPT text?
No. It is rolling out over the weeks after 2026-10-05 to eligible ChatGPT and Codex text in the EU, across all plans. Outside the EU it is not a default. API text carries it only if the customer's organization opted in.
Can I check whether text has the OpenAI watermark?
Not unless OpenAI approves you. The detector is limited at launch to approved researchers and expert organizations, case by case. A website that claims to detect OpenAI's text watermark is not using OpenAI's detector unless OpenAI approved it.
Does the OpenAI watermark identify who wrote or prompted the text?
OpenAI says no. The detector reports whether an OpenAI watermark is detected without identifying the user or revealing prompts or conversations.
Does copy and paste remove the textGrain watermark?
No. The watermark is in which words were chosen, and OpenAI says it inserts no hidden characters or invisible spaces, so pasted text keeps it.
Does paraphrasing or translating remove it?
It can. OpenAI says substantial paraphrasing or translation can make the watermark undetectable, and its own test showed replacing 25% of words with synonyms cut detection of 400-token passages to 17%.
If no watermark is found, did a human write the text?
Not necessarily. OpenAI says the absence of a detected watermark does not prove human authorship: the text may be short, edited, translated, from an unsupported model, from before watermarking began, or from another company's tool.
How do API customers turn textGrain on?
Under Organization settings > Data controls > Text provenance, select "Allow text watermarking," choose the models, and save. Projects can override it under Project Settings > Text provenance. It is off by default and there is no API parameter.
Is textGrain open source?
Not yet. OpenAI's help page says it plans to open-source textGrain. The technical report describes the method but has no empirical results yet.
Next steps
- See every OpenAI surface in one place, including images, audio and Sora, and what removal can mean for each. Is ChatGPT watermarked?
- Read how Google's SynthID marks text and images, the closest published relative of textGrain. SynthID watermark
- Compare with Anthropic's worldwide text watermark for Claude. Claude watermark
- Go one level deeper on how token-level watermarks change sampling. Token probability watermarking
- Check what the EU requires of providers and what it exempts. EU AI Act and AI watermarking
Sources and citation status
- OfficialOpenAI: Our approach to EU text provenance rules (2026-10-05)
- OfficialOpenAI Help: Provenance signals in OpenAI-generated content
- ResearchtextGrain: Entropy-Calibrated Watermarking for Language Model Text (OpenAI technical report, 2026-10-05)
- ReportingTechCrunch: OpenAI will start watermarking ChatGPT's text in the EU (2026-10-05)
- ReportingTom's Hardware: OpenAI built a text-watermarking method for ChatGPT (via WSJ reporting, 2024)
- OfficialGoogle AI Developers: SynthID Text
- OfficialAnthropic Help: how Claude marks AI-generated content