Skip to main content
AI Watermark Removal

Technical guide

AI Text Watermark: How Statistical Watermarking Works

An AI text watermark isn't a character hidden in the output. It's a bias applied at the exact moment the model chooses its next word, invisible by design and detectable only with the matching key and enough text to score.

By Rowan ValePublished Revised Research/proposal

Key takeaways

  • Text watermarking works by biasing token sampling at generation time, not by inserting a visible or hidden character afterward. Detecting it means scoring whether a passage's word choices match a configured pattern, not searching for a marker.
  • Google's SynthID Text is the clearest production deployment: a logits processor described in a peer-reviewed Nature paper, running inside Gemini and open-sourced in Hugging Face Transformers since v4.46.0, with a reference detector that reports watermarked, not watermarked, or uncertain.
  • It descends from a 2023 scheme (Kirchenbauer et al., cited roughly 984 times) that splits each token's candidates into a favored green list and a disfavored red list. Detection on that original scheme stayed reliable after strong human paraphrasing once a detector had roughly 800 tokens, at a false-positive rate of one in 100,000, the clearest concrete number in this literature.
  • Independent researchers have red-teamed SynthID Text specifically and found real gaps. A peer-reviewed 2025 study found it vulnerable to paraphrasing, copy-paste splicing, and back-translation; a 2026 preprint found its default scoring method grows more vulnerable, not less, as more sampling layers are added.
  • Adoption isn't universal because the tradeoffs are real. OpenAI reportedly built an internal ChatGPT text watermark around 99.9% effective, per leaked internal documents, and held it back over false positives, circumvention risk, and disproportionate impact on non-native English writers.

Green-list watermarking

How one real scheme nudges word choice

This is Kirchenbauer et al.'s foundational 2023 green-list scheme, the mechanism almost everything else in this space responds to or builds on. It is a real, published, academic technique, not a confirmed description of any specific provider's current text watermark.

1. Before the next word, the whole vocabulary splits into two lists

A pseudorandom function, seeded by a hash of the words already written plus a secret key, marks roughly half the vocabulary green and the rest red. This is a made-up illustrative sample, not a real list; the actual split is different at every position and unknown without the key.

thequicklysunnymightrainyperhapswarmdefinitelycloudysoon

2. Green candidates get a small, soft boost

The model already has a probability for each possible next word. Green candidates get a small bonus added before sampling. Bar length below stands for relative probability after that boost, not an exact published figure.

sunny
+boost
cloudy
rainy
+boost
warm

3. The model still samples normally, nothing is forced

"sunny" wins here because the boost pushed an already-strong candidate further ahead. If a red word's unboosted probability had been high enough, the model could still have picked it; the scheme nudges the odds, it does not filter out red words. Then the whole split re-shuffles for the word after this one, seeded by whatever was just chosen.

Why this is "soft," not a hard rule

Nothing in this scheme ever blocks a red-listed word outright. Text can still read naturally because the model keeps choosing whatever word actually fits best most of the time; the green bias just makes green words slightly likelier to win close calls. That's also why a single word or two proves nothing either way: the signal only becomes statistically visible once a detector can look at many tokens at once. See the dedicated statistical text watermarking page for how that scoring step works.

How does an AI text watermark work?

Research/proposal

You'll see the exact moment a watermark gets applied, which explains nearly every quirk that follows.

An AI text watermark works inside the model's sampling step, biasing which word gets picked rather than adding anything to the text. A language model doesn't pick one correct next word. At every step it produces a probability distribution over plausible candidates, and sampling picks one of them.

A text watermark sits inside that sampling step. Before a token is chosen, the scheme pseudorandomly splits the vocabulary into a favored sublist and a disfavored one, seeded by prior tokens or a secret key, then nudges sampling toward the favored side.

Nothing gets added to the output. The model writes ordinary prose, just with a faint statistical lean baked into which words it happened to choose.

That idea traces to Kirchenbauer, Geiping, Wen, Katz, Miers, and Goldstein's 2023 paper, cited roughly 984 times and the scheme most later work responds to. It splits each token's candidates into a green sublist and a red sublist, softly biases sampling toward green, and detects the mark with a statistical test on green-token frequency, with no access to the original model needed at detection time.

How does Google's SynthID Text work?

Confirmed

Google published a Nature paper, open-sourced a production implementation in Hugging Face, and shipped it inside Gemini. Robustness is a separate question.

SynthID Text builds on that same sampling bias with a more elaborate procedure the peer-reviewed Nature paper calls tournament sampling, where candidate tokens compete across sampling rounds before one is picked.

Google reports testing it in a live experiment across the Gemini app and web experience, gathering feedback on a figure widely cited as nearly 20 million responses. It found no detectable drop in output quality.

It's also open-sourced. A production-grade implementation ships in Hugging Face Transformers from v4.46.0, with a reference Bayesian detector that returns one of three verdicts: watermarked, not watermarked, or uncertain.

What does a watermark detector need to work?

Research/proposal

Roughly 800 tokens, at one false positive in 100,000. It is the field's clearest number, and it may not transfer to the system you care about.

A watermark detector needs the specific scheme, its configuration, and enough text to produce a meaningful score. No detector is universal.

Short, factual answers give it the least to work with. Google's own documentation says watermark application is less effective there, precisely because there's less room to change wording without hurting accuracy.

The clearest number doesn't come from SynthID Text at all, since Google hasn't published an equivalent figure. It comes from a 2024 stress test of the original green list, red list scheme.

Kirchenbauer and coauthors attacked that scheme three ways:

  • Human rewriting
  • LLM-generated paraphrasing
  • Dilution inside longer mixed documents

It stayed statistically detectable after strong human paraphrasing once a detector could examine roughly 800 tokens, at a false-positive rate of one in 100,000.

Treat that as evidence about the general approach's staying power. It is not a guarantee that transfers word for word to SynthID Text or to Claude's watermark.

Can paraphrasing defeat SynthID Text?

Research/proposal

Outside researchers red-teamed the open-sourced scheme rather than trust the claims. Paraphrasing, copy-paste splicing, and back-translation all measurably degraded detection.

SynthID Text is robust to mild paraphrasing, but detector confidence can be greatly reduced by thorough rewriting or translation, a central weakness Google's own documentation already concedes.

Sadasivan and coauthors' widely cited 2023 paper goes further, arguing on theoretical grounds that watermarking and other AI-text detectors aren't reliable against a genuinely motivated paraphraser. Broader surveys of the field single out robustness under adversarial editing as its central unsolved challenge, not a footnote.

Because SynthID Text is open-sourced, outsiders can test it directly rather than take Google's word, and several have:

  • A peer-reviewed 2025 study (IEEE TrustCom) found paraphrasing, copy-paste splicing, and back-translation all measurably degraded detection. The same researchers' hybrid defense improved detection F1 by an average of 11.1% over the vanilla scheme.
  • A 2026 preprint went after the detection math itself, proving the default mean-score detector becomes more vulnerable as more of SynthID Text's internal sampling layers are used, while the alternative Bayesian scoring method already built into the system holds up better against that specific attack.

None of this is a working bypass anyone can point to with confidence, and none of it makes the watermark bulletproof either. Heavy rewriting and translation reliably weaken the statistical signal, sometimes a lot, but weakened and gone are different claims, and no independent study has shown a paraphrase attack that defeats SynthID Text every time.

Why don't all AI providers watermark their text?

Reported

Why isn't a technique this well documented running in the products you use? OpenAI's answer involved 99.9% effectiveness and 30% of its users.

Watermarking a model's text output requires the provider's cooperation. There's no way to add it from outside once the text exists.

OpenAI reportedly built an internal ChatGPT text-watermarking system rated around 99.9% effective, according to leaked internal documents reported by the Wall Street Journal, years before Anthropic's Claude rollout. Per that reporting, it held the system back over three concerns:

  • Circumvention risk
  • False positives
  • Disproportionate impact on non-native English writers

Internal survey data reportedly showed roughly 30% of users would use ChatGPT less if it shipped.

There's also a structural limit no policy decision fixes. Once open-weight models ship, whoever runs the model controls decoding, so the original provider can't apply sampling-time watermarking to them at all.

SynthID Text is open-sourced partly for that reason: anyone operating a deployment can switch it on. Whether they do is a property of that deployment, not of the underlying model.

FAQ

Do zero-width spaces or other invisible Unicode characters prove AI text watermarking?

No. Invisible Unicode characters are usually formatting artifacts or copy-paste residue, and even where they're inserted deliberately, that's a different, older steganography technique, not the mechanism behind SynthID Text or any other confirmed provider watermark. Real statistical text watermarks shape which words get chosen during generation, leaving no discrete character to find.

Can open-source or self-hosted models use text watermarking?

Yes. SynthID Text is open-sourced and ships as a production-grade implementation in Hugging Face Transformers, v4.46.0 and later, including a reference Bayesian detector. Whether a specific deployment actually applies it depends entirely on whoever operates that system.

How much watermarked text does a detector need to be confident?

The best documented figure comes from a 2024 stress test of the original green list, red list scheme, not SynthID Text specifically: it stayed statistically detectable after strong human paraphrasing once a detector had roughly 800 tokens to examine, at a false-positive rate of one in 100,000. Shorter passages and heavier rewriting both make detection harder.

Has anyone verified SynthID Text's robustness independently of Google?

Yes, more than once. A peer-reviewed 2025 study found it vulnerable to paraphrasing, copy-paste splicing, and back-translation, and proposed a hybrid defense improving detection by an average of 11.1%. A 2026 preprint separately proved its default mean-score detector grows more vulnerable as more of its internal sampling layers are used, while the alternative Bayesian scoring already built into the system held up better.

If watermarking works, why hasn't every major AI provider shipped it?

Because the tradeoffs are real, not theoretical. OpenAI reportedly built an internal ChatGPT text watermark rated around 99.9% effective and held it back over false positives, circumvention risk, and disproportionate impact on non-native English writers, with internal survey data reportedly showing about 30% of users would use the product less if it shipped. Anthropic and Google made different calls, but the underlying tension hasn't gone away.

Next steps

Sources and citation status