Skip to main content
AI Watermark Removal

Text artifacts

Zero-Width Space Watermark: Real Technique, Wrong Mechanism

Encoding a hidden message in invisible characters is a genuine, decades-old steganography trick. It is not the mechanism behind any confirmed AI provider watermark, because SynthID Text and Claude's watermark both live in word choice rather than in characters. Strip every zero-width character from a passage and you destroy the first kind completely while leaving the second entirely untouched.

By Rowan ValePublished Revised Sources verified Community discussion

Key takeaways

  • Zero-width steganography encodes bits as characters: a zero-width joiner (U+200D) standing in for one value, a zero-width non-joiner (U+200C) for the other, building a payload one bit at a time with nothing rendering on screen.
  • It's fragile in a way real watermarks aren't. Delete the characters, or run the text through almost any cleaner, and the hidden message is gone completely.
  • Deployed watermarks work the opposite way. SynthID Text runs as a logits processor during generation, and Anthropic says Claude weaves an imperceptible watermark directly into the text itself, so there's no character to find and delete in the first place.
  • Zero-width characters show up in ordinary, non-adversarial text constantly through copy-paste, emoji rendering, and Arabic-script formatting, so finding one proves nothing about a document's origin.
  • Commenters on r/singularity floated at least seven possible mechanisms for the Claude watermark, hidden Unicode among them. Anthropic settled it on August 14, 2026: its watermark is a SynthID-Text variant, and "Nothing is added to the text and there are no hidden characters."

Invisible character audit

Four characters explain nearly all of them

(invisible)

Zero-width space

U+200B

(invisible)

Zero-width joiner

U+200D

(invisible)

Byte-order mark

U+FEFF

(invisible)

Non-breaking space

U+00A0

Strip these characters

Destroys any zero-width steganography encoded into a passage. That technique is real and decades old.

Does it touch the watermark?

No. SynthID Text and Claude's mark both live in word choice, not characters, so this leaves them completely untouched.

The one exception worth a second look

Homoglyphs (look-alike characters swapped for the real ones) are the one category that isn't obviously unrelated to watermarking, and even there the only concrete claim is a single unverified report about a discontinued Claude Code behavior, not a confirmed provider mechanism.

How does zero-width space steganography hide a message?

Confirmed

You'll get the real encoding scheme, in enough detail to see why it's both clever and brittle.

A zero-width message is built one bit at a time out of invisible characters that render nothing on screen. Pick a character with zero visual width. The zero-width joiner (U+200D) and zero-width non-joiner (U+200C) are the usual choices.

Then let one stand for a 1 and the other for a 0. Slot them between ordinary letters and you build a payload one bit at a time.

Nothing renders on screen. A program that knows the scheme reads the sequence straight back out.

This is general-purpose text steganography, and it long predates modern AI. Hiding information in imperceptible formatting is an old idea that happens to work unusually well in Unicode.

Do SynthID Text and Claude use zero-width characters?

Confirmed

Both shipping watermarks live in word choice during generation, which is nowhere near the invisible characters people keep scanning for.

Neither one does: both shipping watermarks live in word choice during generation, not in any zero-width character. SynthID Text doesn't add anything to finished text. It runs as a logits processor during generation itself.

After ordinary Top-K and Top-P sampling narrows the field of candidate next tokens, it reshapes which of them the model is likely to pick. The signal ends up spread across dozens of word choices instead of sitting in any single character.

Anthropic no longer describes Claude's watermark just in spirit. Its explainer of August 14, 2026 calls it "a version of the SynthID-Text approach published by Google DeepMind," with a secret key and the preceding words determining which eligible next word gets picked, and states that "Nothing is added to the text and there are no hidden characters."

Neither system touches zero-width Unicode at all. That's the whole reason stripping invisible characters and defeating a watermark are different activities.

Why did people think Claude's watermark was hidden Unicode?

Community discussion

Anthropic named no mechanism for twelve days, and seven theories grew in the gap. On August 14, 2026 it named one, a SynthID-Text variant, with hidden characters ruled out.

Claude's watermark shipped with no published mechanism, and seven theories grew in the gap, hidden Unicode among them. When Claude's watermark rolled out, an r/singularity thread summarizing the announcement drew the same follow-up question over and over: how does a text watermark even work?

The theories commenters offered ran wide.

  • hidden Unicode or invisible characters
  • statistical word-choice patterns
  • overrepresented n-grams
  • first-letter or sentence-position patterns
  • token-probability nudges
  • something SynthID-like
  • a hybrid of several signals at once

Anthropic's support article named none of them, and real curiosity meeting an unpublished mechanism is exactly the soil zero-width theories grow in. The explainer Anthropic published on August 14, 2026 then ruled the zero-width theory out directly: "Nothing is added to the text and there are no hidden characters."

One adjacent claim is worth naming precisely because it describes a different technique. An independent blogger has written that Claude Code once used Unicode homoglyphs, characters that look identical to ordinary ones but carry different code points, inside date strings as an internal flag that was later discontinued.

That's homoglyph substitution, not zero-width steganography. It comes from a single source, and the follow-up article that might have corroborated it returned an error when checked.

So it stays interesting and unresolved, and it isn't evidence about zero-width spaces either way. Until a provider ties a specific character to a specific scheme in writing, an invisible character in copied AI text is evidence of ordinary text handling and nothing more.

What does removing zero-width characters actually fix?

Confirmed

A cleaner fixes broken search-in-page and corrupted pastes. One rule separates a good one from a text mangler: U+00A0 gets replaced, not deleted.

Removing zero-width characters fixes broken search-in-page and corrupted pastes, and it leaves a statistical watermark untouched. This site's text cleaner strips zero-width spaces, zero-width joiners and non-joiners, byte-order marks, and their relatives directly in your browser. Nothing is uploaded anywhere.

If a stray invisible character is breaking search-in-page, corrupting a paste into code, or tripping a duplicate-content check, that fixes it in seconds.

What none of this does is touch a real statistical watermark. If SynthID Text or Claude's watermark is present, it lives in which words were chosen, so cleaning invisible Unicode stays a hygiene job, not a removal one.

FAQ

Does removing zero-width characters count as removing an AI watermark?

No. It's useful text hygiene, fixing broken search indexing and copy-paste artifacts, but no major provider has documented zero-width characters as its watermarking mechanism. Google's SynthID Text and Anthropic's Claude watermark both work by shaping which words a model picks during generation, not by inserting an invisible character afterward.

How does zero-width steganography actually hide a message?

By treating the presence or absence of an invisible character as a binary digit. A common version inserts a zero-width joiner (U+200D) after certain letters to mean one bit and a zero-width non-joiner (U+200C) to mean the other, building a payload character by character with no glyph rendering on screen. It's a decades-old idea that predates modern AI entirely.

Has any AI provider ever used invisible Unicode characters as part of a real watermark?

None has documented doing so, and Anthropic has documented the opposite: its explainer of August 14, 2026 states "Nothing is added to the text and there are no hidden characters." Speculation on r/singularity had floated hidden Unicode as one guess among several, alongside statistical word-choice patterns, overrepresented n-grams, and token-probability nudges, before that explainer settled it. One blogger has separately claimed Claude Code once used homoglyphs in date strings as a since-discontinued internal flag, which is uncorroborated and describes a different technique.

Is stripping zero-width characters good for anything besides watermark paranoia?

Yes. They can break search-in-page, mess up pastes into code or spreadsheets, and occasionally trip duplicate-content checks, so clearing them is reasonable hygiene independent of any watermarking question.

Next steps

  • Paste a passage into the free in-browser cleaner to see which invisible characters it contains, with a count for each. Clean your text
  • See why finding an invisible character in AI output identifies almost nothing about its source. Hidden Unicode, explained
  • Read how statistical watermarking bends token probabilities during generation, which is the mechanism zero-width theories keep getting mistaken for. Statistical text watermarking
  • Read Anthropic's own wording on what Claude embeds and what it doesn't. Anthropic's support article

Sources and citation status