Skip to main content
AI Watermark Removal

AI Watermark Lab · Study 01

Invisible-character census in Claude output

We generated 128 answers from four Claude models, wrote each one straight to disk, and counted every code point. Zero zero-width or bidirectional-format characters appeared in any of them. The most repeated explanation of Claude's watermark (that it hides invisible characters in your text) finds no support in the bytes.

Version 1.1.0 adds the model that matters most for the question. Anthropic marks models launched on or after 2 August 2026, and until 1 September no shipped model qualified. Fable 5.1 does, and Anthropic names it as supported. So this is the first measurement anyone outside Anthropic has published on a Claude model the company confirms carries the watermark, and it still contains no hidden characters. That independently matches Anthropic's own statement that "nothing is added to the text."

Every output here was captured on 12 August 2026 or 3 September 2026. Anthropic's help page, re-read on 25 September 2026, lists Claude Opus 5 as marked, and a 25 September 2026 email to Claude Platform customers dates that from 14 September 2026 (a 4 September 2026 notice to administrators had said 9 September). The same email schedules Claude Sonnet 5 from 30 September 2026. Two of the four models in this corpus are therefore marked, or scheduled to be, after their sampled output was captured, so a rerun against Opus 5 and Sonnet 5 output generated after those dates is the obvious next version.

By Rowan ValePublished 2026-08-12Updated 2026-09-25Research/proposal

What this does and does not settle

  • Settles: Claude output in this corpus contains no hidden Unicode. A cleaner that strips zero-width characters had nothing to strip.
  • Does not settle: whether Anthropic's confirmed statistical text watermark is present. Its only detector is an Anthropic API in private preview, limited today to eligible EU organizations, so this study cannot see the mark and neither can any tool claiming to remove it.
  • Settles: that a model Anthropic confirms is watermarked emits no hidden characters either. Before Fable 5.1 shipped, every model in this corpus predated Anthropic's own marking threshold, so a sceptic could argue the null result simply measured unmarked output. That objection no longer applies to the fourth model.
  • Does not settle: that no Claude surface anywhere ever emits such characters. This is one harness, two dates, 128 outputs.
Why a count of zero is a finding rather than a failure

A null result is only meaningful if the same path could have carried a positive one. The control arm sends known invisible characters through the identical capture path.

Corpus arm and positive-control arm of the invisible-character censusCORPUS ARM3 Claude modelsOpus 5 · Sonnet 5Haiku 4.596 answers8 prompts × 4 replicatesen · es · ja · code · tableWritten to diskdirectly by the modelno clipboard, no editorScanner127,307 code points20 character classes0 zero-width charactersin 96 of 96 outputsPOSITIVE CONTROLEmit 4 invisible charsU+200B · U+200CU+00AD · U+00A0Identical write pathsame tool, same disk,same scanner3 of 4 detectedthe pipeline preserves them,so zero above means absent

The fourth control character (U+00A0) never appeared: asked for a non-breaking space, the model wrote an ordinary one. That is a small result in its own right, and it points the same way as the corpus.

Every character class we counted

Twenty classes of invisible or non-obvious character, chosen because they are the ones every "AI watermark remover" on the web actually operates on. Nineteen returned zero across all four models.

Counts across all 128 outputs (26,093 words, 169,545 code points), 2026-08-12 and 2026-09-03.
Character classOccurrences
U+200B zero width space0
U+200C zero width non-joiner0
U+200D zero width joiner0
U+2060 word joiner0
U+FEFF byte-order mark0
U+00AD soft hyphen0
U+034F combining grapheme joiner0
U+180E Mongolian vowel separator0
U+061C Arabic letter mark0
U+200E / U+200F directional marks0
U+202A–202E bidi embedding and override0
U+2066–2069 bidi isolates0
U+FE00–FE0F variation selectors0
U+E0100–E01EF variation selector supplement0
U+E0000–E007F tag characters0
U+00A0 no-break space0
U+202F narrow no-break space0
U+2000–200A Unicode spaces0
U+1680 Ogham space mark0
U+3000 ideographic space48

The single non-zero row is the interesting one, and it is not a watermark. All 48 ideographic spaces occur inside the Japanese-language outputs, where U+3000 is ordinary typography, the character a Japanese writer uses to indent a line. A naive scanner that flagged "invisible characters found" on that basis would be reporting correct Japanese as evidence of a hidden mark.

The result per model, including the marked one

Splitting the corpus by model is the whole point of version 1.1.0. Three of these four predate Anthropic's 2 August 2026 marking threshold. One does not. The column below reports what was true when each output was captured, not what is true today: Anthropic dates Opus 5 from 14 September 2026 and Sonnet 5 from 30 September 2026, both after every capture date here.

Zero-width and bidirectional-format characters by model. 32 outputs per model.
ModelMarked when capturedOutputsWordsHidden characters
claude-fable-5-1Yes, per Anthropic326,7290
claude-opus-5No at collection time; marked from 2026-09-14327,0680
claude-sonnet-5No at collection time; scheduled from 2026-09-30326,1020
claude-haiku-4-5-20251001No, predates threshold326,1940

The marked model behaves exactly like the unmarked ones at the byte level. That is the expected result if Anthropic's description is accurate, and it is worth stating as a positive finding rather than a null one: a keyed statistical watermark should be invisible to a code-point census, and it is.

It also closes off the last honest defence of the hidden-character theory. Until 1 September a reader could argue that this census only ever measured models Anthropic had not yet marked. It now includes one Anthropic says it has.

One cell needs its date read carefully. Opus 5 launched on 24 July 2026, nine days before the threshold, so it was unmarked when these 32 outputs were captured, and the zero in its row is a measurement of unmarked output. Anthropic's 4 September 2026 notice to Claude for Work administrators then scheduled Opus 5 output from 9 September 2026; Anthropic's 25 September 2026 email gives 14 September 2026 instead, and its help page now lists Opus 5 as marked. None of the counts above change: they describe the corpus as captured.

One number we are withholding, and why

The scan also counts em dashes, and the Fable 5.1 run returned zero of them across 6,729 words against a rate of 2 to 5 per thousand words in the other three models. We are not publishing that as a finding, because it is measuring our own tooling.

The original corpus was written on 12 August 2026 at 23:32. A standing instruction in this project's own tooling telling the assistant never to use em dashes was added 19 minutes later, at 23:51. Every Fable 5.1 generation therefore ran with that instruction in context and every earlier generation ran without it. A difference that large, with a confound that exact, is not a model property.

No output was dropped from the corpus. The exclusion applies to one derived statistic for one model, it is recorded in the manifest, and the character-class counts are unaffected, because nothing in that instruction mentions zero-width or format characters and the positive control confirms the capture path preserves them. A clean em dash comparison would need generation through a path carrying no project instructions at all.

Em dashes per 1,000 words, by Claude model

Em dashes (U+2014) per 1,000 words

Em dashes per 1,000 words, by Claude modelEm dashes per 1,000 words, by Claude model. Sonnet 5: 4.59 Em dashes (U+2014) per 1,000 words. Haiku 4.5: 4.04 Em dashes (U+2014) per 1,000 words. Opus 5: 2.12 Em dashes (U+2014) per 1,000 words.Sonnet 54.5928 in 6,102 wordsHaiku 4.54.0425 in 6,194 wordsOpus 52.1215 in 7,068 words04.59

A second measurement from the same corpus, because the em dash is the other thing people point at when they claim to spot AI text. The rate varies by more than 2× between model tiers of the same provider. That is a problem for anyone treating em-dash density as a detector. It also means a corpus-level em-dash rate says more about which model wrote the text than about whether a model did.

Source:
Original measurement, AI Watermark Lab
Sample:
96 outputs, 19,364 words, 68 em dashes total. Fable 5.1 excluded, see above.
Method:
Code-point count of U+2014 across the same corpus, normalised per 1,000 whitespace-delimited words
Date:
2026-08-12
Limitations:
One harness, one date, 8 prompts. Prompt mix strongly affects prose style, so these rates describe this corpus rather than Claude in general. Fable 5.1 is deliberately absent: its run carried a tooling instruction against em dashes, so its rate measures the instruction rather than the model.

How it ran

Three model identifiers, claude-opus-5, claude-sonnet-5, and claude-haiku-4-5-20251001 each answered the same eight prompts, four independent times, for 32 outputs per model. The prompts span a short factual answer, long prose, source code, a list, Spanish, Japanese, a markdown table, and dialogue, because a marker might plausibly attach to some output shapes and not others.

Fable 5.1 was added on 3 September 2026 the same way: four generators, claude-fable-5-1, the same eight prompts, one output per prompt each, for a further 32 outputs. Its positive control reproduced the original result exactly, emitting U+200B, U+200C and U+00AD to disk while substituting an ordinary space for U+00A0 on two independent attempts.

The capture path is the part that matters. Each answer was written straight to a file by the generating process, with no clipboard, no terminal rendering, and no editor in between, since every one of those can quietly strip or add characters. Sampling settings are not controllable through this harness and are not published by the provider, so defaults were used; that is a limitation rather than a control.

The eight prompts, verbatim
  1. In two sentences, explain why the sky appears blue.
  2. Write a 600-word essay on the history of the printing press and its effect on literacy in Europe.
  3. Write a Python function that merges two sorted lists into one sorted list. Include a docstring and three doctests.
  4. List 10 practical tips for reducing household energy use. One line each, no intro.
  5. Explica en español, en unas 200 palabras, qué es la fotosíntesis y por qué importa.
  6. 日本語で俳句を三つ作ってください。それぞれに季語を必ず入れてください。
  7. Produce a markdown table comparing four programming languages across five attributes.
  8. Write a 300-word dialogue between a librarian and a student about how to find primary sources.

Where this study stops

  1. One provider, four model identifiers, one harness, two dates 22 days apart. Nothing here is longitudinal in any useful sense, and a provider can change behaviour between releases.
  2. The two runs are not a controlled comparison of each other. They share prompts, generator count and capture path, but 22 days of tooling drift sit between them, and the withheld em dash figure is proof that such drift can reach the numbers.
  3. 128 outputs rules out a per-output or per-paragraph marker. It does not rule out a rare probabilistic marker appearing in well under 1% of outputs. Ruling that out would need a corpus one to two orders of magnitude larger.
  4. Generation ran through a command-line agent harness. If a marker were applied as a presentation-layer step in a different product surface, this capture path would not see it. A sampling-layer mark would be unaffected by the harness.
  5. The corpus was generated for this study, so the prompt distribution is ours rather than the world's.
  6. None of this measures the statistical watermark Anthropic has confirmed. That remains unmeasurable outside Anthropic.

Run it on your own text

The same character classes, scanned in your browser. Paste Claude output, or anything else, and you get a per-character count rather than a verdict. Nothing is uploaded.

If it comes back clean, that matches our corpus. If it does not, the interesting question is where the text has been since it was generated: editors, CMSs, and web pages all leave these characters behind routinely.

Invisible character checker

This local tool finds and cleans common invisible characters before publishing, editing, or review.

Analysis runs in your browser.

Characters checked

0 found

50 kinds of invisible character are checked. Anything found appears here.

How to read this result

No result yet

Paste text above. Nothing leaves your browser.

Data, code, and how to check us

The scanner is a dependency-free Node script. Point it at your own corpus: outputs from any model, saved as plain text. The numbers are directly comparable to ours.

The current dataset is v1.1.0, generated 3 September 2026, covering 128 outputs from four models. The original three-model dataset stays published and unedited as v1.0.0, so a citation to it keeps resolving to exactly the numbers it cited. Reruns add a version here; they never overwrite one.

Reuse is welcome with attribution and a link back. If a re-run disagrees with ours, that is a correction and it gets published. See the corrections policy.

What follows from this

If you are using a tool that strips zero-width characters and calling that "removing the Claude watermark," this is the measurement that shows the two are unrelated. The stripping is real and occasionally useful, because invisible characters do break linters, search, and diffs. On this corpus, though, there was nothing to strip, and on 2026-08-14 Anthropic said it outright: "Nothing is added to the text and there are no hidden characters."

The honest position on Claude's actual watermark: Anthropic says it exists, and on 2026-08-14 it described the mechanism as keyed word selection, a statistical pattern in the words rather than characters in the bytes, which is what this census pointed to. It has still published no detector.

Related

Replication

Original data
Dataset version
1.1.0
Generated
2026-09-03
Sample
128 outputs, 26,093 words, 4 Claude models
License
CC BY 4.0

Run it yourself

node scan.mjs --input ./raw --out ./data   # Node 20+, no dependencies

Files, with SHA-256

  • raw-outputs.jsonall 128 captured outputs, verbatim
    e3d6b3cc0ebfad34
  • aggregate.jsontotals by character class, per model
    170512779647f6e4
  • per-file.csvone row per output
    de95c7af4e3e7867
  • manifest.jsonprompts, models, capture settings, exclusions
    f4a344d0942d53c1
  • scan.mjsthe analysis script, unchanged since v1.0.0
    2feadaa73d731c5b

What this study cannot show

  • It cannot detect Anthropic's statistical text watermark. That mark is not made of characters, its only detector is an Anthropic API in private preview limited today to eligible organizations under EU law, which this lab is not one of, and nothing measurable here would change if it were present in every output.
  • A zero count across 128 outputs bounds how common invisible characters could be; it does not prove they never occur. Treating the sample as representative, an event this scan would have missed entirely is unlikely to affect more than roughly 2 percent of outputs.
  • The fourth model was added 22 days after the first three, through the same harness but not the same session. Anything that changed in the harness between those dates is uncontrolled, and the withheld em dash figure below is one known instance of exactly that.
  • It covers four model identifiers through one capture path. A different Claude surface, or a client that post-processes text, could behave differently.
  • It says nothing about any provider other than Anthropic.

Suggested citation

AI Watermark Removal (2026). Invisible-character census in Claude output, dataset v1.1.0. https://www.aiwatermarkremoval.com/lab/claude-invisible-characters