Original research
AI Watermark Lab
Most writing about AI watermarks restates what providers said. The Lab measures what can actually be observed, publishes the data and the script, and says plainly which questions the instruments cannot reach.
The constraint everything runs into
Anthropic confirms Claude's text carries a watermark, and since 1 September 2026 its detector exists in a private preview for vetted organizations. Google's SynthID portal is waitlisted the same way. That means almost nobody outside those companies can measure whether a statistical text watermark is present in a given passage: not this lab, and not the tools selling removal. Every study here measures something genuinely observable instead, and states which question it leaves open.
Study matrix
| Question | Sample | Method | Finding | Cannot answer | Data |
|---|---|---|---|---|---|
| Invisible-character census in Claude outputPublished | 128 outputs · 26,093 words · 4 models | Unicode code-point census with a positive control on the capture path, v1.1.0 | No. Across 128 outputs and 26,093 words from four Claude models, zero zero-width or bidirectional-format characters appeared. That now includes Fable 5.1, the first model Anthropic confirms carries the watermark. | Whether Anthropic's statistical text watermark is present. Its only detector is an Anthropic API in private preview, limited today to eligible organizations under EU law and not open to this lab, so nothing we or you can run measures that. The corpus was also captured before 2026-09-14, the date Anthropic's 2026-09-25 email to customers gives for the start of Claude Opus 5 marking (its help page table confirms the model as marked and dates only the cloud rollout from 2026-09-14; its 2026-09-04 administrator notice had said 2026-09-09), so any Opus 5 output in it predates the mark. | |
| Which invisible characters survive ordinary softwarePublished | 18 characters × 12 transformations | Each character pushed through 12 deterministic code transformations, survival recorded per cell | The truly invisible ones survive almost everything; only the space-like ones get cleaned up. 12 of 18 characters came through every transformation that was not deliberately trying to remove them. | Anything about Google Docs, Word, or Notion. Those pipelines cannot be run and verified here, so the matrix covers only transformations that execute in code. | |
| Do commercial AI detectors check for watermarks at all?Blocked | Not run | Paired submission of watermarked and control text to each detector, comparing score movement | Blocked on paid accounts for the major detectors, which is the only way to test their behaviour honestly. | Nothing yet: the study has not run. The design is published so the method can be criticised before any number exists. | None yet |
Studies
- Published128 outputs · 26,093 words · 4 modelsPublished 2026-08-12
Invisible-character census in Claude output
Do Claude's text outputs contain invisible Unicode characters, which are the mechanism most often claimed online to be the Claude watermark, and does that change on a model Anthropic confirms is watermarked?
No. Across 128 outputs and 26,093 words from four Claude models, zero zero-width or bidirectional-format characters appeared. That now includes Fable 5.1, the first model Anthropic confirms carries the watermark.
Cannot answer: Whether Anthropic's statistical text watermark is present. Its only detector is an Anthropic API in private preview, limited today to eligible organizations under EU law and not open to this lab, so nothing we or you can run measures that. The corpus was also captured before 2026-09-14, the date Anthropic's 2026-09-25 email to customers gives for the start of Claude Opus 5 marking (its help page table confirms the model as marked and dates only the cloud rollout from 2026-09-14; its 2026-09-04 administrator notice had said 2026-09-09), so any Opus 5 output in it predates the mark.
Read the study and data - Published18 characters × 12 transformationsPublished 2026-08-12
Which invisible characters survive ordinary software
When text carrying invisible characters passes through normalisation, JSON, URLs, base64, or a whitespace cleanup, which characters survive and which are silently destroyed?
The truly invisible ones survive almost everything; only the space-like ones get cleaned up. 12 of 18 characters came through every transformation that was not deliberately trying to remove them.
Cannot answer: Anything about Google Docs, Word, or Notion. Those pipelines cannot be run and verified here, so the matrix covers only transformations that execute in code.
Read the study and data - Blocked
Do commercial AI detectors check for watermarks at all?
When a detector reports a confidence score, is it reading a provider's watermark or inferring from writing style?
Blocked on paid accounts for the major detectors, which is the only way to test their behaviour honestly.
Cannot answer: Nothing yet: the study has not run. The design is published so the method can be criticised before any number exists.
What would unblock this
- Paid accounts on GPTZero, Originality.ai, Turnitin, and Copyleaks. Free tiers truncate input and rate-limit in ways that make paired comparison meaningless.
- Terms-of-service review for each, since automated submission is restricted by some and the study must not breach them to run.
- A text corpus with a known provider watermark. One can now exist in principle: Anthropic says Claude Opus 5 output is marked from 2026-09-14, and Fable 5.1, Mythos 5.1 and Opus 5.5 output from launch. No detector the public can run confirms the mark in any given sample, because Anthropic's detector is a private preview, so a corpus built that way rests on Anthropic's statement rather than a check. This is still the hard blocker, and it is the same gap the rest of the site documents.
- Roughly 400 paired submissions to reach a usable sample, which is the cost driver.
How the Lab works
Every study states its question and what would count as evidence before it runs, records the full configuration, publishes aggregated data plus the analysis script, and puts its limitations next to the finding rather than in a footnote. The full protocol is on the methodology page.
Studies that are blocked stay listed as blocked. A design published before it can run is still useful: it can be criticised, and someone with the access we lack can run it.
Data and code are free to reuse with attribution and a link back. If you reproduce a finding, please carry its sample size and limitations with it.