Skip to main content
AI Watermark Removal

Detector

AI Text Watermark Detector

Claude's text watermark has a detector as of 1 September 2026, in a private preview almost nobody qualifies for, and Google's is runnable code whose applicability to your text is disputed by Google's own sources. The tools most people actually reach for aren't checking for a watermark at all, and they say so themselves.

By Rowan ValePublished Revised Sources verified Research/proposal

Key takeaways

  • Anthropic states in writing that supported Claude models watermark generated text, and separately that marking applies at launch to models released on or after August 2, 2026. Fable 5.1 and Mythos 5.1, released 1 September 2026, are the first models to meet that threshold and Anthropic names them. Its detector shipped the same day but only to eligible organizations under EU law, so for almost everyone the mark stays documented rather than verifiable. Anthropic's help page, re-read 2026-09-25, now carries a per-model table that also marks Opus 5.5 and Opus 5. Anthropic's 2026-09-25 email dates Opus 5 from September 14, where its 2026-09-04 notice to administrators had said September 9, and schedules Fable 5, Sonnet 5 and Opus 4.8 from 2026-09-30.
  • Gemini's detection method, SynthID Text, is thoroughly documented and open-sourced, but a Google-affiliated forum reply says it doesn't apply to Gemini API text at all, only, per Google's marketing page, to the app and web experience.
  • The underlying statistics are real and well studied. The original green-list/red-list scheme uses a one-proportion z-test on token frequency, and a 2024 follow-up found the signal survived heavy human paraphrasing once a detector had roughly 800 tokens, at odds low enough (1 in 100,000) to trust.
  • Even a real watermark detector can be attacked. Researchers approximated a provider's secret key from public queries for under $50 in 2024, then used it to forge the signal onto text the provider never wrote, or erase it from text it did.
  • Generic AI detectors are a fundamentally weaker tool. GPTZero, Originality.ai, and ZeroGPT all confirm, in their own published materials, that they score writing style rather than any provider's watermark.

Detection model

Watermark checker vs AI detector

Watermark checker

Looks for an intentional provider signal such as Claude text marks, SynthID Text, C2PA, or another known provenance layer.

Generic AI detector

Estimates whether text looks model-generated using style, probability, or classifier signals.

Why this distinction matters

A watermark-checker verdict is stronger evidence, because it's checking for a specific, intentionally embedded signal. A generic detector score is a probabilistic style guess that can misfire on short text, heavy editing, translation, or simply unusual human writing. Treat the two as different categories of evidence, not interchangeable confidence scores.

How does a real watermark detector work?

Confirmed

A real detector runs a one-proportion z-test and hands back a p-value. That statistic is the whole difference between it and an AI-likelihood score.

A genuine text watermark detector isn't scanning for a mark. It's running a statistical test on the sequence of words a model chose, asking whether that sequence matches a pattern only someone holding the right key could have produced.

The foundational version is Kirchenbauer and colleagues' 2023 green-list/red-list scheme. It splits the vocabulary into two pseudorandom groups before every token, reseeded from whatever came before, then nudges sampling toward the green group.

Detection is a one-proportion z-test: count how often green tokens actually got picked, then check whether that rate could plausibly have happened by chance. The output is an interpretable p-value, not a vague AI-like score.

SynthID Text builds on the same basic idea with something more elaborate, tournament sampling, where candidate tokens compete across sampling rounds before one is picked. Google describes two scoring approaches its own reference detector can use, a mean score and a Bayesian score.

Google also documents three ways a detector could be deployed:

  • Fully private, where only the provider can check anything
  • A semi-private API with limited third-party access
  • Fully public

Its documentation doesn't say which of the three, if any, Google actually runs for Gemini in production.

Has Anthropic released a Claude watermark detector?

Confirmed

Anthropic explained the mechanism on August 14, 2026 and announced a detection API first-party. Rechecked August 15: the API is not yet released.

Anthropic's support article states that compatible Claude models have woven an imperceptible watermark into generated text since August 2, 2026, worldwide. That covers the API, Claude, Claude Code, Claude Cowork, Claude Tag, and Claude models accessed through AWS, Google Cloud, or Microsoft Foundry, because, in Anthropic's own words, watermarking is applied at the model level and is present no matter which Claude product or surface the text comes from.

The same article commits Anthropic to letting users and third parties detect that mark. The technical explanation arrived on August 14, 2026: an Anthropic post describing the mechanism as "a version of the SynthID-Text approach published by Google DeepMind" and stating "We will soon be offering a watermark detection API." It shipped on September 1, 2026, in private preview for eligible organizations under EU law rather than to the public, so no detector a general reader can run exists.

What Anthropic has confirmed is narrower than a working tool. A detected mark would only indicate content may have been processed by Claude, not proof that it was.

Anthropic also lists the conditions where marks may not be recoverable at all:

  • Models released before the rollout
  • Heavily edited or translated text
  • Very short passages
  • Files whose metadata was stripped by re-saving, format conversion, or a screenshot

What does a clean SynthID Text result on Gemini output prove?

Community discussion

Run Google's own reference detector on Gemini API text and a clean result may mean the text was never marked in the first place.

Google's documentation of SynthID Text is, if anything, more technically detailed than Anthropic's. The method is open-sourced, with a production-grade reference implementation shipping in Hugging Face Transformers since version 4.46, including a Bayesian detector that reports watermarked, not watermarked, or uncertain.

That's real, usable software anyone can run today. What's contested is whether it's checking anything when pointed at ordinary Gemini API output.

Google DeepMind's own SynthID page says the Gemini app and web experience carry the watermark. A Google-affiliated account on the Google AI Developer Forum, replying on August 5, 2026 to a question about EU AI Act Article 50(2) compliance for two named Gemini models, wrote that API text is "NOT SynthID-watermarked" and that native text watermarking "is not planned at the moment." That reply was retracted by its own author on August 19, 2026, in favour of the opposite claim, and DeepMind's page has not changed to match.

Nothing published reconciles the two. So a clean result on Gemini API output could mean the text was never watermarked in the first place, not that detection failed.

Google's own hosted checker doesn't settle it either. The SynthID Detector portal launched at I/O on May 20, 2025 with only image detection live and text detection promised in the coming weeks.

Access has stayed waitlist-gated for journalists, media, and researchers, with no public API. An independent check in November 2025 found it still waitlist-only.

Can a real watermark detector still get it wrong?

Research/proposal

Roughly 800 tokens before the signal holds, under $50 to approximate the secret key, and one shipped Google checker that misfired with no attacker at all.

The statistics behind these detectors are well studied, and so are the ways around them. A 2024 follow-up to the original green-list/red-list scheme found the signal stayed detectable after strong human paraphrasing once a detector had around 800 tokens to analyze, at a false-positive rate of roughly 1 in 100,000.

Shorter passages carry a much weaker signal. That's part of why Google separately notes watermarking is less effective on short, factual responses: there's less room to nudge word choice without hurting accuracy.

The more unsettling finding is that a secret key isn't necessarily safe just because it's secret. ETH Zurich researchers showed in 2024 that a provider's green-list/red-list rule can be approximated from public API queries for under $50.

They then used that approximation two ways: forging a convincing watermark onto text the provider never generated, with over 80% spoofing success in their tests, and scrubbing a genuine watermark out, pushing removal success from near 0% to over 85% in settings assumed safe beforehand.

A 2025 follow-up from the same lab found a statistical way to catch that specific kind of forged text. So this is an active arms race, not a settled failure, and both a clean result and a confirmed match deserve some skepticism.

Do GPTZero and Originality.ai check for watermarks?

Confirmed

GPTZero, Originality.ai, and ZeroGPT each publish a description of their own method, and not one of those descriptions mentions a watermark.

Most of the detectors people actually reach for don't check any of the above:

  • GPTZero describes its own tool as scoring seven components covering perplexity and burstiness, how predictable and varied the writing is, plus other style signals, with no mention of SynthID, C2PA, or any provider watermark.
  • Originality.ai describes its detector as "a modified version of the BERT model," a statistical classifier.
  • ZeroGPT calls its method "DeepAnalyse."

All three are confirmed, from their own published descriptions, to be inferring style rather than verifying an embedded signal.

That's a meaningfully weaker kind of evidence. A style-based score can be wrong in either direction for reasons that have nothing to do with who wrote the text, and OpenAI has pointed to exactly this as a reason these tools disproportionately misjudge non-native English writers.

A real watermark detector, even an imperfect one, is at least checking for something a provider deliberately put there. A style detector is checking for a vibe.

FAQ

Is an AI detector a watermark detector?

A true watermark detector checks an intentional signal a provider built in at generation time. Generic AI detectors, including GPTZero, Originality.ai, and ZeroGPT, infer style instead and are less authoritative, by their own published descriptions of their own methods.

Will a Claude watermark detector ever become public?

Partly already, and the rest is only a stated intention. Anthropic announced the API on August 14, 2026 and shipped it on September 1 into a private preview for regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, EU civil society groups, and enterprises with their own duty to verify marking. It says it plans to expand access over time, without naming a date or criteria. Anthropic's 2026-09-04 notice to Claude for Work administrators compresses that same gate to "eligible EU organizations" and adds that it does not see or store text submitted to the API, neither of which its public help page includes when re-read 2026-09-25. So vetted third parties can check for the mark today and the public still cannot.

Can I run Google's SynthID Text detector myself?

The underlying method is open-sourced with a reference implementation in Hugging Face Transformers, so technically yes, for text you know came from a SynthID Text-marked source. But a clean result on ordinary Gemini API output doesn't necessarily mean detection failed. Per a Google-affiliated forum reply, that surface may never have been watermarked to begin with.

Can a watermark detector be tricked into confirming a fake match?

ETH Zurich SRI Lab researchers have shown it under lab conditions: approximating a provider's secret watermarking rule from public queries, then using it to forge a convincing match onto text that model never generated. It's an active area of both attack and defense research, not something documented as happening outside a lab.

Next steps

Sources and citation status