Methodology
Methodology and source policy
Every factual claim on this site carries an evidence label, a primary source, and a date. This page defines exactly what those labels mean, what qualifies as a source, and how experiments are run and reported.
It is written so that a reader who distrusts us can check the work, and so that a researcher who wants to reproduce an experiment has the protocol rather than a summary of it.
Updated 2026-08-12
The evidence ladder
Claims are labelled by the strength of the source behind them, not by how confident the sentence sounds. A badge appears on every article and, where a single page mixes evidence levels, on individual sections.
| Label | What it means | What we require |
|---|---|---|
| Confirmed | The provider or standards body states it directly, in its own current documentation | A live primary URL and a quotable sentence |
| Official announcement | Stated in an official post or policy, but not yet reflected in product documentation | The announcement URL and its date |
| Reported | Established by credible journalism or a named investigator, not by the provider | The original report, not an aggregator's rewrite of it |
| Research/proposal | Demonstrated in academic work, including preprints, against a stated setup | The paper, its venue, and whether it tested a production system or a reimplementation |
| Community discussion | Circulating among users or practitioners with partial evidence | A link to the actual thread and an explicit statement of what is unverified |
| Rumor/speculation | Widely repeated, traceable to no primary source | Where it came from, and why it does not hold up |
Each label, with a live example from this site rather than a hypothetical:
| Label | Live example | What earned that label |
|---|---|---|
| Confirmed | Grok images and video carry a visible watermark | xAI's own Grok FAQ states it, and adds that no setting removes it. Quotable, current, primary. |
| Official announcement | Claude marks models launched on or after 2026-08-02 | Anthropic's help centre states the policy, but its own release notes date every generally available model earlier, so the product has not caught up to the announcement. |
| Reported | Grok Imagine images carry no C2PA Content Credentials | Established by reporting, not by xAI, and directly contradicted by two vendors who sell removal. Comparable authority on both sides, so neither is Confirmed. |
| Research/proposal | A watermark-stealing attack raised scrubbing success past 85 percent | Published at ICML, with a stated setup. It targeted a research reimplementation, not a production system, and the label says so. |
| Community discussion | Em dashes indicate AI authorship | Widely believed, partially measurable. We measured our side of it and published the counts; no human baseline exists to compare against. |
| Rumor/speculation | The EU mandated a 24x24px icon in PMS 286C blue | Circulates on marketing sites, traceable to no Commission document. The real EU icon set exists and carries none of those attributes. |
What counts as a source
Primary sources first, in this order, and the ordering is not a formality. It decides which of two conflicting statements gets quoted first.
- Provider documentation, support articles, model cards, and terms that the provider maintains.
- Official announcements, transparency reports, and regulatory filings.
- Legislative and standards text: the AI Act articles themselves, the C2PA specification, national standards.
- Peer-reviewed research, then preprints, with the distinction stated rather than blurred.
- Credible reporting that does original work, used to locate primaries.
- Community threads, used as evidence of demand, terminology, and belief. They are never evidence of mechanism.
Some provider domains block automated fetching. Where a claim was retrieved through a text mirror or an archive capture rather than the live page, the page says so instead of implying direct access.
Marketing pages from tools that sell watermark removal are not sources. They are the thing this site exists to check.
Dating and re-verification
Provider watermarking status changes month to month, so an undated claim about it is worthless within a quarter.
How often a claim gets reopened depends on how fast that kind of fact moves:
| Claim type | Cadence | Why |
|---|---|---|
| Provider watermarking status | Monthly | Help-centre pages change without notice and without a changelog. |
| Public detector availability | Monthly | The single fact most likely to change, and the one readers act on. |
| Regulation in force | Quarterly | Legal text is stable; guidance and enforcement posture are not. |
| Published research | On citation | A paper does not change. Its interpretation and its citation count do. |
| Our own study data | On rerun only | A dataset is fixed once published. A new run is a new version, never an edit. |
| Tool behaviour | On code change | Anything the site ships is re-checked when the code behind it changes. |
- Every page shows a published date and an updated date.
- Volatile provider claims carry a last-verified date, meaning someone opened the primary source on that date and confirmed the wording still says what we say it says.
- When a primary source changes wording, the change itself becomes a fact worth reporting, with both versions where we can evidence them.
- Where an archive capture is the only evidence of a previous wording, that is stated rather than presented as a direct observation.
If a page's last-verified date is old and the topic is fast-moving, treat it as stale and tell us. That is a legitimate correction even when nothing on the page is wrong yet.
How experiments are run
Original testing follows the same protocol every time, and the protocol is published with the results rather than described afterwards.
- State the question and what result would count as evidence either way, before running anything.
- Record the full configuration: model identifier, interface or API, date, prompt, output length, language, and any sampling settings that were controllable.
- Publish the aggregated data, and the raw data where it can be released without republishing someone else's copyrighted text.
- Publish the analysis code, or enough of the procedure that someone else can write it in an afternoon.
- State the sample size next to every number, including in charts.
- State the limitations in the same place as the finding, not in a footnote.
Dataset versioning
Study data is versioned rather than edited, so a citation stays valid after a rerun.
- Every published file carries a SHA-256, shown on the study page next to its download link.
- A dataset version changes only when the data changes. Rewording the page around it does not bump the version.
- A rerun produces a new version alongside the old one. Previous versions are not deleted, because a citation to them has to keep resolving.
- The generation date, the sample size, and the script version are published together. Any one of them alone is not enough to reproduce a result.
- Exclusion criteria are stated before the numbers. If an output was dropped from a sample, the page says how many and why.
- Scripts are published in full and carry no dependencies where that is achievable, so running them does not require trusting a lockfile.
Handling contradictions
Providers contradict themselves and each other, and the contradictions are often the most useful thing on the page.
The rule is to report both sides with their labels and dates, name the source of each, and say plainly that nothing reconciles them. We do not silently pick the more recent one, the more authoritative-sounding one, or the one that makes a cleaner sentence.
Where a contradiction has a plausible resolution that no source confirms (for example, that a consumer app and a developer API behave differently), the hypothesis is stated as a hypothesis.