SkillEnsure

Blog

The Watermark Problem: Why AI Labels Miss the Point
Artificial Intelligence

The Watermark Problem: Why AI Labels Miss the Point

Detectors and disclosure badges measure writing style, not who actually made the decisions behind a text.

The Watermark Problem: Why AI Labels Miss the Point

As generative language models have become embedded in everyday writing, a regulatory and cultural response has followed close behind: mark the output. Watermarks, detector scores, and platform badges now promise to tell readers whether a given text was "AI-generated." The premise seems reasonable if a machine helped write something, surely that fact should be visible. But the premise rests on an assumption that does not survive scrutiny: that a finished text carries a stable, recoverable signature of its own origin. It does not. What these mechanisms actually measure is a statistical property of surface style, and conflating that property with authorship produces a badge that is at once technically defensible and substantively misleading.

Three Mechanisms, Frequently Confused

Public debate tends to collapse three distinct mechanisms into a single "AI or not" judgment.

The first is a model-side watermark: a statistical signal embedded in generation, of the kind Anthropic introduced for Claude output under the EU AI Act's Code of Practice on Transparency of AI-Generated Content. Such a watermark is explicit about its own limits it can be edited or translated away, and it says nothing about whose idea the text expresses.

The second is the third-party consumer detector tools like ZeroGPT, unaffiliated with any model provider, whose reliability has never been established at scale.

The third is the platform-level editorial decision: a badge displayed, a ranking penalty applied. This is a policy choice, not a mechanical consequence of the first two.

Treating these as one causal chain the model marks, so the detector will catch, so the platform will punish is a category error. Each link requires an independent justification that the others do not supply.

An Empirical Illustration

A useful way to see the instability of detection is to run the same text through the same tool twice. One documented case: an article written in April 2021, before any consumer-facing large language model existed, scored 97% "AI probability" on ZeroGPT roughly a year ago and 8.6% on a recent re-test identical text, identical tool, incompatible verdicts. Inspection of which passages the detector flagged in the second pass reveals a pattern: it is not random sentences but the most neutral, most structurally regular ones definitions, procedural explanations that draw suspicion, while passages carrying distinctive authorial voice go unflagged.

This is not proof of how detection works in general, but it is suggestive of what it is actually sensitive to: stylistic regularity, not causal origin. The trouble is that regularity long predates language models. A detector trained to recognize a model's imitation of clear, well-structured prose ends up flagging the human writing that the imitation was modeled on. The inference runs backward.

Assisted, Generated, Produced

If detection cannot reliably locate origin, a more useful question is not whether a model participated but who retained control of the decisions. Three categories are worth distinguishing:

  • Assisted — a human sets the idea, angle, and structure, and retains final editorial judgment throughout, regardless of how many rounds of machine-assisted rewriting, translation, or fact-checking occurred along the way.
  • Generated — a human supplies a prompt but not sustained editorial control; the resulting text, as it stands, is the model's.
  • Produced — an automated pipeline operating at scale, without meaningful human oversight at all.

These categories carry entirely different degrees of editorial responsibility, yet a binary "processed by AI" badge treats them identically. What distinguishes them is not the quantity of model involvement but the locus of decision-making — who kept their hand on the choices that mattered.

The Cost of Collapsing the Distinction

Flattening this spectrum into a single label has a predictable social cost. A writer who spent hours drafting, verifying, and revising gets filed in the same bucket as a content farm publishing unreviewed volume, because a casual reader responds to the badge rather than to the process behind it. The asymmetry compounds: the same readers who reject a labeled article as "AI slop" routinely accept unlabeled AI-generated search summaries without scrutiny, even though independent measurements suggest such summaries substantially reduce click-through to original sources. The rejection, in other words, tracks visibility of the label far more than it tracks the actual involvement of a model, AI is resisted when disclosed and voluntarily claimed, and absorbed without comment when invisible and imposed by default.

There is also a structural risk in attaching consequences to an imperfect proxy at all. Once a badge affects visibility, ranking, or reputation, the incentive shifts toward optimizing against the badge rather than against the behavior it was meant to discourage transparent users have the least reason to conceal machine involvement and the most to lose from disclosing it, while actors with something to hide face the strongest incentive to learn what the detector responds to and route around it. A signal introduced to reward honesty can end up taxing it instead.

Conclusion

The instinct behind AI-transparency badges is not wrong: readers have a legitimate interest in knowing how a text was produced. But the mechanisms currently deployed to satisfy that interest watermarks, detectors, platform flags measure statistical regularities in finished text, not the chain of decisions that produced it. A more defensible standard would ask not "did a model touch this?" but "who can explain, defend, and take responsibility for what was published?" That question cannot be answered by inspecting the artifact after the fact. It can only be answered by the person who made the decisions which is, after all, the thing the badge was trying to identify in the first place.

by: L&D Team

Published on: Sep 7, 2026