Skip to main content

Jay Shah

50 min readInteractive Laboratory

How does a text watermark work?

A first-principles investigation of how generation-time text watermarks work, from a weighted coin to Gemma, and where the evidence stops.

Plain text leaves nowhere obvious to hide a watermark. That is what makes the problem interesting.

An image has millions of pixel values that can shift without changing what a person sees. Video repeats that grid across thousands of frames. Text has no comparable slack. Every character is visible, and copy-paste discards file metadata. If a mark is going to survive, it has to live in the model's choices rather than in an invisible layer around the text.

I started pulling on that question after Anthropic announced that future Claude models would watermark their text output.1 Three days later, the company said Claude uses "a version of the SynthID-Text approach" that changes randomness among acceptable next-word choices.2 Anthropic plans to provide a detection API, but it has not published the algorithm, keying scheme, threshold, or production evaluation data.

The timing matters. Article 50 of the EU AI Act requires providers of systems that generate synthetic text to make their output machine-readable and detectable where technically feasible, starting August 2, 2026.3 Anthropic signed the corresponding Code of Practice on transparency.

Theo Browne framed the same constraint in his video about Claude's announcement: images and video carry massive numerical redundancy, while text operates under discrete constraints.4 Dr. Mike Pound's Computerphile explanation led me to the green-list method introduced by Kirchenbauer and colleagues. A secret key partitions the vocabulary into favored and unfavored groups at each generation step.510

The difference in hiding space is structural. An image can distribute a small change across a field of pixels. A video can distribute it across pixels and time. Text has one narrow place to put the mark: the sequence of choices made as it is generated.

That is easier to see when the three media are placed side by side:

Video
spatial + temporal
◂time▸
Image
spatial
Text
sequential
The
small
boat
reached
the
quiet
shore
Frames × pixels. Each frame is a 2D grid. Time adds a third axis. Marks scatter across all three.
Pixels. One 2D grid. A few shifted values vanish among millions.
Tokens. One axis: the sequence. Each choice is visible. The mark is which token the model picked.
A text watermark can't scatter marks across a pixel grid. It can only lean on which word the model chose when several words would have worked.

The mark appears at generation time. A language model writes one token at a time, and at many positions several continuations are reasonable. A secret rule can give some candidates a small advantage without forcing the model to choose them every time.

The copied text carries no label. The signal is in the pattern of choices the model made, accumulated across many positions.

A detector that knows the rule can walk through the copied text, rebuild the favored group at each position, and count how often the model selected it. One position says almost nothing. Evidence accumulates over many positions, although the required length depends on the watermark profile, the tokenizer, and how many plausible choices the text offered the model.

At a high level, the mechanism has three steps:

1Generate

Bias pseudorandom sampling toward keyed subsets of valid tokens.

2Check

Reconstruct those subsets from context and count observed choices.

3Judge

Test whether the green count exceeds ordinary baseline chance.

I worked outward from the smallest test I could inspect. Weighted coins made the statistics visible. A 20-word vocabulary made the keyed partition possible to trace by hand. From there, I moved the same idea into a local language model, reproduced it with Gemma under paired controls, tested the checker on natural-web text, and measured what editing removed.

Start with the statistical engine

Start with two weighted coins

Begin with two coins. One lands heads 25% of the time; the other lands heads 40% of the time. Flip each coin 40 times and count the heads. If I hand you one unlabeled sequence, can you tell which coin produced it?

The 25% coin averages 10 heads in 40 flips. The 40% coin averages 16. Forty flips are not many, though, and the ranges overlap. The 25% coin can get lucky; the 40% coin can have a cold run. That overlap is the first lesson: the expected count is useful only when we compare it with the amount of ordinary variation.

Baseline coinNo added preference
p = 25%
H
About 1 head in 4
Nudged coinA persistent preference for heads
p = 40%
H
About 2 heads in 5

Watch the counts, not one particular flip.

Every run produces a different sequence, but the long-run averages stay near one head in four and two heads in five. A single unlabeled batch remains difficult to identify because chance can reverse the expected order. Detection works by accumulating a small preference across many observations, not by finding a telltale individual choice.

From head count to z score

Take a single observed batch: 32 heads in 80 flips.

The baseline coin produces an expected average of 20 heads in 80 flips (80 × 0.25 = 20). The batch contains 12 more heads than that expectation. A raw difference of 12 lacks meaning without a scale: 12 extra heads in 20 flips is extraordinary, while 12 extra in 20,000 is negligible.

Under the baseline model, 80-flip batches fluctuate by a standard deviation (σ) of about 3.87 heads around their average. Here T is the number of flips and p is the baseline probability of heads:

σ = √T × p × (1 − p) = √80 × 0.25 × 0.75 ≈ 3.87 heads

Dividing the observed excess by the expected baseline movement gives the z-score:

z = 12 extra heads3.87 heads of ordinary standard deviation = 3.10

The z-score measures how many baseline standard deviations separate the observed count from its expected value. It gives us a common scale for the excess. It does not, by itself, say that a watermark exists. That conclusion requires a background distribution, a decision rule, and controls.

Turning a score into an alarm requires a decision threshold chosen before the result is inspected. I set the experiment's cutoff at z > 3 before collecting data, so 3.10 crosses it. Under the exact binomial model, 32 or more heads in 80 flips occur about 0.224% of the time. The familiar one-sided normal tail beyond z = 3 is about 0.135%, but this discrete 80-flip example does not match that approximation exactly. Either rate is small enough to look persuasive in one trial and common enough to produce false alarms at scale.

More flips, clearer signal

The coin experiment gives us the intuition, but it also marks the boundary of the analogy. Coin flips are independent. Real text tokens are conditional on their predecessors, so the same variance calculation no longer has a guaranteed fit. Context, repeated tokens, and tokenizer behavior can inflate or suppress the green count.

The nudged source accumulates an expected excess of 0.15 × T heads. Baseline random fluctuations grow as √T × 0.25 × 0.75. The watermark signal grows linearly with length (O(T)), while random noise grows only with the square root of length (O(√T)).

At 40 flips, the expected excess is 6 heads and baseline noise is 2.74 (z = 2.19). At 400 flips, the expected excess reaches 60 heads while baseline noise reaches only 8.66 (z = 6.93). The length control holds both probabilities fixed while increasing T, making the signal separate from baseline variation.

Number of flips
Watch excess outrun scatter as T grows.
Baseline average10heads
Nudged average16heads
Persistent excess6heads
Expected z2.19movement units
Baseline movement
2.74
Persistent excess
6.00

A single z-score says nothing about error rates. Simulating 2,000 batches from each coin produces two distributions. They overlap heavily at short lengths and separate as the sample grows. The cutoff control shows how each threshold trades false alarms against missed detections.

Baseline and nudged score distributionsTwo score histograms and a movable cutoff.
Baseline over line0%
Nudged over line0%

The hills overlap. Move the cutoff and watch both rates change.

The same idea in code

Coins are tokens

A language model choosing its next token has the same basic counting structure, with one important qualification. Each eligible position is one flip. A token landing in the favored set counts as heads. The implementation reuses the z formula and the configured 25% null rate. Their interpretation does not transfer unchanged because token choices are dependent, which is why the later experiment measures an empirical background distribution.

Independence does not carry over. Coin flips do not depend on one another. Token choices depend on context, can repeat within a sentence, and pass through a tokenizer that may split a visible word into several tokens. Those dependencies can inflate or suppress the green count in ways a coin never would. This is why the later experiments need empirical background text rather than a coin formula alone.

Coin versionOne weighted flip
H

Heads is a hit. Tails is not.

becomes
Text versionOne next-token choice

The small boat was...

quietoldblueready

The key marks quiet and blue favored for this context.

CoinTextOne flipOne eligible token positionHeadsChosen token is in the favored setCount and score headsCount and score favored tokens

The first experiment, in code

The repository is small. src/watermark_lab/stats.py owns the statistics. labs/01_biased_coin.py reads the frozen configuration, simulates both sources, scores every batch, and writes raw rows.

The scorer is six lines:

def green_hit_z_score(*, hits: int, trials: int, null_probability: float) -> float:
    expected = trials * null_probability
    variance = trials * null_probability * (1.0 - null_probability)
    return (hits - expected) / math.sqrt(variance)

The simulator uses a local seeded generator. It never touches module-global random state:

def simulate_hit_counts(
    *, trials: int, hit_probability: float, replicates: int, seed: int
) -> tuple[int, ...]:
    generator = random.Random(seed)
    return tuple(
        sum(generator.random() < hit_probability for _ in range(trials)) for _ in range(replicates)
    )

The lab runs every configured length under both conditions:

for length in config.lengths:
    for condition in ("null", "biased"):
        probability = _probability(config, condition)
        seed = derive_group_seed(
            base_seed=config.base_seed,
            condition=condition,
            trials=length,
        )
        hit_counts = simulate_hit_counts(
            trials=length,
            hit_probability=probability,
            replicates=config.replicates,
            seed=seed,
        )
        for hits in hit_counts:
            z_score = green_hit_z_score(
                hits=hits,
                trials=length,
                null_probability=config.null_hit_probability,
            )

That loop produced 10,000 baseline batches and 10,000 nudged batches at each of five lengths.

Flips Nudged batches above cutoff Baseline batches above cutoff
40 21.33% 0.16%
80 54.20% 0.13%
160 88.62% 0.21%
200 95.23% 0.17%
400 100.00% 0.20%

The null rates wobble rather than falling smoothly. They are finite Monte Carlo estimates. The 100% at 400 means all 10,000 nudged batches crossed under this coin setup, not a promise about real model output. The chart shows both lines diverging:

Recorded Stage 1 detection rates by lengthNudged and baseline rates at five lengths.
Nudged batches above cutoff Baseline batches above cutoff10,000 batches per point, cutoff z > 3

The simulation shows why a fixed preference becomes easier to detect as the number of observations grows. It also makes the trade-off visible: move the cutoff and false alarms fall or missed detections rise. The coin experiment still leaves one problem unsolved. It tells us how to score heads, but not which text choices should count as heads. That requires a key tied to context.

The key mechanism

The coin had no way to decide which token choices should count as heads. The text version needs a rule that looks at a position and a candidate token, then answers the same question every time: does this candidate belong to the favored group for this context?

I built one with 20 visible words and one sentence small enough to trace by hand:

Early one morning Jack went up the hill.

The four tabs trace selection, score adjustment, generation, and checking on this sentence.

Four words used for this choice
SelectionThe key ranks all 20 candidates.
fixed draw 0.30
Green hits0
Checked positions0
Expected hits0
z score--

Selection. The program hashes the teaching key, the four context token IDs (Early one morning Jack), and each candidate ID. Sorting the 20 hashes gives a stable ranking. The first five become green: Early, went, walked, snow, and trail. Some are poor continuations because the selector reads token IDs, not grammar. The comparison key changes eight memberships while preserving the five-of-20 fraction.

Score increase. The program adds 2 to the five green logits before normalization. The relative odds of each green word therefore multiply by exp(2), about 7.39. In this step, went rises from 22.85% to 46.51%, while the unmodified ran falls from 27.91% to 7.69% because all candidates share the new probability total.

Generation. The saved draw of 0.30 selects walked from the original distribution and went after the increase. The program appends went, drops Early from the four-token window, and rebuilds the favored set from one morning Jack went. The first two generated words land in green; the last two win despite falling outside the favored set. The watermark changes the odds without dictating every token.

Checking. The checker replays the selection rule through the copied sentence. It rebuilds the favored group before each of the four generated words without access to generation scores or random draws. The result is G=2 hits across T=4 positions, z = 1.155. The comparison key produces zero hits.

The key printed here is for teaching. Anyone can read it and study how the pattern works or how edits disrupt it. A production service would keep the generation key out of prompts, browser code, and public logs. A production service would keep the generation key out of prompts, browser code, and public logs.

What happens in a real model

One real model step

The coin experiment is clean because we control every probability. The next question is whether the same score increase produces a detectable excess inside a real language model, where scores span five orders of magnitude and the sampler applies temperature, top-k, and top-p before the watermark takes effect?

I loaded mlx-community/LFM2-350M-4bit at revision 18dc72abf3b2337f9123cfd6eeeb58dfa7947066 on an Apple GPU with MLX-LM 0.31.3 and MLX 0.32.0. I then ran one autoregressive loop from the prompt "Early one morning Jack went up the hill. At the top he". The control and marked paths shared the same model state and seed. The figure compares the raw distribution with the distribution after adding 2 to green logits at the first generated position.

Model input ends here

Early one morning Jack went up the hill. At the top he

Control continuation

Marked continuation

Marked text, comparison key

The copied text stays fixed. Only the checker key changes.

The table makes the first-step effect concrete. The score increase does not replace the model's distribution; it reshapes it. He begins with the highest raw score but falls behind four green candidates after the processor runs.

Candidate Raw score Green Score after increase Final marked chance
As 15.1875 yes 17.1875 34.715%
he 14.8125 yes 16.8125 21.724%
Jack 14.6875 yes 16.6875 18.582%
The 13.5000 yes 15.5000 4.211%
He 15.3125 no 15.3125 3.332%

Jack rose from 11.642% to 18.582%. The saved draw chose Jack in both paths, a non-event that matters: two distributions can return the same token. A later draw split the continuations, after which each path conditioned on its own history.

The marked path began, "Jack climbed slowly, his boots sinking slightly into the soft snow-covered earth." I scored the copied continuation across 39 eligible positions. The generation key produced 21/39, z 4.160. The comparison key on the same text produced 7/39, z -1.017. The paired control with the generation key produced 8/39, z -0.647.

All three fixed marked passages scored higher than their paired controls. That is useful as a smoke test because it shows that generation and detection are connected. It is not an accuracy study. Three passages cannot estimate detection performance or text quality.

Operation order

Recording delta=2 doesn't fully specify a watermark. My first loop applied the increase before temperature and filtering. Transformers 5.14.1 applies them differently:

Temperature→Top-K→Top-P→Watermark Processor→Softmax→Sample

My earlier teaching loop used:

Watermark Bias→Temperature→Top-P→Top-K→Softmax→Sample

I replayed both sequences on the same 50,257 saved GPT-2 scores. The operation control exposes the candidate count and selected-token probability after each step.

A separate compatibility checkDoes a repeated pair count once or every time it appears?

The Transformers route kept 40 after top-k and 19 after top-p. My earlier route kept 11 after top-p. For token was, final probabilities: 8.643% vs 8.826%. Only 0.18 percentage points apart, but the structural difference matters. Temperature changes the effective increase. Filtering can remove a candidate before the watermark reaches it. The operation sequence is part of the watermark profile.

A six-token compatibility fixture exposed another mismatch. It alternated token IDs 373 and 21272, producing five pair occurrences but only two distinct pair values. Both ignore_repeated_ngrams settings in Transformers returned 3/5, z 1.807. My explicit distinct-value count returned 1/2, z 0.816. Maintained behavior must be inspected directly rather than inferred from an option name.

Gemma end-to-end

The watermark core should not depend on how a specific model formats chat. The model adapter and the watermark logic should be separable so that each can be inspected on its own. I kept the shared interface small: pass a watermark profile into generation, extract copied assistant text, build the matching checker. A Gemma-specific adapter owned prompt rendering, tokenization, and generated-ID slicing.

I pinned google/gemma-4-E2B-it at revision 3e22461f65e89153144f8adb70e3b8c2cc9845a7 in BF16 on one Modal NVIDIA L4 with Transformers 5.14.1. Both the control and watermarked calls share every argument except one:

CONTROL
model.generate(
    input_ids=encoded.input_ids,
    attention_mask=encoded.attention_mask,
    do_sample=True,
    temperature=0.8,
    top_k=40,
    top_p=0.95,
)
WATERMARKED
model.generate(
    input_ids=encoded.input_ids,
    attention_mask=encoded.attention_mask,
    do_sample=True,
    temperature=0.8,
    top_k=40,
    top_p=0.95,
    watermarking_config=profile.to_transformers(),
)

The component trace shows which data cross each boundary and where the private key enters.

The model loaded in 5.8 seconds. The three marked smoke outputs generated at 18.422, 18.747, and 19.259 tokens per second. Each appears beside its paired control below.

CONTROL

G / T
z
WATERMARKED

G / T
z
Control z
Watermarked z
z = 3

A later natural-length ladder produced 12 marked and 12 paired control outputs. Eight marked rows crossed z > 3; no control did:

Length ladder: 12 paired control and watermarked z scores

Prompt content and length changed together, so the ladder doesn't isolate length as the cause. The committed evidence uses a public key so anyone can verify it. A private service would keep key material inside the host process and expose a version identifier, not the key itself.

Test the detector before trusting it

Score outside text before trusting the cutoff

A detector score is only useful if we know how often ordinary text produces the same score. I therefore tested the checker on text that did not receive this experiment's watermark. Before treating a crossing as evidence, I ran the same checker on text that did not receive this experiment's watermark.

I scanned the pinned C4 realnewslike validation shard in file order. A passage needed at least 500 Gemma tokens, at least 65% Unicode letters among non-whitespace characters, and no duplicate text, code dump, or obvious list structure. The selector scanned 2,479 rows, rejected 1,451 as too short and four as obvious lists, froze the first 1,000 passing rows for calibration, and reserved the next 24 for paired generation. Detector scores played no role in that split.

C4 is natural-web text scraped from Common Crawl. The corpus contains no verification that any passage was written by a human.

With the public Gemma key and all-pair counting, the 1,000 scores had a median of 0.029, a 99th percentile of 2.457, and a maximum of 3.729. Four rows crossed strict z > 3. A thousand rows can't validate one-in-100,000 behavior, but four crossings in a declared negative set are enough to distrust the cutoff in isolation.

The maximum row exposed the counting rule. Counting every adjacent-pair occurrence gave 132/399, z 3.729. Counting each pair value once on the same token sequence gave 114/358, z 2.990. The second rule removed 41 observations, including 18 green hits, and moved the row below the cutoff.

Fernandez and colleagues showed that standard asymptotic tests underestimate false positives on short or repetitive text.12 The z formula stays useful because every term is visible, but it must travel with the empirical background and the repetition policy that produced it.

The sorted calibration view compares the first 100 rows with the full 1,000. The larger set reveals the background distribution and all four crossings.

Median0.0289
99th percentile2.4568
Strict crossings4 / 1,000
Maximum3.7286
Sorted natural-web z scoresEvery frozen all-pair score sorted from low to high with the strict z greater than three line.
Keep the text fixedThe counting rule moves the maximum row across the cutoff.
Observed green hits132
Eligible pair observations399
z score3.7286
Strict z > 3Crosses

Every adjacent-pair occurrence counts, including repeated values.

Distinguishing the watermark from model artifacts or domain noise requires three paired controls alongside the marked score.

One high score needs three controls

I froze 24 paired prompts before generation. Each pair shared its 50-token source prefix, prompt-derived seed, model revision, sampler, and 400 generated-token safety cap. Only the marked call received the watermark configuration.

A high score on the marked text under the generation key still leaves three explanations open: ordinary Gemma output might score high under that key, natural text from the source domain might score high, or the marked text might score similarly under a different key. Rank 1000 gives all four checks at 160 copied tokens:

Checked text and key Green hits Eligible checks z
marked text, generation key 58 159 3.3424
paired model control, generation key 47 159 1.3278
natural-web continuation, generation key 43 159 0.5952
marked text, comparison key 29 159 -1.9688

Only the marked-correct-key condition crosses, while the three controls stay below the line. Rank 1001 complicates the picture: its marked and control paths shared every token ID through the first 80 copied tokens, both scoring 26/79, z 1.624. The watermark changed probabilities, but these seeded draws followed the same early path.

The four branches hold the text, key, or source condition apart so each alternative explanation can be checked separately.

Frozen prompt rank 1000One source prefix, paired generation seed, and model profile
The inconvenient rowRank 1001 followed the same early path.
Marked, generation key26/79z 1.6239
=
Model control, generation key26/79z 1.6239

The marked and control token IDs were identical through 80 copied tokens. The watermark changed probabilities, but these seeded draws did not split yet.

A single row confirms the path runs without estimating accuracy. The full 24-row cohort shows whether the separation holds across documents.

Read every paired row before the average

All 24 pairs completed at 80 copied tokens. I plotted each document's difference on its own line before calculating a mean.

Contrast Mean paired z difference 95% paired bootstrap interval
paired model control 1.8296 [1.3424, 2.3276]
natural web 1.7538 [1.3100, 2.1977]
comparison-key replay 2.0461 [1.6131, 2.4792]

The averages separate cleanly, but the individual rows are less tidy. That distinction matters because a detector will eventually be applied to one document at a time. Some differences are small, rank 1001 is exactly zero against its model control, and at least one row points against the mean. Three marked rows crossed z > 3; none of the three control families did at this prefix. The intervals summarize 24 frozen documents, not a population.

The matched cohort shrank with prefix length. Counts at 40, 80, 160, 200, and 400 copied tokens were 24, 24, 21, 17, and 0. A 400-token generation cap meant both the marked and control outputs terminated before reaching 400 copied tokens, so no pair completed at that length. At 200 tokens, four of 17 marked rows crossed, along with one natural-web row. Because the documents changed as the prefix grew, that sequence can't support a clean causal claim about length.

The contrast controls preserve row order, allowing each document to be tracked across comparisons.

Twenty-four labeled paired z differencesEach frozen selection rank on its own row, with zero, mean, and paired interval.
Matched rows by copied-token prefix

These scores come from unedited text. Editing rewrites the checker's history.

What editing does to the signal

Edits rebuild the checker history

Rank 1000 scored 28/79, z 2.144, at its first 80 copied token IDs. A deterministic 10% word deletion scored 25/79, z 1.364. A Gemma paraphrase scored 26/79, z 1.624.

An edit does not peel a label off the text. It changes the string itself, which changes the history the checker uses to reconstruct the mark.

That distinction is easy to miss. The watermark is not a hidden label attached to a sentence. It is a pattern in the sequence of token choices, and editing changes that sequence. Tokenization creates a different ordered history, so the checker rebuilds different keyed groups at every later position.

I carried the first 12 marked outputs through eight conditions:

Edit Mean z change Mean copied-token length ratio
normalization 0.0000 0.9986
homoglyph 1% -0.0217 1.0448
homoglyph 5% -0.9311 1.2183
deletion 10% -0.3248 0.8980
deletion 30% -0.9960 0.7021
mixing 25% -0.6712 1.0012
mixing 50% -1.3424 1.0088
paraphrase -1.7105 0.9636

Homoglyph substitution swapped ASCII letters for visually identical Unicode code points. For example, replacing Latin 'a' (U+0061) with Cyrillic 'а' (U+0430) looks the same to a reader, but the tokenizer splits the word into unknown byte sequences, scrambling the context hashes for every later position.

Deletion and mixing can damage grammar or change claims. A lower detector score alone says nothing about whether the meaning survived, so meaning preservation has to be measured separately.

All 12 paraphrases passed the declared length, decimal-number, and embedding-cosine screens. A non-independent assistant review marked ten pass, two uncertain. Every passed rewrite reduced z, and no paraphrase crossed the cutoff.

James Padolsey's Declaude explainer reports known-key rewrite tests against open KGW and EXP implementations.13 Padolsey reports that full rewrites left about 0.5% of the original windows intact and reduced detection on those implementations to roughly chance. Those measurements do not cover Claude's private SynthID configuration or replace this project's results.

The edit view follows rank 1000 through deletion and paraphrase. It holds the key and checker profile fixed while the visible string, token history, and keyed decisions change.

Unedited rank 1000

All 12 frozen rowsSelect an edit to see every paired change.
Twelve labeled paired z changes after editingEach frozen selection rank on its own row for the selected edit.

A lower z score does not prove that an edit preserved meaning. Paraphrase length, number checks, embedding cosine, and assistant review were recorded separately; two reviews remained uncertain.

A stronger mark costs something

The default experiment added delta=2 to green scores. I held the eight prompts, model, key, sampler, and 400-token safety cap fixed, then changed delta. The generated continuations reached different lengths, so achieved copied length remains a separate measurement rather than a controlled constant.

Delta Mean z Strict crossings Mean conditional NLL Mean repeated-pair fraction Mean copied tokens
1 0.2923 0/8 0.5004 0.0373 232.50
2 2.1761 1/8 0.5415 0.0471 265.75
3 2.4684 3/8 0.5783 0.0483 271.75

Mean z, conditional NLL, repeated adjacent-pair fraction, and achieved copied length all rose with delta. Higher conditional NLL means the pinned Gemma checkpoint assigned less probability to the recorded continuation. It does not show that a reader would find the text less fluent or useful. Ranks 1004 and 1006 had lower z at delta 3 than at delta 2, so individual paths don't climb monotonically.

NLL and repetition are model-based proxies. They cannot establish factual accuracy, reader preference, or a universal setting from eight prompts.

Each line in the delta sweep represents one frozen prompt. The coral paths mark ranks 1004 and 1006, whose scores fell from delta 2 to delta 3.

Eight labeled z-score paths across delta one, two, and threeEvery frozen prompt path with two non-monotonic rows labeled.

Conditional NLL and repeated pairs are model-based proxies. They do not replace a human quality or factuality study.

The delta sweep tested KGW-style green lists only, and Claude's production watermark belongs to a different family.

Claude, other detectors, and the law

Claude uses a SynthID-Text variant

The experiments above implement a KGW-style green-list watermark, while Anthropic says Claude's production mark uses a version of SynthID-Text. These experiments explain the statistical idea without reproducing Claude's system.

KGW uses the key and recent context to select a vocabulary subset, raises the scores of tokens in that subset, and later counts the excess selected tokens.10

SynthID-Text uses keyed tournament sampling.11 It draws candidates from the model distribution, lets keyed scoring functions select tournament winners, and measures the resulting correlation during detection. The paper also describes repeated-context masking and settings with different distortion guarantees. Its live quality evaluation compared about 20 million watermarked and unwatermarked Gemini responses, using voluntary thumbs feedback as a proxy for quality rather than as a test of removal attacks or detector accuracy.

Anthropic calls Claude's watermark "a version of the SynthID-Text approach."2 The company says its implementation changes the randomness used to choose among suitable words and adds no hidden characters or extra tokens. It also says factual passages, light editing, and code often provide too few choices to carry much evidence. A detection API is planned.

Anthropic reported no practical quality loss in its internal tests. It has not published Claude's tournament settings, key construction, context masking, scorer, threshold, model coverage, or evaluation data needed to verify that result independently. The KGW experiments here explain a related idea, but they do not validate Claude's production system.

Other systems answer different questions. Ghostbuster and DNA-GPT detect machine-generated text without a generation-time key.78 EditLens estimates how much an AI system edited a document.9 The comparison below separates these tasks from keyed watermark detection.

Keep unlike detectors separateFour methods answer different questions.Select a family to compare what it changes, what it measures, and what its score can support.
No method shown here establishes authorship, intent, ownership, or misconduct.The green-list experiment supports one narrower statement: consistent with this configured watermark and key.

Article 50 requires machine-readable marking

Article 50(2) helps explain why providers are developing marking systems.3 It requires providers of AI systems, including general-purpose AI systems, that generate synthetic audio, images, video, or text to mark their outputs in a machine-readable format so detectors can identify them as artificially generated or manipulated. The duty applies as far as technically feasible and accounts for content limits, implementation costs, and the state of the art. It also exempts standard editing and uses that do not substantially alter the supplied input or its meaning.

The law sets a transparency duty. It does not prescribe SynthID-Text, turn a detector score into proof of authorship, or tell a school or employer how to judge a person.

Claude’s mark does not identify a user

Anthropic says Claude's watermark and key contain no information about a user, organization, or chat.2 Under that design, the detector asks whether the text carries evidence associated with Claude's watermark. It cannot recover who requested the text.

That claim concerns Claude's disclosed design. It does not establish that every possible text watermark carries only one bit or that a provider holds no separate account records.

Detection does not settle authorship

A positive result can indicate Claude involvement, according to Anthropic, without showing that Claude wrote every word. Authorship asks who supplied the ideas and accepted responsibility. A policy decides which assistance was allowed, and a disciplinary process must weigh evidence beyond one score.

A negative result is weaker still. Short answers and exact code may offer few choices to mark. Light editing may leave too few Claude-selected tokens, while a wrong key, an older model, or later revisions can also suppress the score. The edit experiments above show how a rewritten token history can weaken the signal even when the meaning survives.

Passive classifiers create a related risk. Sean Goedecke reports that students he knows rewrite their own prose or record their drafts because they fear false accusations.6 His examples are anecdotal, but the policy error is clear: an uncertain detector score cannot carry the burden of proof by itself.

The evidence in this article supports one narrow statement:

Consistent with this configured watermark and key.

What the experiment actually shows

The experiment supports a narrower claim than the word "watermark" often suggests. With enough matching text and the correct profile, a detector can recover evidence of the token-selection bias used in this KGW-style experiment. That evidence gets weaker when the passage is short or when an edit changes the token history.

At 80 copied tokens, all 24 Gemma pairs completed generation. The marked outputs scored higher on average than the paired model controls, natural-web continuations, and comparison-key replays. The 95% paired bootstrap interval for each contrast excluded zero. The rows were not perfectly tidy: rank1001 matched its control exactly, and at least one difference pointed against the cohort mean.

The cutoff was less trustworthy than a single score might imply. Four of 1,000 C4 passages crossed z > 3 when the checker counted every adjacent-pair occurrence. One crossed when it counted each pair value once. The background corpus and the repetition policy are part of the detector. They cannot be treated as implementation details.

Editing weakened the mark by different amounts. Normalization left the mean score unchanged, while deletion, mixing, homoglyph substitution, and paraphrasing reduced it. All 12 paraphrases passed the declared automatic screens and scored lower than their source passages. The assistant review rated ten pass and two uncertain, so the meaning-preservation result is limited to that review and its stated uncertainty.

Raising δ increased the cohort's mean detection score, conditional NLL, and repeated-pair fraction. Two of eight prompt paths still scored lower at δ = 3 than at δ = 2. These are model-based proxies, not human judgments of fluency, factuality, or preference.

I would not use these measurements to estimate production false-alarm rates, evaluate Claude's private watermark, or judge authorship. The honest conclusion is smaller: this kind of detector can recover a statistical trace under controlled conditions. It cannot, by itself, tell us who wrote a passage, why they wrote it, or whether a policy was broken.

The code, frozen traces, generation scripts, and verification artifacts are available in the text-watermarking-lab repository.14

Sources and evidence

  1. Anthropic, How Claude marks AI-generated content, published 2026-08-11. This support page announced model-level text watermarking before the fuller method post on August 14.
  2. Anthropic, How Claude's text watermark works, published 2026-08-14. Provider claims about Claude's method, quality, rollout, detection API, code, editing, and user identity come from this announcement.
  3. European Union, AI Act Article 50, and European Commission, Code of Practice on Transparency of AI-generated Content. These sources establish the transparency requirement and its technical qualifications, not a legal assessment of any provider's compliance.
  4. Theo Browne, Claude watermarks your code now, 2026-08-14, 31:58.
  5. Computerphile with Dr Mike Pound, Ch(e)at GPT?, 2023-02-16. Used for green-list intuition.
  6. Sean Goedecke, AI detection tools cannot prove that text is AI-generated, 2025-12-05. His student examples are personal reports rather than a prevalence study.
  7. Verma et al., Ghostbuster: Detecting Text Ghostwritten by Large Language Models, submitted 2023-05-24; NAACL 2024.
  8. Yang et al., DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text, submitted 2023-05-27.
  9. Thai et al., EditLens: Quantifying the Extent of AI Editing in Text, submitted 2025-10-03.
  10. Kirchenbauer et al., A Watermark for Large Language Models, ICML 2023.
  11. Dathathri et al., Scalable watermarking for identifying large language model outputs, Nature 634, 818-823, 2024.
  12. Fernandez et al., Three Bricks to Consolidate Watermarks for Large Language Models, version inspected 2023-11-08.
  13. James Padolsey at NOPE, How AI text watermarking works, inspected 2026-08-17. Its open-model removal measurements are self-reported and do not describe Claude.
  14. Jay Shah, text-watermarking-lab.

Comments