Modern AI Engineering

Lesson 15.3 · 25 min

LLM Watermarking: Embedding Invisible Signatures in AI Text

Can we hide an invisible signature in AI-written text using nothing but the choice of words, so that only someone with a secret key can find it?

In short: LLM watermarking hides a statistical signal in generated text. At each step, a secret key and the previous token pick a random “preferred” (green) subset of the vocabulary, and the model's scores for those tokens are nudged up slightly. Any single word looks normal, but over hundreds of words the text contains far more green tokens than chance would give, which a detector with the key can measure with a simple z-score. It works well on long, unedited text and weakens with short texts and heavy paraphrasing.

What is a watermark, and why put one in LLM text?

A watermark is a hidden or subtle mark embedded in something to show where it came from, ideally without spoiling it. Banknotes carry watermarks visible against the light; photos can carry invisible digital watermarks in their pixels. An LLM watermark is a hidden pattern in the words a language model chooses, which a detector can later check.

Why would we want that? As AI-written text becomes common and hard to tell apart from human writing, several groups want a reliable way to answer “Did our model write this?”

  • Provenance and transparency: platforms and regulators increasingly expect AI-generated content to be identifiable. The EU AI Act, for example, includes transparency obligations for marking AI-generated content in a machine-readable way.
  • Misinformation and spam: spotting mass-produced AI text in reviews, comments or fake news campaigns.
  • Education: helping (carefully) with questions of AI-assisted homework.
  • Training data hygiene: model builders may want to filter AI-generated text out of future training data.

Think of it like a secret dice-rigged word game Imagine a writer who, whenever two words fit equally well, flips a secret coin that only they and a friend can predict, and picks the word the coin favours. Readers notice nothing; both words fit. But the friend, who can replay the coin flips, sees that the writer “won” the coin far more often than luck allows. That excess of wins is the watermark.

How does an LLM write text and choose the next word?

An LLM writes one token (a word or word piece) at a time. At each step it outputs a score, called a logit, for every token in its vocabulary (often 30,000 to over 200,000 tokens). The softmax function turns these scores into probabilities that sum to 1, and a sampler picks the next token at random according to those probabilities (temperature and top-p settings reshape them). The chosen token is appended, and the process repeats.

The hidden freedom that makes watermarking possible

Take the sentence “The storm caused ___ damage.” Plausible next words include severe, major, extensive, significant, serious, widespread. Each might get 10–20% probability. Any of them gives a perfectly good sentence. This happens constantly in natural language: at most positions there are several good choices.

That freedom is the hiding place. If we gently bias which of the good choices gets picked, in a pattern only we can predict, the text stays natural but carries a signal. At positions with no freedom (after “Barack”, the next token is almost surely “Obama”), the watermark simply cannot act, and that is fine.

This idea, in the form described below, comes from Kirchenbauer and colleagues' 2023 paper A Watermark for Large Language Models, often called the green-list or red/green watermark.

The secret key, preferred tokens, and nudging the probabilities

Generating one watermarked token

  1. Seed from key + previous token: Combine a secret key with the previous token (for example, hash them together) to get a random seed. The same key and previous token always give the same seed.
  2. Split the vocabulary: Use the seed to randomly mark a fraction γ (gamma, say 25%) of the vocabulary as green (preferred tokens). The rest are red (other tokens).
  3. Nudge green logits: Add a small constant δ (delta, say 2.0) to the logit of every green token. Red tokens are unchanged.
  4. Softmax and sample as usual: Green tokens are now more likely, but red tokens can still be chosen when they are clearly the best fit.
  5. Repeat: The next position uses the newly chosen token as its “previous token”, so it gets a completely different green list.

A small worked example. Four candidate words have equal logits, so each has 25% probability. Two of them happen to be green. After adding δ = 2 to the green ones, each green word's weight becomes e² ≈ 7.39 versus 1 for each red word. Green probability per word = 7.39 / (2 × 7.39 + 2) ≈ 0.44, red ≈ 0.06. The text still uses a fitting word, but very likely a green one.

Pause and think: Pause and predict: why does the green list depend on the previous token instead of being one fixed list of words for the whole text?

A fixed list would make the text overuse the same set of words (noticeable, and harmful to quality), and an attacker could discover the list just by counting word frequencies across many watermarked texts. Re-randomizing at every position spreads the bias evenly over the vocabulary and keeps the pattern invisible without the key.

One token vs thousands of tokens: how detection works

One green token proves nothing: even human text hits the green list about γ = 25% of the time by chance. The evidence comes from counting over many tokens. The detector does not need the model at all, only the text, the tokenizer and the key.

Detecting the watermark

  1. Tokenize the text: Split the suspect text into tokens with the same tokenizer.
  2. Recompute each green list: For each position, hash the key with the previous token to rebuild that position's green list.
  3. Count green hits: Count how many tokens fall in their position's green list: |s|_G out of T tokens.
  4. Compute a z-score: Compare the count with what chance predicts (γ·T) in units of standard deviation.
  5. Decide: If z exceeds a threshold (for example 4), flag the text as watermarked. A high threshold keeps false accusations of human text extremely rare.

green_list_watermark.py

import hashlib
import numpy as np
V, GAMMA, DELTA, T = 1000, 0.25, 2.0, 200    # vocab, green share, nudge, length
rng = np.random.default_rng(0)
def green_list(prev_token, key=b"secret"):
# Hash (secret key, previous token) -> seed -> a fresh 25% "green" subset
h = hashlib.sha256(key + prev_token.to_bytes(4, "little")).digest()
seed = int.from_bytes(h[:8], "little")
return set(np.random.default_rng(seed).permutation(V)[: int(GAMMA * V)])
def generate(watermark):
tokens = [0]
for _ in range(T):
logits = rng.normal(0, 2, V)             # stand-in for the model's scores
if watermark:
logits[list(green_list(tokens[-1]))] += DELTA   # nudge green tokens up
p = np.exp(logits - logits.max())
p /= p.sum()
tokens.append(int(rng.choice(V, p=p)))
return tokens
def z_score(tokens):
# Count green hits; compare with what chance (GAMMA) would give
n = len(tokens) - 1
hits = sum(t in green_list(prev) for prev, t in zip(tokens, tokens[1:]))
return hits / n, (hits - GAMMA * n) / np.sqrt(n * GAMMA * (1 - GAMMA))
def edit(tokens, frac):
out = list(tokens)
for i in rng.choice(np.arange(1, len(out)), int(frac * (len(out) - 1)), replace=False):
out[i] = int(rng.integers(V))            # swap a word for a random one
return out
wm, plain = generate(True), generate(False)
for name, toks in [("unwatermarked", plain), ("watermarked", wm),
("wm, 30% edited", edit(wm, 0.3)), ("wm, 60% edited", edit(wm, 0.6)),
("wm, first 16", wm[:17])]:
frac, z = z_score(toks)
print(f"{name:15} tokens={len(toks) - 1:3}  green={frac:.2f}  z={z:5.2f}  flagged={z > 4}")

Output:

unwatermarked   tokens=200  green=0.20  z=-1.63  flagged=False
watermarked     tokens=200  green=0.67  z=13.72  flagged=True
wm, 30% edited  tokens=200  green=0.48  z= 7.51  flagged=True
wm, 60% edited  tokens=200  green=0.33  z= 2.45  flagged=False
wm, first 16    tokens= 16  green=0.62  z= 3.46  flagged=False

How is this different from an AI text detector?

Why the quality of the text does not break

  • The nudge is small. δ shifts preferences between good options; it does not force a green token. A red token that is clearly best still wins.
  • It only acts where there is freedom. In low-entropy positions (one obvious next token), the bias changes nothing. In high-entropy positions, any choice was fine anyway.
  • The green list changes every step, so no word gets systematically overused.
  • Strength is tunable. Smaller δ means less effect on text but weaker detection; production systems tune this trade-off.
  • Some schemes aim to be distortion-free, choosing tokens with keyed randomness so that, averaged over keys, the output distribution matches the original model exactly.

Google DeepMind's SynthID Text, described in a 2024 Nature paper and used in Gemini, uses a related keyed scheme (a “tournament” over candidate tokens). The authors reported a large live experiment in which users did not rate watermarked responses noticeably differently from unwatermarked ones. Exact quality impact still depends on the scheme, the settings and the task; code and very factual text offer less freedom and therefore weaker watermarks.

What happens when someone edits the text?

Editing replaces tokens, and each replaced token is green only by chance (25%). Because each green list depends on the previous token, changing one token can also disturb the check for the token after it. So edits pull the green fraction down toward γ and the z-score toward zero.

  • In our run, editing 30% of the tokens dropped z from 13.72 to 7.51: still clearly flagged.
  • Editing 60% dropped z to 2.45: below the threshold, the watermark is effectively gone.
  • A 16-token snippet of perfectly watermarked text gave z = 3.46: not enough evidence on its own.

Watermarks can be removed Paraphrasing with another model, translating to another language and back, or tricks like asking the model to insert and later delete filler characters can wash the signal out. Mixing a little AI text into a long human document dilutes it. Watermarks are evidence that is cheap to check when present; their absence does not prove a human wrote the text, and their presence in a short snippet is weak evidence.

Pause and think: From the output: unwatermarked text scored z = −1.63 and the 60%-edited text scored z = 2.45. Should we conclude the edited text is human-written?

No. A z below the threshold only means we lack strong evidence of the watermark. The text may still be AI-generated and edited, generated by a model without a watermark, or simply too short. Watermark detection can support a “yes”, but a “no” is not proof of human authorship.

Where LLM watermarking is used, and its advantages and disadvantages

Real-world use Google DeepMind applies SynthID watermarking to Gemini text output and released an open-source implementation of SynthID Text (for example via Hugging Face Transformers) in 2024; the SynthID family also covers images, audio and video. OpenAI publicly said in 2024 that it had built a text watermarking method but had not released it, citing concerns such as easy circumvention and effects on some user groups. Image generators widely combine invisible watermarks with C2PA content credentials. For text, deployment remains limited and differs across providers.

AdvantagesDisadvantages
Invisible to readers; little or no visible quality cost when tuned wellOnly detects text from models that embed it; open-weight models can simply skip it
Detection needs only the key and tokenizer, not the modelNeeds enough tokens; short texts give weak evidence
Controllable, very low false-positive rate on human textParaphrasing, translation and heavy edits weaken or remove it
Robust to light edits, cropping and copy-pasteLow-entropy text (code, facts, lists) carries a weaker signal
Keys can be kept private to resist forgeryKey management and who may run detection raise trust and privacy questions

Going one level deeper

We used a threshold of z = 4 without saying where it comes from. The threshold is a choice about how often we are willing to accuse text that carries no watermark. For such text, the z-score behaves roughly like a standard bell curve, so each threshold maps to a chance of a false flag (using the normal approximation).

What a threshold means when many texts are checked (normal approximation)
Threshold zChance that one unwatermarked text is flaggedExpected false flags in 1,000,000 texts
2about 1 in 44about 22,750
4about 1 in 31,600about 32
6about 1 in 1 billionabout 0.001

A threshold of 2 sounds strict for a single text, but at the scale of a platform it wrongly flags tens of thousands of writers. That is why detectors use high thresholds, and why a high threshold in turn needs longer texts.

Dilution: AI text mixed into human text

  1. Set up: A 1,000-token document. Human-written tokens are green 25% of the time (pure chance); watermarked tokens 67% of the time, as in our run. Chance predicts 250 green tokens with a standard deviation of √(1000 × 0.25 × 0.75) ≈ 13.7.
  2. 300 AI tokens, 700 human: Green count ≈ 0.67 × 300 + 0.25 × 700 = 201 + 175 = 376. z = (376 − 250) / 13.7 ≈ 9.2. Still clearly flagged.
  3. 100 AI tokens, 900 human: Green count ≈ 67 + 225 = 292. z = (292 − 250) / 13.7 ≈ 3.1. Below the threshold: the watermark is there, but it is drowned out.
  4. What a detector can do: Score sliding sections of the document instead of the whole. The 100 AI tokens alone would give z = (67 − 25) / √(100 × 0.1875) ≈ 9.7.

So a low score for a whole document does not rule out a watermarked passage inside it, and scoring many sections brings back the first problem: every extra test is another chance of a false flag, so the threshold has to rise with the number of sections checked.

Practice: try it yourself

We will build a small detector calculator using only the z-score formula. It answers three practical questions: how strong is the evidence for a given green count, how often does a threshold accuse unwatermarked text, and how many tokens do we need before a watermark of a given strength can be detected?

practice_watermark_calculator.py

import math
GAMMA = 0.25                      # share of the vocabulary that is green
def z_score(green_hits, total):
return (green_hits - GAMMA * total) / math.sqrt(total * GAMMA * (1 - GAMMA))
def false_positive_rate(z):
# chance that unwatermarked text scores above z (normal approximation)
return 0.5 * math.erfc(z / math.sqrt(2))
def tokens_needed(green_rate, threshold=4.0):
# smallest T whose expected z-score reaches the threshold
t = 1
while z_score(green_rate * t, t) < threshold:
t += 1
return t
for hits, total in [(9, 20), (45, 100), (90, 200)]:      # all are 45% green
print(f"{hits:>3}/{total:<3} green -> z={z_score(hits, total):5.2f}")
for z in [2, 4, 6]:
print(f"threshold z={z}: about 1 in {1 / false_positive_rate(z):,.0f} "
f"unwatermarked texts flagged")
for rate in [0.70, 0.50, 0.40, 0.30]:
print(f"green rate {rate:.2f}: about {tokens_needed(rate)} tokens to reach z=4")

Output:

  9/20  green -> z= 2.07
45/100 green -> z= 4.62
90/200 green -> z= 6.53
threshold z=2: about 1 in 44 unwatermarked texts flagged
threshold z=4: about 1 in 31,574 unwatermarked texts flagged
threshold z=6: about 1 in 1,013,594,692 unwatermarked texts flagged
green rate 0.70: about 15 tokens to reach z=4
green rate 0.50: about 48 tokens to reach z=4
green rate 0.40: about 134 tokens to reach z=4
green rate 0.30: about 1200 tokens to reach z=4

Now change it:

  • Set GAMMA = 0.5. Predict the sign of the z-score for 45/100 before running, and explain it.
  • Call tokens_needed(0.40, threshold=5.0). Predict whether it needs a few more tokens or many more than the 134 needed for z = 4.
  • Add the case (25, 100) to the first loop. Predict its z-score without computing anything.

Pause and think: At a 70% green rate, 15 tokens are enough to reach z = 4. At 30%, it takes about 1,200. Why does a weaker watermark cost so many more tokens?

What counts is the excess over chance. At 70% the excess is 45 points; at 30% it is only 5 points, 9 times smaller. The z-score grows with the excess times √T, so the length needed grows with 1 / excess². A 9 times smaller excess needs 81 times more tokens: 15 × 81 ≈ 1,200.

Pause and think: A platform scores 1,000,000 human-written posts a day with a threshold of z = 2. Roughly how many are wrongly flagged, and what is the trade-off in fixing it?

About 1 in 44, so roughly 22,750 posts a day. Raising the threshold to 4 cuts that to about 32. The trade-off is sensitivity: a higher threshold needs more tokens or a stronger watermark to be reached, so short or heavily edited watermarked texts will more often go undetected.

Key takeaways

  • LLM watermarks exploit the freedom of choosing among several good next tokens.
  • A secret key plus the previous token selects a fresh green list; green logits get a small boost δ.
  • Detection counts green tokens and computes a z-score; evidence grows with the square root of length.
  • Unlike AI-text classifiers, watermarks give controllable false-positive rates but only for cooperating models.
  • Short texts, paraphrasing and heavy editing weaken the signal; absence of a watermark proves nothing.

Key terms

  • Watermark: A hidden mark embedded in content to indicate its origin.
  • Logit: The raw score a model assigns to each vocabulary token before softmax.
  • Green list: The keyed, per-position subset of preferred tokens whose logits are boosted.
  • γ (gamma): The fraction of the vocabulary placed on the green list at each step.
  • δ (delta): The amount added to green tokens' logits; the watermark strength.
  • z-score: How many standard deviations the observed green count is above what chance predicts.
  • SynthID Text: Google DeepMind's text watermarking scheme, used in Gemini and released as open source.

← 15.2 Prompt Injection: Attacks Against LLM-Powered Systems · 16.1 Multimodal AI: Perceiving Text, Images, and Audio Together →