Understanding AI Text Watermarking

Summary of Understanding AI Text Watermarking

by Ben Jaffe and Katie Malone

29m•August 24, 2026

Overview of Understanding AI Text Watermarking

This episode of Linear Digressions explains how Anthropic’s new AI text watermarking works, using Google DeepMind’s 2024 SynthID text approach as the technical model. The hosts break down how an invisible watermark can be embedded directly into an LLM’s word-choice process, how it can later be detected with a secret key, and why the method is effective but not foolproof—especially for short text or heavily edited output.

What AI Text Watermarking Is

  • The goal is to make it possible to tell whether text was generated by an AI model.
  • The watermark is meant to be invisible in the final text:
    • Not a visible label.
    • Not hidden Unicode characters.
    • Not a secret acrostic or obvious tag.
  • Instead, the watermark is built into the token selection process itself.

How the Watermark Is Generated

Standard LLM generation

  • An LLM predicts the next token based on the preceding text.
  • It samples from a probability distribution over possible next tokens.
  • The chosen token is appended, and the process repeats.

Watermarked generation

  • Anthropic uses:
    • A private watermarking key
    • The recent context / preceding tokens
  • These feed a random seed generator.
  • That seed creates a series of random functions that influence token selection.
  • The model still samples valid next tokens, but the sampling is slightly biased in a secret, structured way.

Tournament-style sampling

  • The hosts use a “March Madness” / knockout tournament analogy:
    • Candidate tokens are compared in rounds.
    • Random binary vectors determine which candidates get an advantage.
    • Tokens that align with the watermark function are more likely to advance.
  • This process repeats across multiple levels until one token is selected.

How Detection Works

  • Detection requires:
    • The text to test
    • The same watermarking key
  • A scoring function estimates whether each token looks like it came from the watermarking process.
  • Individual tokens provide only a weak signal.
  • The signals are summed across the full text.
  • If the total score crosses a threshold, the text is classified as watermarked.

Key Properties and Limitations

Fast and scalable

  • Both generation and detection are designed to be computationally efficient.
  • That’s important if watermarking is going to be usable at scale.

Works best on longer text

  • Short snippets do not provide enough signal for reliable detection.
  • Detection becomes much stronger as the text gets longer.
  • The episode cites results showing strong performance around ~400 tokens, with high true-positive rates at low false-positive rates.

Best for high-entropy text

  • Watermarking is strongest when the model has many plausible next-token choices.
  • It is weaker in low-entropy situations where only one completion makes sense.
    • Example: “The capital of France is ___” is too predictable to carry much watermark signal.

Not a style detector

  • It does not simply flag “LLM-sounding” language.
  • It won’t necessarily eliminate common AI phrasing or stylistic habits.
  • The overall output distribution remains natural-looking, whether or not watermarking is applied.

Can be weakened by editing

  • Heavy human editing or synonym replacement can disrupt the watermark.
  • The hosts frame this less as a bug and more as a consequence of changing the original AI-generated text.

Practical Implications

  • Anthropic plans to release a detection API in the future so users could test whether text was watermarked.
  • This could matter for:
    • AI-generated content moderation
    • Disclosure and provenance
    • Combating spam or “AI slop”
  • The hosts speculate that easy detection may change how people use LLMs to present AI-written text as human-authored.

Notable Takeaways

  • The watermark is embedded in token choice, not hidden formatting.
  • Detection relies on a secret key plus statistical analysis over many tokens.
  • It is robust enough to be useful, but not perfect.
  • Longer, higher-entropy outputs are much easier to identify than short or heavily edited ones.

Extra Notes from the Episode

  • The hosts note that Google had already implemented a similar approach in Gemini before Anthropic’s announcement.
  • The episode ends with a promo for the show’s Substack and YouTube presence, where they provide written summaries and visual explainers.