Claude's Invisible Watermark Signals AI Involvement but Fails to Address the Real Problem of Reader Deception
Source: Oren Etzioni. "Etzioni on AI: Claude is marking its text — Caveat Promptor! – GeekWire." August 13, 2026. www.geekwire.com
The Gist
Anthropic now secretly marks text written by its new Claude AI models so it can be identified later, mainly because a new EU law requires it. But the author argues this watermark is a weak fix: it can't tell the difference between someone using AI to write lazy fake content and someone just polishing their own writing, and people who really want to hide AI use can simply switch to AI tools that don't watermark. So instead of trusting the watermark, we should judge AI-assisted writing by how good and honest the final result is.
Conclusion
Anthropic's new invisible watermarking of Claude-generated text is a technically real but practically limited and easily circumvented measure that cannot solve the actual problem of AI-generated content deceiving readers, so people should judge AI-assisted work by its outcome rather than by whether it carries a watermark.
Premises
- The watermark works by statistically biasing word choices in a way that survives copy-paste and can be detected with a computable error rate, but only in sufficiently long text (single sentences or minor edits like punctuation fixes leave no room for a mark).
- A watermark only indicates that a model modified text, not that it wrote the substantive content — for example, summarizing or lightly editing a user's own memo will trigger a mark even though every idea is the user's.
- The absence of a watermark proves nothing, since older models, non-signatory models (e.g., Grok/xAI), and open-weight models produce unmarked text regardless of AI involvement.
- Watermarks are difficult to remove — paraphrasing degrades but rarely erases the statistical signal, especially once there is enough text.
- Anthropic has not published key technical details (false positive rate, minimum word count, or a public detector), and the marks persist indefinitely, creating long-term consequences with unclear reliability.
- The real harm worth addressing — 'Claudefishing,' or the mismatch between a reader's expectation of human effort and the reality of AI-generated content with no human thought behind it — is not something a watermark can detect, since it cannot distinguish careless slop from careful AI-assisted work.
- Anyone determined to hide AI use can still do so via non-watermarking models, so the regulation (driven by the EU AI Act's Article 50) has limited practical impact on the actual bad actors it aims to deter.
Assumptions
- That the primary social harm from AI-generated text is reader deception about the amount of human effort/thought involved, rather than concerns about misinformation, labor displacement, or other issues.
- That watermarking is intended (by regulators and Anthropic) primarily as a disclosure/detection mechanism rather than for other purposes like model attribution or copyright enforcement.
- That readers and evaluators are capable of and willing to 'judge the outcome, not the tool,' implying quality can be reliably assessed independent of production method.
- That the EU AI Act's transparency code, despite being written for EU compliance, will have global effects because major AI companies apply watermarking uniformly across users.
- That technical limitations described (short text, paraphrase resistance) are stable and representative of how the watermarking will function in practice, despite Anthropic not having released full technical specifications.