Anthropic details how Claude’s SynthID-Text watermarks will work

Anthropic published a blog post on Friday outlining how it will add invisible watermarks to text produced by its chatbot Claude to comply with the EU AI Act’s Transparency Code.

How the watermarking works

The company said Claude will make subtle, low-stakes choices—such as selecting between synonyms—that create a pattern in responses that is “undetectable to the reader, but is detectable to anyone who has a key that encodes it.” Anthropic added that “Watermarking does not impact the quality of Claude’s output,” and that a watermarked reply should appear indistinguishable from an unwatermarked one to human readers.

Anthropic plans to implement the SynthID-Text approach outlined by the Google DeepMind team in 2024 and told users it will provide a watermark detection API. The company contrasted watermark checks with other AI-detection methods from firms like Pangram, saying those approaches look for stylistic “tells” while watermark verification relies on a different, encoded signal.

Limits: edits, proofreading and code

The company acknowledged that editing can affect detectability. Anthropic said light editing “probably won’t remove the watermark completely,” whereas “a complete rewrite where every word is replaced will.” It also noted that whether a piece of text edited or proofread by Claude carries a watermark depends on the length of the text and the extent of Claude’s edits.

Code is expected to carry less of a watermark because producing functional code restricts the model’s choices. Anthropic said the watermark could appear in places with arbitrary choices—such as comments within code—but it will have “a negligible effect on the actual code produced.”

Reception and rollout

Users have debated the change on platforms such as Reddit, where some described it as a conspiracy and others defended it; Business Insider reported that dozens of users on X said they canceled Claude subscriptions over the decision. Anthropic also said other major model developers have signed the same Code of Practice and will implement their own watermarks. The company will release the detection API as part of its rollout.


Original source: TechCrunch AI

Leave a Comment