Claude

Anthropic Claims it Can’t Track Individuals With New Watermarks in Texts

Claude by Anthropic. © Unsplash
Claude by Anthropic. © Unsplash

Set Trending Topics as a preferred source on Google.

Anthropic has published a detailed blog post spelling out how the new watermark in the text output of its AI model Claude actually works. The company is responding to a sharp reaction from its own user base: after the plan became known via an updated support page, numerous paying customers announced on X and Reddit that they were cancelling their subscriptions (TechCrunch). Two groups pushed back the loudest: people who use Claude only to proofread their own writing, and developers worried that the generated code would degrade.

The watermark sits in the choice between synonyms

The core of the explanation: a language model writes word by word, and at many points it has a choice between several synonyms that would continue the sentence equally well. In a sentence starting “The weather today was cold and …”, it hardly matters whether “overcast” or “grey” comes next – the meaning stays the same, and the choice is normally settled by a random number. Those interchangeable synonyms are exactly where the watermark lives: instead of an arbitrary random number generator, a secret key and the preceding words determine which of the equivalent options is picked. Across a longer text, those decisions add up to a pattern. Anyone holding the key can check whether a sequence of words fits that pattern and derive a probability that Claude was involved.

According to Anthropic, the word choice stays inconspicuous. The model is not permanently nudged towards “overcast” or “grey”; depending on the preceding text, it still decides differently each time. Nor does it reach for exotic synonyms it would otherwise never use – only the options that were plausible candidates anyway are in play.

Nothing is added to the text itself: no hidden characters, no additional tokens, and therefore no higher costs and no noticeable loss of speed. Conclusions about individual people, organisations or chats are not possible either, the company says. The method is a variant of SynthID-Text, developed by Google DeepMind and published in a Nature paper. In a live test on a share of Gemini traffic, no statistically significant quality difference was found between marked and unmarked responses.

Limits: proofreading, facts, code

Anthropic concedes clear gaps of its own accord. The watermark only works where there is genuine freedom of choice. In factual passages where only one formulation is correct, nothing sticks. The same applies to code, which has to be exact – there the pattern can effectively only sit in comments and has no effect on the code itself. And anyone who hands Claude their own text merely for grammar and punctuation fixes leaves behind so few AI-chosen words that detection often fails. The reverse also holds: the more Claude writes, the stronger the signal. Translations carry the watermark in full, because every word comes from the model. A complete rewrite removes the mark again, light editing usually does not.

For files such as .png, .jpg or .svg, Anthropic is not using a watermark but signed metadata following the open C2PA standard, which camera manufacturers and image editing software use as well. A detection API that lets users check texts for the watermark has been announced but is not yet available.

Regulatory trigger, commercial upside

Anthropic points to the EU AI Act as the reason: like around 190 other signatories, the company has signed the EU Code of Practice on transparency of AI-generated content, which requires machine-readable marking of such content. Because there is currently no reliable way to limit the feature by region, it is being rolled out worldwide. Other major model providers have signed the same code and are working on watermarks of their own.

Beyond compliance, though, there are tangible self-interests at play. A reliable marker lets Anthropic recognise its own output on the open web and filter it out of the training data for future models. That addresses a problem discussed under the labels AI cannibalism and model collapse.

Here is what those mean. Large language models are trained to a significant extent on text scraped from the open web. But that web is filling up with content that was itself generated by AI. Anyone collecting training data today is increasingly feeding a model the output of other models – the AI is, in effect, eating its own produce. That phenomenon is known as AI cannibalism.

AI Cannibalism Can Cause Model Collapse

Model collapse is the consequence. A Nature study by a team around Ilia Shumailov showed that models degenerate when they are trained over several generations on such synthetic data without controls. The tails of the distribution disappear first – rare phrasings, exceptions, unusual facts, because they are already underrepresented in AI output. The effect compounds with each generation until the model produces an ever narrower, more error-prone picture of reality. The authors describe the process as irreversible and see it as inherent to all generative models, not just language models. The more AI text circulates online, the more valuable the ability becomes to mark it and sort it out again when assembling training data.

On top of that sits a possible line of business: the announced detection API can be charged for. Anyone wanting to check whether a text came from Claude – publishers, universities, platforms, HR departments – would be paying to verify an output whose creation has already been paid for. That is precisely what draws criticism in the debate: a free detector would double as a circumvention tool, because it would let anyone keep tweaking a text until the watermark disappears. A paid one, in turn, makes independent verification harder. How Anthropic will structure access, and what it will cost, remains open.

Rank My Startup: Erobere die Liga der Top Founder!
Advertisement
Advertisement

Specials from our Partners

Top Posts from our Network

Deep Dives

© Wiener Börse

IPO Spotlight

powered by Wiener Börse

Europe's Top Unicorn Investments 2023

The full list of companies that reached a valuation of € 1B+ this year
© Behnam Norouzi on Unsplash

Crypto Investment Tracker 2022

The biggest deals in the industry, ranked by Trending Topics
ThisisEngineering RAEng on Unsplash

Technology explained

Powered by PwC
© addendum

Inside the Blockchain

Die revolutionäre Technologie von Experten erklärt

Trending Topics Tech Talk

Der Podcast mit smarten Köpfen für smarte Köpfe
© Shannon Rowies on Unsplash

We ❤️ Founders

Die spannendsten Persönlichkeiten der Startup-Szene
Tokio bei Nacht und Regen. © Unsplash

🤖Big in Japan🤖

Startups - Robots - Entrepreneurs - Tech - Trends

Continue Reading