Anthropic desires to make AI-generated textual content simpler to determine, and on paper, I’ve little or no purpose to complain. The corporate is experimenting with an invisible watermark that may be baked instantly into textual content generated by Claude.
It appears like a wise thought. AI-generated textual content is in all places, and figuring out the place one thing got here from might actually assist. Furthermore, Anthropic isn’t merely hiding a marker someplace inside a doc. Its method modifications how Claude selects phrases to create a statistical sample that may later be detected.
However there’s one element that bothers me. Anthropic is testing simply how persistent that watermark will be, even after the textual content has been modified.
That’s the place I can already odor hassle.
Claude touched my writing. Did it really write it?
Take into consideration translation for a second. Let’s say somebody writes a whole essay themselves in Spanish and asks Claude to translate it into English. The concepts are theirs. The analysis is theirs. The arguments are theirs. Claude’s solely job is translation.
But the ensuing textual content might nonetheless carry Claude’s watermark.
The identical query applies to proofreading. What if somebody writes one thing themselves and asks Claude to repair the grammar? What about shortening a paragraph, altering its tone, cleansing up dictated textual content, or just making an ungainly sentence simpler to learn?
These aren’t fringe makes use of for AI anymore. Folks more and more flip to assistants like ChatGPT, Gemini, and Claude for on a regular basis duties which have little to do with producing unique work. A watermark can inform you that Claude was concerned with a bit of textual content. It can not inform you whether or not Claude really wrote it. Anthropic makes the identical level, saying the watermark reveals Claude’s involvement, not who created the unique work.
Now think about explaining that distinction to a professor after their detection software program has simply flagged your essay.

We already know the way messy AI detection can get
I wouldn’t fear almost as a lot if our observe report with AI detection have been significantly good. It isn’t.
MIT Sloan’s steerage is kind of easy about present AI detectors. It says they’ve excessive error charges and might lead instructors to falsely accuse college students of misconduct.
We’ve already seen what that appears like in apply. College students have discovered themselves defending work they are saying they wrote themselves after automated programs recognized it as AI-generated. In a single case documented by The Guardian, a scholar’s essay was flagged as completely AI-generated regardless of the scholar saying they’d solely used accredited spelling and grammar help. The attraction was ultimately accepted.
To be clear, Claude’s watermark is essentially completely different. Standard AI detectors take a look at writing and primarily estimate whether or not an AI may need produced it. Anthropic is intentionally planting a detectable sign in Claude’s output. In principle, that ought to make its system significantly extra dependable. However reliability isn’t the one downside right here. Interpretation is.

We’re utilizing AI to show we didn’t use AI
Issues have already reached a barely ridiculous level.
College students fearful about AI detection are turning to so-called AI humanizers, which rewrite textual content particularly to make it much less more likely to set off detectors. Some college students are even utilizing these instruments on work they wrote themselves as a result of they’re fearful about false positives. Detector corporations, naturally, are growing methods to determine humanizers.
Learn that once more.
A human can write one thing, fear that an AI will suppose an AI wrote it, feed it by way of one other AI to make it look extra human, after which have yet one more system decide whether or not the AI made it look human.
It’s a technological ouroboros.
Making Claude’s watermark resilient sufficient to outlive modifying and translation is technically spectacular. Earlier analysis has proven that translation can defeat some text-watermarking methods, so fixing that weak spot would symbolize significant progress.
I simply don’t suppose making the sign tougher to take away solves the extra vital downside.

A watermark wants context
There are good causes to watermark AI-generated content material. It might assist determine mass-produced misinformation, undisclosed artificial textual content, or AI-written materials that later leads to coaching datasets.
The issue is that AI assistants now do excess of generate content material from scratch. Folks use them to translate textual content, proofread paperwork, summarize analysis, assist with code, enhance accessibility, or just clear up an e-mail earlier than sending it. In that context, detecting AI involvement doesn’t routinely inform you who really created the work.
All of these interactions contain AI to wildly completely different levels. If Claude writes an essay from scratch, figuring out that’s helpful. If Claude interprets an essay somebody spent three weeks researching and writing themselves, figuring out Claude was concerned tells you significantly much less.
The watermark could also be completely able to answering “Did Claude contact this?” My concern is what occurs when individuals begin treating the reply as proof of “Did Claude write this?”
Anthropic can construct the neatest watermark on the planet. Except the individuals utilizing it perceive that distinction, I think we’re going to have some issues.


