Up until recently, it’s been pretty easy for people to get away with using AI to write. Not because we can’t tell that it’s been AI-generated, but because it’s been virtually impossible to get a definitive answer. Indeed, something may look a lot like AI and it may fail multiple AI detection tests, but using these techniques, we still can’t definitively assert that something was 100% written by AI and not a human.
And for that reason, people have been using it to help write essays, draft emails, create blog posts and even produce entire books, all while remaining largely invisible. Not free from scepticism and doubt (that’s a whole different issue), but they were able to get away with it for the most part. Indeed, unless somebody explicitly admitted they had used AI, it was often impossible to know for sure.
But Anthropic and plenty of its supporters think that this may be about to change. This week, Anthropic announced that Claude has begun embedding invisible, machine-readable watermarks into AI-generated text, as well as provenance metadata into supported files and images. They’ve done this as a result of the company’s intention to align itself with the new transparency requirements under the EU AI Act that came into force at the beginning of this month. And according to Anthropic, the changes will be applied globally, not just within Europe.
And this update is far from merely technical. In reality, it could mark one of the biggest shifts yet in how AI-generated content is created, shared and identified online. But at the same time, it could also have absolutely no real effect at all if it turns out to be less definitive than it claims to be.
So, What Is Claude Actually Doing?
Unlike the visible watermarks you might see on stock photos, Claude’s watermark won’t be something users can see with the naked eye. Instead, Anthropic says it will be an “imperceptible” watermark embedded directly into the text itself. The content will read exactly the same, but hidden signals within the output will allow detection tools to determine whether Claude was involved in generating or processing it. Anthropic is also adding signed provenance metadata to supported file formats and images.
But, did you catch the important bit? Or did you skim over it just like most other people when they first read the news? Don’t worry, we’ll reiterate.
The watermark will indicate whether Claude was involved in generating or processing the content.
That phrasing and seemingly small little detail is actually the most important one in the whole announcement, because it means that the watermark doesn’t necessarily confirm Claude wrote something from scratch. Not at all.
According to Anthropic’s own documentation, a detected watermark simply indicates that Claude processed, edited, translated or otherwise worked on the content. In other words, finding a watermark won’t automatically prove that an entire article, report or essay was generated by AI. It simply flags the fact that it’s had some involvement with the text, and that’s a very different situation.
More from Artificial Intelligence
- When Law Meets AI: Meet Kyra Law
- What Is Fabless Manufacturing?
- Quite Contrary: The Real Reason Enterprise AI Pilots Fail, According To Chase W. Hughes
- Is Spotify’s AI Persona Badge Protecting Real Artists Or Discriminating Against A New Creative Medium?
- Intel’s $20 Billion Stock Offering: Is The Chip Race Becoming More About Money Than Tech?
- Does The AI Industry Now Need Financial Engineering To Continue To Grow?
- Meta Just Released A New Open-Weight AI Model – Is This A Direct Assault On OpenAI And Anthropic?
- Could Sounding Human Become The Best Marketing Strategy on LinkedIn?
How Will Claude’s Watermark Work?
Anthropic hasn’t released the full technical details yet, but according to TechRadar, experts seem to believe that the watermark is likely created through subtle patterns in word selection during the generation process.
The basic idea is that the model makes tiny choices between equally valid words or phrases, creating a statistical fingerprint that detection tools can later identify. So, to human readers, the text looks completely normal, but if it’s run through the appropriate detection tool, the pattern will be identified.
Anthropic says the watermark should survive common actions such as copying, pasting and light editing, but if the text is heavily rewritten, translated, or extensively paraphrased, it may be weakened or removed entirely.
So from here, it’s easy to see where things start to get complicated. First, there’s the fact that if it’s edited slightly, the watermark loses its magic. And second, the “watermark” is less of a watermark and more of a supposedly specific arrangement of words attached to the statistical likelihood of them being grouped in certain ways. And perhaps I’m just a diehard sceptic, but if this new shiny “fix-all” tool is really just another pattern detection tool, isn’t it going to have all the same issues every other AI detection tool has?
We’ll probably need to wait and see what Anthropic releases regarding the mechanics of the watermarking, but at this stage, I’m not convinced.
The End of “Invisible AI” Or More of the Same?
The announcement has sparked debate because it touches on a question that’s been hanging over generative AI for years. That is, should people know when AI was involved?
For educators, publishers and employers, the answer is often a resounding “yes”. A watermark could make it easier to identify content that was generated or heavily assisted by AI, helping organisations enforce disclosure policies or investigate cases where AI use may have been hidden. Business Insider also noted that the move could have implications for everything from publishing to education, where concerns about undisclosed AI use continue to grow and wreak havoc in education, both in high school and tertiary institutions.
But some users view the new tool a little differently. They see it as more of a technique for collaboration rather than a replacement for their own work. If somebody uses Claude to improve grammar, brainstorm ideas or restructure a document, should that content carry the same marker as something generated entirely by AI? And that’s a new question the industry is now grappling with (as if we didn’t have enough to stew on already).
Will Watermarking Actually Work?
That’s perhaps the biggest “who knows?” of all. Watermarking has long been seen as one possible solution to AI transparency, but researchers have repeatedly pointed out its limitations. One of the biggest challenges is that the more invisible a watermark becomes, the easier it may be to remove through rewriting, paraphrasing or translation. Much like how AI detection becomes less and less effective when AI writing improves.
According to The Verge, even Anthropic acknowledges that watermark detection won’t be perfect. The company has said that the absence of a watermark also shouldn’t be treated as proof that content was created by a human.
Ultimately, detection tools are still being developed and the full technical specifications have yet to be released publicly. So, in other words, this isn’t a magic solution to AI detection.
This May Be a Sign Of What’s Coming…Or It Might Not
The most interesting thing about Anthropic’s announcement isn’t the watermark itself, it’s more about what it represents and why it’s even an issue in the first place.
For years, AI companies competed to make their outputs as human-like as possible, and now they’ve become pretty good at that, regulators are increasingly asking them to stop. Kind of like a, “please mimic human writing” followed by, “wait, but don’t do it so well!”
Whether Claude’s watermark proves to be effective remains to be seen, but the introduction of the tool seems to indicate more of a shift towards a broader desire for transparency and traceability in AI-generated content. And if other major AI companies follow suit, the era of AI quietly working behind the scenes may start to come to an end.
If it’s effective, that is. Because if it’s not, it’ll just be more of the same mistrust, accusations and poor writing.
