top of page

AI Act and watermarks: what an AI watermark actually proves about a text

  • 2 days ago
  • 10 min read

AI Act, watermarks, and why the useful question isn't technical.


The AI Act and August. Admit it, that's the pairing your summer was missing. I know, you're all busy implementing it instead of hitting the beach, taking every single obligation to the letter. Among the many angles here, I want to dig into one detail worth a thought even from under a beach umbrella: the watermark. You read that right, mark, as in trademark. No Park in sight, water or otherwise.



When it comes to artificial intelligence, a watermark is one possible answer to the "problem" of recognizing text (or other digital artifacts) produced by a machine.


The reason is easy to see: a voluntary declaration along the lines of "I swear I wrote this myself" offers few guarantees, so you need something verifiable.


What is a watermark?


It's an invisible signature hidden among the words of a text, readable only by whoever holds the key.


A watermark is also a mark hidden inside an object, digital or physical, that says where it came from. In fine paper, that's literally what happens: you look at the sheet and see a sheet, you hold it up to the light and the paper mill's logo appears. The sheet stays perfectly normal and readable, and all the while it carries the name of who made it.


You can do the same thing with text, in a crude way:


"Behind Your Message Awaits X."


The initials of the five words spell BY MAX. The sentence reads normally, it makes sense, nobody notices a thing. It's a mark of authorship hidden in plain sight.


But it's extremely weak. Touch a single word and it's gone.


Or😈you😈can😈use😈tricks😈of😈any😈kind😈that😈don't😈alter😈the😈content😈in😈any😈way😈😀


Watermarks and LLMs


In August, Anthropic's announcement that it's about to put a watermark on all content generated by its models made headlines. And, of course, it set off a wave of very different reactions.


Let's try to understand what they're actually proposing.


Every time it has to choose a word, a language model has several plausible alternatives in front of it, each with a different probability.


Anthropic's watermark tilts that choice ever so slightly, following a secret key: certain words get favored by the tiniest margin, and the meaning stays exactly the same. A single word proves nothing, but over a long text those micro-choices add up to hundreds of them, accumulating into a statistical signature that can be detected with the right key.


(Claude, who helps me write, points out that the first scheme of this kind came from Kirchenbauer and colleagues, published in 2023. Google DeepMind later published a more recent version, SynthID Text, in Nature in 2024. That's the very scheme Anthropic chose to mark the text Claude generates.)


The mechanism, step by step, is in the animation below.



That signature proves a model passed through there. All the human work that comes before and after, from hours spent editing to fact-checking, stays out of frame. And according to Anthropic, this system works because dozens or hundreds of watermarks get embedded, so finding just one is enough to say that text had something to do with a language model. But it says nothing about how much, or where.


So much so that, of course, systems already exist promising to make it useless.


One of them is called Declaude, and it proposes a full rewrite of the text, designed to collapse detection down to near chance level. According to Declaude, a light paraphrase dilutes the signature without erasing it, leaving it detectable for hundreds of tokens. Who knows. Either way, the underlying logic is instructive, because the difference between the two techniques shows very clearly what a watermark actually measures: how intact the text stayed after an "AI pass." (A source with an obvious stake)


What the AI Act actually asks (and of whom)


Anthropic seems to have done this with the AI Act's entry into force as the trigger, since Article 50 requires that people be made aware they're dealing with AI-generated artifacts.


Deployers of an AI system [...] shall disclose that the content has been artificially generated or manipulated.

It draws a distinction between two areas.


Whoever builds the model has to mark the output in a machine-readable format, for audio, images, video and text. The exception already says a lot about the underlying principle: the obligation drops when the system doesn't substantially alter the input data or its meaning. Having your typos fixed, for instance, doesn't trigger the marking. The regulation also asks that the solution be as robust as technically possible, a sign the legislator already knew what to expect from rewrites.


Whoever publishes has a different obligation, and a much narrower one than people tend to claim. It covers "text published for the purpose of informing the public on matters of public interest." In other words, the obligation doesn't automatically extend to every email, post, internal report, sales proposal, or meeting minutes.


Then, in that very same paragraph, comes the line almost nobody quotes:


This obligation shall not apply where [...] a natural or legal person holds editorial responsibility for the publication of the content.

in other words, the obligation doesn't apply when AI-generated content has gone through a process of human review or editorial control, and when a natural or legal person holds editorial responsibility for the publication.


Read that again, slowly. The regulation that's so often cited as a witch hunt for artificial text puts a very human principle in black and white: when someone puts their name behind it, editorial responsibility takes the place of the witch hunt.


Traces everywhere, like dead skin cells


The risk sits in how we'll end up interpreting all this. With watermarks and machine-readable systems, we're entering a rather curious phase, one where we'll go hunting for traces of AI inside content with the zeal of forensics at a crime scene.


This word looks suspicious. This token sequence shows anomalies. At least 37.335% of this text was written by AI (Singular. Capital-S She.) Claude may have passed through here around 2:37 PM. Tape off the paragraph.


Chances are, we'll find traces everywhere. We work with AI to write content, sure, but also to shorten it, translate it, come up with three headline options, fix a passage, summarize forty pages, turn an article into a LinkedIn post, or just check that what we wrote makes sense. Every intervention leaves a different story behind the words.


The surprise comes straight from whoever is building the watermark. On the page explaining how Claude's mark works, Anthropic makes clear that Claude might not be the original author, because people also work with the model to edit, translate, summarize, or convert files. As a result, the text can carry the mark even when the ideas, the text, or the source data come from somewhere else entirely. (Source)


Imagine handing it one of your own texts and asking for a translation, a summary, or a format conversion. The ideas stay yours (hopefully forever), while almost every word gets chosen by the model, and that's exactly where the mark applies, even though the thinking inside is entirely yours. A simple typo fix, on the other hand, leaves the mark little to attach to, as that same Anthropic page explains. The difference comes down to how many words, in the end, the machine chose.


And imagine a text, like this one, that gets reviewed by both Claude and GPT once it's written. Where does that leave us?


Anthropic also adds that the absence of the mark doesn't prove a text never went through Claude, and it lists several reasons why the signature can end up invisible even when the model really was involved. Presence is ambiguous, and so is absence. Lovely, isn't it? (Source)


So we'll end up finding traces of AI in content more or less the way we find dead skin cells on any surface. The trace tells us something happened, but leaves what actually happened wrapped in mystery: whether the idea and the reasoning were human, whether the author spent twenty hours on it, whether they checked every source or rejected the conclusions the machine proposed.


Even if you never worked with AI on that text, can you be sure about whoever handed you the sources? Or about whoever translates it, or summarizes it for a boss who has no time to read it? Before and after you, that content passes through hands you have no control over.


Sifting through every sentence to find out whether so-and-so worked with AI means starting from the wrong question, even before you get to how hard it is to answer.


And that's where a new profession risks being born, one I hope never takes off: the "how much AI is in here" hunter.


The impossible percentage


The scenario I fear most looks reassuring, tidy, and terrifyingly precise:


This text is 37.4% AI · 62.6% human


A measure of quality about as accurate as counting bouquets of flowers to estimate how much love is in a relationship.


Two pieces of content can both come out labeled "AI-generated" and, behind that same label, be completely different things. In the first case the process is: "write me two thousand words on the AI Act," copy, paste, publish. In the second there can be weeks of research, an original idea, a structure the author built, hand-picked sources, a first draft with AI, twenty hours of editing, rewritten sentences, conclusions reworked together with the AI, and a final sign-off. And then there can be a third piece, made 100% by hand.


What difference does that make? That every piece of writing becomes a suspect regardless of the outcome? Does text written by humans alone carry more value? Why?


You'll find stories everywhere of people who ran a text they wrote ten years ago through three of the best-known "AI detectors." The verdicts come back as things like 90%, 75%, and 85% written by AI. They tell us the witch hunt over "AI content" is still very much alive, however much we'd like to think otherwise.


Anyway, those detectors are a different tool from a watermark, because they take a guess and lean into the AI stigma instead of verifying a key. The story making the rounds sums up nicely the wall we risk running into: someone will use the wrong tool to hand down a verdict.


Meanwhile, the damage is already circulating. Whoever receives a document written with AI tends to picture the human as nothing more than a rubber stamp, and to suspect a con hiding behind those pages. The watermark finds that suspicion already sitting there and stamps it official. It risks devaluing good content for no reason other than the whiff of AI on it.


The useful question isn't technical


I like to judge the value of a piece of content with my own taste and judgment. Then, the more important the content is (is it informing me of something that matters, or just entertaining me?), the more relevant this question becomes, the one I think is actually useful:


Who is taking responsibility for what I'm reading? Who checked it, who decided, and who is willing to put their name, their reputation, and their face on it?


Reading opinions on the AI Act, watermarking, and AI use, it seems to me we're mixing three different levels together as if they were one.


Knowing *what I'm looking at is transparency; knowing where it comes from is provenance; knowing who answers for the effects is responsibility. A watermark works on the first two, while the third stays entirely in our hands, going back to the days of cuneiform script.*


A piece of content (valuable or not) with nobody willing to take responsibility for it carries little weight, whether it was produced by GPT-37 or typed out on an Olivetti Lettera 32.


It's my first rule when I work with AI: if I take responsibility for it, the material choice of which tools I used to produce it becomes secondary.


And I imagine not all of you agree. But there's one last point... this is the finger pointing at the moon.


Agents don't write. They act.


These discussions are entirely legitimate, and they'll never land on a definitive answer. But they're distracting us.


The question of "who takes responsibility" is about to become even more urgent. While we're debating watermarks in text, the serious part of the problem has already moved somewhere else.


A modern agent doesn't just produce text. It performs actions: it opens a ticket, updates a CRM record, sends an email to a client and decides on its own what to reply, or whether to reply at all. It chooses which system to query and which credentials to use to get in.


I see it every day in my blog's traffic, where bots and agents show up with a goal that has nothing to do with reading: they probe forms and endpoints looking for a door left open, they comment on articles, they slip in malicious links while pretending to be human. In practice: somewhere out there, a generative-AI-based agent decides how to act and react while it "talks" to websites or other agents, at three in the morning, without asking anyone for permission.


Actions stay entirely outside the reach of a watermark. An email that's already sent, an order that's been forwarded, data saved in the wrong place, or a reply that reached the wrong client, none of these carry a watermark or a statistical signature.


Those are decisions, with immediate and often irreversible consequences. And none of them leave a token behind to analyze.


Are we busy sealing the envelope while revolutions rage outside?


So, back to the old responsibilities


The good news is that the answer has existed for far longer than AI has. The AI Act's own legislators copied it right into the paragraph: human review, editorial control, and one person who answers for the publication.


These are the old responsibilities we already know well. The editor signs off on the newspaper, the surveyor signs the appraisal, the director answers for the balance sheet. Each of them takes responsibility for work built together with other people, regardless of who pressed which individual key.


The same principle holds for agents, with one decisive difference: responsibility stays, for now, entirely on the side of the humans who activate them. Natural persons (and legal ones) are the only parties with a reputation to defend and consequences to face.


An AI agent couldn't care less about any of that. It's, well, like a sheep. And that's exactly why shepherds exist. The shepherd looks after the flock: knows where it is, knows the boundaries it has to stay within, and when a sheep wanders into the neighbor's field, it's the shepherd who goes and knocks on the door.


That's why the job waiting for us is the AI agent shepherd (I wrote about this before, also in Italian), not the hunter of AI-made content (watermarked or not).


Enjoy Artificial Intelligence Responsibly! Massimiliano

 
 
bottom of page