Anthropic began marking Claude’s output on August 2, 2026. The commentary since has been almost entirely about the fact of it. AI watermark detection is the part nobody is working through, and it is the only part that changes anyone’s job.
Here is the problem in one line: the obligation to mark content arrived before any public ability to check a mark, and the technique’s failure modes run precisely opposite to how editorial policies are written.
AI watermark detection does not exist yet, and that is not a detail
Start with the thing that makes every other question theoretical.
In its explanation of how the technique works, Anthropic says it "will soon be offering a watermark detection API" and is "in the process of working out the details of its implementation." No public tool exists today. The key is held by the vendor, which is inherent to the design rather than a temporary gap, so third-party detectors cannot appear independently the way plagiarism checkers did.
Meanwhile the legal clock runs. Article 50 of the EU AI Act became applicable on August 2, 2026, with systems already on the market given until December 2, 2026 to satisfy the marking and detection duty. So there is a marking regime in force, a deadline approaching, and no way for anyone receiving content to act on either.
If you commission writing, that means the current state of AI watermark detection gives you nothing operational. Not less than you hoped. Nothing.
The asymmetry that breaks editorial policy
Now the part that will still be true after a detector ships, and which almost no coverage has connected.
Anthropic’s own summary of robustness is that "light editing probably won’t remove the watermark completely; a complete rewrite where every word is replaced will." Read that as an editor rather than an engineer.
A writer who takes a Claude draft and tidies it, fixes a few facts and adjusts the tone produces text that is still detectable. A writer who takes the same draft and rewrites it thoroughly, thinking hard about every sentence, produces text that is not detectable.
The second writer did substantially more work and added more judgment. Under any sane definition of authorship they have a stronger claim to the output. The mark says the opposite. It flags the diligent light-editor and clears the person who barely engaged, since someone who never used AI at all also produces an unmarked result and is indistinguishable from the heavy rewriter.
Most disclosure policies say something like "declare AI assistance." Detection sorts writers by how much they changed, not by how much they used, and those are different axes.
Absence proves nothing, which cuts both ways
Anthropic’s support documentation is blunt: "lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed."
Count the ways an unmarked passage arrives on your desk. It came from a different model. It came from an older Claude that has not been retrofitted yet. It is a file type outside .svg, .png and .jpg. Its metadata was stripped by a screenshot, a format conversion or an upload pipeline. It was heavily rewritten. It is too short. Or a human wrote it.
That list has one honest human on it and six other explanations. So a negative result is not evidence of human authorship, and any policy that treats "no mark found" as a clearance is misreading the tool. A positive result is better but still hedged: a detected mark "provides a signal that content was processed by Claude, but is not fully conclusive."
Your most checkable content is the least markable
There is a second asymmetry, and it is arguably worse for a publisher than the first.
Anthropic notes that watermarking "is sparser on factual passages where there are fewer choices that can be made without decreasing the accuracy of the text." The mark needs free word choices to encode a signal. Factual writing does not have them, because the wording is pinned by correctness.
Now list the content a publisher most wants to verify: news copy, technical documentation, financial summaries, product specifications, research abstracts. Every one is the constrained, factual register where the mark is thinnest. The genres that carry the signal best are personal essays, opinion and long discursive prose, which are also the genres where AI involvement matters least to a reader.
Add the length problem, since detection "doesn’t work well on small samples," and short factual copy is close to unmarkable by construction. A product description, a headline, a photo caption or a summary paragraph will not carry a reliable signal no matter what generated it.
The one thing you can verify today
Everything above concerns text. The file side is different, and it is the only part that works right now.
A .png, .jpg or .svg from Claude carries signed C2PA provenance metadata, and that manifest is cryptographically verifiable with public tooling today. No vendor API required. It makes a narrower claim than a watermark, describing how a file was produced rather than proving what generated some text, but it is a real claim you can check.
Its weakness is separability. Screenshot the image, re-encode it, or run it through a pipeline that strips metadata and the credential is gone while the pixels are unchanged. It proves presence, never absence.
Where AI watermark detection will still help
Being fair to the technology, there are real uses once the API exists, and they are narrower than the discourse implies. This is the honest list.
Long-form prose that arrives unedited. A 2,000-word essay pasted straight from a model is the best case for the technique.
Confirming an admission. If a contributor says they used Claude for a draft, a positive result corroborates it. Corroboration is a legitimate function even when discovery is not.
Spot-checking at volume. Across hundreds of submissions, a positive rate that shifts tells you something about a contributor pool even when no single result is conclusive.
What it will not support is policing. A thorough rewrite defeats it, and thorough rewriting is exactly what a motivated person does.
What to put in your policy instead
Five positions that hold whether or not a detector ships in December.
Do not write "we detect AI" into a policy. You cannot, today, and the version you will get is probabilistic and one-directional. A policy that promises detection is a policy you will breach.
Ask about process, not output. "Which tools did you use and for what" is answerable, auditable against a contributor’s own account, and does not degrade when someone edits carefully. It is also the question you actually care about.
Treat provenance metadata as worth preserving. If your pipeline strips metadata on upload, you are destroying the one cryptographically verifiable signal in the chain. That is a fixable engineering decision and most teams have never looked at it.
Do not treat an unmarked file as cleared. Given six non-human explanations for a missing mark, absence should change nothing in your process.
Watch December 2 rather than August 2. The transitional window for the marking and detection duty runs to December 2, 2026, and the detection API is the thing to wait for. Until it exists, nothing in this area is operational.
The wider pattern is one we keep running into: a capability ships, the verification tooling trails it by months, and the gap gets filled with confident claims nobody can check. That was the story with vendor benchmark scores, and with the gated release of an offensive security model whose evaluation numbers only its approved partners can reproduce. If you want the underlying mechanics rather than the editorial consequences, our explainer on how AI content watermarking works covers the two mechanisms and where each one breaks.
Frequently Asked Questions
Can I check whether text was written by Claude?
Not today. Anthropic says it will soon offer a watermark detection API and is still working out the implementation. The verification key is held by the vendor by design, so independent detectors cannot appear the way plagiarism checkers did. Any service claiming to detect Claude watermarking right now is not using Anthropic’s key and is guessing.
If no watermark is found, was it written by a human?
No, and this is the most consequential misunderstanding. Anthropic states that lack of a detected mark does not mean the content was not AI-generated. A different model, an older Claude, an unsupported file type, stripped metadata, heavy rewriting or a passage that is simply too short all produce an unmarked result. Absence of a mark is not evidence of human authorship.
Why would a lightly edited draft be detectable when a heavily edited one is not?
Because the mark lives in the model’s word choices, and rewriting replaces them. Anthropic’s own summary is that light editing probably will not remove the watermark while a complete rewrite will. The practical effect runs against intuition: the contributor who did more work and added more judgment produces the cleaner result, while the one who barely engaged gets flagged.
Does this work on news copy and documentation?
Less well than on other writing. Watermarking is sparser on factual passages because correctness constrains word choice and leaves little room to encode a signal. Combined with the fact that detection performs poorly on short samples, the tightly written factual content publishers most want to verify is the content where the mark is weakest.
What is actually verifiable right now?
Signed C2PA provenance metadata on supported image files, where it has not been stripped, is cryptographically checkable and makes a specific claim about origin. Everything on the text side is waiting on the detection API. Until then the honest answer is that you can verify file provenance in some cases and nothing about text.
Should my disclosure policy require contributors to declare AI use?
Yes, and it should ask about process rather than promise detection. Which tools, at which stage, for what purpose. That is answerable, checkable against a contributor’s own account, and does not degrade when someone edits carefully. A policy claiming you will detect undisclosed use is one you cannot currently enforce.
Does the EU AI Act require me to detect AI content?
No. Article 50 places the marking obligation on providers of generative AI systems, not on publishers receiving content. Deployers have separate disclosure duties, notably around deepfakes. The relevant date for the marking and detection duty on systems already on the market is December 2, 2026, and penalties reach 15 million euros or 3 percent of worldwide annual turnover.
What is the single change worth making now?
Check whether your publishing pipeline strips metadata on upload. Signed provenance is the one cryptographically verifiable signal in this whole area, many content management systems and image processors discard it by default during resizing or re-encoding, and almost nobody has looked. Preserving it costs little and is the only part of this you control.