Guide
AI content provenance, explained without the jargon
Provenance is the question of where a piece of content came from. Three different technologies attempt an answer, and they fail in different ways.
Published 4 September 2026, 5 min read
Three mechanisms, three different bets
A watermark puts the signal inside the content. For text that means biasing word choice according to a secret pattern, so the words themselves carry the signature. The bet is that the signal survives ordinary handling, and it does survive copying and re-formatting, because there is no separate metadata to lose.
Content credentials, standardised by the C2PA, do the opposite: they attach cryptographically signed metadata to a file recording what produced it and what edited it afterwards. The bet is that the record travels with the file. The signature makes tampering detectable, which is a genuine strength.
Fingerprinting stores a compact representation of known content and compares new material against it. It is how platforms match uploads against a catalogue. It only recognises what is already in the database, so it says nothing about content it has never seen.
Where each one breaks
Watermarks degrade when content is edited, because the signal lives in the choices that editing changes. Paraphrasing, translation or heavy rewriting weakens it, and short passages may never carry enough signal to measure. Detection is probabilistic, so results come as likelihoods.
Content credentials survive editing but not stripping. Metadata is separate from the content, so anything that re-encodes a file — a screenshot, a copy-paste, a platform that discards metadata on upload — can remove the record entirely. Absence of credentials therefore means very little: it may mean the content was not signed, or simply that it passed through something careless.
Fingerprinting is blind outside its database, and text presents a further difficulty: matching text against known material is closer to plagiarism detection than provenance, and legitimate quotation looks much like copying.
The asymmetry that matters most
Across all three, positive and negative findings carry very different weight, and this is the single most misunderstood point.
A positive finding is meaningful evidence. A detected watermark, or a valid signed credential, is a real signal about how content was produced, though even then vendors are careful: Anthropic describes a detected Claude mark as a signal that the model may have processed the content, not proof of authorship.
A negative finding is nearly worthless. No watermark may mean the generator did not watermark, that the model predates the feature, that the passage is too short, or that ordinary editing degraded the signal. No credentials may just mean a screenshot. Treating absence of evidence as evidence of human authorship is the mistake underlying most false accusations.
What to do with this in practice
If you publish, keep your own records. Drafts, notes, version history and commit logs are more convincing than any detector output, and they are within your control. Institutions increasingly accept process evidence precisely because detector scores are unreliable.
If you evaluate other people's work, do not treat a tool's output as a verdict. Detection scores are probabilistic and known to misfire, particularly against non-native English writers. Use them to start a conversation, never to conclude one.
And separate the two questions people tend to merge. 'How was this made' is a provenance question that technology partially answers. 'Is this any good, and is it accurate' is an editorial question that no provenance system addresses at all, and it is usually the one that actually matters.
Frequently asked
What are C2PA content credentials?
A standard for attaching cryptographically signed metadata to a file, recording what produced it and what edited it. The signature makes tampering detectable, but the metadata can be removed entirely by re-encoding, screenshotting or uploading to a platform that strips it.
Does no watermark mean a human wrote it?
No. It may mean the generator does not watermark, that the model predates the feature, that the text is too short to carry a measurable signal, or that editing degraded it. Absence of a mark is not evidence of human authorship.
How can I show my own work is genuinely mine?
Keep process evidence: drafts, notes, outlines, version history, commit logs. Records you control are more convincing than a detector score, and institutions increasingly accept them precisely because detectors are unreliable.
Keep reading
Related guides.
Free, no account