AI Content Detector
Read this first. The AI Act watermark that Claude and Gemini embed in text is statistical — woven into word choice, readable only with the provider's key. No public tool detects it, including this one. What this page checks is real and verifiable: leftover invisible characters in text, and signed C2PA provenance in files.
How it works
This tool answers two different questions about a piece of content, and it is honest about the one it cannot answer.
What it cannot do. Since 2 August 2026, Article 50 of the EU AI Act requires providers of generative AI to mark their output in a machine-readable format. For *text*, Anthropic and Google implement this as a statistical watermark: the model is nudged in its low-stakes word choices (which synonym to use, how to order a clause) so the finished text carries a pattern. That pattern is invisible to a reader and, crucially, detectable only by someone holding the key that encodes it. No public tool can read it. Any site claiming to detect a Claude or Gemini text watermark is guessing, and its guesses will accuse real human writers. This tool does not pretend otherwise.
Text tab: invisible characters. What *is* verifiable is the Unicode residue that survives a copy-paste out of a chat interface, a CMS, or a word processor. The scanner walks your text codepoint by codepoint and flags four classes:
• Invisible: zero-width space (U+200B), zero-width joiner and non-joiner, word joiner, BOM, soft hyphen, plus variation selectors (U+FE00–FE0F) and Unicode tag characters (U+E0000–E007F). The last two matter for security: they render as nothing at all but can encode an arbitrary hidden payload, a technique known as ASCII smuggling that has been used to hide instructions inside text pasted into an AI agent. • Homoglyphs: Cyrillic or Greek letters that render identically to a Latin one, so that 'password' can contain a Cyrillic 'а'. These are flagged only inside a word that already contains Latin letters, so a genuine Russian or Greek paragraph is not flagged as tampered. • Odd spaces: no-break space (U+00A0), narrow no-break space (U+202F, frequent in ChatGPT output around punctuation and inside numbers), thin, figure, punctuation and ideographic spaces. They look like a space, sort like a space, and break your regex like nothing else. • Smart typography: em dash, en dash, curly quotes, ellipsis character. These are the weakest signal on the list, because Word, Pages and every CMS with smart-quote substitution produce them too. They are shown for completeness, not as evidence.
Each finding is listed with its codepoint and count, then shown in position in a highlighted preview where invisible characters get a visible stand-in. The cleaned output strips the invisible ones, normalises odd spaces to a plain space, and folds homoglyphs back to Latin. Smart typography is left alone unless you tick the box, because flattening it is an editorial choice rather than a fix.
File tab: C2PA content credentials. For files, the AI Act marking is concrete: a signed manifest embedded in the file itself, following the C2PA (Coalition for Content Provenance and Authenticity) standard. Drop a PNG, JPEG, WebP, SVG or PDF and the tool locates the manifest in its container (the caBX chunk in PNG, APP11/JUMBF segments in JPEG, the C2PA RIFF chunk in WebP), then reads the declared fields out of it: the claim generator, the software agent, and the IPTC digital source type that says whether the content was produced by a trained algorithm.
Two limits, stated plainly. First, this reads the manifest, it does not cryptographically verify the signature: full validation requires walking a certificate chain against the C2PA trust list, which is a different and much heavier job. A manifest shown here is what the file *declares*, not what has been *proven*. Second, absence proves nothing at all: a screenshot, a re-encode, a resize, or an upload to almost any social platform strips this metadata cleanly. C2PA can confirm provenance; it can never rule it out.
Everything runs locally. Your text and your files are read in the browser and never uploaded.
Frequently Asked Questions
Can this detect the AI Act watermark in text from Claude or ChatGPT?
- No, and neither can anything else that is publicly available. The text watermark providers use is statistical: the model biases its own word choices to encode a pattern that is readable only with the provider's key. Anthropic has said it plans to offer detection, but at the time of writing no public detection tool or API exists. What this page checks instead is the invisible Unicode residue that copied text often carries, which is verifiable and reproducible.
Does finding invisible characters prove a text was written by AI?
- No. It proves the text passed through something that inserts them, which includes chat interfaces but also word processors, CMSs, PDF extraction and translation tools. Treat a zero-width space as a fingerprint of copy-paste, not a confession. The only class here that is genuinely suspicious on its own is a hidden payload, meaning variation selectors or Unicode tag characters, which nothing legitimate puts in prose.
What is ASCII smuggling and why does it matter?
- It is a way of hiding readable instructions inside ordinary-looking text. Unicode tag characters (U+E0000 to U+E007F) and variation selectors render as absolutely nothing, but each one can carry a byte, so someone can encode a full instruction inside a sentence that looks perfectly normal. If you then paste that sentence into an AI assistant, the assistant reads the hidden bytes along with the visible ones. Scanning pasted text before handing it to an agent is the practical defence, and it is why this tool flags those ranges as hidden rather than merely odd.
Why does the em dash count as only a weak signal?
- Because it is the most over-claimed tell on the internet. LLMs do use em dashes heavily, but so does every writer with a decent editor, and Word converts a double hyphen into one automatically. Accusing someone of using AI because their text has em dashes is pattern-matching, not evidence. It is listed here so you can flatten it when you need plain ASCII, not so you can point at it.
What does a C2PA manifest actually tell me?
- It tells you what the file declares about its own history: which tool generated or edited it, when, and whether the IPTC digital source type marks it as produced by a trained algorithm. That is the machine-readable marking Article 50 of the AI Act requires for generated images. This tool reads those declared fields; it does not verify the signature against the C2PA trust list, so treat what you see as a claim rather than a proof.
A file shows no provenance. Is it therefore human-made?
- No, and this is the single most important limitation. C2PA metadata is fragile by design: taking a screenshot, re-saving, resizing, or uploading to most social platforms removes it entirely. So a manifest is meaningful when present and meaningless when absent. Provenance can confirm origin, never rule it out.
Is my text or my file uploaded anywhere?
- No. Both tabs run entirely in your browser: text is scanned in JavaScript, files are read with the File API into memory and never sent anywhere. There is no backend and no logging. Open your browser's DevTools Network tab while you use it and you will see no outbound request carrying your content.