Find the characters you cannot see

Text can carry invisible characters, hidden instructions and letters that only look like the ones you read. Scan any text, web page or document, see exactly what is in there, and get a clean copy back. And in the AI check tab, inspect predictability patterns across the full text. Spanish screening has passed a limited reserved test; English remains experimental. Results support editorial review and do not establish authorship.

URLs are fetched by our server to read the page (browsers block cross-site reads). The page content is public; nothing you paste in the Text or Files tabs is ever sent.

Drop files here or click to choose
Word (.docx), PDF, PowerPoint (.pptx), .txt, .md, .csv — parsed in your browser

This one is different. Measuring predictability needs a language model, so this text is sent to our server, scored and discarded. Nothing is stored. If that is not acceptable for this text, use the Text tab instead — that one never leaves your browser.

Spanish and English. Natural prose only.
We fetch each page, extract the article text and score it.
Private by design. Your text never leaves this page.

Why hidden characters matter

They can carry instructions meant for a machine, not for you. A zero-width character renders as nothing on screen, but it is still text. Someone can hide a line inside a CV, a web page or an email that a person will never see and an AI assistant will read and act on. If any part of your workflow lets a model read documents or pages you did not write, this is worth checking before it reaches the model.

Lookalike letters are a security problem, not a typo. A Cyrillic "а" is a different character from a Latin "a" and looks identical. That gap is what makes phishing domains work, what slips words past moderation filters, and what quietly breaks an exact match in a database. We flag them only inside mixed-script words, so genuine Cyrillic or Greek text is left alone.

And most text is dirtier than it looks. We scanned 136 real web pages and 74% of them carried non-standard spaces — over a thousand in total. Nobody put them there on purpose: they come from copying between a word processor, a PDF and a CMS. They are invisible until a table collapses, a CSV import fails or a search returns nothing.

About the AI watermark you probably came here for

It is real, and it is not a hidden character. Since 2 August 2026, Anthropic "weaves an imperceptible watermark directly into the text itself" for new Claude models. They have not published how it works, but their own documentation gives it away: the mark may go undetected when "the passage is very short". An inserted character works the same in three words as in three pages — it is there or it is not. Needing length only makes sense for a statistical mark, one that lives in which words the model picks. Google's SynthID works exactly that way, by nudging word probabilities as the text is generated.

Nobody outside the provider can read it today. Anthropic has published no detector, no specification and no date — only a promise to enable third-party detection, which the EU AI Act requires. SynthID's text detector needs Google's own key, and its public portal is still limited to testers. So any tool claiming it detects or removes the provider watermark right now is doing one of two things: looking for hidden characters that are not there, or paraphrasing your text with another model and charging you for it. Paraphrasing does degrade these marks — independent research puts it near 100% — but that is rewriting, not removal.

One thing you can actually check. For images, SynthID is verifiable today: upload the image to the Gemini app and ask whether it was generated by Google AI. It is the only part of this whole subject that works right now, free and without keys. For text, the wall stands — and the moment a public detector ships, we will connect to it here.

What this does not tell you

It does not tell you whether a text was written by AI. We measured this rather than guessed: across 925,000 characters of AI-generated articles, and across raw API output from five different models, we found zero invisible characters. Models do not mark their text this way, so a clean result here is not evidence a human wrote it — and a site that claims otherwise is selling you something it cannot deliver.

Two more limits, stated plainly. PDF text extraction can drop zero-width characters that carry no glyph, so a PDF scan is best effort — the text and file tabs are stronger. And we only read what is in the text: if a mark was never there, we will not invent one.