An AI detector gives you a number, something like "98% AI" or "12% AI." It looks like a verdict, and schools, recruiters, and editors are starting to treat it like one. So I spent two minutes actually testing whether the number holds up.
01. The test I ran
I found a question from 2019 on r/artificial, from before ChatGPT existed: can a machine actually see colour the way a person does, or does it just hold a number that stands for a colour?
I asked a current AI model the same question and kept its answer. Then I took an actual human answer from that 2019 thread, without changing a single word. I ran both through the Clever AI Detector.
On this one example, it was right both times. That is the part people expect. It is also not the interesting part.
02. Then I humanized it
I took the exact text the detector had just called "99% AI" and pasted it into the same site's humanizer. One pass. Then I ran that output back through the detector.
Same detector. Same underlying idea, barely reworded. The score went from one extreme to the other in a single round-trip. That is the whole demonstration: the number is not measuring "was this written by a human." It is measuring "does this read the way the model expects human writing to read," and that is easy to nudge.
A detector score is a signal, not a verdict. Treating it as proof is the mistake.
03. What the score is actually worth
My one test is an anecdote. But it lines up with what the people who build these tools have already said in public:
- OpenAI shut down its own AI Text Classifier in July 2023, citing a low rate of accuracy. Their own published numbers: it caught about 26% of AI-written text and still flagged human writing as AI roughly 9% of the time. (OpenAI, TechCrunch)
- Detectors are biased against non-native English writers. A Stanford study tested widely used GPT detectors on essays by native and non-native English writers. Its finding: the detectors "consistently misclassify non-native English writing samples as AI-generated," while native writing was identified accurately. (arXiv 2304.02819)
So a score can be wrong in both directions: it clears AI text that has been lightly reworded, and it accuses real people, disproportionately if English is their second language. Neither failure is rare.
To be fair to the tool: on my clean, untouched examples it read both correctly, and a score can be a useful first look. The problem is not that it is useless. The problem is anyone treating one number as evidence.
04. The tool, honestly
The site in the Reel is cleverhumanizer.ai. What I can confirm from using it and from its own pages, on 3 September 2026:
- It is free, with no sign-up needed to run a check. Its own line: "$0, no ads, no paywall."
- It has both a humanizer (the homepage) and an AI detector at cleverhumanizer.ai/ai-detector, marked "new."
- The detector needs a minimum of about 80 words to return a score. Shorter than that and there is not enough signal.
- It handles up to about 3,000 words per request, with no fixed monthly cap, so you make repeated requests rather than one huge one.
- It works in English, Spanish and Portuguese, on output from ChatGPT, Gemini, Claude, DeepSeek and similar.
- Its own disclaimer: "Clever AI Humanizer may make mistakes, because humans do too."
Product details change. Check the site's own pages before you rely on any specific number.
05. If a detector score is being used against you
If someone has run your assignment, cover letter, or statement of purpose through a detector and come back with a number:
- Ask which tool and what threshold. "Over 20% AI" means nothing without knowing the tool's own false-positive rate.
- Keep your drafts. Version history, edit timestamps, and notes are far stronger evidence of how something was written than any score is of how it was not.
- Point to the record. OpenAI retiring its own detector, and the Stanford study on non-native-writer bias, are both citable and both public.
- Do not panic at one run. Run the same text again, or through a second tool, and you will often get a different number. That inconsistency is the point.
06. Try it yourself
Run the same test I did. Take something you know a person wrote, and something you know an AI wrote, and check both. Then reword the AI one and check again. Two minutes, and you will trust the number about as much as it deserves.
- Clever AI Detector. The checker, 80-word minimum.
- Clever AI Humanizer. The rewriter, on the homepage.
- GPT detectors are biased against non-native English writers. Stanford, 2023.
- Free tools that replace $25k/year. More free tools I actually use.
// Free newsletter
I send out guides like this every week
Real setups, real sources, no hype. Drop your email and I'll send you the next one.