Search for a free AI detector and you will find twenty sites, each offering a percentage score in exchange for your text. The scores disagree with each other constantly. That is not a bug in one tool — it is what these tools are, and understanding why changes both which one you pick and what you can honestly do with the result.
This page compares the free tiers on what they actually give you, then covers the two pieces of published evidence that matter more than any of the comparisons, and finally the practical playbook for the situation most people are actually in: being flagged for text you wrote yourself.
1. What an AI detector actually measures
A detector is not comparing your essay against a database of ChatGPT outputs. There is no such database — generated text is not published anywhere as a corpus to match against, and the models produce different text every time.
What a detector does instead is measure statistical properties of your writing and compare them to what its training data says human writing usually looks like. The two properties most often cited are:
- Perplexity — how predictable each word is given the ones before it. Language models pick likely words, so generated text trends toward low perplexity. Fluent, plain, well-edited human writing also trends toward low perplexity, which is the entire problem.
- Burstiness — how much sentence length and complexity vary. Human writing swings: a long clause, then a four-word sentence. Generated text is more even. Again, a writer who has been taught to write cleanly and consistently looks like the machine.
Both measures are statistical tendencies with overlap between the two populations. A detector turns that overlap into a percentage, and the percentage is presented as confidence — but there is no ground truth it was ever checked against for your document. A 97% score means "this text is statistically unusual in this particular way", not "this text was written by AI".
2. The evidence that matters more than the tool comparison
Two documented events should be in front of anyone choosing a detector.
OpenAI retired its own detector over accuracy
OpenAI shipped a tool specifically designed to identify AI-written text, and then withdrew it. Their own notice is blunt: as of 20 July 2023, the AI classifier is no longer available due to its low rate of accuracy. The company that makes one of the models being detected concluded that detecting its output reliably was not something it could do.
Detectors systematically misclassify non-native English writers
The second is worse, because it has victims. A 2023 study published in Patterns (Liang et al., Stanford) tested seven widely used detectors on 91 human-written TOEFL essays by non-native English speakers, alongside 88 essays by US eighth-graders as a control.
- Across the seven detectors, a mean of 61.22% of the human-written TOEFL essays were flagged as AI-generated.
- 89 of the 91 essays (97.8%) were flagged as AI-generated by at least one of the seven.
- 18 of the 91 (19.78%) were flagged as AI-written by all seven working together — unanimous agreement on text a human wrote.
- On the native-speaker control essays, false positives were close to zero.
Read that last line again, because it is the finding. The detectors were not simply inaccurate at random. They were accurate on the control group and wrong on non-native writers, at scale. The study's explanation is that limited vocabulary and simpler sentence construction — which is what developing English proficiency looks like — are exactly the features the detectors treat as machine-like.
If you are writing in a second or third language, or if you write plainly on purpose, any score you get from a detector is drawing on a signal that was measured to be wrong for a majority of people in your position. That does not make the tools worthless. It makes a bare percentage worthless as proof.
3. The free tiers, compared on what they give you
All of these have a usable free mode. The differences that matter are how much text you can check without paying, whether they publish anything about accuracy, and whether a free result is a percentage or just a pass/fail label.
GPTZero
One of the two tools that study tested, and the one most often cited in academic-integrity discussions, so it is the result a teacher or editor is most likely to be holding. The free tier gives a limited monthly word allowance and a document-level verdict with sentence highlighting. It reports a probability and its own error rate rather than presenting a verdict as fact. Worth using because it is the one others cite, not because it is the most accurate.
ZeroGPT
Free with a genuinely generous character allowance and no account required, which makes it the fastest one to just try. It returns a percentage and highlights "AI" sentences. It publishes no accuracy figures, and unaccountable percentages from a tool with no stated methodology are the weakest form of evidence on this page — useful as a second opinion, never as a decision.
QuillBot AI Detector
Free, short-document focused, and paired with a paraphraser, which makes it the obvious pick if you are already using QuillBot for editing. The commercial relationship is worth naming: the same product sells a tool to rewrite text so that detectors do not flag it. A detector and a detector-evasion product from one vendor is not an independent check.
Copyleaks AI Detector
Offers a free scan with a limited word count, aimed at educators, and is one of the few that publishes accuracy and false-positive figures in its documentation. The free tier is capped tighter than the others, but a tool that states its error rate is more useful than one that implies it has none.
Sapling AI Detector
Free, no account, short inputs, fast. Returns a probability per sentence. It is a reasonable quick sanity check and a poor basis for an accusation, for the reasons above.
Winston AI
Free tier with a small monthly word allowance; the paid tiers are what the product is built around. If you need many documents checked, the free allowance runs out quickly, which is worth knowing before you build any process on it.
Scribbr's free AI detector
Aimed at students writing academic work, with a free word allowance and a clear statement that the result should not be used as evidence on its own — a caveat that speaks well of it.
Originality.ai
Widely used by publishers and SEO teams. Free checking is minimal; the product is subscription-oriented, and it is the one that also runs plagiarism checks side by side. Include it when the question is "will a client's tool flag this", which is a different question from "is this AI".
4. Why two detectors give you two different answers
If your paragraph scores 12% on one and 88% on another, you have not made an error and neither has the software necessarily. They differ because:
- Different features. Some weight perplexity, some weight burstiness, some add their own signals like repetition of rare tokens. A paragraph can look human on one axis and machine-like on another.
- Different training data. Each vendor assembled its own human sample. If your writing style is under-represented in that sample — technical documentation, second-language English, academic register, dialect — you sit near the edge of what the tool calls normal.
- Different thresholds. One tool calls 60% "likely AI", another calls it "mixed". The percentage is not a calibrated probability, so it is not comparable across tools.
- Length. Short texts give the statistics less to work with, so scores wobble far more. A 150-word abstract is the worst case for every one of these tools.
- They keep changing. Models and thresholds are updated, so a result you saved last month may not reproduce today. Never treat a saved screenshot as a stable measurement.
The practical consequence: collect results, not a result. Two detectors agreeing is weak evidence. One detector at 99% with no stated methodology is no evidence at all.
5. If you have been wrongly flagged: the playbook
This is the situation that brings most people to this page, and it has an answer that does not involve arguing about percentages.
- Ask which tool, and ask for the report. A verdict handed down without the tool, version, input length and score is not a finding you can respond to. Ask in writing, politely, for the specifics.
- Show the process, not the score. Version history is the strongest thing you can produce: a Google Doc or Word file with edit history shows the text being built over days, with revisions, reordering and deletions. Detectors cannot see that, and humans find it far more convincing than any counter-percentage.
- Bring the false-positive evidence. The Stanford study is published, peer-reviewed, and specifically measured this failure on second-language writers. If that is your situation, a finding has to explain why it is exempt from a documented bias rather than the other way round.
- Offer to be assessed on your actual work. An oral defence, a supervised in-class paragraph on the same topic, a follow-up assignment. This reframes the question from "what does the software say" to "can this person do the work", which is the question that matters and the one you can answer.
- Do not defend yourself with a second detector. A different tool saying 4% does not refute 88%, because you would be arguing that this particular machine is the authority after all. It concedes the premise.
- Escalate to a human decision-maker. Ask for the policy on file, ask who owns the final decision, and ask for it in writing. Institutions have appeals processes; use the one that exists.
6. If you are assessing other people's writing
Given the measurement above, a detector score is not sufficient grounds for a misconduct finding, and treating it as one exposes you to a complaint you will lose. What holds up instead:
- Require process artifacts up front — drafts, outlines, version history, or a short in-class or oral component. Cheaper than investigating after the fact.
- Assess for understanding, not for polish. A viva, a follow-up question about a specific paragraph, or applying the same reasoning to a new example all discriminate between the writer and the tool.
- Publish the policy before the assignment. State whether AI assistance is permitted, and if so, how it must be disclosed. Most confusion here is a policy gap, not a cheating problem.
- Be explicit about the second-language bias. If any of your students write in a second language, you now know that a detector flag correlates with that as strongly as with AI use.
7. If you are writing with AI assistance
The honest position is simpler than the tooling suggests. Ask what your course, employer or client actually permits and follow that, and disclose when it is required. Two things worth doing regardless: keep your own drafts (the version history is your evidence), and use AI for the parts where you can check the output — outlining, restructuring, tightening a sentence you wrote — rather than generating a whole document you would not be able to defend in conversation. A detector score is unreliable either way; being able to explain your own text is not.
FAQ
What is the best free AI detector?
There is no best, because they measure the same statistical properties on different training data and disagree with each other. GPTZero and Copyleaks are the most useful free options to know about: GPTZero because it is the one academic discussions most often cite, Copyleaks because it publishes its accuracy and false-positive figures, which most do not.
Are AI detectors accurate?
Less than their confidence scores suggest. OpenAI withdrew its own AI text classifier on 20 July 2023 citing a low rate of accuracy. A 2023 Stanford study in Patterns found that across seven widely used detectors, a mean of 61.22% of human-written TOEFL essays by non-native English speakers were flagged as AI-generated, and 89 of 91 were flagged by at least one detector.
Why do different AI detectors give different answers?
They weight different signals (perplexity, burstiness, their own features), were trained on different human writing samples, and apply different thresholds. The percentage is not a calibrated probability, so scores are not comparable between tools, and short texts make every one of them wobble more.
Why was my own writing flagged as AI?
Because the detector is measuring style, not origin. Plain, consistent, well-edited writing, or writing in a second language, both look statistically machine-like. The Stanford study showed this bias hits non-native English writers hardest while a native-speaker control group produced almost no false positives.
What should I do if I am falsely accused of using AI?
Ask which tool and for the report; produce your version history, which is far stronger evidence than any counter-score; cite the documented false-positive research if you write in a second language; and ask for an oral or supervised assessment of your actual work. Do not try to rebut one detector with another — that concedes the premise that the software is the authority.
Can AI detectors detect ChatGPT reliably?
No, and the strongest evidence is that the company behind ChatGPT built a detector and retired it in July 2023 over its low accuracy. Detectors catch some generated text and misclassify a substantial amount of human writing, with the errors concentrated on non-native English speakers.
Do free AI detectors have word limits?
Yes, and they differ a lot. ZeroGPT and Sapling allow short text with no account, QuillBot and Copyleaks are capped more tightly, and Winston AI and Originality.ai are built around paid tiers with only a small free allowance. Short inputs also produce the least stable scores, so a long document checked in fragments gives you several different unreliable numbers.

