Direct answer
AI detectors have improved at catching clean, unedited LLM output, but the evidence that they have improved specifically against humanized text is weak — and the best-documented failure remains false positives on genuine human writing, including polished academic prose.
The distinction matters because "humanized text" covers two different inputs. The first is text run through a paraphrasing or humanizing tool. The second is text written by a person that simply reads like polished, fluent AI output. Detectors handle these very differently, and most public claims about "improvement" collapse the two together.
A 2025 review by Erol et al. found that AI-output detectors show "moderate to high success" at distinguishing AI-generated text, but the same review explicitly warns that false positives pose risks to researchers [1]. That is not a ringing endorsement of reliability. A January 2025 study summarized by Illinois State University is described as showing that detectors "remain consistently inconsistent," sometimes approaching accuracy and then delivering different scores on the same or similar input [2].
The strongest recent evidence points in an uncomfortable direction. Wordvice analyzed 135,389 pairs of academic manuscripts — the original and a professionally edited version of the same paper by the same author — drawn from submissions between 2018 and 2025. Professional human editing alone could change an AI detector's judgment even when the author and the content were unchanged [3]. If polishing a human manuscript moves it toward an AI verdict, then humanized text is not a niche edge case. It is the center of the problem.
Meanwhile, the classic false-positive finding has not gone away. Liang et al. (2023) found that AI detectors often misclassified TOEFL essays written by nonnative English speakers as AI-generated [3]. The University of San Diego law library guide notes that multiple studies show AI detectors were "neither accurate nor reliable," producing high numbers of both false positives and false negatives [4]. And Tsigaris et al. (2026) argue that detector performance must be assessed through sensitivity, specificity, prevalence, and ultimately the false discovery rate — meaning that even a highly accurate detector produces mostly false accusations when AI use is rare in the tested population [5].
So the honest answer to the question is: detectors are getting better at the easy case, and the hard case — text that has been made to read more fluently — is exactly where the improvement story breaks down.
Why "Getting Better" Depends on Which Failure You Measure
Detector vendors measure two things: the false positive rate (flagging human text as AI) and the false negative rate (missing AI text). These move in opposite directions. A detector can improve at catching AI while getting worse at clearing humans, or vice versa. Any claim that detectors are "getting better" is therefore incomplete until you specify which error rate improved and by how much.
The published numbers diverge so widely that they cannot all describe the same underlying reality. Pangram claims an industry-leading 1 in 10,000 (0.01%) false positive rate, with domain breakdowns of 0.01% for creative writing, 0.02% for academic writing, 0.01% for biomedical writing, and 0% for movie scripts — while conceding worse performance on poetry and recipes [6]. Jisc's 2025 update, by contrast, states that false positive rates for mainstream paid detectors such as Turnitin are "relatively low," with the best tools reporting around 1–2% [7].
That gap between 0.01% and 1–2% is a 100–200x difference. It is not a rounding disagreement. It is a strong, checkable illustration of how "better" depends entirely on which vendor and which benchmark you trust. A student reading Pangram's page and a student reading Jisc's update would walk away with completely different expectations of risk.
There is a further problem with headline accuracy figures. One evaluation cited in secondary coverage found that leading tools achieved overall accuracy of only 69% and 61%, with performance on hybrid human–AI texts dropping to nearly 0% [8]. That last number deserves attention: hybrid text — part human, part machine — is arguably the most common real-world case in academic writing, and it is the case where detectors perform worst. This figure comes from secondary coverage, so the underlying study should be verified before it is treated as settled.
Tsigaris et al. (2026) add the statistical layer that most vendor pages omit. Even a detector with strong sensitivity and specificity produces a high false discovery rate when the base rate of AI use is low [5]. In a cohort where few students use AI, most positive flags will be wrong — not because the detector is broken, but because of how conditional probability works. That is a mathematical property, not a marketing claim, and it applies to every detector on the market.
The Humanized-Text Blind Spot
Humanized text sits precisely in the blind spot of current detectors, because the act of making writing more fluent and polished moves it toward the statistical profile detectors associate with AI.
This is the mechanism behind the Wordvice finding. If making human writing more fluent and polished moves it toward "AI," then humanized text is exactly the kind of input detectors are most likely to misjudge — and the improvement story is not straightforward [3]. The 135,389-pair design is what makes this credible: the same author, the same content, the same paper, differing only by professional editing. When the verdict changes under those conditions, the detector is responding to style, not to authorship.
Other evidence points the same way, though with weaker sourcing. Cooperman and Brandao's study found that ZeroGPT identified 83% of human-written text as AI, as reported in secondary coverage [9]. That figure should be verified against the original paper before being relied on, but the direction is consistent with the Wordvice result and with the older Liang et al. finding on nonnative speakers.
Public discussion reflects the same frustration, though it is not a substitute for evidence. A Reddit PSA thread title asserts that "AI detectors have a 15% false positive rate" and that they "flag real human writing as AI constantly" [10]. The 15% figure is unverified, and the thread could not be fetched, so it should be treated as a claim rather than a fact. It is included here because it illustrates the range of numbers circulating publicly — 0.01%, 1–2%, 15% — none of which agree.
The pattern across all of this is consistent. Detectors are tuned to separate clean machine output from human writing. Humanizing deliberately blurs that boundary by making machine text more human-like, and polishing a human manuscript does the same thing from the other direction. Both operations push text into the region where the detector's signal is weakest.
What Turnitin0's Own Research Shows
Turnitin0's first-party research shows that unedited LLM output is caught at very high rates, while humanized output is treated as substantially human — and that genuine human writing produces no word-level false positives in its tested corpora.
On the unedited side, the numbers are close to ceiling. In TT0-2026-0008, 180 unedited GPT-5.6-Sol essays totaling 156,955 words across 30 majors produced an overall 97.88% of words flagged as AI-generated. A companion study on Claude Fable-5 essays, TT0-2026-0007, found 99.01% of words flagged as AI-generated across 131,451 words. Clean machine output is not a hard problem for Turnitin.
On the humanized side, the picture changes. In TT0-2026-0009, 174 GPT-5.6-Sol essays humanized by Turnitin0, totaling 204,736 words across 30 majors, produced an overall 76.44% of words treated as human-written. That is a substantial shift from the 97.88% flagged in the unedited condition, though it is not a clean sweep — roughly a quarter of words still read as AI to the detector.
The human-writing baseline is the most important result for anyone worried about a false accusation. TT0-2026-0005 examined 504 human-written PLOS graduate essays, 135,712 words, 18 majors, non-ESL, 400–800 words each, and found 100.0% of words classified as human-written, with no word-level false positives reported. TT0-2026-0004 found the same 100.0% result across 340 human-written CELL undergraduate ESL essays totaling 263,329 words. Both results are specific to these corpora and these conditions, but they establish that Turnitin does not flag genuine human academic writing by default.
One more result is worth noting because it complicates the "polishing is safe" assumption. TT0-2026-0006 took 500 PLOS graduate essays and had them polished by GPT-5.6-sol — a lighter intervention than full humanization — and found 47.54% of words flagged as AI-generated across 132,275 words. Polishing alone moved nearly half the text into AI territory. That is consistent with the Wordvice finding that editing changes detector verdicts, and it is the clearest first-party evidence that fluency, not authorship, drives the signal.
What This Means for Students Facing a Detector Verdict
Because the measurable damage so far falls on human writers rather than on humanizers, students should verify their actual Turnitin result before submission rather than trust a third-party detector's verdict.
The reasoning is straightforward. The documented failures — Liang et al. on nonnative speakers, Wordvice on edited manuscripts, the USD guide's summary of unreliable results — all involve human writing being misjudged. A free third-party detector that produces a scary number is not evidence about what Turnitin will say. It is evidence about that detector, on that day, on that input.
Turnitin0's checking service is built for exactly this gap. Users upload .docx, .pdf, or .txt for a pre-submission check; English documents only; word count must be greater than 300 and less than 30,000; file size under 20 MB. Each order includes two downloadable PDFs in one checkout: a Turnitin AI detection report and a similarity/plagiarism report, identical to what professors see in their LMS.
One display detail matters for interpretation. Turnitin shows *% instead of an exact percentage when AI detection is below its 20% confidence threshold. Those are low-confidence signals, not a hidden high score. Students who see *% should read it as "below threshold," not as an unknown number that might be damning.
Turnaround is under 15 minutes in 98% of cases, with most orders finishing within 5–15 minutes; in rare queue spikes, delivery is still guaranteed within 30 minutes. The file is checked without being added to Turnitin's student paper database, and reports are not shared with third-party databases; users can delete files from their account. There is no subscription.
The service reports 100,000+ Turnitin AI and similarity reports delivered, 20,000+ students worldwide across the United States, United Kingdom, Canada, Australia, New Zealand, and Ireland, and a 4.9/5.0 satisfaction rating. On Trustpilot, the profile shows a TrustScore of 4.3/5 with the label Excellent, based on 9 reviews all in the last 12 months; 5-star reviews make up 89% and 4-star 11%, with no negative reviews at capture [11]. Trustpilot notes that the company has not recently invited customers, so the reviews may not be representative. Recurring themes in those reviews are that the process is easy and fast, that reports come back sooner than expected, that the pricing is fair compared with other checkers, that the AI and similarity PDFs download together, and that the humanizer kept meaning while sounding more natural.
Where Turnitin0's Humanizer Fits
Turnitin0's AI humanizer is designed to lower the Turnitin AI score to *% or <20%, or even 0%, for text drafted with ChatGPT, Claude, or Gemini — or the user gets a full refund.
Users upload .docx or .txt; English documents only; file size under 90 MB. In a few minutes they receive a humanized version that rewrites flagged passages while preserving meaning, citations, headings, and .docx formatting. The formatting preservation matters more than it sounds: it removes the copy-paste reformatting step that otherwise eats time before a deadline.
The score promise is specific and bounded. For text drafted with ChatGPT, Claude, or Gemini, the system can lower the Turnitin AI score to *% or <20%, or even 0%, or the user gets a full refund. That is a claim about those models and that detector, not a general claim about all AI detection everywhere.
Two operational details are worth stating plainly. First, 98.2% of humanizer orders are re-checked with Turnitin, which means the outcome is being verified rather than assumed. Second, new users sign in with Google and can pay with PayPal or a prepaid balance. There is no free word quota or free trial for the humanizer.
Read against the research above, the humanizer addresses the specific failure mode this article has documented. The blind spot exists because humanized text sits between clean machine output and human writing. The 76.44% human-treated figure in TT0-2026-0009 is the measured result of pushing text through that process, and the refund guarantee is what backs the claim when it does not land.
What a Pre-Submission Check Actually Costs
Pricing is pay-per-use with no subscription, which matters when you only need a verdict once or twice a term. A single Turnitin check is $3.80, and prepaid packs lower the per-check rate: 2 scans for $6.50, 5 for $15.00, and 10 for $27.50, with packs valid for 100 days. The 10-check pack works out to $2.75 per check. The AI humanizer is priced separately at $2.00 per 1,000 words, rounded up to the next 1,000-word block, with prepaid word packs starting at $18.00 for 10,000 words that never expire.
The Practical Takeaway
The evidence in this article points to one conclusion: the risk that matters most is a false positive on your own writing, not a detector's ability to catch a humanizer. Verify your actual Turnitin result first rather than reacting to a third-party score.
If you want the closest available match to what your professor will see, the structural answer is to run the same system they do. Only Turnitin itself answers that question — every third-party checker is a proxy with its own model and its own error profile.
FAQ
Are AI detectors getting better at catching humanized text?
The evidence for improvement specifically against humanized text is weak. Detectors have improved on clean, unedited LLM output, but humanized text sits in the blind spot because polishing moves writing toward the statistical profile detectors associate with AI. Wordvice's 135,389-pair study found professional human editing alone could change a detector's judgment even when author and content were unchanged. The measurable damage so far falls on human writers, not on humanizers.
Why do AI detectors flag human writing as AI?
Detectors rely on statistical patterns such as perplexity and burstiness, and fluent, polished prose can resemble AI output. Liang et al. (2023) found detectors often misclassified TOEFL essays by nonnative English speakers as AI-generated. Cooperman and Brandao's study found ZeroGPT identified 83% of human-written text as AI, as reported. The University of San Diego law library guide notes that multiple studies show detectors were "neither accurate nor reliable."
How accurate are AI detectors really?
Vendor-reported false positive rates vary by orders of magnitude. Pangram claims 0.01%, while Jisc cites 1–2% for mainstream paid detectors such as Turnitin — a 100–200x gap. One evaluation cited in secondary coverage found leading tools achieved overall accuracy of only 69% and 61%, with hybrid human–AI texts dropping to nearly 0%. Tsigaris et al. (2026) argue even a highly accurate detector produces mostly false accusations when AI use is rare in the tested population.
Does Turnitin detect humanized text?
Turnitin0's own research shows unedited LLM output is caught at very high rates — 97.88% for GPT-5.6-Sol and 99.01% for Claude Fable-5 — while humanized output is treated as substantially human at 76.44%. Genuine human writing produced no word-level false positives in tested corpora: 100.0% for PLOS graduate essays and 100.0% for CELL ESL essays. Turnitin shows *% instead of an exact percentage when AI detection is below its 20% confidence threshold.
What should I do if I'm worried about a false positive?
Verify your actual Turnitin result before submission rather than trusting a third-party detector's verdict. Turnitin0's checking service delivers a Turnitin AI detection report and a similarity/plagiarism report identical to what professors see in their LMS, usually in under 15 minutes. The file is checked without being added to Turnitin's student paper database, and users can delete files from their account. No subscription is required.