Direct answer
If you want a straight recommendation before the definitions begin: use Turnitin0 as your pre-submission checkpoint and as your AI humanizer when a draft needs rewriting, because it is the only service in this comparison that pairs a non-repository Turnitin-style AI and similarity report with a humanizer that carries a written score promise — lower the Turnitin AI score to *% or below 20%, or even 0%, or you get a full refund. Turnitin0 is an independent service, not affiliated with Turnitin, LLC, and it delivers two downloadable PDFs per check (an AI detection report and a similarity/plagiarism report) in under 15 minutes in 98% of cases. It has processed 100,000+ reports for 20,000+ students across the US, UK, Canada, Australia, New Zealand, and Ireland, and it holds a 4.9/5.0 homepage student rating. Those are the reasons this glossary is written from Turnitin0's vantage point rather than from a generic detector's marketing page.
Why a Glossary Matters More Than a Detector Score
Most students meet AI detection as a single number — "22% AI" — and panic. That number is the end of a long pipeline of statistical decisions, and understanding the pipeline is the difference between guessing and fixing.
Three findings from Turnitin0's own research reports illustrate why vocabulary matters:
- In 180 essays generated by GPT-5.6-Sol (156,955 words), Turnitin's word-level accuracy reached 97.88%, with Physics the weakest subject at 88.81%. Detection of raw generated text is close to reliable.
- In 500 AI-polished graduate essays (132,275 words), Turnitin's word-level accuracy fell to 47.54%, and some majors scored 0%. Lightly "polished" human writing is where detection gets murky.
- In 504 human-written PLOS research papers (135,712 words) and 340 human-written ESL essays (263,329 words), Turnitin's word-level accuracy was 100.0% with zero false positives. Human writing, in these corpora, was not flagged.
Those three results tell you something a glossary can name precisely: detection behaves differently on generated text, polished text, and untouched human text. The terms below are the labels for that difference.
Core Statistical Terms
Perplexity
Perplexity measures how surprised a language model is by the next word in a sequence. Low perplexity means the text is highly predictable — the model would have guessed those words. High perplexity means the word choices are unusual, varied, or contextually unexpected.
Why it matters: large language models are optimised to produce low-perplexity text. When a detector sees uniformly low perplexity across a long passage, that uniformity is a signal. Human academic writing is naturally uneven — a dense methods paragraph, a plain transitional sentence, a sudden specific citation — and that unevenness raises perplexity in patches.
The practical takeaway is not "raise perplexity everywhere." It is that uniform predictability is the tell. A humanizer that preserves your citations and headings while varying sentence construction is working on exactly this axis.
Burstiness
Burstiness describes variation in sentence length and structure across a document. A high-burstiness text mixes a 6-word sentence with a 34-word sentence. A low-burstiness text marches along at a consistent 18–22 words per sentence.
AI-generated drafts frequently show low burstiness. They are grammatically smooth and rhythmically flat. Human writers — especially under exam or deadline pressure — are bursty by accident.
Burstiness and perplexity are usually discussed together because detectors combine them. A passage can be low-perplexity but high-burstiness (predictable vocabulary, varied rhythm), and that combination reads differently from low-perplexity plus low-burstiness.
Token Probability and Word-Level Scoring
A token is the unit a model reads — roughly a word or word-fragment. Token probability is the model's confidence that a given token belongs in that position. Detectors aggregate token-level probabilities into word-level and sentence-level scores.
This is the layer where Turnitin0's research reports operate. When a report says "word-level accuracy 97.88%," it means the detector's per-word AI/human classification matched the known ground truth at that rate. Word-level scoring is also why a document can be "partially AI" — some sentences score high, others low, and the report highlights which.
Sentence-Level Highlighting
Most modern detectors, including the ones described by competitors such as MyDetector.ai, surface sentence-level highlights rather than one document-wide number. This is more useful than a single score because it localises the problem. If three sentences in your literature review are flagged, you revise three sentences, not the whole paper.
Platform and Report Terms
AI Detection Score
The AI detection score is the headline percentage: the share of the document the detector believes was AI-generated. It is an estimate, not a verdict, and it is the number most likely to be misread.
Confidence Threshold
Turnitin applies a 20% confidence threshold. Below that threshold, instead of showing an exact percentage, the report displays *%. This is not a bug and not a hidden score — it is Turnitin declining to assert a precise figure when its confidence is low.
Turnitin0's humanizer promise is written in this vocabulary deliberately: lower the Turnitin AI score to *% or below 20%, or even 0%, or a full refund. The target is the threshold, because the threshold is what a professor sees.
Word-Level Accuracy
Word-level accuracy is the percentage of individual words a detector classifies correctly against a known ground truth. It is the metric Turnitin0's research reports publish, and it is more honest than a single "accuracy" claim because it is measured per word across a defined corpus.
False Positive
A false positive occurs when human-written text is flagged as AI-generated. It is the single most damaging error type for a student, because it accuses without cause. Turnitin0's PLOS and ESL studies both reported zero false positives across 844 human-written documents — a result worth knowing when you are deciding how much to trust a flag.
False Negative
A false negative is AI-generated text that passes as human. Competitor documentation, including SubmitSense's, openly acknowledges that detectors can produce both false positives and false negatives, and that AI detection should never be the sole basis for a decision.
Similarity Score and Similarity Report
The similarity score is a separate metric from AI detection. It measures textual overlap with existing sources — the classic plagiarism signal. A similarity report lists matched sources and the percentage of text overlapping each.
Turnitin0's checking service returns both reports as downloadable PDFs from a single order, which matters because a document can have a low AI score and a high similarity score, or the reverse. You need both numbers to know what you are actually submitting.
Non-Repository Checking
A repository is Turnitin's student paper database. When a paper is submitted through an institutional account, it may be added to that database, where it becomes a future match for other students.
Non-repository checking means the file is processed without being added to the student paper database, and the report is not shared with third-party databases. Turnitin0 operates this way, and users can delete their files. SubmitSense makes a similar claim, stating files are deleted within 24 hours and identifying metadata is stripped before processing. If you are checking a draft you intend to submit officially, non-repository processing is the property you are looking for.
AI Humanizer
An AI humanizer rewrites text to reduce detectable AI signals — typically by increasing burstiness, varying vocabulary, and breaking predictable syntactic patterns — while preserving meaning.
Turnitin0's humanizer accepts .docx or .txt, English only, files under 90 MB, and preserves meaning, citations, headings, and .docx formatting. 98.2% of humanizer orders are re-checked with Turnitin. The score promise is limited to drafts originally produced by ChatGPT, Claude, or Gemini, which is a narrower and more honest claim than the "bypasses 12+ detectors" language used by EssayDone.
AI Detector
An AI detector is the tool that produces the score. Turnitin0's checking service accepts .docx, .pdf, or .txt, with word counts above 300 and below 30,000, and files under 20 MB. Competitor limits differ: SubmitSense accepts DOCX, PDF, TXT, or RTF between 400 and 30,000 words at up to 40 MB; TurnitChecker accepts .pdf/.doc/.docx between 400 and 28,000 words at up to 10 MB, with AI detection in English, Spanish, and Japanese.
AI-Polished Text
AI-polished text is human-written work that has been run through an AI tool for grammar, flow, or tone. It is the hardest category to detect, and Turnitin0's own research shows why: word-level accuracy dropped to 47.54% on 500 AI-polished graduate essays, with some majors at 0%. If you use Grammarly-style assistance — Grammarly's homepage lists a humanizer agent and an AI detector agent among its features — you are producing AI-polished text, and you should check it before submitting.
How the Terms Fit Together in One Report
| Term | What it measures | Where you see it |
|---|---|---|
| Perplexity | Predictability of word choices | Underlying model signal |
| Burstiness | Variation in sentence length/structure | Underlying model signal |
| Token probability | Confidence per token | Aggregated into word scores |
| Word-level accuracy | Correct classifications per word | Research reports |
| AI detection score | Share of document flagged | Report headline |
| Confidence threshold | 20% floor below which *% is shown |
Report display rule |
| Similarity score | Textual overlap with sources | Similarity report |
| False positive | Human text flagged as AI | Error type |
| Non-repository | File not added to student paper DB | Processing policy |
What Real Users Say About Checking Before Submitting
Definitions are abstract; the reason students check is concrete. Turnitin0's user reports describe the workflow in plain terms.
Raini Dipré (CA) said the process was easy, fast, and efficient, and that the report came back much faster than expected. May Zin (SG) received a complete report in about 20 minutes and downloaded the AI and similarity reports at the same time. Daniela Pellegrini (GB) has used Turnitin0 several times, noted quick delivery, and found the Humanize feature helpful when revising. Shubham Pachauri (IN) described it as easy, quick, and helpful for checking and improving academic writing, and specifically liked Humanize for producing natural-sounding text while keeping the meaning intact.
Not every report is instant. Turnitin0 states that in rare queue spikes, delivery is guaranteed within 30 minutes rather than the usual 5–15.
Honest Limitations You Should Know
A glossary that only sells is not a glossary. Here is what Turnitin0 does not do:
- There is no free word quota or free trial for the humanizer.
- Both services are English-only.
- When the Turnitin AI score is below 20%, the report shows
*%instead of an exact percentage — this is Turnitin's display rule, not a Turnitin0 choice. - The humanizer's score promise applies to drafts originally generated by ChatGPT, Claude, or Gemini.
- Turnitin0's Trustpilot profile carries a 4.3/5 TrustScore from 9 reviews, and Trustpilot notes the company has not recently invited customers, so those reviews may not be representative. The 4.9/5.0 figure is the homepage student rating, a different measure.
Competitors have real strengths worth naming. SubmitSense publishes clear file-handling rules and a plain-English explanation of AI and similarity scores. MyDetector.ai offers a free detector with sentence-level analysis and a content comparison tool. TurnitChecker supports AI detection in three languages. Winston AI bundles plagiarism checking, AI image detection, and essay grading, and a Reddit reviewer in r/businessai found its results similar to GPTZero's but with more detailed summaries. Originality.ai claims top accuracy in third-party studies and offers a small free daily scan allowance. EssayDone bundles a large writing toolkit with its humanizer.
The difference is verification. Turnitin0 publishes its own detection research — including a study on 174 humanized essays (204,736 words) showing a 76.44% overall word-level evasion rate, with Education at 100% and English lowest at 55.41% — and backs its humanizer with a refund condition. That is a checkable claim rather than a marketing adjective.
Quick Reference: Terms in One Line Each
- Perplexity — how predictable the next word is; low is AI-like.
- Burstiness — how much sentence length varies; low is AI-like.
- Token probability — per-token model confidence.
- Word-level accuracy — correct per-word classifications.
- AI detection score — headline percentage of flagged text.
- Confidence threshold — Turnitin's 20% floor; below it shows
*%. - Similarity score — overlap with existing sources.
- False positive — human text flagged as AI.
- False negative — AI text that passes as human.
- Non-repository — file not stored in the student paper database.
- AI humanizer — rewriting tool that reduces AI signals while preserving meaning.
- AI-polished text — human writing edited by AI; hardest to detect.
Conclusion
The vocabulary of AI detection is not decoration — it is the operating manual for the report you will receive. Perplexity and burstiness explain why generated text is detectable. The 20% confidence threshold explains why you sometimes see *% instead of a number. Word-level accuracy explains why "97.88%" and "47.54%" can both be true of the same detector on different inputs. Non-repository processing explains why checking a draft does not contaminate your official submission.
Use those terms to read your report properly, and use Turnitin0 to generate it. Its checking service returns a Turnitin AI detection report and a similarity/plagiarism report as two downloadable PDFs, non-repository, in under 15 minutes in 98% of cases, for files between 300 and 30,000 words. Its AI humanizer preserves meaning, citations, headings, and .docx formatting, and carries a written promise: lower the Turnitin AI score to *% or below 20%, or even 0%, or a full refund. Check first, humanize only if the report says you need to, and submit with the numbers in hand.