Research report
Does Turnitin Detect Gemini 3.5 Flash-Generated Essays?
A Benchmark Experiment on Turnitin's Detection of Gemini 3.5 Flash-Generated Essays
Overview
This report presents a benchmark experiment evaluating Turnitin's AI detection accuracy on essays generated by the Gemini 3.5 Flash language model. We analyzed 180 essays across two datasets (undergraduate and graduate levels), spanning 30 majors and five domains. The primary metric is word-level accuracy—the proportion of words correctly flagged as AI-generated. Overall, Turnitin achieved 98.35% word accuracy, correctly identifying 144,688 out of 147,117 words.
- Overall word accuracy: 98.35% — Turnitin correctly flagged 144,688 of 147,117 words across all 180 essays.
- Domain performance: Humanities (99.04%) and Business (99.02%) showed the highest accuracy; STEM (97.70%) was the lowest.
- Major-level variation: Accuracy ranged from 94.36% (Information Technology) to 99.82% (Business Administration and International Relations).
- Level comparison: Undergraduate essays (98.55%) were detected slightly more accurately than graduate essays (98.14%).
- Word count effect: 800-word essays achieved the highest accuracy (99.16%), while 1200-word essays had the lowest (97.88%).
- Datasets: Two datasets of 90 essays each, generated by Gemini 3.5 Flash:
- Undergraduate essays across 30 majors (Business, Education, Humanities, Psychology, STEM).
- Graduate-level essays across the same 30 majors.
- Essay specifications: All essays were non-ESL, written in the essay genre, with word counts of 400, 800, or 1200 words.
- Detection tool: Turnitin's AI detection system.
- Evaluation metric: Word-level accuracy = (correctly flagged words / total words) × 100. No humanizers were applied.
- Scoring unit: accuracy is measured in words. First identify correct paragraphs under the expectation below, then sum the words in those paragraphs for the correct-word count; total words are summed over all counted paragraphs.
- Correct paragraph: a counted paragraph that was flagged as AI. Correct words are the words in those paragraphs.
- References excluded: from a References/Bibliography-style heading onward, those paragraphs are left out of word counts.
- Domain and major sample sizes: Some domains (e.g., Education, Psychology) had only 6 essays, limiting statistical power.
- Word count buckets: The 400-word bucket had 61 articles, while 800 and 1200 had 59 and 60, respectively—slight imbalance.
Overall Performance
- Total articles: 180
- Correct words: 144,688 out of 147,117
- Word accuracy: 98.35%
Performance by Domain
- Business: 99.02% accuracy (24,608 correct of 24,852 words, 30 articles)
- Education: 98.63% accuracy (4,753 correct of 4,819 words, 6 articles)
- Humanities: 99.04% accuracy (38,565 correct of 38,939 words, 48 articles)
- Psychology: 98.94% accuracy (4,866 correct of 4,918 words, 6 articles)
- STEM: 97.70% accuracy (71,896 correct of 73,589 words, 90 articles)
Performance by Major
- Highest accuracy: Business Administration and International Relations (99.82% each)
- Lowest accuracy: Information Technology (94.36%), Mathematics (95.10%), Physics (96.17%)
- Other notable majors:
- Accounting: 98.74%
- Biology: 96.40%
- Computer Science: 98.20%
- Nursing: 98.23%
- Psychology: 98.94%
Performance by Academic Level
- Undergraduate: 98.55% accuracy (73,426 correct of 74,505 words, 90 articles)
- Graduate: 98.14% accuracy (71,262 correct of 72,612 words, 90 articles)
Performance by Word Count
- 400 words: 98.38% accuracy (22,829 correct of 23,205 words, 61 articles)
- 800 words: 99.16% accuracy (44,623 correct of 45,002 words, 59 articles)
- 1200 words: 97.88% accuracy (77,236 correct of 78,910 words, 60 articles)