Research report
Can Turnitin0's Humanization of GPT-5.6-Sol-Generated Essays Bypass Turnitin AI Detection?
Benchmark Experiment: Evaluating Turnitin0's Humanization Against Turnitin AI Detection
Research ID: TT0-2026-0009
Published July 30, 2026
Cite this research
What journalists and bloggers can cite from this report.
Overall word accuracy : 76.44% (156,497 correct words out of 204,736 total words).
According to Turnitin0 Research, overall word accuracy : 76.44% (156,497 correct words out of 204,736 total words).
- 174 articles tested
- Checker: turnitin0726
- LLM: gpt-5.6-sol
- Humanizer: turnitin0
- 1 linked dataset
- Detection screenshots available
- Published July 30, 2026
Peng, J. (2026). Can Turnitin0's Humanization of GPT-5.6-Sol-Generated Essays Bypass Turnitin AI Detection?. Turnitin0 Research. https://www.turnitin0.com/research/reports/can-turnitin0-s-humanization-of-gpt-5-6-sol-generated-essays-bypass-turnitin-ai-detection
Licensed under CC0 1.0 Universal. You may copy, modify, distribute, and cite this research—including text, figures, tables, and datasets—without asking permission.
Overview
This benchmark evaluates whether Turnitin0's humanization of GPT-5.6-Sol-generated essays can bypass Turnitin's AI detection. A dataset of 174 humanized essays (204,736 words) across 30 majors, five domains, and two academic levels was tested. The primary metric is word-level accuracy—the percentage of words not flagged as AI-generated. An overall word accuracy of 76.44% indicates that Turnitin0 successfully evades detection for the majority of text, though performance varies by dimension.
- Overall word accuracy: 76.44% (156,497 correct words out of 204,736 total words).
- Domain performance: Education achieved 100% accuracy (6,750/6,750 words), while Humanities had the lowest at 70.83% (38,231/53,976 words).
- Major-level variation: Education (100%), Environmental Science (98.04%), and History (94.34%) showed the highest accuracy; English (55.41%) and Political Science (55.94%) the lowest.
- Academic level: Undergraduate essays (80.63% accuracy, 84,646/104,980 words) outperformed graduate essays (72.03%, 71,851/99,756 words).
- Word count buckets: 1,200-word essays had the highest accuracy (78.30%, 132,985/169,830 words), while 800-word essays had the lowest (63.80%, 4,999/7,836 words).
- Dataset: 174 humanized essays originally generated by GPT-5.6-Sol, then processed through Turnitin0's humanizer. Essays span 30 majors across five domains (Business, Education, Humanities, Psychology, STEM) and two academic levels (graduate, undergraduate). All are non-ESL, essay genre, with word counts in buckets of 400, 800, and 1,200 words. Total word count: 204,736 (excluding references/bibliography).
- Detection tool: Turnitin AI detection was run on each essay to identify AI-generated words.
- Metric: Word-level accuracy = (correct words not flagged as AI) / (total words) × 100. Only word-level counts are reported; no article-level or paragraph-level accuracy is used.
- Dimensions analyzed: Overall, domain, major, academic level, and word count bucket.
- Scoring unit: accuracy is measured in words. First identify correct paragraphs under the expectation below, then sum the words in those paragraphs for the correct-word count; total words are summed over all counted paragraphs.
- Correct paragraph: a counted paragraph that was not flagged as AI. Correct words are the words in those paragraphs.
- References excluded: from a References/Bibliography-style heading onward, those paragraphs are left out of word counts.
- Dataset scope: Essays are limited to 30 majors and two academic levels; performance may vary for other subjects or writing styles.
- Total articles: 174
- Correct words: 156,497
- Total words: 204,736
- Word accuracy: 76.44%
This overall result shows that Turnitin0's humanization successfully bypasses Turnitin AI detection for the majority of words across the entire dataset.
- Business: 71.67% accuracy (25,896/36,133 words)
- Education: 100.00% accuracy (6,750/6,750 words)
- Humanities: 70.83% accuracy (38,231/53,976 words)
- Psychology: 90.81% accuracy (7,036/7,748 words)
- STEM: 78.48% accuracy (78,584/100,129 words)
Education and Psychology domains show the highest evasion rates, while Humanities and Business are lower.
- Top 3 majors: Education (100.00%), Environmental Science (98.04%), History (94.34%)
- Bottom 3 majors: English (55.41%), Political Science (55.94%), Health Sciences (62.28%)
- Range: From 55.41% (English) to 100.00% (Education)
Accuracy varies widely across majors, suggesting that subject-specific writing styles influence detection evasion.
- Graduate: 72.03% accuracy (71,851/99,756 words)
- Undergraduate: 80.63% accuracy (84,646/104,980 words)
Undergraduate essays show higher evasion rates than graduate essays, possibly due to simpler language patterns.
- 400 words: 68.39% accuracy (18,513/27,070 words)
- 800 words: 63.80% accuracy (4,999/7,836 words)
- 1,200 words: 78.30% accuracy (132,985/169,830 words)
Longer essays (1,200 words) achieve higher accuracy, while shorter essays (800 words) are more likely to be flagged.
Detection report screenshots
Turnitin marks suspected AIGC text in cyan.