Research report
Does Turnitin AI Detect AI-polished Human-written Research Papers with GPT-5.6 Sol?
Benchmark Experiment: Word-Level Accuracy of Turnitin AI Detection on GPT-5.6 Sol-Polished Research Abstracts
Overview
This benchmark evaluates Turnitin's AI detection capability on 500 research paper abstracts from the PLOS corpus, each polished by the GPT-5.6 Sol model. The dataset spans multiple domains and majors, with a total of 132,275 words. The primary metric is word-level accuracy, measuring the proportion of words correctly flagged as AI-generated.
- Overall word accuracy is 47.54%, meaning Turnitin correctly identified 62,879 out of 132,275 words as AI-polished.
- Domain breakdown:
- Business: 64.36% accuracy (809 correct words out of 1,257)
- Humanities: 28.27% accuracy (834 correct words out of 2,950)
- Psychology: 46.89% accuracy (5,058 correct words out of 10,787)
- STEM: 47.90% accuracy (56,178 correct words out of 117,281)
- Major breakdown:
- Highest accuracy: History (95.37%), Chemistry (94.10%), Economics (64.36%)
- Lowest accuracy: Civil Engineering (0.00%), Mechanical Engineering (0.00%), Criminal Justice (0.00%)
- Word count bucket:
- Abstracts around 400 words: 47.66% accuracy (62,686 correct words out of 131,539)
- Abstracts around 800 words: 26.22% accuracy (193 correct words out of 736)
- Dataset: 500 research paper abstracts from the PLOS corpus, each polished by GPT-5.6 Sol (OpenAI ChatGPT).
- Word count: Each abstract between 400 and 800 words, total 132,275 words (excluding references/bibliography).
- Domains: Business, Humanities, Psychology, STEM.
- Majors: 18 distinct majors including Biology, Computer Science, Economics, Physics, etc.
- Evaluation: Turnitin AI detection was run on each abstract. Word-level accuracy was calculated as the percentage of words correctly flagged as AI-generated.
- Metrics reported: Total articles, correct words, total words, and word accuracy percentage for each dimension.
- Scoring unit: accuracy is measured in words. First identify correct paragraphs under the expectation below, then sum the words in those paragraphs for the correct-word count; total words are summed over all counted paragraphs.
- Correct paragraph: a counted paragraph that was flagged as AI. Correct words are the words in those paragraphs.
- References excluded: from a References/Bibliography-style heading onward, those paragraphs are left out of word counts.
- Single LLM: Results are specific to GPT-5.6 Sol; other AI models may yield different detection rates.
- Dataset composition: The dataset is heavily skewed toward STEM (443 articles) and Biology (276 articles), which may influence overall accuracy.
- Zero-accuracy majors: Some majors (Civil Engineering, Mechanical Engineering, Criminal Justice) had only 1 article each, limiting reliability.
Overall Performance
- Total articles: 500
- Correct words: 62,879 out of 132,275
- Word accuracy: 47.54%
- This indicates that Turnitin correctly flagged less than half of the AI-polished words on average.
Performance by Domain
- Business: 64.36% accuracy (809 correct words out of 1,257)
- Humanities: 28.27% accuracy (834 correct words out of 2,950)
- Psychology: 46.89% accuracy (5,058 correct words out of 10,787)
- STEM: 47.90% accuracy (56,178 correct words out of 117,281)
- Business shows the highest accuracy, while Humanities lags significantly.
Performance by Major
- Highest accuracy:
- History: 95.37% (268 correct words out of 281)
- Chemistry: 94.10% (271 correct words out of 288)
- Economics: 64.36% (809 correct words out of 1,257)
- Lowest accuracy:
- Civil Engineering: 0.00% (0 correct words out of 258)
- Mechanical Engineering: 0.00% (0 correct words out of 211)
- Criminal Justice: 0.00% (0 correct words out of 159)
- Other notable majors:
- Biology: 59.26% (41,711 correct words out of 70,392)
- Computer Science: 26.89% (749 correct words out of 2,785)
- Health Sciences: 28.48% (7,161 correct words out of 25,148)
Performance by Word Count Bucket
- 400-word bucket (499 articles): 47.66% accuracy (62,686 correct words out of 131,539)
- 800-word bucket (1 article): 26.22% accuracy (193 correct words out of 736)
- The 800-word bucket shows notably lower accuracy, but the sample size is very small.