Turnitin0 - A trusted institution providing third-party Turnitin checking services to students worldwide.

Research report

Does Turnitin Detect Essays Generated by Claude-Fable-5?

Benchmark Experiment: Turnitin AI Detection Accuracy on Claude-Fable-5 Generated Essays

James Peng

James Peng Turnitin0.com | Lead Editor

3 years of writing experiences on Turnitin AI detection and humanizer.

Overview

This benchmark evaluates Turnitin's AI detection capabilities on a corpus of 170 essays (131,451 words) generated entirely by the Claude-Fable-5 language model. The essays span five domains and 30 majors, at graduate and undergraduate levels, with word counts of 400, 800, and 1,200 words. The goal is to measure word-level accuracy in identifying AI-generated text.

  • Overall word accuracy: 98.53% (129,515 correct words out of 131,451 total words).
  • Domain performance: Accuracy ranged from 97.45% (Business) to 98.87% (STEM).
  • Major performance: Highest accuracy was Criminal Justice (99.80%), lowest was Business Administration (92.96%).
  • Level performance: Undergraduate essays (98.92%) slightly outperformed graduate essays (98.05%).
  • Word count buckets: Accuracy was consistent across 400-word (98.32%), 800-word (98.59%), and 1,200-word (98.56%) essays.
  • Dataset: 170 essays generated by Claude-Fable-5, totaling 131,451 words (excluding references).
  • Domains: Business, Education, Humanities, Psychology, STEM.
  • Majors: 30 distinct majors, including Accounting, Biology, Computer Science, Nursing, and Psychology.
  • Levels: 80 graduate essays and 90 undergraduate essays.
  • Word count buckets: 60 essays of 400 words, 60 of 800 words, and 50 of 1,200 words.
  • Detection tool: Turnitin AI detection, measuring word-level accuracy (correctly flagged words as AI-generated vs. total words).
  • Scoring unit: accuracy is measured in words. First identify correct paragraphs under the expectation below, then sum the words in those paragraphs for the correct-word count; total words are summed over all counted paragraphs.
  • Correct paragraph: a counted paragraph that was flagged as AI. Correct words are the words in those paragraphs.
  • References excluded: from a References/Bibliography-style heading onward, those paragraphs are left out of word counts.
  • Limited domains: While five domains were covered, some majors had only 5-6 essays, reducing statistical power.

Overall Performance

  • Total articles: 170
  • Correct words: 129,515
  • Total words: 131,451
  • Word accuracy: 98.53%

Performance by Domain

  • Business: 97.45% accuracy (22,229 correct / 22,811 words)
  • Education: 98.77% accuracy (4,659 correct / 4,717 words)
  • Humanities: 98.53% accuracy (34,139 correct / 34,650 words)
  • Psychology: 98.83% accuracy (3,451 correct / 3,492 words)
  • STEM: 98.87% accuracy (65,037 correct / 65,781 words)

Performance by Major

  • Highest accuracy: Criminal Justice (99.80%), Pre-Med (99.75%), Communications (99.58%)
  • Lowest accuracy: Business Administration (92.96%), History (93.96%), Biology (96.97%)
  • All other majors ranged from 98.00% to 99.54% accuracy.

Performance by Academic Level

  • Graduate: 98.05% accuracy (58,388 correct / 59,551 words)
  • Undergraduate: 98.92% accuracy (71,127 correct / 71,900 words)

Performance by Word Count

  • 400 words: 98.32% accuracy (24,174 correct / 24,586 words)
  • 800 words: 98.59% accuracy (46,827 correct / 47,498 words)
  • 1,200 words: 98.56% accuracy (58,514 correct / 59,367 words)

Detection report screenshots