Learn how Turnitin’s AI detection analyzes perplexity and burstiness, what a flagged score really means, and practical steps to take if your paper is incorrectly flagged.
How Does Turnitin Tell AI Writing From Human Writing?
If you have ever stared at a Turnitin report wondering how a machine decided your essay “sounds like AI,” you are not overthinking it. The system is genuinely doing something different from a normal plagiarism check, and it is worth understanding exactly what that something is. Turnitin does not compare your essay to a database of known AI outputs. Instead, it studies the statistical fingerprints of how your sentences are built and guesses, sentence by sentence, whether a language model probably generated them.
That distinction matters for students, for teachers, and for anyone trying to make sense of a confusing percentage on a report. Let’s walk through the mechanics, the limits, and what you should actually do if your name ends up next to a high AI score.
Similarity Checking and AI Detection Are Two Separate Systems
A lot of people assume Turnitin’s “AI score” and its plagiarism score are the same thing. They are not. The Similarity Report compares your submission against a database of web pages, journal articles, and previously submitted papers, then flags matching text. The AI writing indicator runs independently and asks a completely different question: does this sentence read like something a language model produced, regardless of whether it matches any existing source. Turnitin’s own FAQ page notes that its AI writing detection covers long-form English, Spanish, and Japanese submissions, and the two reports live in separate parts of the dashboard.
So a paper can score low on similarity and still get flagged for AI writing, or the reverse. They measure different things, and your instructor is meant to look at both.
Perplexity and Burstiness: The Core of the Model
Underneath the interface, Turnitin’s detector leans on two statistical properties that show up across most AI detection tools, not just Turnitin’s.
Perplexity measures how predictable your word choices are. Large language models generate text by picking the most statistically likely next word at each step. That habit produces prose with low, consistent perplexity: smooth, expected, rarely surprising. Human writers wander more. You might reach for an odd word, restructure a sentence halfway through, or make a choice a probability model would not have picked. That unpredictability shows up as higher, more uneven perplexity.
Burstiness looks at variation in sentence length and rhythm across a document. Real writers are bursty. You’ll write a long, layered sentence and then follow it with something short and blunt. AI-generated text tends to even itself out, producing paragraphs where every sentence sits in a similar length and complexity range.
Turnitin does not read your paper as one block. It breaks the submission into smaller segments, scores each one, and aggregates those scores into the overall percentage you see at the top of the report. Each sentence gets its own likelihood estimate before anything gets combined.
What the Percentage Actually Represents
Here’s where a lot of confusion starts. The percentage on your report is not “20% of your paper was plagiarized” or even “20% chance this is AI.” It reflects the amount of qualifying text within the submission that Turnitin’s detector has identified as likely AI-generated, where qualifying text means the prose sentences inside a long-form document like an essay or dissertation. Reference lists, bibliographies, and non-prose content get excluded from that calculation.
Turnitin also built in a buffer for uncertainty. Scores between roughly 1% and 20% show up as an asterisk instead of a hard number, specifically to reduce the chance that low-confidence detections get treated as reliable findings. That asterisk is a built-in admission that the model isn’t fully confident at low percentages, and it is worth pointing out to an instructor if your report shows one.
Since 2024, the report also separates two categories of flagged text. One category covers text likely generated by a large language model that may have also been run through an AI bypasser tool, while a second category covers AI-generated text that appears to have been run through a paraphrasing tool such as a word spinner. In practice, this means Turnitin has specifically trained its model to notice when AI text has been reworded to sound more human, and that detection layer keeps getting refined with each update.
Where the Model Gets It Wrong
No detector is airtight, and Turnitin has never claimed otherwise. Turnitin’s own documentation states plainly that false positives, meaning human-written text incorrectly flagged as AI, are a possibility with any AI detection model.
A few writing styles are more prone to this than others:
- Formulaic or highly structured writing. Lab reports, legal definitions, and methods sections in scientific papers follow rigid templates by design. That rigidity looks statistically similar to AI-generated prose, even when a person wrote every word.
- Non-native English writing. Writers working in a second or third language tend to lean on common vocabulary and simpler sentence structures because those are the words they know most confidently. Detectors read that as low perplexity, the same signal they associate with AI text.
- Heavily edited or paraphrased human drafts. If you write cleanly and consistently, without much stylistic variation, your natural voice can occasionally resemble the smoothness a model produces.
Because of this, Turnitin repeatedly tells institutions not to treat the AI score as a verdict on its own. A flag is meant to start a conversation, not end one. Several major universities, including some UC campuses and other large institutions, have restricted or turned off the AI detection layer entirely over the past year, citing exactly these reliability concerns while keeping traditional similarity checking active.
What a High Score Actually Means for a Student
If your paper gets flagged, panic is the least useful reaction. Here is a more practical sequence:
- Read the report itself, not just the number. Open the highlighted sentences and see where the flags cluster. Scattered flags across formulaic sections tell a different story than a flag covering your entire argument.
- Pull together your process evidence. Draft versions, research notes, and any documented writing history give you something concrete to point to.
- Talk to your instructor early. Explain your process and, if relevant, why your writing style might resemble the patterns the detector is trained to catch.
- Ask for a look at your submission history. If your flagged paper matches the tone and structure of your earlier work, that consistency is meaningful evidence on its own.
Turnitin Report Delays and File Issues
Separately from AI scoring, a lot of students run into a report that just won’t load or process. This usually comes down to file formatting, file size, or a temporary processing backlog rather than anything about your writing. If your report gets stuck in a loading state or throws an error, Turnitin’s own support documentation walks through the common causes and fixes.
Where This Is Headed
Detection and generation are locked in a constant back and forth. Every time detectors get better at spotting AI patterns, generation tools adjust their output to look more human, and the cycle continues. Turnitin has already added detection for text that has passed through AI paraphrasing and bypasser tools, and that layer keeps improving with each model update. None of that changes the underlying advice, though: write in your own voice, keep evidence of your process, and treat a flagged report as a starting point for a conversation rather than a final judgment.
Frequently Asked Questions
Can Turnitin tell which AI tool wrote a paper, like ChatGPT versus Gemini? No. It estimates the likelihood that a segment of text was AI-generated based on writing patterns. It does not identify a specific tool or model.
Does a low AI score guarantee a paper is fully human-written? Not entirely. Detection focuses on prose sentences in long-form writing and can miss AI-assisted content in tables, code, or heavily edited passages. A low score is a good sign, not an absolute guarantee.
Is the AI score the same as the plagiarism or similarity score? No, they are separate. Similarity checking looks for matching text against existing sources. AI detection analyzes writing style and statistical patterns, independent of whether the text matches anything else.
Why does my report show an asterisk instead of a percentage? Scores in the low range, generally under 20%, are shown with an asterisk rather than a number. This signals reduced confidence rather than a hard result.
What should I do if I believe my paper was flagged incorrectly? Contact your instructor, share draft history or research notes, and ask for a review of the specific flagged sentences rather than just the overall percentage.
Beat Plagiarism & AI Detectors with Confidence
Worried about plagiarism scores or AI detection flags? Get Plagiarism reports, expert humanization, and original papers written from scratch — all in one place and before your deadline.
100% Confidential • Turnaround as fast as 3 hours • Money-back guarantee

