AI detectors call 100% human writing “AI-like.” Doug Llewellyn on what these tools actually measure, and why they get it wrong so often.

Why AI Detectors Call Human Writing “AI-Like”

AI writing detectors are being asked to draw a hard line between “human” and “AI” that doesn’t match how most people actually write anymore, and the tools keep proving they can’t draw that line reliably, even when the honest answer is 100% human. When a detector calls a fully human piece of writing “AI-like,” it’s revealing what the tool is actually measuring, and that usually isn’t authorship at all.

That story stuck with Doug Llewellyn, CEO of Data Society, enough that he brought it up unprompted: a writer posted something on LinkedIn she’d written entirely herself, ran it through an AI detector out of curiosity, and got a score of 80% AI. When she pushed back, the only explanation she got was that her writing style was “very AI-like.”

“I did see someone post, I think it was on LinkedIn the other day, that she’d written something 100% herself and put it in, and it said it was 80% AI. And the explanation was, ‘Well, your writing style is very AI-like.’ What does that mean?”

Doug Llewellyn, Chief Executive Officer, Data Society

It’s a fair question, and most people who ask it don’t get a real answer. “AI-like” describes the tool’s own uncertainty about a text, not anything about who actually wrote it.

AI detectors don’t read for meaning, intent, or authorship. They score text against statistical patterns, things like word predictability and sentence-length variation, that tend to differ between large language model output and human writing on average. “Tend to” and “on average” are doing a lot of work in that sentence, and the track record shows it.

A 2023 evaluation by researcher Debora Weber-Wulff and colleagues tested 14 AI detection tools and found every one scored below 80% accuracy, with only five clearing 70%. Turnitin itself reports a false positive rate under 1% at the document level, but when The Washington Post ran its own test on the tool, more than half the sample essays it checked came back with at least one call that was wrong, whether that meant flagging human writing as AI or missing AI writing altogether. Stanford researchers found the problem isn’t evenly distributed either: across seven GPT detectors, non-native English writers’ essays were flagged as AI-generated at an average false positive rate of roughly 61%, far above the rate for native English writers using similar language. A 2024 analysis from Common Sense Media found racial disparities in the same direction, with false positive rates of 20% for Black students, 10% for Latino students, and 7% for white students. Several Russell Group universities, including Cambridge, have declined to turn on Turnitin’s AI-detection feature at all, citing exactly this kind of unreliability.

None of that describes who’s secretly using AI. It describes whose natural writing style happens to score closer to whatever pattern the detector was trained to flag. A “very AI-like” writing style is often just a clear, structured, unsurprising one, the kind plenty of careful human writers already had before any of these tools existed.

The detector’s binary, human or AI, also assumes writing happens in one mode at a time. Llewellyn’s own process is a useful example of why that assumption doesn’t hold up.

“I personally, when I write long emails or texts, I usually write it all myself, but I’m a little less careful in the first draft. It’s all my thinking. I throw it through one of the tools, it comes back, it’s still 85, 90% me, just differently formatted. I make a few edits, so my final work product is somewhere between 80 and 90% me.”

Doug Llewellyn, Chief Executive Officer, Data Society

By his own account, an AI tool touches nearly everything he sends, and the result is still, by his estimate, 80 to 90% his own thinking. That’s what AI-assisted writing looks like for a large and growing share of professional communication: a human draft, a pass through a tool for formatting and polish, a human edit on the way out the door. Somewhere in that process, the question a detector is built to answer, did a person write this, stops making sense as a yes-or-no question. The detector doesn’t know that. It returns a percentage anyway, and lets the reader assume the percentage means what they think it means.

Detector false positives get discussed mostly in academic settings, but the same tools, and the same weaknesses, are showing up in workplaces: content review, brand compliance, even HR conversations about whether an employee’s writing is really their own. A detector score used to gate a performance conversation, a content approval, or a policy decision carries the same double-digit false positive rates the research above describes. Treating that score as settled fact, rather than a noisy signal from a tool with a well-documented bias problem, is how organizations end up penalizing their most careful writers for writing carefully.

This is the same governance gap that shows up anywhere AI tools get adopted faster than anyone builds real understanding of what they’re doing. An agent nobody’s tracking and a detector score nobody’s questioning are two versions of the same failure: treating an AI tool’s output as a decision instead of an input that still needs a person’s judgment.

Whether a detector can catch AI writing is, increasingly, the wrong question. The more useful one is whether “AI-generated” versus “human-written” is still a distinction worth building policy around at all. Data Society’s AI upskilling programs help teams build the judgment to evaluate AI-assisted work fairly, and AI Advisory services help organizations set policies that don’t rely on unreliable detector scores.

Don’t wanna miss any Data Society Resources?

Stay informed with Data Society Resources—get the latest news, blogs, press releases, thought leadership, and case studies delivered straight to your inbox.

Data: Resources

Get the latest updates on AI, data science, and our industry insights. From expert press releases, Blogs, News & Thought leadership. Find everything in one place.

View All Resources
  • Why AI Detectors Call Human Writing “AI-Like”

    September 28, 2026

    Read more

  • Cost Per Outcome vs. Cost Per Token: How to Actually Measure AI ROI

    September 25, 2026

    Read more