Think about how absurd this situation is. Someone uploads a paper written long before generative AI was around, only to get hit with an AI flag. Right next to the result is a paid button offering to clean up the writing. Panic created, money made.
This is a standard experience for many. A 2026 University of Florida study tested five commercial AI detectors and found false positive rates ranging from 0.05% to 68.6%. Research from Stanford found that non-native English speakers face substantially higher false positive rates across major platforms. A PNAS Nexus study found more than half of all TOEFL essays were wrongly flagged as AI-generated across every detector tested.
These platforms regularly produce untrustworthy findings, and the companies selling them know it.
The Conflict Of Interest Built Into The Product
What makes this scenario so telling is the sequence that follows the flag itself.
Many of the companies selling AI detection tools also sell “humanisation” products, sometimes called AI removers or rewriters, designed to make flagged text score as human-written. The conflict of interest is clear, and built into the product architecture. A detector that flags aggressively creates more anxiety, more urgency and more conversions for the paid tool sitting one click away.
Turnitin has publicly claimed a false positive rate of under 1%. Independent testing suggests otherwise, with reported rates as high as 50% in some samples. The UF study found the range across commercial detectors so wide as to make the tools functionally unreliable for anything beyond a rough indication. One detector in the study produced a false negative rate of 99.6%. Another produced a false positive rate of 68.6%. These aren’t calibration issues; they’re numbers that would disqualify a tool from use in any other evidential context.
The question this raises is whether the detection business model is designed to fail in order to sell the fix. We asked a range of voices across education, technology and AI ethics to weigh in.
More from Artificial Intelligence
- Are AI Detectors Doing More Harm Than Good In Universities?
- Claude Users Found Their Private Chats Online – What Does This Say About AI And Privacy?
- Can AI Save The Insurance Industry From The Wildfire Crisis?
- Why Are OpenAI and Anthropic Secretly Lobbying Against Open Source?
- Andy Burnham Launches New AI Taskforce To Drive Economic Growth
- Meta Has Pledged To Be Net Zero By 2030. So Why Is It Walking Away From RE100?
- New Data Reveals How Much AI Is Costing UK Business Operations
- SaaS Founders Are Panicking About AI But A $21 Billion Bet Says They’re Wrong
Our Experts
- Martin Harris, Head of Digital, Tank
- Dr. Morissa Schwartz, AI Evaluator, Educator and Founder, GenZ Publishing
- Michelle Edge, Partner, Eleven Hundred Agency
- Harpal Singh, AI SEO and GEO Consultant and Founder, Blimpp
Martin Harris, Head of Digital, Tank

“Fifteen years of assessing marketing tech has taught me to start with one question: who profits from this tool’s verdict? And with AI detectors, the answer is uncomfortable. The same company that flags your writing will happily sell you the fix. If your humaniser revenue depends on people getting flagged, why would you ever work hard to reduce false positives? You don’t need a conspiracy theory here. It’s just a business model, doing what business models do.
“There’s a more basic problem underneath. These tools don’t actually detect AI. They detect predictable writing. Clear, well-structured prose scores badly on them, which is absurd when you think about it, and it’s why Stanford researchers found the majority of essays by non-native English speakers were falsely flagged. The detectors are also losing the arms race, because language models improve faster than the tools chasing them. When universities like Yale quietly dropped these detectors, that wasn’t caution. It was an admission that the evidence was never there.
“The bit I find strange is the misplaced anxiety. Brands are fretting over whether their content reads as AI while ignoring the question that actually affects revenue in 2026: can ChatGPT, Perplexity and Google’s AI Overviews find your content, understand it and cite it? A detection score has never made anyone a penny. Don’t use these tools as evidence for decisions about people. And whatever you do, don’t pay the company that flagged you to make the flag go away.”
Dr. Morissa Schwartz, AI Evaluator, Educator and Founder, GenZ Publishing

“AI detectors should be treated as weak screening signals, not verdicts. They estimate statistical patterns in text; they do not prove authorship. Polished human prose, formulaic academic writing, heavily edited copy and work by multilingual writers can all be falsely flagged.
“A company that sells both detection and humanisation has an obvious conflict-of-interest risk, but that structure alone doesn’t prove intentional over-flagging. The right questions are whether the company publishes its thresholds, false-positive rates, validation datasets and commercial incentives, and whether those claims have been independently audited.
“In my work evaluating AI-generated and human-authored writing, the reliable approach is evidence triangulation: review version history, drafts, citations, source notes and the writer’s ability to explain their reasoning. A student or employee should never be penalised based on one detector score. If a vendor markets detection as certainty while also selling the cure, its product design deserves scrutiny.”
Michelle Edge, Partner, Eleven Hundred Agency

“It’s certainly fair to question the motivations behind these tools. It’s a bit like a mechanic that diagnoses a fault with your car and immediately offers to fix it. It doesn’t necessarily mean the diagnosis is wrong, but it does mean a certain amount of scepticism is warranted.
“The bigger issue is that the jury is still out on whether these tools are even accurate. If you take a piece of well-edited human writing and put it through a few AI detectors, there’s a good chance at least one will flag it. Good human writing often shares some of the patterns these tools associate with AI-generated text, such as clear, consistent structure and precise grammar.
“But if writers start to worry that clarity and polish will get them flagged as machines, we risk incentivising uneven or unedited writing just to pass the test. As AI keeps evolving, hunting for the right list of tells is a losing battle. We should be looking for ways to identify original thinking and real human insight instead.”
Harpal Singh, AI SEO and GEO Consultant and Founder, Blimpp

“It is not that all AI detectors intentionally exaggerate scores. It is most concerning that the business models are based on anxiety over issues like a lack of accuracy. When faced with the problem, users are offered a paid ‘humaniser’ option with a false definitive warning by the AI-detection systems. Urgency leads to a problem-based design; false positives create urgency. While there may not be intentional malice, the design, warning systems and scoring systems are most likely unreasonably aggressive.
“AI detection systems should never be viewed as proof of violation or authorship. As a minimum, all significant cases must be reviewed and the writer engaged personally. An analysis of version history and notes must be included.
“For detection systems, the initial question is not ‘how accurate is your detection system?’ It must also include: ‘do you conduct independent evaluations of false positive rates and disclose the results?'”
For any questions, comments or features, please contact us directly.
![]()
