Universities are relying on AI-detection software to catch cheating. How well do the programs work?
- CAREER FEATURE
- Anna McKie
Anna McKie is a freelance journalist based in London.
Search author on: PubMed Google Scholar
- [Bluesky](https://bsky.app/intent/compose?text=Universities+are+relying+on+AI-detection+software+to+catch+cheating.+How+well+do+the+programs+work%3F https%3A%2F%2Fwww.nature.com%2Farticles%2Fd41586-026-01358-2)
- [Whatsapp](https://wa.me/?text=Universities+are+relying+on+AI-detection+software+to+catch+cheating.+How+well+do+the+programs+work%3F https%3A%2F%2Fwww.nature.com%2Farticles%2Fd41586-026-01358-2)
- X
Some AI-detection tools used to analyse students’ work have been flagged as untrustworthy. Credit: Paul Guzzo/Getty
Last November, Lauren Jager, a chemistry undergraduate student at Idaho State University in Pocatello, was applying to PhD programmes when she noticed that some application portals warned students about using generative artificial-intelligence tools for their personal statements. They informed students that they would use detectors to sniff out applications that contained AI-generated text. The portals weren’t specific about which detectors they were using. But they were clear on one thing: “They said that if they felt that the personal statement had been written with AI, then they would disregard your entire application,” Jager says.
She didn’t think much of it — she hadn’t used AI at all — but a friend said they’d run their own statements through an AI detector on the Internet, just for safety. Jager decided to do the same with a few detectors she’d found online.
“They all came back at almost 100% AI,” she says. “I started freaking out.”
AI writing tools could hand scientists the ‘gift of time’
Looking back, Jager wonders why her essays were flagged. “I’m a chemistry major, but I was almost an English major. I grew up writing a lot and I used to study grammar books,” she says. “I think of myself as a good writer, but I write very much to the book, following all the rules specifically. Maybe that’s why it thought it had been written by AI.”
She ended up rewriting her statement entirely. And, instead of trying to write to the best of her ability, she says she wrote in a way to make sure it wasn’t flagged by the AI detector. “I was making it less perfect,” Jager recalls. When she ran her essay through the checkers again, they predicted that it was 30% AI-written. “I called that good enough and sent it in.”
When we last spoke to Jager, she had been accepted to do her PhD at the University of Utah in Salt Lake City.
Jager’s situation and cases like it are happening around the world. Universities that are struggling with a surge of written material that might have been generated, or heavily shaped, by AI are turning to AI-detector tools to help them grapple with the problem. And, just like other AI technology, such tools are far from perfect.
Educators and copycats have been trying to outsmart each other since the first Sumerian student copied their classmate’s cuneiform. In the modern Internet era, software provided by companies such as Turnitin, based in Oakland, California, detects text similarities on the basis of a vast corpus of previously published work. This has made it much harder for people to plagiarize text directly. In response, cheaters went online to find ghostwriters, paying others to produce their work for them. Many institutions tried in-person or timed exams to prevent this. But these come with their own set of issues, potentially disadvantaging certain groups, encouraging rote memorization and preventing students from demonstrating their ability to conduct deep research.
Shadow scholars: inside Kenya’s multibillion-dollar fake-essay industry
But large language models (LLMs) that power chatbots, such as ChatGPT, have made it even cheaper and faster to produce written work. Many of those interviewed for this article noted the irony that even though generative AI is built on a corpus of previous work, the writing it produces isn’t easily detected by plagiarism tools, which compare full sentences with previously published work.
Cath Ellis, an academic integrity professional at Western Sydney University in Australia, says AI use represents a fundamental change, both in the scale of its use and how to understand the written word. “Up until now, we’ve been mostly able to rely on a written document,” Ellis says. But “we’re now starting to see this massive volume of fraudulent or at least heavily fabricated processes,” she says. “The volume of stuff coming through has just rocketed.”
Fighting fire with fire
A growing number of companies say that the tools they have made can identify text that has been written by another AI system. On-the-market products include Copyleaks, GPTZero and ZeroGPT, and those created by Grammarly, QuillBot and Turnitin.
Many AI detectors rely on a measure known as perplexity, which estimates how predictable each word in a sequence is likely to be. Because AI-generated text tends to follow more statistically predictable patterns than does human writing, passages with lower perplexity scores are more likely to be flagged as machine-generated, whereas less predictable phrasing is taken as a signal of human authorship (see ‘The telltale signs of AI’).
Source: M. Suvanto et al. Preprint at arXiv https://doi.org/rdhv (2026)
The question is: do these tools actually work? And should they be used at all if there is any chance of unjustly accusing a student, like Jager, of cheating?
Clear evidence of AI detectors’ difficulties in correctly assessing human-written text was highlighted by several users on the social-media platform Reddit — they discovered that the US Declaration of Independence is often flagged as AI-written. Nature ran part of the 1776 text through ZeroGPT a number of times and was told it was between 95% and 100% AI-generated.
Finding your academic voice: faster, healthier writing with AI speech recognition
Other detectors might show more promise. Pangram Labs, based in New York City, says that its tool has a near-zero false-positive rate. Its approach involves training a model on a large corpus of human-written, then AI-rewritten, text. The model thus comes to understand how each newly released chatbot writes. So far, this approach, which avoids using perplexity as the only direct measure, has stood up to scrutiny: independent assessments have judged the product to be among the most accurate available.
Mike Perkins, who researches the impact of AI on academia at the British University Vietnam in Hanoi, says even when detectors perform reasonably well in controlled tests, their results should not be used as evidence in high-stakes decisions. “The short answer is no, they don’t [work reliably],” he says. “The long answer is yes, they can work — but the fact that there are so many concerns about false positives means they shouldn’t really be used when it comes to anything that’s sensitive for a student.” Otherwise, students like Jager can be caught in the net.
Part of the problem, argues Perkins, is that teachers have grown accustomed to accepting automated scores from plagiarism software. “People see a score and trust it,” he says. “Similarity tools worked because they could show you exactly where text matched something else. With AI-detection tools, that evidence just isn’t there.”
Evading the detectors
Hybrid texts are another problem for those hoping to catch cheaters. “Work that I, and others, have done shows that if you just test a piece of AI text against a detector, it’s going to be pretty good at identifying it,” says Perkins. “But if you start to manipulate that text in different ways, then the detection really starts to break down.”
People don’t even need to edit passages themselves: they could ask another AI system to rewrite them, or run the text through ‘humanizer’ tools that are designed to lower AI-detection scores. Detection companies are now trying to identify the use of such tools, adds Perkins, but new systems that evade detectors are emerging quickly. “It becomes a huge arms race that doesn’t really help anyone,” he says.
AI FOMO: everyone is mastering AI except me — or are they?
But Karpinska cautions that although this information might be useful for assessing how much content is probably generated by AI on a large scale, it does not mean that the results from Pangram — or from other tools trying to catch up to its abilities — should be taken at face value in individual cases. In other words, it can reveal trends, but not the guilt of any one author. “We certainly cannot mass-reject people because of it,” she says.
The issue of bias
Expert-level test is a head-scratcher for AI
Enjoying our latest content?
Log in or create an account to continue
- Access the most recent journalism from Nature's award-winning team
- Explore the latest features & opinion covering groundbreaking research
doi: https://doi.org/10.1038/d41586-026-01358-2
References
- Dik, S., Erdem, O. & Dik, M. Preprint at arXiv https://doi.org/10.48550/arXiv.2506.23517 (2025).
- Elkhatat, A. M., Elsaid, K. & Almeer, S. Int. J. Educ. Integr. 19, 17 (2023). Article
- Walters, W. H. Open Inf. Sci. 7, 20220158 (2023). Article
- Russell, J. et al. Preprint at arXiv https://doi.org/10.48550/arXiv.2510.18774 (2026).
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E. & Zou, J. Preprint at arXiv https://doi.org/10.48550/arXiv.2304.02819 (2023).
- Kobak, D., González-Márquez, R., Horvát, E.-A. & Lause, J. Sci. Adv. 11, eadt3813 (2025). Article
- Liang, W. et al. Nature Hum. Behav. 9, 2599–2609 (2025). Article
Related Articles
- Finding your academic voice: faster, healthier writing with AI speech recognition
- AI writing tools could hand scientists the ‘gift of time’
- Research integrity: Don’t let transparency damage science
- This robot can beat you at table tennis
- Expert-level test is a head-scratcher for AI
- Can AI tools assess coding assignments?
- LLMs behaving badly: mistrained AI models quickly go off the rails
Subjects
Latest on
- [‘Humanizer’ tool can erase signs of AI-written text — alarming scientists
News 07 JUL 26](https://www.nature.com/articles/d41586-026-02105-3)
- [AI tools can speed up thinking, but evidence still comes from the lab bench
Correspondence 30 JUN 26](https://www.nature.com/articles/d41586-026-02069-4)
- [AI systems devise hypotheses and ways to test them
News & Views 30 JUN 26](https://www.nature.com/articles/d41586-026-01873-2)
- [What’s behind China’s historically high counts of corresponding authors?
Nature Index 04 JUN 26](https://www.nature.com/articles/d41586-026-01623-4)
- [White House proposes vast overhaul of US science funding: what you need to know
News 03 JUN 26](https://www.nature.com/articles/d41586-026-01779-z)
- [First and last authors more likely to be men in leading science journals
Nature Index 02 JUN 26](https://www.nature.com/articles/d41586-026-01495-8)
- [AI can cause harm: safeguards must catch up
Correspondence 07 JUL 26](https://www.nature.com/articles/d41586-026-02109-z)
- [AI tools can speed up thinking, but evidence still comes from the lab bench
Correspondence 30 JUN 26](https://www.nature.com/articles/d41586-026-02069-4)
- [Trump has big AI and quantum ambitions: this scientist’s job is to make them reality
News 29 JUN 26](https://www.nature.com/articles/d41586-026-02023-4)
Jobs
You will contribute to the success of the BMC Series by supporting editorial handling of content in BMC Sports Science, Medicine and Rehabilitation an New York City, New York (US) Springer Nature Ltd
AI and Data Science for Bioengineering | Translational Pediatric Bioengineering Barcelona (Localidad), Cataluña (ES) Institute for Bioengineering of Catalonia (IBEC)
The University invites individuals from diverse backgrounds to apply for faculty positions in this field Dongguan, Guangdong, China City University of Hong Kong (Dongguan)
The University invites individuals from diverse backgrounds to apply for faculty positions in this field Dongguan, Guangdong, China City University of Hong Kong (Dongguan)
The University invites individuals from diverse backgrounds to apply for faculty positions in this field Dongguan, Guangdong, China City University of Hong Kong (Dongguan)
Related Articles
- Finding your academic voice: faster, healthier writing with AI speech recognition
- AI writing tools could hand scientists the ‘gift of time’
- Research integrity: Don’t let transparency damage science
- This robot can beat you at table tennis
- Expert-level test is a head-scratcher for AI
- Can AI tools assess coding assignments?
- LLMs behaving badly: mistrained AI models quickly go off the rails
Subjects
Sign up to Nature Briefing
An essential round-up of science news, opinion and analysis, delivered to your inbox every weekday.
