AI detectors fail the cheating test they were built to pass, and students know it

The market for AI detection software rests on a promise: that an algorithm can tell whether a student wrote an essay or a machine did. A hands-on test published by CNET this week suggests that promise has collapsed, with every major detector defeated by free and low-cost workarounds and detectors contradicting one another on the same images.

CNET contributor Rachel Kane, a journalism professor at Arizona State University, fed real PDF assignments into ChatGPT, Claude, Gemini, and NotebookLM, then ran the resulting essays through AI-humanizing tools. Every humanized text passed detection with flying colors, according to the piece. The test also found the industry has a structural conflict: some companies that sell AI detection also sell the humanizing tools designed to defeat it.

The contradictions were starkest in image detection. One tool flagged a graphic made before generative AI existed as 100 percent AI-generated, while another judged a fully AI-created infographic to be only 1 percent AI. Students facing accusations can exploit such disagreements, the article notes, since two detectors will often produce opposite verdicts on identical work.

Even assignments designed to be “AI-proof” failed the test. Given a single prompt to complete an assignment requiring visual elements, Claude produced professional-looking infographics in a standardized format without any further instruction. The author never reviewed the assignment or the output, undercutting the argument that students absorb material when they delegate work to models.

Support independent reporting built on evidence, transparency, and scientific rigor.

Support independent reporting

Independent research backs the anecdotal findings. A systematic evaluation published in the International Journal for Educational Integrity found widely used commercial detectors performed inconsistently on GPT-4-generated content, producing both false negatives and uncertain classifications. Brandeis University’s AI steering council has compiled studies showing detection tools carry high false-positive rates and can be biased against non-native English speakers, and reporting since 2023 has documented detectors flagging human-written work as machine-generated.

The economics of the ecosystem point in one direction. Homework-helper apps that complete multi-page assignments on demand are sold as study aids, with subscription pricing from about $7 to $12 per month; exam-specific tools advertise undetectable experiences at $15 to $192 per year. The CNET piece concludes the primary customer of the entire industry is the student who wants to get the work over with, not the educator trying to enforce integrity.

Universities have begun responding structurally rather than technologically. Institutions are moving toward question randomization from larger pools and more in-person testing, recognizing that software arms races favor whoever controls the writing tool, not whoever controls the checker. For now, the practical takeaway for educators is that a detector score alone is not evidence, and for students, that the enforcement machinery is easier to defeat than it is to trust.

Sources: I Used Every AI Cheating App and Detector and Came to One Conclusion (CNET, July 31, 2026); Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text (International Journal for Educational Integrity, Springer); Limitations of AI Detection Tools (Brandeis University AI Steering Council); Detecting AI may be impossible. That’s a big problem for teachers (Washington Post, June 2023)

Scroll to Top