Sadly enough, not very. Here’s a look at how these detectors work and which are the best ones.
We have come a long way from identifying whether the viral dress was blue and black or white and gold. Today, there’s a new battle underway: distinguishing between human and machine-generated content. This question is much trickier to answer, seeing as the machine—in this case, artificial intelligence (AI)—is trained on human input.
The rise of AI-generated content has made life much more difficult for academicians, journalists and others, sparking concerns about cheating, fake news and fabricated web reviews. Researchers at Cornell University found that, about 66% of the time, people found phony news articles generated by GPT-2 credible.
Consequently, many startups cropped up to detect artificial intelligence (AI)-generated content, be it text, video or audio. With the generative AI market expected to be valued at over US$109 billion by 2030, it has spotlighted developing solutions that combat plagiarism and deepfakes.
AI content detectors—an antidote to AI-generated content?
AI content detectors were born from concerns surrounding the prevalence of AI-generated content. These tools analyze text for features like fluency, frequency of certain words, punctuation patterns, sentence length and more. Google Brain’s Senior Research Scientist, Daphne Ippolito, expounds, “If you have enough text, a really easy cue is the word ‘the’ occurs too many times.”
While the detection method might appear straightforward, AI content detectors tend to suffer from high false-positive rates, erroneously labeling human-written content as AI-generated. Plus, it is easy to evade such detectors by paraphrasing content to avoid repetition of words or sentence structures. As a result, research has found that AI detectors are not yet reliable in practical scenarios.
Which AI content detector is the best?
Considering the mixed success rate of AI detectors, how do you determine which to use? Here are the most popular ones and how they fare:
GPTZero
Created by 22-year-old Princeton student Edward Tian during his Christmas break, GPTZero is arguably the most promising AI content detector. Tian developed the software to help educators and journalists fight the battle of AI-related plagiarism and fake news. This tool distinguishes AI authorship by considering two key factors: perplexity and burstiness. Perplexity measures how complex a text is, and burstiness refers to the variation between sentences. Lower values for these two factors suggest a greater likelihood that AI created the text.
When we tested it in March 2023, GPTZero inaccurately identified a piece of content wholly written by ChatGPT as human-written. Almost eight months later, the software appears to have improved (albeit not reaching 100% accuracy). Upon inputting the same ChatGPT-curated content into the tool again, GPTZero said there was an 86% chance that it was entirely AI-generated (see below).
Turnitin
In February 2023, Turnitin unveiled an AI writing detector that it claimed could identify up to 97% of content authored by ChatGPT and GPT-3. Plus, it said that the detection had a low false positive rate of less than 1/100. However, a few months later, the company admitted that its detector software might have a “reliability” issue.
The Washington Post conducted an in-depth investigation to find out if Turnitin’s AI content detector is as accurate as it says it is. Turns out, it is not. The study revealed that Turnitin’s AI detector software often errs, mistakenly flagging essays composed entirely by humans as AI-generated.
Copyleaks
In 2022, Copyleaks secured US$7.75 million in funding to enhance its anti-plagiarism offerings for schools and universities, with a special focus on detecting AI content within student submissions. When we tried it out with the same ChatGPT-curated content we used for GPTZero, Copyleaks failed to identify it as AI-generated. It confidently surmised that the whole thing was human-generated (see below).
While its success rate in identifying content created by GPT-3.5 is rather limited, Copyleaks demonstrates 93% sensitivity when dealing with GPT-4 generated content. Sensitivity, in this context, refers to the tool’s ability to accurately identify AI-generated content as such. Conversely, GPTZero shows less proficiency in handling GPT-4 content, with a sensitivity rate of only 27%.
Special mention: FakeCatcher
Intel’s FakeCatcher claims to be able to identify deep fake videos with 96% accuracy, partly by analyzing pixels for subtle signs of blood flow in human faces. However, when BBC tested the tool on different videos—real and fake—the tool struggled to detect which was which. Since it doesn’t analyze audio and cannot work with super-pixelated videos, it is difficult to employ this software in real-world scenarios.
Will AI content detectors ever be perfect?
It is unlikely—although it does appear that GPTZero is improving. As AI detectors evolve, so too do AI content generators. For example, ChatGPT is constantly upgrading, Bard will move beyond its testing phase, and new, more sophisticated generators will crop up. In the race between generators and detectors, the odds are that the latter might not catch up. But that doesn’t mean you should surrender and give up. You might just need to get creative.
For instance, academicians have found ways in which AI content generators can become a learning tool instead of a hindrance. Teachers can use it to create syllabi and generate more interesting educational content, and it can help students develop layouts for their assignments and presentations.
AI content detectors leave much to be desired. But, soon enough, another college student, frustrated with AI content, may develop the antidote.
Also read:
- Is Using AI for Academic Writing Cheating?
- The AI Language Model Avengers: Meet the Top Emerging AI Heroes Reshaping Communication
- AI Detector for Educators: What is GPTZero?
Header Image by Freepik





