DetectGPT: a new tool to classify AI generated text | Generative ai Google | Microsoft Generative ai Tools | Best Free Generative ai Tools | Turtles AI
DetectGPT: a new tool to classify AI generated text
DukeRem18 February 2023
In recent months, researchers have been exploring new ways to detect whether a given text was generated by a language model such as ChatGPT. A new tool, DetectGPT, shows a 95% accuracy in this task.
A graduate student, Alexander Khazatsky, posed a question to his colleague, Eric Anthony Mitchell, on whether there was a way to classify an essay as being written by ChatGPT. This question led Mitchell to delve deeper into the matter.
Previous approaches have included training a model using both human- and LLM-generated text and then asking it to classify whether another text was written by a human or an LLM. However, this approach may require a vast amount of data for training in order to be successful across various subject areas and languages.
A second existing approach involves using the LLM that likely generated the text to detect its own outputs. This approach asks an LLM how much it “likes” a text sample, and this liking of a piece of text is a shorthand way to say “scores highly.” It involves a single number, which is the probability of that specific sequence of words appearing together according to the model. Mitchell explains that if the model likes a text sample a lot, it is likely to be generated by the model, and if it doesn't like it, it's probably not from the model. This approach works reasonably well and is much better than random guessing.
However, Mitchell had an initial intuition that even powerful LLMs have subtle, arbitrary biases for using one phrasing of an idea over another. Therefore, the LLM will tend to “like” any slight rephrasing of its own outputs less than the original. By contrast, even when an LLM “likes” a piece of human-generated text, the model's evaluation of slightly modified versions of that text would be much more varied. In other words, if a human-generated text is perturbed, it is roughly equally likely that the model will like it more or less than the original.
To test his intuition, Mitchell used popular open-source models, including those available through OpenAI’s API. He realized that calculating how much a model likes a particular piece of text is essentially how these models are trained. These models give a number automatically, which turns out to be really useful in detecting machine-generated text. This approach, called DetectGPT, involves measuring the curvature of the probability function of the model as a function of the amount of perturbation to the original text. The results were impressive, with the method achieving an F1 score of 94.9% for English text and 92.6% for German text.
