The research, conducted by Sydney Sears and Dr. Deena Skolnick Weisberg at Villanova University, tested over 2,500 adults aged 18 to 81 recruited via Prolific. It used three original short stories by published human authors and three corresponding stories generated by ChatGPT, all matched for genre and narrative elements .
The study was structured in three parts, each designed to isolate different aspects of reader perception:
Study 1 (1,682 participants): Each person read a single story and was told—either correctly or incorrectly—whether it was written by a human or AI. They then rated it for quality and narrative absorption .
Study 2 (424 participants): Participants read one human-written and one AI-generated story without being told which was which. They then had to guess the author of each .
Study 3 (481 participants): A similar blind-reading design, again requiring participants to identify whether each story came from a human or ChatGPT .
The results were striking and consistent across the experiments:
Detection failure. In Study 2, correct identification was just 39.93%—worse than random chance. In Study 3, it was 51.97%, statistically no different from chance .
AI stories rated higher. Across all studies, AI-generated stories received higher ratings for both quality and how absorbing readers found them .
A 'pro-human' attribution bias. The highest-rated stories were those participants were told were written by humans—even when they were actually AI-generated . The researchers argue this reveals a deep-seated bias toward narratives perceived as authentically human .
AI literacy helps; literary expertise does not. Each one-point increase in self-reported AI expertise was associated with a 14% increase in the odds of correct identification. Each point on a validated AI Literacy Scale (AILS) corresponded to a 33% increase in detection odds .
Dr. Weisberg noted why AI writing may appear to outperform human work on these metrics: AI writing "tends to be clearer, more direct and easier to process," while human-written stories are "often more subtle and complex." She added that "people generally prefer predictability, because difficult or subtle material requires more brain power" .
The study builds on a growing body of evidence about how AI-generated text is perceived. AI-written prose tends to avoid ambiguity, follows predictable narrative structures, and delivers resolutions efficiently—all qualities that make reading feel effortless. Human fiction, by contrast, often embraces complexity, subtlety, and unresolved tension, which can feel less immediately satisfying even when it is artistically richer .
The 'pro-human' bias uncovered in the study is particularly revealing. Readers penalized stories they believed were AI-written, even when the content was identical to stories they praised under a 'human' label. This attribution bias creates a paradox: students who transparently disclose AI use may face a credibility penalty even when their work is high quality, while those who hide AI use may benefit from readers' preference for perceived human authorship .
The study's findings pose fundamental challenges for universities and their academic integrity policies:
Detection is unreliable. Since even trained readers perform near chance at distinguishing AI from human writing, traditional plagiarism detection tools and instructor judgment alone cannot reliably identify AI-authored submissions .
AI literacy is the most effective countermeasure. The study found that familiarity with AI systems improved detection ability, suggesting that universities should invest in AI literacy training for both students and faculty rather than relying solely on enforcement .
Assessment design needs rethinking. If AI can produce writing that readers judge as equal or superior to human work, take-home essays and standard writing assignments may no longer be valid measures of individual student ability. Universities must reconsider what skills they are actually assessing and how .
Academic integrity policies must evolve. The study suggests that a more nuanced approach is needed—one that defines acceptable versus unacceptable AI use, emphasizes process and reflection over final product, and incorporates AI literacy into curricula .
Dr. Weisberg stressed that the findings do not make human authors obsolete: "Humans write as a form of creative self-expression or to challenge ourselves, or to make sense of our experiences. The fact that AI can generate human-like stories doesn't change any of that" .