Readers Prefer ChatGPT Stories but Don't Know It
OpenAI
A study by Villanova University researchers found that readers rated short stories generated by ChatGPT higher than human-written ones, yet they also rated stories labeled as human-written more favorably regardless of true authorship. Participants often failed to distinguish between the two, highlighting the influence of perceived authorship on judgment.
In a study published in Judgment and Decision Making, researchers Sydney Sears and Deena Skolnick Weisberg from Villanova University investigated how readers evaluate short stories based on actual content and perceived authorship. The first experiment involved 1,682 American adults recruited via Prolific, who each read a story of about 1,000 words. Half of the stories were written by humans and previously published, while the other half were generated by GPT-4, with each AI story matching a human one in theme and elements. Participants were told the author was either a human or ChatGPT, but this information was accurate only half the time. The results showed that AI-generated stories received higher ratings for absorption (1.42 vs. 1.00) and perceived quality (1.54 vs. 0.97) compared to human stories. However, regardless of the true origin, stories labeled as human were consistently rated more favorably, and this effect was stronger among participants with negative views of AI. In two further experiments with 905 participants who had to identify the origin of paired stories, only 40% and 52% respectively succeeded, showing that readers could not reliably tell AI from human text. Self-reported literary expertise did not help, though familiarity with AI tools gave a slight advantage. The researchers suggest that AI-generated texts might be more fluid and easier to comprehend, which could explain their better evaluations. However, the study is limited to three pairs of short realistic stories and an online American sample, and does not imply that AI writes better in general. The results underscore that expertise does not ensure accurate discrimination and that perceived authorship significantly influences text reception.
Source: Numerama —
original
