1.54 versus 0.97. That’s the quality gap between short stories generated by ChatGPT 4.0 and short stories written by humans, scored by readers on a scale running from minus 3 to plus 3. The machine won.
Immersion went the same way, 1.42 for the AI stories against 1.00 for the human ones. And across three experiments with more than 2,500 participants, nobody did better than chance at telling which was which.
The setup left little room to bluff
In the first experiment, 1,682 participants each read one of six short stories, each about 1,000 words. Three came from well-known literary magazines and short story collections. Three were generated with ChatGPT 4.0, prompted to match the theme, style and narrative perspective of the human originals.
Then came the trick. Half the participants were told a human wrote the story. The other half were told ChatGPT did. That label was accurate for only half the people in each group, according to researchers Sydney Sears and Deena Skolnick Weisberg, whose study was published in the journal Judgment and Decision Making.
The label moved the score more than the author did
Regardless of who had written the thing, participants rated it higher when they were told a person was behind it. The byline was doing work the prose wasn’t.
Attitudes toward AI bent the results further. People who came in positive about AI gave higher ratings generally, and higher still when they were told ChatGPT wrote what they’d just read. Among skeptics, the effect ran the other way. An earlier study on AI-generated poems turned up the same bias.
Reading both side by side didn’t help
Two further experiments with 905 total participants made the task easier on paper. Each person read one human story and one AI story, then had to say which was which. A direct comparison, no memory involved, 50-50 odds.
Performance stayed at chance.
One thing did predict success, and it isn’t the one you’d guess. Self-reported experience with AI systems correlated positively with picking the right origin. Self-reported experience with fiction didn’t help at all. Reading a lot of novels apparently trains you to enjoy prose, not to audit it.
Easy to read isn’t the same as good
The authors have a plain explanation for the scores, and it isn’t that the machine writes better. AI-generated text tends to be smoother, easier to read and more emotionally upbeat than human writing. People prefer material that’s easier to process, so those traits can lift ratings without any literary gain underneath.
High-quality literary fiction often works the opposite way. It’s deliberately hard to get into and makes you do the work. A story can be high quality and not very engaging, and the reverse, the researchers suggest.
Length matters too. Holding 1,000 words together is a different job than holding hundreds of pages together, and the short story format plays to the model’s strengths. The researchers’ conclusion still lands: AI can produce creative work people rate as at least equal to human work, while those same people don’t believe AI can do it.
Train the model on one writer and the experts flip
Expertise buys you something, but less than you’d hope. A study from last October by Stony Brook University and Columbia Law School found that with simple prompts, professional readers clearly preferred the human-written texts.
Then the models were trained on individual authors’ styles. The experts preferred the AI-generated texts eight times more often for style imitation and twice as often for writing quality.
All data and materials from the Sears and Weisberg study are freely available on the Open Science Framework, which means you can run the comparison on yourself rather than take anyone’s word for it. Do that before you decide you’d have spotted the machine. Everyone thinks they would have.