> ## Content Index
> Fetch the complete content index at: https://leventbulut.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Can Readers Tell AI Stories From Human Stories? Bulut Doctrine
- URL: https://leventbulut.com/can-readers-tell-ai-stories-from-human-stories/
- Published: 2026-09-22T16:43:03.000Z
- Updated: 2026-09-22T16:43:03.000Z
- Description: Readers spot AI stories at chance level; a classifier does it at 93%. The difference is not in style but in narrative architecture. What new research shows.
- Author: Levent Bulut
- Tags: Computational Narratology

**Conflict of Interest and Academic Grounding Statement (COI):** This article examines whether readers can distinguish stories written by AI from those written by humans, through the findings of two independent studies published in 2026 (*Judgment and Decision Making* and StoryScope), framed by the distinction between style and narrative architecture. The [Levent Bulut](https://leventbulut.com/graph/) is the founder of the Objective Projection and Narrative Engineering methodology. The formalisations, dataset and evaluation files referred to in the text are openly published so that independent researchers can audit them to academic standards, and they are open to audit. Figures from external studies are reproduced exactly from their sources; Summarization Bias is a registered protocol that has not yet been tested. 

### Abstract

Two studies published in 2026 appear to contradict each other. In one, readers guessed whether a short story was written by a human or an AI at chance level — and rated the AI stories as higher quality. In the other, a classifier made the same distinction with better than 93 per cent accuracy. This article proposes where the contradiction resolves: readers look at **style**, the classifier looks at **narrative architecture**. AI does not always sound like a machine; but when it builds a story, it can behave like one — in its characters' choices, in the ordering of causes, in the flow of time, in moral ambiguity and in how much it explains.

Imagine reading two short stories. One was written by a human. One was generated by AI.

Could you tell which is which?

Most likely not. And the interesting part of this article is that the difference is real even though you cannot see it.

## Readers cannot tell

A study published in summer 2026 in *Judgment and Decision Making* (Cambridge University Press) asked the question directly. Researchers at Villanova University paired three human-written stories, published in reputable literary journals or story collections, with three ChatGPT stories matched to each in topic and style.

In the first experiment, 1,682 adult readers each read one of these stories. Readers were told who wrote the story — but half the time they were told wrongly. The result showed two effects stacked on each other: readers found the AI stories **higher in quality and more absorbing**, yet gave higher ratings to any story labelled as human-written, whoever had actually written it. The highest-rated stories of all were AI stories labelled as human.

In two follow-up experiments, 905 readers read stories of both kinds and were asked to guess their origin. They were right **40 per cent and 52 per cent** of the time — performance indistinguishable from chance.

In short: readers do not see the difference, but they respond to the label.

## Machines can tell — by looking somewhere else

The StoryScope study (Russell et al., 2026), posted to arXiv the same year, asked the question differently: can AI stories be distinguished from human ones *without using stylistic cues at all*, by looking only at how the narrative is built?

The researchers assembled 10,272 writing prompts. For each prompt there was one story by a human author and stories from five language models (Claude, DeepSeek, Gemini, GPT, Kimi). In total 61,608 stories, each around 5,000 words. From each, 304 narrative features were extracted across ten dimensions including character, plot, setting and temporal structure.

Using those narrative features alone — with every stylistic feature, such as sentence rhythm and density of figurative language, held out — the human-versus-AI distinction was made at **93.2 per cent**. That retains more than 97 per cent of the performance of a model that also uses style. On the task of identifying which model wrote a story, performance was 68.4 per cent. In the study's framing, AI stories cluster in a shared region of narrative space, while human stories are far more diverse.

*(Interpretation — marked separately from the data.)* Placed side by side, the two studies turn the contradiction into something else: **whatever carries the difference is not where the reader is looking.**

## Style versus narrative architecture

A story can be read at two layers.

**Style** is the surface of the sentence: word choice, rhythm, image, voice. It is the layer a reader meets first and usually evaluates consciously. Current language models now produce fluent and often careful text at this layer, which is what the Cambridge readers' preference for the AI stories reflects.

**Narrative architecture** is the set of decisions beneath the sentences: when a character chooses and why, how events connect, whether time runs straight or turns back, whether the story states its own meaning or leaves it to the reader. Readers *experience* this layer but usually do not name it. The feeling that "something was missing but I can't say what" often comes from here.

The distinction matters for [computational narratology](https://leventbulut.com/computational-narratology/): most text-detection tools work at the level of style and are easily defeated when a text is rewritten. Narrative architecture is far more robust to rewriting, because changing it requires rebuilding the story.

## The AI doesn't always sound like a machine. It may behave like one.

The differences StoryScope found are concentrated on exactly five axes of narrative architecture.

### 1\. Character agency

Whether a character's choices actually drive the story. StoryScope names character agency explicitly among the narrative choices it shows to be discriminating without stylistic cues. In AI stories, characters more often *pass through* events; in human stories they more often *make* them.

### 2\. Causal structure

How events connect. According to the study, AI stories favour **tidy, single-track plots**. Each event joins neatly to the next; there are fewer loose ends, parallel threads, or details that only acquire meaning later.

### 3\. Temporal complexity

Flashbacks and nonlinear structure are markedly more frequent in human stories. AI mostly lets time run in a straight line.

### 4\. Moral ambiguity

Human authors frame their protagonists' choices as **more morally ambiguous**. A choice in which right and wrong are cleanly separated leaves the reader less to think about.

### 5\. Explanatory density

Perhaps the most striking finding: AI stories **over-explain their themes**. Rather than leaving what the story means to the reader, they say it in the text.

The study also reports model-specific fingerprints: notably **flat event escalation** in Claude, over-reliance on dream sequences in GPT, and external character description in Gemini.

## Explanatory density and the told–shown axis

The fifth finding touches directly on a question [Levent Bulut](https://leventbulut.com/graph/) has been working on. The [Objective Projection](https://leventbulut.com/objective-projection-definition/) framework places narrative between two modes: in *told* mode, content is stated on the surface; in *shown* mode, it is encoded in physical cues and the reader reconstructs it.

A hypothesis registered under the name Summarization Bias proposes that language models err on this axis not randomly but in **one direction** — toward told mode. StoryScope's finding that AI stories over-explain their themes is **consistent** with that hypothesis.

Care is needed here: consistent with is not evidence for. StoryScope was not designed to test this hypothesis; it measured different features by a different method. Summarization Bias remains a protocol registered before data collection and has not yet been tested. But an independent study arriving, by its own route, at a finding that points the same way is a sign of a question worth testing. A scene-level view of the same problem was examined separately: [Why AI Cannot Maintain Micro-Focus](https://leventbulut.com/why-ai-cannot-maintain-micro-focus-macro-narrative-blindness/).

## An open question: who labelled the features?

There is an important detail in StoryScope's method. The values of the 304 features for each of the 61,608 stories were assigned by **a language model** (Gemini 3 Flash).

The detail is interesting because Bulut's own published reliability study measured how well language models label inferential narrative features, and found it low: on one rule, with a human marking 9 of 100 scenes as present, five machine raters returned 0, 1, 40, 72 and 78, and agreement was at chance level for four of them.

*(Interpretation.)* This does not invalidate StoryScope's finding. If the labels were random noise the classifier could not reach 93 per cent; the labels carry a systematic signal. The open question is subtler: is that signal the narrative architecture a human would see if they labelled the same features, or partly the reading habits of the labelling model itself? Because the two studies used different features, different languages and different texts, no direct inference runs from one to the other. But the sentence "the machine sees narrative architecture" should be read with the rider *through a machine's eyes*.

## So what does this mean?

A few things can be said, none of them with a claim to certainty.

First, "can I tell AI text apart?" is probably the wrong question for a reader. Readers look at style, and at the level of style the gap is closing. The difference sits in the layer readers usually do not name.

Second, that layer is **designable**. A character's choice, the line of causation, the ordering of time, moral tension, how much explanation is left to the reader — these are decisions. As proposed in the first article of this series ([Who Is the Author When AI Writes the Story?](https://leventbulut.com/who-is-the-author-when-ai-writes-the-story/)), the question of authorship can itself be asked by looking at exactly these layers.

Third, the Cambridge study's label effect is a warning in its own right: readers rate according to who they believe the author is. That finding cuts both ways in the debate over disclosing AI use.

[Levent Bulut](https://leventbulut.com/About/)'s approach offers no verdict here; its convergence claim is statistical, not deterministic. Objections to the framework are collected separately: [Bulut Doctrine — Critiques](https://leventbulut.com/bulut-doctrine-critiques/).

### Limitations

- **The two studies used different texts.** The Cambridge study used short stories; StoryScope used stories of around 5,000 words. The StoryScope authors note that this length enables extraction of fine-grained features that shorter texts cannot support. Part of the gap between reader and classifier may come from text length.
- **No comparison on the same stories.** Readers and classifier did not evaluate the same texts. "The reader does not see it, the machine does" is a reading of two separate studies side by side, not the result of a single experiment.
- **The labeller is a language model.** StoryScope's feature values were assigned by a language model. This article does not claim that is a problem; it notes only that it is a question worth asking.
- **Summarization Bias is untested.** The hypothesis is a protocol registered before data collection; StoryScope's finding is consistent with it but does not validate it.
- **Models change quickly.** The findings reported belong to particular model versions; narrative tendencies may change in later versions.
- **No biometrics.** Reader responses mentioned here are self-report data from the studies cited; there is no physiological measurement.

## Frequently Asked Questions

### Can readers tell an AI story from a human one?

On current evidence, largely no. In a study published in 2026 in *Judgment and Decision Making*, readers guessed the origin of short stories correctly 40 and 52 per cent of the time — performance indistinguishable from chance. They also rated the AI stories as higher quality.

### Then is an AI story the same as a human story?

No. The difference lies not in style but in how the narrative is built. The StoryScope study separated human and AI stories at 93.2 per cent using only narrative features, with no stylistic cues. AI stories over-explain their themes and favour single-track plots; human stories show more moral ambiguity and more temporal complexity.

### Why do AI detection tools so often get it wrong?

Most of them work at the level of style — word choice, sentence rhythm, the frequency of particular patterns — and that layer changes easily when a text is rewritten. Narrative architecture is far more robust, because changing it requires rebuilding the story itself. How narrative features are themselves labelled, however, is a separate question worth asking.

### References

- *Bot or not: Can people tell the difference between stories written by a human or by an AI system?* (2026). *Judgment and Decision Making*. Cambridge University Press. Villanova University; senior author D. Weisberg. [Cambridge Core](https://www.cambridge.org/core/journals/judgment-and-decision-making/article/bot-or-not-can-people-tell-the-difference-between-stories-written-by-a-human-or-by-an-ai-system/45E6DC0BB90AA648654D5AE243F6C667?ref=leventbulut.com)
- Russell, J., et al. (2026). *StoryScope: Investigating idiosyncrasies in AI fiction*. arXiv:2604.03136\. [arxiv.org/abs/2604.03136](https://arxiv.org/abs/2604.03136?ref=leventbulut.com)
- Bulut, L. (2026). *The Bulut Doctrine: Architectural Framework*. Zenodo. [10.5281/zenodo.18689179](https://doi.org/10.5281/zenodo.18689179?ref=leventbulut.com)
- Bulut, L. (2026). *Why AI Cannot Write Emotional Scenes: Objective Projection as a Framework for Auditable Narrative Generation*. Zenodo. [10.5281/zenodo.22466696](https://doi.org/10.5281/zenodo.22466696?ref=leventbulut.com)
- Bulut, L. (2026). *Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models* (v1.1). Zenodo. [10.5281/zenodo.22817289](https://doi.org/10.5281/zenodo.22817289?ref=leventbulut.com) · arXiv:2609.20712
- Bulut, L. (2026). *Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus*. Zenodo. [10.5281/zenodo.21740239](https://doi.org/10.5281/zenodo.21740239?ref=leventbulut.com) · arXiv:2609.13936
- Bulut, L. (2026). *Objective Projection Dataset* \[Dataset\]. Hugging Face Datasets. [10.57967/hf/8960](https://doi.org/10.57967/hf/8960?ref=leventbulut.com)
- Complete registered record: [Levent Bulut — Corpus](https://leventbulut.com/corpus/)

### Citation

@misc{bulut2026tellapart,
  author       = {Bulut, Levent},
  title        = {Can Readers Tell AI Stories From Human Stories?},
  year         = {2026},
  howpublished = {leventbulut.com},
  url          = {https://leventbulut.com/can-readers-tell-ai-stories-from-human-stories/},
  note         = {Objective Projection / Bulut Doctrine. ORCID: 0009-0007-7500-2261}
}