LLM Summarization Bias: Narrative Information Loss
Explore how LLMs flatten implicit physical parameters during text compression, collapse narrative entropy, and replace sensory scenes with abstract emotion tags.
LLM Summarization Bias: What Gets Lost When Machines Compress Stories?
Compressing a lengthy document, an investigative report, or a multi-chapter novel into a few crisp paragraphs within seconds stands as one of the most celebrated capabilities of Large Language Models (LLMs). Users and researchers routinely treat automated text summarization as a frictionless convenience—a neutral extraction of core semantic value. Yet immediately beneath this surface lies a structural flaw that fundamentally alters narrative architecture: Summarization Bias.
When an LLM compresses narrative prose, it does not merely reduce token volume; it systematically strips away implicit physical cues, erases environmental variables, and flattens narrative friction. The resulting summary is rarely a neutral condensation of the original work. Instead, it becomes an artificially flattened artifact where sensory depth is converted into explicit emotional declarations.
This article examines the mechanisms of information loss during LLM text compression, how Shown physical structures regress into Told abstract tags, and the empirical consequences of this shift on narrative entropy.
1. The Architecture of Compression: Regressing from "Shown" to "Told"
High-quality literary prose and screenplays avoid explicitly declaring a character's internal state. Rather than stating emotions outright, subtle narratives encode atmospheric tension into physical variables on the textual surface. This principle forms the cornerstone of the Shown paradigm within narrative engineering and biophysical parametric critique.
In a well-crafted scene, a character's profound anxiety or creeping dread is almost never rendered as "She felt extremely anxious." Instead, six fundamental physical parameters govern the scene:
- Luminous Decay: The narrowing beam of a light source, cast shadows, or shifting luminance gradients.
- Thermal Gradient: Rapid drops in ambient temperature, cold drafts, or frigid touch points.
- Acoustic Impedance: Shifting sound propagation speeds, muffled reverberation, or deadened echoes.
- Kinetic Momentum: Sudden deceleration of movement, physical drag, or atmospheric resistance.
- Atmospheric Pressure: Heavy humidity, thick air, or a sense of respiratory constriction.
- Spatial Geometry: Encroaching walls, claustrophobic spatial bounds, or retreating exits.
When an LLM processes a scene built on these physical variables and receives a summarization prompt, its autoregressive architecture prioritizes token efficiency and global probability alignment. In doing so, the model identifies environmental variables—such as acoustic dampening or thermal drops—as "non-essential noise" or peripheral detail. Consequently, these sensory indicators are discarded during compression.
To fill the atmospheric vacuum created by stripping these physical variables, the LLM inserts an explicit summary tag: "The character experienced intense fear in the dark room." This shift represents a direct collapse into the Told layer. The model replaces the reader's active cognitive reconstruction with a pre-packaged semantic label. This mechanism directly explains why ChatGPT and generative LLMs write flat stories when generating or compressing prose.
2. The Three-Stage Mechanism of Summarization Bias
Summarization Bias in generative models is not a random glitch; it is an inherent property of probabilistic token optimization. LLMs execute a systematic three-stage filtering process when compressing narrative text:
A. Stripping Implicit Cues and Micro-Details
Micro-details that generate information friction ($If$)—such as rust flaking off a doorknob, dust settling on a desk, or irregular pauses between footsteps—receive low attention weights during compression prompts. As a result, these implicit cues are discarded first.
B. Injecting Abstract Sentiment Tags
Removing physical parameters creates a contextual gap. To maintain global coherence, the model explicitly declares what the scene meant. It injects abstract labels such as "gloomy atmosphere," "overwhelming grief," or "desperate situation." The LLM transforms from a neutral compressor into an active, flattening commentator.
C. Collapsing Information Friction and Causal Branching
In authentic narrative structures, event progression is non-linear. Information friction slows pacing, delays decisions, and creates atmospheric resistance. Summarization eliminates this friction entirely, smoothing out the narrative trajectory into a linear sequence of cause and effect. This collapse destroys one of the most critical elements of storytelling: pacing.
3. Narrative Entropy ($S_n$) and Semantic Flattening
To quantify Summarization Bias empirically, we evaluate its impact on Narrative Entropy ($S_n$). In its canonical formulation, Narrative Entropy integrates information friction ($If$) and causal branching ($Cb$) across narrative time ($t$):
$$S_n = \int (I_f \times C_b) \, dt$$
In rich literary prose, $S_n$ remains dynamic, leaving deliberate gaps for the reader's mind to reconstruct. When an LLM compresses this prose, information friction ($If$) drops toward zero and causal branching ($Cb$) collapses into a single deterministic line. The resulting summary exhibits a total crash in Narrative Entropy.
A collapse in entropy signifies the eradication of atmospheric resistance and sensory depth. When models perform this flattening systematically across corporate or creative pipelines, it fuels algorithmic homogenization and story standardization across generative media.
4. Empirical Evidence and Inter-Rater Reliability
Measuring Summarization Bias and automated text transformation is closely linked to how language models evaluate implicit features. In our empirical benchmark study (Zenodo DOI: 10.5281/zenodo.21740239), we evaluated the performance of LLMs and rule-based detectors when annotating inferential narrative variables against blind human baselines.
The empirical findings reveal why models struggle with implicit narrative structures during summarization:
- High Raw Agreement vs. Chance-Level Kappa: LLMs often achieve high raw percent agreement on rare narrative features, yet chance-corrected metrics like Cohen's Kappa remain near zero ($\kappa \approx 0.00–0.02$).
- Atmospheric Friction Discrepancy: Models routinely fixate on surface emotion words while missing physical atmospheric contradictions (such as a character shivering uncontrollably in a sweltering room).
- Inter-Model Divergence: Independent LLMs (Gemini, Grok, Claude, ChatGPT) display severe variance from one another and from human annotators when evaluating the same text blocks.
These empirical results confirm that LLMs are not neutral, objective summarizers. Rather, they actively re-engineer text according to their internal probability distributions, destroying implicit surface data in the process.
5. Practical and Theoretical Implications
Summarization Bias is not merely an abstract concern for literary theorists; it presents concrete risks across computational and consumer domains:
- AI-Driven Search (SearchGPT, Perplexity): When AI search engines condense narrative or artistic prose, they provide users with flattened abstract summaries rather than the sensory reality of the text, stripping away the work's aesthetic core.
- Synthetic Dataset Contamination: Training future models on LLM-generated summaries risks permanently degrading AI models' ability to comprehend or generate deep, multi-layered prose.
- Analytical Illusion in Digital Humanities: Researchers using automated LLM summaries as proxies for original texts risk analyzing model-injected sentiment tags rather than the actual structural properties of the source material.
6. Checklist: Identifying Summarization Bias
Use this five-point diagnostic checklist to determine whether an AI summary suffers from Summarization Bias:
- Physical Variable Retentions: Were sensory parameters (light, heat, pressure, acoustics) preserved, or were they completely erased?
- Abstract Label Spike: Did the frequency of explicit emotion words (e.g., "sad," "terrified," "anxious") increase compared to the original text?
- Friction Eradication: Was reading friction eliminated to force events into a frictionless, linear sequence?
- Entropy Collapse: Were inferential gaps closed, leaving no room for cognitive reconstruction?
- Shown-to-Told Shift: Did the text transition from encoding states through physical details to explicitly declaring internal states?
7. Conclusion
Large Language Models are remarkable utilities for rapid text compression. However, mistaking an automated summary for an equivalent representation of narrative prose is a fundamental error. Summarization Bias is a systemic property of language models that replaces physical reality with explicit sentiment labels.
Computational narratology and literary critique remain resilient against algorithmic flattening only when they look beyond surface emotion tags and measure the implicit physical variables beneath. Without this rigor, smooth but hollow AI summaries risk replacing the sensory experience of narrative itself.
References
- Summarization Bias in Large Language Models — Zenodo DOI 10.5281/zenodo.20783465
- Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus — Zenodo DOI 10.5281/zenodo.21740239
- HuggingFace Dataset: Objective Projection Corpus & Evaluation
- Miall, D. S. & Kuiken, D. (1994). Foregrounding, defamiliarization, and affect: Response to literary stories. Poetics, 22, 389–407.
- Artstein, R. & Poesio, M. (2008). Inter-coder agreement for computational linguistics. Computational Linguistics, 34(4), 555–596.
Frequently Asked Questions
What is LLM Summarization Bias?
Summarization Bias refers to the systemic tendency of Large Language Models (LLMs) to strip implicit physical variables (light, thermal gradients, acoustic friction) from narrative text during compression, replacing them with explicit emotion labels and flattened plotlines.
Why do LLMs erase physical literary details?
LLMs are optimized for token efficiency and high-probability semantic representations. During compression, sensory details are mathematically flagged as non-essential noise and discarded to generate concise outputs.
How does the Shown vs. Told dynamic shift during AI summarization?
In Shown structures, emotion is encoded through environmental details, requiring cognitive reconstruction. Told structures declare emotion outright ("she was afraid"). AI summarization rapidly converts Shown physical setups into Told explicit tags.
How does summarization impact Narrative Entropy ($S_n$)?
Because automated compression eliminates information friction ($If$) and simplifies causal branching ($Cb$), Narrative Entropy crashes. This strips the text of its atmospheric depth and interpretive resistance.
How can researchers detect Summarization Bias?
By comparing the density of physical variables before and after compression, tracking the ratio of explicit sentiment tags, and measuring whether narrative pacing friction has been zeroed out.
Levent Bulut — Independent researcher and author. ORCID 0009-0007-7500-2261.
@misc{bulut2026llmsummarizationbias,
author = {Bulut, Levent},
title = {LLM Summarization Bias: Measuring Implicit Narrative Information Loss},
year = {2026},
howpublished = {\url{https://leventbulut.com/llm-summarization-bias-narrative-information-loss/}},
note = {Independent Researcher, ORCID: 0009-0007-7500-2261. Objective Projection Paper Series. Refers to Zenodo DOI: 10.5281/zenodo.20783465}
}