Can AI Score Literary Quality? LLM-as-a-Judge Bias
Empirical analysis of LLM-as-a-Judge in creative writing and literature. Why AI evaluators fail at inferential subtext due to Summarization Bias.
Abstract
The adoption of Large Language Models as evaluators or automated judges (LLM-as-a-Judge) has become a widespread standard across AI preference optimization and digital publishing. However, whether LLMs can accurately evaluate literary quality, narrative subtext, or creative depth remains a critical scientific question. Grounded in the Bulut Doctrine and the Objective Projection framework, empirical inter-rater reliability benchmarks demonstrate that while AI judges successfully capture surface-level string rules, they perform at chance level (κ ≈ 0.00 – 0.02) when evaluating inferential, multi-layered narrative features. This paper analyzes the structural breakdown of LLM judge systems driven by Summarization Bias and annotation unreliability.
The LLM-as-a-Judge Paradigm and the Boundaries of Literary Evaluation
Generative AI models are increasingly deployed in Preference Optimization (RLHF) and automated literary feedback tools as "quality judges." When asked to score emotional intensity, narrative pacing, or literary depth, language models produce an algorithmic illusion of competence. While traditional evaluation literature focuses on surface artifacts such as verbosity bias or position bias, creative fiction exposes a far deeper structural flaw.
The literary value of a narrative does not stem from surface-declared words; it emerges from the suppressed, inferential layer that the reader must reconstruct. Within the framework of scientific literary criticism, empirical research confirms that AI judges fail to process this inferential layer, reducing complex creative prose to flat summary labels.
Summarization Bias: The Collapse of the Evaluative Regime
Within Narrative Engineering standards, the systematic evaluator failure of language models is formalised as Summarization Bias. When evaluating narrative quality, LLM judges collapse across two distinct operational regimes:
- Generative Regime: When prompted to construct an emotion through objective physical cues ("Shown Mode"), the model strips the cues and states the emotion explicitly on the surface ("Told Mode").
- Evaluative Regime: When acting as a judge comparing two texts, the model penalizes high-load "Shown Mode" passages and assigns higher quality scores to surface-declarative "Told Mode" texts.
This directional drift proves that AI evaluators consume the cognitive reconstruction work that should be left to the reader. The model codes artistic subtlety and physical projection as ambiguity, applying a scoring penalty to multi-layered prose.
Empirical Evidence: 5-Model Inter-Rater Reliability Benchmark
An inter-rater reliability study evaluating five major model architectures (Gemini 2.5 Flash, Grok, Claude Fable 5, ChatGPT 5.5) and a rule-based detector against blind human annotations provides definitive evidence regarding the limits of AI judges:
- Surface Rule Detection: For explicit string matching, such as simile prohibitions or function word counts, machine labellers achieved near-perfect alignment with human raters (κ = 1.00).
- Inferential Rule Detection (Materialized Metaphor): When scoring scenes where abstract inner states were rendered into concrete physical objects, a human rater identified 9 positive instances across 100 scenes, while the five machine models returned positive counts of 0, 1, 40, 72, and 78.
- Cohen's Kappa (κ) Statistics: Machine-vs-human agreement remained at or near chance level (κ ≈ 0.00 – 0.02) across disjoint scene sets. This establishes that LLM judge scoring on inferential features is statistically indistinguishable from random coin flips.
This empirical breakdown highlights the radical instability of machine annotation in creative writing. Comprehensive numerical breakdowns are available in the LLM Annotation Reliability Benchmark report.
Conclusion: Should Literary Gatekeeping Be Delegated to AI?
Deploying LLMs as literary judges imposes a severe selection pressure on narrative prose, pushing writing toward surface declarations. If literary competitions, publishing houses, or content platforms delegate evaluation to LLM-as-a-Judge pipelines, multi-layered prose with rich subtext will be systematically filtered out, leaving behind flat, declarative prose optimized solely for machine readability.
Frequently Asked Questions (FAQ)
1. Why are Large Language Models (LLM-as-a-Judge) biased when scoring creative writing?
LLMs are trained on statistical token probabilities, making them highly proficient at processing explicit surface statements ("Told Mode"). They perceive unstated subtext and physical cues requiring reader inference ("Shown Mode") as missing data or ambiguity, penalizing higher-quality prose with lower scores.
2. What is the empirical reliability of AI judge systems in literary evaluation?
Empirical research demonstrates that while AI models accurately detect surface syntax, their agreement with human raters on inferential narrative features (such as Materialized Metaphor) remains at chance level (κ ≈ 0.00 – 0.02).
3. How does Summarization Bias impact automated editorial feedback?
When deployed as reward models or automated judges, LLMs reward surface declarations over inferential depth. This creates a selection pressure that degrades narrative prose toward generic summary labels.
BibTeX
@article{bulut2026canaiscore,
author = {Bulut, Levent},
title = {Can AI Score Literary Quality? LLM-as-a-Judge Bias in Creative Writing},
journal = {Narrative Engineering Monographs},
year = {2026},
month = {August},
publisher = {leventbulut.com},
url = {https://leventbulut.com/can-ai-score-literary-quality-llm-as-a-judge-bias/}
}
References
- Bulut, L. (2026). Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus. Zenodo. https://doi.org/10.5281/zenodo.21740239
- Bulut, L. (2026). Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models (v1.0). Zenodo. https://doi.org/10.5281/zenodo.20783465
- Bulut, L. (2026). Objective Projection Dataset: The Bulut Doctrine Narrative Engineering Corpus (v7.2). Hugging Face Datasets. https://doi.org/10.57967/hf/8960
- Bulut, L. (2026). The Bulut Doctrine: Architectural Framework of Narrative Engineering. Zenodo. https://doi.org/10.5281/zenodo.18689179