Scientific Literary Criticism: What It Can and Cannot Do
A century of attempts to make literary criticism empirical formalism, stylometry, reading experiments, distant reading and the test each had to pass.
Scientific literary criticism: what it can and cannot do
Roughly once a generation, someone announces that literary criticism is about to become a science. Roughly once a generation, the announcement is followed by a period of real progress in a narrow area, an overreach into questions the method cannot touch, a backlash, and a long quiet.
This has happened often enough that the pattern itself is worth studying. What follows is a map of five serious attempts, the single test each of them had to pass, and an honest account of where the current wave — mine included — stands against that test.
1. What the phrase actually means
"Scientific literary criticism" is not one thing, and much of the confusion around it comes from collapsing three quite different ambitions:
- Descriptive. Measure properties of texts — word frequencies, sentence structures, the shape of a plot across a corpus. Nothing here is contested in principle. Texts have countable features.
- Causal. Show that a formal property of a text produces a measurable effect in readers. This is where the real claims live, and where the difficulty is.
- Evaluative. Decide which works are good. No empirical method has ever done this and, in my view, none will. A method that tells you Hamlet scores 84 has not measured value; it has measured whatever it measured and attached a prestigious name to it.
Most disputes about whether criticism can be scientific are actually disputes about whether someone has quietly slid from the first ambition to the third.
2. Five attempts, and what each got right
Russian Formalism (1910s–1920s)
The Formalists made the first move that matters: they proposed that literary effect comes from devices in the text, not from the author's biography or the reader's soul. Shklovsky's notion of defamiliarization — that art works by making the habituated strange again, slowing perception down — is a mechanism claim. It says this feature produces that effect.
Propp went further and did something recognisably empirical: he took a corpus of Russian folktales and extracted a finite inventory of recurring narrative functions in a fixed order. Whatever one thinks of the result, the method was corpus-based and its output was checkable against new tales.
What they got right: effect is produced by structure, and structure can be described precisely.
What was missing: no reader was ever measured. Defamiliarization was a hypothesis about minds, argued entirely from texts.
Practical criticism (1929)
I. A. Richards did something almost nobody had done: he ran an experiment on readers. He handed his Cambridge students poems with the authors' names removed and collected their written responses.
The results were unflattering to everyone involved. Competent, trained readers produced wildly divergent readings, missed obvious features, and were heavily influenced by preconceptions about what poetry should sound like. Strip the author's name and the consensus evaporates.
What he got right: if you want to claim something about how readers respond, ask readers.
What it revealed: reader response is far more variable than criticism's confident tone implies — a finding that keeps being rediscovered.
Stylometry (1964 onwards)
The cleanest success story in the whole field. Mosteller and Wallace took the twelve Federalist Papers whose authorship was disputed between Hamilton and Madison, and settled the question using Bayesian analysis of function-word frequencies — words like and, to, upon, which carry no subject matter and which a writer deploys unconsciously at a stable personal rate.
Their conclusion attributed the disputed papers to Madison, and the method has been re-run with newer techniques many times since. It launched authorship attribution as a quantitative field.
Why it worked: the question had a fact of the matter. Someone wrote those papers. There was a right answer independent of the method, and the method could be tested against known cases.
What it does not tell you: anything about whether the papers are any good, or why they persuade.
The empirical study of literature (1980s–)
This is the tradition that took Formalism's hypothesis and finally measured it. In a much-cited set of four studies, Miall and Kuiken presented readers with a short story segment by segment and recorded reading times alongside ratings. Segments with heavier stylistic foregrounding were read more slowly and rated as more striking and more affecting — and the effect did not depend on the reader's literary training.
That is a real result. Shklovsky's claim about slowed perception, made in 1917 from the armchair, turned out to be measurable seventy-odd years later, in seconds.
What this tradition got right: it closed the loop. Text feature → reader measurement → statistical test.
Its limitation: effects are modest, variance between readers is large, and generalising from one short story to literature is exactly the move the method forbids.
Computational and distant reading (2000s–)
The most recent wave scaled up: instead of close-reading a canon, analyse thousands of books at once and look at the shape of the aggregate — genre boundaries drifting, sentiment arcs across a plot, vocabulary shifting by decade.
It produces things no human reader could see, because no human reader can hold ten thousand novels in mind. It also produces the field's most common failure mode, which is worth naming precisely: an operation performed on a corpus is not automatically a finding about literature. The gap between "the model clustered these books together" and "these books constitute a genre" is a full interpretive argument, and it does not come free with the graph.
3. The test all five had to pass
Strip away the vocabulary differences and every attempt above faces the same three obligations. They are unglamorous and they are where most work dies.
Operational definitions
A concept has to be defined so that two independent people applying it to the same text reach the same answer. Not roughly the same. The same.
This sounds like bookkeeping. It is the hardest part of the entire enterprise, and it is where the humanities-facing end of this field is weakest, because critical vocabulary evolved for a different purpose — to be productive and suggestive rather than to be applied identically by strangers.
Falsifiability
The claim has to be able to lose. "This novel enacts a tension between memory and place" cannot lose; there is no observation that would count against it. "Readers will pause longer at foregrounded segments" can lose, and did not.
Replication
Someone with no stake in the outcome has to get the same result. This is the obligation nobody can satisfy alone, and the one that finally decides whether a framework was real or was one person's private vocabulary with citations attached.
4. Where reliability keeps breaking
Of those three, the first is the quiet killer, and it is worth understanding why.
Suppose you propose that a certain narrative device produces a certain effect. Before you can test that, you need to identify which passages contain the device. If you do it yourself, you have measured your own consistency, not the device. If two people do it and disagree, you have no data — you have two incompatible datasets.
The statistic for this is chance-corrected agreement, usually Cohen's kappa. And it comes with a trap that catches even careful work: raw percentage agreement is nearly meaningless when the feature is rare or near-universal. If a device appears in 96 of 100 passages, a rater who marks every passage positive scores 96% agreement while contributing exactly zero information. Reporting that number without the class distribution overstates performance dramatically.
This is not a hypothetical failure mode. It is the ordinary one.
5. A worked example, including the part that failed
Since a page arguing for empirical discipline should submit to it, here is a case from my own work.
I maintain a corpus annotated for six craft features — two prohibitions (the narrator may not name an emotion, may not use a simile) and four positive techniques. The annotations were generated by a rule-based detector and shipped with the dataset. Whether a human would agree with them had never been tested.
Across three studies, five automated labellers were scored against blind human labels. On the feature requiring the most inference — an abstract state rendered as a concrete physical detail — they marked 0, 1, 40, 72 and 78 of a hundred passages as positive. The human marked 9. Cohen's kappa was at or indistinguishable from chance for five of the six labellers.
Two further details, both of which cut against me:
- The labellers diverged from each other as much as from the human. On one rule, two of them marked 9 and 82 passages positive respectively, against a human count of 96 — same written definition, same hundred passages.
- In the first study the human rater was me, and my own criterion drifted partway through. I had to relabel the rule entirely.
Two readings survive, and I cannot separate them with one rater per study. Either the feature is genuinely inferential — reconstructed by the reader rather than present in the text, in which case no automatic detector should find it. Or my definition is too loose for anyone to apply consistently, including me. Distinguishing these needs a second independent human rater.
The reason for putting this on a page about the field rather than burying it in a methods appendix: this is the normal outcome, not an unusual one. Most annotation schemes for interpretive features have never been tested this way. Some fraction of them would not survive it.
6. What this field can honestly deliver
Strong ground:
- Authorship and dating. There is a fact of the matter and methods can be validated against known cases.
- Formal description at scale — vocabulary, syntax, structural patterns across corpora.
- Reception and response measurement: reading times, eye movements, physiological signals, recall. Modest effects, but real ones.
- Comparative claims within a controlled set: this version was read more slowly than that version.
Weak ground:
- Any single-number score of literary quality.
- Claims about what a text means, as opposed to how it is built or how it is processed.
- Generalising from one corpus, one language, or one genre to "literature."
- Anything resting on an annotation layer whose inter-rater reliability has not been reported.
The honest position is narrower than the enthusiastic one and much more durable. A method that reliably measures one thing is worth more than a framework that gestures at everything.
7. A checklist
If you are building or evaluating work in this area, five questions do most of the work:
- Could two strangers apply this definition and agree? If it has not been tested, it is not yet a definition.
- Is chance-corrected agreement reported, with the class distribution? Raw percentage alone, on a skewed feature, tells you nothing.
- What observation would make this claim false? If none, it is criticism — which is fine, but it is not a finding.
- Has anyone unconnected to the author reproduced it? Publication volume is not replication; they are different axes and it is easy to confuse them in one's own favour.
- Is a descriptive result being reported as a causal or evaluative one? The slide usually happens in the abstract's last sentence.
8. Why keep trying
Given that record, an obvious question is why anyone bothers.
Because the alternative is not humility, it is unfalsifiability with good prose. Criticism that cannot be wrong accumulates no knowledge; it accumulates positions. The empirical constraint is uncomfortable precisely because it can take things away from you — and a field where nothing can ever be taken away is not one where anything is ever established.
And occasionally it delivers something no amount of insight would have: a claim made in 1917 about slowed perception, confirmed in the 1990s by a stopwatch and a short story. That is worth the discipline it costs.
Further reading
Foundational work in the tradition
- Propp, V. (1928). Morphology of the Folktale.
- Richards, I. A. (1929). Practical Criticism: A Study of Literary Judgment.
- Mosteller, F., & Wallace, D. L. (1964). Inference and Disputed Authorship: The Federalist. Addison-Wesley.
- Miall, D. S., & Kuiken, D. (1994). Foregrounding, defamiliarization, and affect: Response to literary stories. Poetics, 22, 389–407.
- Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46.
- Byrt, T., Bishop, J., & Carlin, J. B. (1993). Bias, prevalence and kappa. Journal of Clinical Epidemiology, 46(5), 423–429.
- Artstein, R., & Poesio, M. (2008). Inter-coder agreement for computational linguistics. Computational Linguistics, 34(4), 555–596.
From this project
- Inter-rater reliability of LLM and rule-based annotation — the three studies discussed in section five
- Is the Physics of Literature really physics? — the same three obligations applied to my own framework
- The Scope Map — what this framework does not claim
- Objective Projection Convergence Test v2.0 — a pre-registered protocol, not yet run
Frequently asked questions
What is scientific literary criticism?
An umbrella term for approaches that study literature using empirical methods — corpus analysis, reader experiments, statistical modelling — rather than interpretation alone. It covers descriptive work on textual features, causal work on reader response, and computational analysis at corpus scale. It does not include, and has never succeeded at, empirical judgements of literary value.
Can literary quality be measured?
No method has done so, and the ambition is generally treated as a category error. What can be measured are formal properties of texts and measurable effects on readers. A score presented as a quality judgement has usually measured something else and renamed it.
What is stylometry?
The statistical study of writing style, most successfully applied to authorship attribution. Its classic case is the 1964 analysis of the disputed Federalist Papers using frequencies of common function words, which attributed them to Madison. It works well because authorship questions have a fact of the matter to be right or wrong about.
What is foregrounding, and was it ever confirmed?
Foregrounding refers to stylistic features that deviate from ordinary usage and draw attention to the text itself — a concept from Russian Formalism and the Prague School. Four studies by Miall and Kuiken in 1994 found that readers spent longer on more heavily foregrounded segments and rated them as more striking and more affecting, independently of literary training.
Why does inter-rater reliability matter so much?
Because most empirical claims about literature require first identifying which passages have a feature. If independent raters cannot agree on that, every downstream statistic inherits the disagreement. Chance-corrected agreement should be reported together with the class distribution, since raw percentage agreement is inflated when a feature is rare or near-universal.
Is distant reading scientific?
Its measurements are reproducible, which is a genuine strength. The interpretive step — from what a model grouped together to a claim about genre, period or influence — is a separate argument that the computation does not supply. Treating the output as self-interpreting is the most common failure in this area.
Levent Bulut — Independent Researcher. ORCID 0009-0007-7500-2261.