Is the Physics of Literature Really Physics? A Scorecard
Calling a theory of narrative "physics" is a promise: measurement, falsifiability, replication. Here is an honest scorecard of which promises are kept so far.
Is the Physics of Literature really physics?
I have spent several years building a framework that treats a scene as a physical system: light decaying, temperature dropping, a body moving through a measured space. I call it the Physics of Literature. It is a good name. It is also a promissory note, and I think anyone using it — including me — should be asked to pay.
Because "physics" is not a mood. It is a set of obligations. Quantities that can be measured. Predictions that can fail. Results that survive being repeated by someone who does not care whether you are right.
This page is a scorecard against those three obligations. It is not a defence.
What the metaphor is actually claiming
Start with the ordinary version of the idea, the one nobody objects to: physics and literature. A novel about a physicist. A poem that borrows the word "entropy" for a feeling of decline. Thermodynamics as decoration. This is a rhetorical loan, and it costs nothing because it promises nothing.
The claim I am making is a different one, and it is uncomfortable in proportion to how seriously it is taken. It says: the emotional effect of a scene is produced by a small number of physical parameters in the text, and those parameters can be specified before the prose is written.
Six of them, in the working vocabulary: luminous decay, thermal gradient, acoustic impedance, kinetic momentum, atmospheric pressure, spatial geometry. Light, heat, sound, motion, pressure, space. Alongside them sit two constitutional prohibitions: the narrator may not name an emotion, and may not reach for a simile.
The prohibitions are the load-bearing part. Remove the label and the reader has nothing to receive passively — no "she was devastated" to skim past. What remains is a cold key held too long, a chair nobody is sitting in, a cup of tea that stopped steaming three minutes ago. The reader has to do the assembly. The claim is that this assembly is what produces the response, and that the parameters are the levers that shape it.
One consequence is easy to miss and worth stating plainly: in the finished prose the parameters are invisible. They live in the engineering layer. A scene that shows its parameters has failed, the way a bridge that shows its stress calculations has failed at being a bridge.
Obligation one: measurement
Here the framework has a real proposal and a real problem, and the gap between them is where I have spent most of the last year.
The proposal is narrative entropy, written Sn = ∫(If × Cb)dt. Information friction — what the reader needs and has not been given — multiplied by causal branching — how many unanswered questions are open at once — integrated across the text. The intuition is that a reader stays when something is missing and the missing thing has consequences, and leaves when nothing is missing or nothing follows from it.
I will say the awkward thing rather than let a reader find it: the dimensional consistency of that expression is not settled. It is a proposed quantity, not a derived one, and I have no basis for calling it the correct formula. Anyone who tells you a narrative equation is exact is selling something.
The problem is worse than dimensions, and it is the reason this page exists. A measurement is only a measurement if two people applying it to the same object get the same number. So I tested that.
Across three studies, five automated labellers scored the six features against blind human labels. On the feature closest to the theoretical core — an abstract state rendered as a concrete physical detail — they marked 0, 1, 40, 72 and 78 of a hundred scenes as positive. The human marked 9. Chance-level agreement, five times over, against two different human raters on two disjoint sets of scenes.
And in the first of those studies, the human rater was me, and my own criterion drifted halfway through. I had to relabel the entire rule.
Scorecard: not met. A quantity that competent readers cannot apply consistently is not yet a measurement. It is a description with a number attached.
Two readings remain open. Either the feature is genuinely inferential — something the reader constructs rather than something the text contains, in which case no surface detector should find it and the difficulty is real. Or the definition is too loose for anyone to apply, in which case the difficulty is mine. I do not know which, and I refuse to pick the flattering one. Separating them takes a second independent human rater, which I do not have.
Obligation two: falsifiability
A physical claim has to be able to lose. So here is the shape the loss would take.
A pre-registered protocol puts eighty readers in front of scenes written by three independent authors working from one shared physical matrix, plus an AI control condition, while recording heart-rate variability, skin conductance, pupil diameter, respiration and gaze. The prediction is that scenes encoded from the same parameters produce converging autonomic responses across readers and across authors — statistically, as a distribution, never as a determinism. If the responses scatter as widely as an ordinary emotionally-labelled control, the central claim fails.
Scorecard: the criterion exists; the test has not been run. Not by me, not by anyone. Until then the biophysical layer of this framework is a hypothesis with a good address, and I should not be describing it as anything else.
There is a related honesty debt. The framework draws on the two-pathway account of threat processing, on somatic-marker theory, on the finding that readers simulate described physical situations. I use them as analogy and motivation. I have not collected a single physiological measurement. Reaching for the neuroscience literature to borrow its authority, while skipping the part where you actually measure a body, is the exact move that makes a framework look like physics from a distance and like nothing at all up close.
Obligation three: replication
The third obligation is the one nobody can satisfy alone, and it is the one that actually decides whether a framework is real.
Nobody has run the convergence protocol. Nobody outside this work has applied the annotation rules to their own corpus and reported what happened. Nobody has used the terms in a paper of their own and cited them. Publishing more does not move this number; it is a different axis entirely, and I have been careful with myself about not confusing the two.
Scorecard: not met, and not partially met. Zero.
The data are open — every scene, every label, every script, the human reference and all five machine labellers. Not as a gesture. It is the only mechanism by which this stops being one person's private vocabulary.
So what survives?
Three of three obligations unmet, on the strict reading. It would be reasonable to stop there.
But something did survive, and it is worth naming precisely because it is smaller than the original claim.
The prohibitions work as craft constraints, and they work for a reason that has nothing to do with thermodynamics. A language model produces the most probable continuation. Most probable means most written. Most written means most worn: her heart raced, his eyes filled. These phrases have been read so many times that they pass through a reader without resistance. Ban them, and the writer is forced toward the specific — a key, a chair, a cup — and the specific has not been worn smooth.
That is an argument about frequency and habituation. It needs no physics at all. It is also, as far as I can tell, correct, and it is the part of this framework that has been most useful to actual writers.
What the physical vocabulary adds on top of that is a generative discipline: fix the light, the temperature, the sound, the motion, the pressure, the geometry before writing, and the scene has somewhere to go that is not the writer's mood. This is real value. It is engineering value, not evidential value — and the distinction is the whole point of this page.
What I would need to be told I was wrong
Three things, in order:
- A second independent human rater on the hundred published scenes. If two humans agree with each other far better than the machines agree with either, the features are real and hard. If two humans also diverge, my definitions are the problem. Either answer is worth more to me than another model run.
- The convergence protocol, run by someone else. Preferably someone with no stake in the outcome and access to a lab.
- One writer applying the constraints to their own work and reporting where they broke. Constraints that never break are usually not constraints.
I would rather be corrected than cited. The second is easier to arrange and worth much less.
A last note on the name
I am keeping it. Not because the obligations are met — they are not — but because the name is what makes them enforceable. Call this a "philosophy of narrative" and nobody ever asks for the numbers. Call it physics and the first honest question is where is the measurement, which is the question that produced everything on this page, including the results that went against me.
A metaphor you can fail is more useful than one you cannot.
Further reading
- Physics of Literature — the definitional page: what the term means and where it came from
- Physics and Literature vs. Physics of Literature — why the distinction in section one matters
- Narrative Entropy (Sn) — canonical definition and version history
- Inter-rater reliability of LLM and rule-based annotation — the three studies behind the measurement scorecard
- Objective Projection Convergence Test v2.0 — the pre-registered protocol and its falsification criteria
- The Habituation Problem — the frequency argument, stated formally
- The Scope Map — what this framework does not claim
- The Bulut Doctrine: Architectural Framework — the primary theoretical reference
Frequently asked questions
What is the Physics of Literature?
An approach that treats a narrative scene as a physical system described by a small set of parameters — light, heat, sound, motion, pressure and spatial geometry — rather than by named emotional states. The parameters are specified in an engineering layer before writing and remain invisible in the finished prose.
Is it actually physics?
Not yet, on the three criteria that matter. Its central quantity cannot yet be applied consistently by independent raters; its falsification test is registered but unrun; and no independent group has replicated any part of it. The name is retained because it makes those obligations explicit rather than because they are met.
What is narrative entropy?
A proposed quantity, Sn = ∫(If × Cb)dt, combining information friction — what a reader needs but has not been given — with causal branching — how many unanswered questions are open simultaneously. Its dimensional consistency is unsettled and it is not offered as an exact formula.
Does the framework claim a deterministic reader response?
No. The predicted convergence is statistical — a distribution across readers, not an identical response in each. Reader state, memory and context all modulate the outcome.
Does it rest on neuroscience?
It draws on work in threat processing, somatic markers and situation-model simulation as analogy and motivation. No physiological measurements have been collected in this work; the pre-registered protocol that would collect them has not been run.
What survives if the physical claims fail?
The prohibitions on naming emotions and using similes function as craft constraints on independent grounds: the most probable phrasing is the most frequently written, and the most frequently written has been habituated. Banning it pushes a writer toward specificity. That argument needs no physics.
Levent Bulut — Independent Researcher. ORCID 0009-0007-7500-2261.