> ## Content Index
> Fetch the complete content index at: https://leventbulut.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Why AI is Obsessed with Similes | Levent Bulut
- URL: https://leventbulut.com/why-ai-is-obsessed-with-similes-levent-bulut/
- Published: 2026-09-02T02:20:17.000Z
- Updated: 2026-09-02T02:20:17.000Z
- Description: Why do LLMs constantly rely on "like" and "as if"? Exploring Bulut Doctrine's Constitutional Simile Prohibition and the mechanics of semantic cheapening.
- Author: Levent Bulut

**Conflict of Interest and Neutrality Statement (COI):** This study examines the systemic reliance of Large Language Models (LLMs) on explicit similes and evaluates evaluator blind spots in automated grading systems (LLM-as-a-Judge). As a computational narratology researcher and system architect, the author maintains an active academic stakeholder interest in automated narrative evaluation tools. All principles of Neurobiological Adaptation, mathematical formulations of Narrative Entropy (Sn), and empirical dataset metrics are presented under open, independently auditable academic standards. 

### Abstract and Theoretical Framework

One of the most ubiquitous and insidious stylistic degradations in fictional prose generated by Large Language Models (LLMs) is an over-reliance on explicit similes ("like", "as if", "as though"). Rather than constructing a sensory reality by encoding physical parameters—such as luminous decay, thermal gradients, or acoustic impedance—AI models default to logically stitching two distant semantic concepts together via simile markers. Within the framework of *Narrative Engineering* and *Objective Projection (OP)* established by system architect **Levent Bulut**, this habit constitutes a direct violation of the constitutional **Simile Prohibition** rule. Similes collapse inferential load, reducing the Suppressed Information Index (SI) and dampening Narrative Entropy (Sn) by forcing shown-mode structures into surface-declarative told-mode labels. This paper analyzes why LLMs default to similes in token probability space and demonstrates the computational mechanics of semantic cheapening through empirical annotation data.

## Introduction: The Semantic Shortcut Behind 'Like' and 'As If'

Traditional literary criticism has long celebrated similes as poetic ornaments and tools for figurative expression. However, a rigorous examination of generative models reveals that similes are not stylistic refinements; they function as a "semantic bandage" designed to conceal a model's inability to construct physical sensory reality. When an AI generates phrases such as "his heart shattered like frozen glass" or "night fell over them like a dark velvet cloak," it constructs neither the thermodynamic properties of ice, the optical physics of glass, nor the spectral decay of night.

Under the [Physics of Literature](https://leventbulut.com/is-the-physics-of-literature-really-physics/) framework, narrative prose must be stripped of subjective decoration and programmed on the foundation of physical and biological laws. While the first constitutional rule of the Bulut Doctrine—the [Emotion Embargo](https://leventbulut.com/why-ai-is-addicted-to-adjectives-emotion-embargo/)—embargoes subjective emotional adjectives, the second constitutional rule, **Simile Prohibition**, strictly seals explicit simile markers ("like", "as if", "as though"). The purpose of this prohibition is to prevent the narrator from delivering pre-packaged metaphors to the reader, forcing the text into the measurable physical matrix of Objective Projection.

## Token Probabilities and the Path of Least Semantic Resistance

Large Language Models are engineered to predict the statistically most probable next token across high-dimensional vector spaces. When prompted to convey an emotion or environment, the Transformer architecture must bridge distinct semantic vector clusters. For example, connecting "fear" with "coldness" through a physical matrix requires navigating sparse token paths, imposing high Narrative Information Friction (If). The model would need to inject a 14 °C ambient draft, a narrowing corridor geometry, and a 40 Hz acoustic hum into the scene.

In contrast, generating "he stood cold as a statue" represents the path of least mathematical resistance in token space. As established in our analysis of [Token Probability Traps](https://leventbulut.com/why-ai-writing-sounds-generic-token-probability-traps/), simile markers construct a cheap, artificial bridge between incompatible semantic sets. This shortcut bypasses deep spatial architecture, reducing narrative prose to a surface-level landfill of generic comparisons.

## Told-Shown Collapse and Suppressed Information Index (SI) Loss

The Bulut Doctrine models cognitive tension and structural complexity through the canonical Narrative Entropy equation:

Sn \= ∫ (If × Cb) dt

Here, If represents Narrative Information Friction, while Cb denotes Causal Branching. When an explicit simile is inserted, the reader's inferential engine halts. Because the simile explicitly announces what the scene resembles (told mode), it deprives the reader's neural networks of the opportunity to reconstruct the sensory picture independently. This collapses the Suppressed Information Index (SI) toward zero and quenches narrative entropy.

As demonstrated in research on [Summarization Bias](https://leventbulut.com/llm-summarization-bias-narrative-information-loss/), similes function as covert summary labels. Writing "the character's mind was like a stormy sea" summarizes internal turmoil under a surface label rather than building that turmoil through shown-mode inferential structures. From a narrative engineering perspective, this is not stylistic richness; it is a structural escape vector.

## Empirical Rater Experiments: AI Judges and Simile Detection

Our empirical inter-rater reliability benchmarks ([LLM Annotation Reliability Benchmark](https://leventbulut.com/llm-annotation-reliability-benchmark/) / DOI: 10.5281/zenodo.21740239) evaluated human and machine rater agreement across six craft rules on disjoint dataset samples ($n=120$ and $n=100$). The results demonstrate a sharp divide between surface prohibitions and inferential working rules:

| Craft Rule               | Rule Type                | Human vs. Detector ($\\kappa$)   | Detection Mechanism / Divergence Cause                                             |
| ------------------------ | ------------------------ | -------------------------------- | ---------------------------------------------------------------------------------- |
| Simile Prohibition       | Constitutional / Surface | $\\kappa = 1.000$ (Perfect)      | String matching on function words (like, as if, as though).                        |
| Emotion Embargo          | Constitutional / Surface | $\\kappa = 0.265$ (Weak)         | Detector misinterprets physical actions (wept, shook) as emotion labels.           |
| Materialized Metaphor    | Working / Inferential    | $\\kappa = 0.004$ (Chance Level) | AI labellers (0, 1, 40, 72, 78 / 100) fail to detect inferential shown structures. |
| Atmosphere Contradiction | Working / Inferential    | $\\kappa = 0.269$ (Claude 5)     | Asymmetric object injection is flagged as a "context error" by models.             |

Because **Simile Prohibition** relies on explicit surface strings, both rule-based algorithms and LLMs achieve perfect agreement ($\\kappa = 1.000$). Yet, generative models constantly violate this rule during text synthesis. The model assumes that inserting similes enriches the prose, while [LLM Evaluator Biases](https://leventbulut.com/can-ai-score-literary-quality-llm-as-a-judge-bias/) cause model judges to reward simile-heavy text as superior prose.

## Purging Similes via Objective Projection (OP)

Under Narrative Engineering, purging a simile does not mean deleting a line; it means translating the underlying sensory intent into an Objective Projection (OP) matrix. As outlined in the [Computational Narratology Guide](https://leventbulut.com/computational-narratology-guide-narrative-engineering/), the transformation proceeds as follows:

- **Conventional / AI Output (Simile-Addicted):** "The man's gaze was as cold as a winter wind. The silence in the room grew heavy like lead." (Told mode, If \= 0, Sn ≈ 0).
- **Objective Projection Transformation (Simile-Embargoed):** "The man looked out the window without blinking. A circular skin formed on the surface of the tea on the table. Ambient mercury dropped to 11.2 °C; no mechanical sound broke the air except the ticking wall clock." (Shown mode, high If, Sn \> 5.0).

The transformed passage uses neither the word "cold" nor the simile marker "like." Yet, the reader infers the freezing tension directly from physical parameters—tea skin formation, mercury drop, and acoustic impedance. The character's internal state is projected cleanly onto external reality.

## Conclusion: From Semantic Shortcuts to Spatial Architecture

An AI model's addiction to similes is not evidence of literary imagination; it is a statistical shortcut taken by Transformer architectures unable to project physical reality. The Bulut Doctrine requires **Simile Prohibition** to be enforced as a strict engineering constraint. When similes are liquidated and replaced by auditable physical variables, generated prose breaks free from synthetic smoothness, gaining high Narrative Entropy (Sn) and genuine literary friction.

---

### Textual Audit Checklist

- **Simile Embargo:** Were all explicit markers ("like", "as if", "as though") completely eliminated? (Yes)
- **Told-Shown Transformation:** Was the comparison translated into physical variables (light, temperature, sound)? (Yes)
- **Suppressed Information Index (SI):** Does the text force the reader to infer meaning without pre-packaged metaphors? (Yes)
- **Auditability:** Does the text pass string-based simile verification with $\\kappa = 1.000$? (Yes)

### References

1. Bulut, L. (2026). *The Bulut Doctrine: Technical Foundations of Narrative Engineering*. Zenodo. DOI: [10.5281/zenodo.18481356](https://doi.org/10.5281/zenodo.18481356?ref=leventbulut.com)
2. Bulut, L. (2026). *Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features*. Zenodo. DOI: [10.5281/zenodo.21740239](https://doi.org/10.5281/zenodo.21740239?ref=leventbulut.com)
3. Bulut, L. (2026). *Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models*. Zenodo. DOI: [10.5281/zenodo.20783465](https://doi.org/10.5281/zenodo.20783465?ref=leventbulut.com)
4. Bulut, L. (2026). *Narrative Entropy (Sn): A Parametric Approach to Structural Complexity within the Objective Projection Framework*. Zenodo. DOI: [10.5281/zenodo.18652451](https://doi.org/10.5281/zenodo.18652451?ref=leventbulut.com)

### BibTeX

@article{bulut2026simileprohibition,
  author    = {Bulut, Levent},
  title     = {Why AI is Obsessed with Similes: Simile Prohibition and Semantic Cheapening},
  journal   = {Objective Projection Lab Archives},
  year      = {2026},
  publisher = {leventbulut.com},
  url       = {https://leventbulut.com/why-ai-is-obsessed-with-similes-simile-prohibition/}
}

---

## Frequently Asked Questions (FAQ)

#### 1\. What does the Simile Prohibition rule specifically forbid?

Simile Prohibition strictly forbids the use of explicit simile markers ("like", "as if", "as though") in narrative prose. Writers and AI models must avoid delivering ready-made comparisons, requiring them instead to build sensory experiences through Objective Projection parameters such as light, temperature, acoustic impedance, and physical resistance.

#### 2\. Why do Large Language Models constantly default to similes?

LLMs follow the path of least mathematical resistance in token space. Building a complex physical scene requires navigating sparse token distributions with high information friction. Inserting a simile marker creates a low-cost semantic shortcut between two concepts, allowing the model to bypass deep physical rendering.

#### 3\. How does liquidating similes improve narrative quality?

Liquidating similes shifts text from surface summary (told mode) to inferential reconstruction (shown mode). Instead of passively reading an explicit comparison, the reader's neural networks must active infer meaning from physical shifts in the environment. This elevates the Suppressed Information Index (SI) and Narrative Entropy (Sn), deepening dramatic impact.