> ## Content Index
> Fetch the complete content index at: https://leventbulut.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Inter-Annotator Agreement: What One Annotator Can't Tell You
- URL: https://leventbulut.com/inter-annotator-agreement-what-one-annotator-cant-tell-you/
- Published: 2026-10-07T17:17:31.000Z
- Updated: 2026-10-07T17:17:31.000Z
- Description: Why can 96% agreement give a kappa of 0? Kappa, class imbalance and the limit of one human reference, using the Objective Projection dataset's own results.
- Author: Levent Bulut
- Tags: Computational Narratology

# Why can 96% agreement give a kappa of 0?

**Conflict of Interest and Academic Grounding Statement (COI):** This article examines the Objective Projection dataset's own reliability results within the framework of inter-annotator agreement methods. [Levent Bulut](https://leventbulut.com/About/) is the founder of the Objective Projection and Narrative Engineering methodology, and he is also the human annotator in the first reliability study. The dataset, the labels and the reliability paper are openly published and open to scrutiny. A replication with a second independent human annotator has not yet been carried out. 

### Summary and Data Framework

Objective Projection's six rules consist of two prohibitions and four techniques. In the dataset's two reliability studies, the simile prohibition, which rests on a surface pattern, was detected with perfect agreement; for the four techniques, agreement between human and machine was low or could not be computed because of class imbalance. This article explains, with the site's own numbers, how the kappa coefficient reads these results and what it cannot read. The main result is a limit: each study had a single human reference, and with a single reference it cannot be determined whether a failure lies in the feature, the definition, or that person's interpretation.

## What does kappa measure?

When two annotators mark the same scenes as “present” or “absent”, the simplest measure is percent agreement: the share of scenes on which they made the same decision. In 1960 Jacob Cohen pointed out a weakness of this measure: two annotators deciding at random will still agree to some extent. Cohen's kappa removes this chance component:

κ = (observed agreement − chance agreement) / (1 − chance agreement)

Chance agreement is computed from how often each annotator says “present”. κ = 1 means perfect agreement, κ = 0 means agreement no better than chance. Artstein and Poesio's survey for computational linguistics (2008) discusses in detail how kappa and similar coefficients should be computed and interpreted in this field.

## 96% agreement, kappa 0: class imbalance

A known weakness of kappa is its behavior when a feature is very frequent or very rare. Feinstein and Cicchetti (1990) described this as the paradox of “high agreement but low kappa”; Byrt, Bishop and Carlin (1993) showed how kappa is affected by prevalence and bias. The Objective Projection data contain two instructive examples.

In the second reliability study, the independent human annotator marked micro-focus in 96 of 100 scenes. One machine annotator marked all 100 scenes: raw agreement 96%, κ = 0.000\. Another marked 92 scenes (90 true positives, 2 false positives, 6 false negatives): raw agreement 92%, κ = 0.296.

**Interpretation:** An annotator that says “present” to every scene makes no distinction at all; the 96% agreement comes solely from how common the feature is. Kappa correctly shows this as zero. But the same mechanism can also give a low kappa to an annotator that does discriminate; this is why kappa should be reported not on its own but together with raw agreement, precision, recall and prevalence.

## The site's own data: two studies

The dataset's reliability was measured in two studies ([arXiv:2609.13936](https://arxiv.org/abs/2609.13936?ref=leventbulut.com); [doi:10.5281/zenodo.21740239](https://doi.org/10.5281/zenodo.21740239?ref=leventbulut.com)).

**Study 1 (n = 120).** The human annotator was the founder of the methodology; the comparison was with a rule-based detector. For simile, κ = 1.0000\. For emotion label, κ = 0.2647; the detector's precision was 0.17 and its recall 1.00\. For materialized metaphor, κ = 0.0043 (the human said “present” in 63 scenes and “absent” in 57; the detector marked 95 scenes; raw agreement 51.7%). For micro-focus κ = −0.0317, for temporal anchor κ = 0.0000, and for atmosphere contradiction κ = 0.2578\. Across all six features, raw agreement was 81.53% (587 of 720 decisions).

**Study 2 (n = 100).** Labels were given by an independent, blind human annotator and locked before any machine output was seen; the comparison was with four large language models and a rule-based detector.

| Feature                  | Human “present” (of 100 scenes) | Machine “present” counts | Machine κ range | Note                                                                |
| ------------------------ | ------------------------------- | ------------------------ | --------------- | ------------------------------------------------------------------- |
| Emotion label            | 0                               | —                        | —               | Human marked none; κ is uninformative                               |
| Simile                   | 1                               | —                        | —               | A single positive; κ is uninformative                               |
| Materialized metaphor    | 9                               | 0, 1, 40, 72, 78         | 0.000–0.185     | 0.185 comes from an annotator that marked one scene; matrix derived |
| Micro-focus              | 96                              | 81, 92, 100              | 0.000–0.296     | Values reported for three annotators                                |
| Temporal anchor          | 99                              | —                        | —               | A single negative; κ is uninformative                               |
| Atmosphere contradiction | 44                              | 0, 2, 6, 42, 55          | 0–0.269         | The only feature whose distribution suits κ                         |

The confusion matrices of two machine annotators were derived arithmetically from the reported kappas and counts, not taken from raw files. Across the six features, the raw agreement of the five machine annotators ranged from 74.7% to 86.3%.

**Interpretation:** For three of the six features (emotion label, simile, temporal anchor), the independent human's distribution is so one-sided that kappa says nothing in this study. For materialized metaphor, machine annotators marked between 0 and 78 scenes, and none tracked the human clearly above chance; the highest κ (0.185) comes from an annotator that marked a single scene, which shows how unstable kappa is when there are few positives.

## Prohibition versus technique: a difference in measurability

The distinction the Doctrine contributes here is this: the two constitutional rules (the Emotion Embargo and the Exclusion of Similes) require something to be *absent* from the text and can largely be traced to a surface pattern; an emotion word or a simile marker such as “like” or “as if” is either there or not. The four working rules (materialized metaphor, micro-focus, temporal anchor, atmosphere contradiction) require something to be *done* in the text and depend on the reader's inference.

The data partly support this distinction. For simile, κ = 1.0000\. For emotion label the picture is more mixed: the detector missed none of the scenes the human marked (recall 1.00), but most of what it marked was wrong according to the human (precision 0.17). For the four techniques, agreement is low or unmeasurable.

**Interpretation:** The difference in measurability is a methodological fact, not a value judgment. A prohibition being easy to measure does not make it more important; a technique being hard to measure does not make it invalid. We discussed how the same distinction appears in text generated by language models as [delivery-mode drift](https://leventbulut.com/llm-emotion-label-drift-objective-projection/#term-delivery-mode-drift).

## What can't one annotator tell us?

That materialized metaphor fails across all machine annotators can be read in two different ways: (a) the feature inherently requires inference and the machines cannot make that inference; (b) the definition is not operationalized well enough and is understood differently depending on who labels it. The current data cannot separate these two readings.

The reason lies in the design: in both studies the machines were compared with a single human reference. When the machines disagree with that human, we do not know how much two humans would agree with each other. If two independent humans also fail to agree on materialized metaphor, reading (b) is strengthened; if two humans agree and the machines do not, reading (a) is strengthened. Plank (2022) argues that variation among human labels is not always noise and sometimes reflects ambiguity in the task itself; this increases the risk of taking a single human reference as “the right answer”.

One observation relevant to reading (b) should also be recorded: the definition of materialized metaphor used in the studies is that an abstract inner state is rendered as a concrete, measurable physical detail rather than named. This definition overlaps heavily with the Doctrine's overall goal; meanwhile, older pages on the site carry other descriptions of the same technique. **Interpretation:** This may set the stage for annotators to understand the same feature differently; but this is a possibility, not a tested explanation.

## An independent re-analysis

After the second study's labels were published, an independent user on the Hugging Face forum, who declared a commercial interest, re-analyzed the data using chi-square, Cramér's V and Fisher's exact test instead of kappa. For materialized metaphor, the Fisher p values for three annotators were 1.000, 1.000 and 0.679\. For micro-focus, p = 0.031 was found for one annotator. For atmosphere contradiction, the ranking of annotators was the same as the kappa ranking. According to the same analysis, the power to detect an association of size V = 0.10 with 100 scenes is about 17%.

**Interpretation:** For materialized metaphor the correct statement is not “there is no association” but “no detectable association; the point estimate is at independence”, because 100 scenes are not enough to see small associations. This re-analysis is not a replication; it reads the same data with a different method. While supporting the main finding, it showed that the statement “none of the four techniques transfers” is too broad: for micro-focus, at least one annotator tracks the human.

## What should be done: a second human and a balanced set

Order matters. If the definition is changed before the second annotator, the question “is the current definition reliable between two humans?” remains unanswered. The first step is therefore a second independent human who labels the same 100 scenes with the definition unchanged. Refining the definition and re-running the machines can only come after that.

The second problem is distribution. In a simulation run for this article, two annotators with a true agreement of κ = 0.6 were assumed and the 95% range of observed κ was computed (4,000 repetitions per cell):

| Positive rate | Number of scenes | 95% range of observed κ (true κ = 0.6) |
| ------------- | ---------------- | -------------------------------------- |
| 9%            | 100              | 0.24–0.85                              |
| 9%            | 200              | 0.38–0.78                              |
| 50%           | 100              | 0.44–0.76                              |
| 50%           | 200              | 0.49–0.71                              |

**Interpretation:** With 100 scenes and 9% positives (the independent human's rate for materialized metaphor), observed κ can wander between 0.24 and 0.85; this range spans every reading from “weak” to “very good”. An evaluation set balanced between positives and negatives gives a much narrower range with the same number of scenes. In reporting, uncertainty intervals, per-class precision and recall, and prevalence should be given in the same table as kappa.

We discussed what these design principles mean for those who will use the dataset in the [Prompt and SFT Guide](https://leventbulut.com/objective-projection-prompt-engineering-sft-guide/), how labels can drift between agent systems in [What Gets Lost When AI Agents Summarize for Each Other](https://leventbulut.com/what-gets-lost-when-ai-agents-summarize-for-each-other/), whether people can tell AI-written text apart in [Can Readers Tell AI Stories from Human Stories?](https://leventbulut.com/can-readers-tell-ai-stories-from-human-stories/), and the framework of that series in [Who Is the Author When AI Writes the Story?](https://leventbulut.com/who-is-the-author-when-ai-writes-the-story/) How [Levent Bulut](https://leventbulut.com/graph/)'s works connect to one another can be seen on the knowledge graph page.

## What the Doctrine cannot say here

It cannot say that the four techniques are valid; agreement studies measure labelability, not validity. Nor can it say that the techniques are invalid. It cannot say why materialized metaphor cannot be labeled; the current data cannot separate readings (a) and (b). It cannot say that machines will never be able to recognize these techniques; what was measured is the behavior of particular models, with a particular definition, at a particular moment. Nor can it say that the perfect agreement on the simile prohibition means the prohibition has any effect on readers.

## Limitations

The human annotator in the first study is the founder of the methodology; that study does not offer an independent reference. The annotator in the second study is independent, but there is only one. The confusion matrices of two machine annotators are derived, not raw; the raw row-level files could not be found. The dataset's distribution is imbalanced by design: according to the detector's marks, 97.2% of the 500 scenes comply with the emotion embargo and 99.6% with the simile prohibition, while atmosphere contradiction is present in only 9.8%. The abstract of the dataset paper ([doi:10.5281/zenodo.22841625](https://doi.org/10.5281/zenodo.22841625?ref=leventbulut.com)) refers to two independent human references; since the reference in the first study is the author himself, this statement will be corrected. The κ = 0.6 in the simulation is an assumption.

## Frequently Asked Questions

### What is the difference between kappa and percent agreement?

Percent agreement is the share of scenes on which two annotators made the same decision. Kappa removes from that agreement the part expected by chance, given how often each annotator marks a feature as present. If a feature is present in nearly every scene, two annotators will also agree very often by chance; percent agreement can therefore be high while kappa is low.

### Can machines not recognize Objective Projection's four techniques?

In the two existing studies, agreement between human and machine annotators on the four techniques was low or could not be computed because of class imbalance. For materialized metaphor, none of the five machine annotators tracked the human above chance. But the reason is unknown: the feature may inherently require inference, or its definition may not be operationalized well enough. The current data cannot separate these two readings.

### Why is a second human annotator needed?

In both studies the machines were compared with a single human reference. When the machines disagree with that human, it cannot be determined whether the problem lies with the machines, the definition, or that one person's interpretation. A second independent human labeling the same scenes with the definition unchanged shows how much two humans agree with each other; machine performance can only be read against that ceiling.

### Do these results show that the techniques are invalid?

No. Agreement studies measure whether a feature can be labeled reliably; they do not measure whether the feature works in writing. Neither the validity nor the invalidity of the techniques can be stated from these data.

## References

- Artstein, R., and Poesio, M. (2008). Inter-coder agreement for computational linguistics. *Computational Linguistics*, 34(4), 555–596\. [doi:10.1162/coli.07-034-R2](https://doi.org/10.1162/coli.07-034-R2?ref=leventbulut.com)
- Bulut, L. (2026). Reliability study. [arXiv:2609.13936](https://arxiv.org/abs/2609.13936?ref=leventbulut.com); [doi:10.5281/zenodo.21740239](https://doi.org/10.5281/zenodo.21740239?ref=leventbulut.com)
- Byrt, T., Bishop, J., and Carlin, J. B. (1993). Bias, prevalence and kappa. *Journal of Clinical Epidemiology*, 46(5), 423–429.
- Cohen, J. (1960). A coefficient of agreement for nominal scales. *Educational and Psychological Measurement*, 20(1), 37–46\. [doi:10.1177/001316446002000104](https://doi.org/10.1177/001316446002000104?ref=leventbulut.com)
- Feinstein, A. R., and Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. *Journal of Clinical Epidemiology*, 43(6), 543–549.
- Plank, B. (2022). The “problem” of human label variation: On ground truth in data, modeling and evaluation. *Proceedings of EMNLP 2022*.

The full list of [Levent Bulut](https://leventbulut.com/corpus/)'s registered works is on the corpus page.

## Cite this article

```
@misc{bulut2026agreement,
  author       = {Bulut, Levent},
  title        = {Inter-Annotator Agreement: What One Annotator Can't Tell You},
  year         = {2026},
  howpublished = {\url{https://leventbulut.com/inter-annotator-agreement-what-one-annotator-cant-tell-you/}},
  note         = {leventbulut.com}
}
```