A Caption Schema for Video Models: Objective Projection

Synthetic captions help video models follow prompts. A caption schema derived from Objective Projection that writes only what is visible, and its limits.

Share
A Caption Schema for Video Models: Objective Projection
How AI Video Captioning Works: Observable Evidence vs Hallucination

Objective Projection for AI Video: How Synthetic Captions Improve Prompt Following

Summary and Data Framework

Image and video generation models learn from image–text pairs in their training data; what a caption says largely determines what the model will learn. OpenAI's DALL-E 3 and Sora work reported that replacing short, noisy captions with detailed synthetic captions improves how well models follow prompts. Starting from that finding, this article proposes a caption schema derived from Objective Projection principles. The schema's basic rule comes from the Doctrine's own principle: write only what is seen and heard in the clip; do not write emotion names, inferred temperatures or inferred times of day. Captioning models are prone to inventing things they do not see.

A model learns what the caption says

A text-to-video model learns from millions of pairings of video clips and the texts that describe them. It learns what kind of image the phrase "a rainy street" corresponds to by looking at which images that phrase appears with in the training data. How detailed and accurate a caption is, and what it emphasizes, therefore directly affects how the model will later respond to prompts.

Scenes with emotion pose a particular problem here. A caption describing such a scene can take one of two routes: an emotion label such as "a sad man waits," or an observable description such as "a man in a wet raincoat sits motionless on a plastic chair." The first tells the emotion; the second shows it. The distinction that Objective Projection, developed by Levent Bulut, proposes for writing is carried here into the language of training data.

Re-captioning: Training with synthetic descriptions

The DALL-E 3 technical report published by Betker and colleagues in 2023 hypothesized that one reason text-to-image models struggle to follow detailed descriptions is noisy and inaccurate captions in the training data. The researchers relabeled the training set with a captioning model trained to describe images in detail, and reported that training on these synthetic captions reliably improves the models' ability to follow prompts.

The Sora technical report that OpenAI published in February 2024 stated that it applied the same re-captioning method to video: a highly descriptive captioning model was trained first, and then used to produce captions for all videos in the training set. The report states that training on highly descriptive video captions improves both text fidelity and the overall quality of the videos. The same report says that users' short prompts are also turned into longer, detailed captions by a language model before being sent to the video model. Sora itself is no longer in use; according to OpenAI, its web and app experiences were discontinued on April 26, 2026. But the method became common ground for much subsequent work in the field. Sora's shutdown and today's tools are discussed in Showing Emotion in AI Video.

More detail, more risk of invention

Detailed captions come at a cost. In 2018, Rohrbach and colleagues showed that captioning models are prone to describing objects that are not actually in the scene, that is, to "hallucination." In one of the paper's examples, a model added a "bench" that was not in the scene to its caption. The researchers found that some of these errors stem from misclassifying the image and some from language priors. They also reported that the models scoring highest on standard captioning metrics were not always those producing the least hallucination.

This finding is a critical warning for a caption schema based on Objective Projection. Some of the Doctrine's six parameters cannot be seen directly in a video frame. Temperature is invisible; only its consequences are visible: a fogged window, sweat on a forehead, breath turning to vapor. The hour is invisible; only a clock on the wall or light that marks the time of day is visible. Sound exists only if the clip has an audio track. Asking a captioning model to "fill in" these parameters may push it to guess what it does not see from language priors, that is, to invent.

The proposed schema: Write only what is seen and heard

The basic rule of the proposed schema is therefore the caption counterpart of Objective Projection's Output Layer principle: a caption contains only what is seen and heard in the clip. Interpretation, emotion names and inferred values do not go into the caption text. The table below shows how the Doctrine's principles are adapted under this rule.

Objective Projection principleGoes into the captionStays out of the caption
Emotion EmbargoVisible action and body: "presses his hand against the rim of the cup"Emotion names: "sad," "anxious"
Exclusion of SimilesThe object itselfSimiles of the "like a ghost" kind
LightVisible light source, color, flicker—
HeatVisible consequences: fog, sweat, vaporInferred temperature value
SoundWhat is heard on the audio trackAny sound if there is no audio track
Temporal AnchorA clock in the frame, light that marks the time of dayInferred hour
Micro-FocusThe object, if the camera closes in on a single object—
Materialized MetaphorThe visible state of the object: "untouched tea with a skin on it"The object's meaning: "a cup symbolizing loss"
Atmosphere ContradictionAn everyday action visible in the backgroundThe interpretation that this action is "indifferent"

Below is an example of what the schema would look like for a single clip. The structured fields make it traceable which observations the caption text was built from. The emotion label is not dropped but moved to a separate field: it is not used in model training, but it is needed later to measure whether the model can build the emotion in the image.

{
  "clip_id": "example_001",
  "caption": "A man in a wet raincoat sits motionless on a plastic chair in a dim hospital corridor. A single fluorescent tube above him flickers. A cleaner pushes a floor polisher past him and exits at the far end. Close-up: his thumb presses a small dent into the rim of a paper cup; a thin skin has formed on the tea.",
  "observable_fields": {
    "light": "single flickering fluorescent tube; rest of corridor dim",
    "visible_heat_cues": null,
    "audio": "constant electrical hum; distant rhythmic beep",
    "time_cues": null,
    "focus_object": "paper cup of tea",
    "background_action": "cleaner pushing a floor polisher"
  },
  "excluded_from_caption": ["emotion labels", "inferred temperature", "inferred time of day", "interpretation of objects"],
  "evaluation_only": {
    "annotator_emotion_label": "<emotion entered by an independent annotator>",
    "annotator_id": "<annotator code>"
  }
}

In this example the visible_heat_cues and time_cues fields are left empty, because the clip contains no cue showing temperature or time. A captioning model should not be forced to "fill in" these fields. An empty field is more accurate than an invented value.

One more summarization risk

Even a detailed caption is never exhaustive; the captioning model itself decides what to write and what to leave out. This is a risk related to the hypothesis the Bulut Doctrine calls Summarization Bias: when summarizing a scene, language models may drift toward conclusion labels rather than observed details. How the same risk might appear in handoffs between AI agents is discussed in What Gets Lost When AI Agents Summarize for Each Other?, and its appearance in fiction generation in Beyond Hallucination. Even a captioning model explicitly instructed not to write emotion labels does not necessarily lose this tendency entirely; a sample of the generated captions should therefore be checked by humans.

How could the schema be tested?

The proposal's central question is this: does a model trained on captions written with observable details instead of emotion labels learn to build emotion in the image more convincingly? Testing this could take two stages.

The first stage checks the captions themselves: two caption sets are produced for the same clip collection, one free and one following this schema. Independent human checkers mark whether the details in each caption are actually present in the clip; this measures whether the schema increases or decreases hallucination. The second stage tests the main question: the same small model is fine-tuned with the two caption sets, the generated videos are shown to viewers who do not know which set was used, and the viewers are asked to identify the emotion in the scene. The second stage requires substantial computing resources; the first can be carried out by a small team.

What the Doctrine cannot say in this area

This article is a schema proposal; the schema has not been used to train any model. There are no data showing that captions written with observable details give better results in emotional scenes. No caption schema can produce "hallucination-free" data; the schema only aims to narrow the room left for invention. The findings in the DALL-E 3 and Sora reports concern descriptiveness in general; they contain no result specific to emotional scenes or to Objective Projection. Criticisms of the Doctrine are collected on a separate page.

Limitations

This article is a proposed method, not an empirical study. The DALL-E 3 and Sora reports are company technical reports and have not gone through peer review; their detailed experimental conditions are not public. The object hallucination study examined 2018 captioning models; current models may have different rates. The field names and structure of the proposed schema are illustrative. Keeping the emotion label only in an evaluation field requires the label to be entered by independent human annotators; the cost of this and inter-annotator agreement must be planned separately.

Frequently Asked Questions

What is re-captioning?

It is the method of rewriting the short, noisy descriptions in an image or video training set with a captioning model trained to produce detailed descriptions. In the DALL-E 3 and Sora technical reports, OpenAI reported that training with this method improves how well models follow prompts.

What is the basic rule of the caption schema based on Objective Projection?

A caption contains only what is seen and heard in the clip. Emotion names, similes, inferred temperatures or hours, and meanings attributed to objects do not go into the caption text. If an emotion label is needed, it is kept in a separate evaluation field that is not used in training.

Do more detailed captions increase hallucination?

They can. Captioning models have been shown to be prone to describing objects not present in the scene, and some of these errors stem from language priors. The schema therefore forbids filling in parameters that cannot be seen, and prefers an empty field to an invented value.

Has this schema been proven to work?

No. The schema has not been used to train any model. First the accuracy of the captions, and then the performance in emotional scenes of models trained on those captions, would need to be tested with independent evaluators.

References

How to cite this article

You can use the BibTeX record below to cite this article. The author's other registered works are listed on the Levent Bulut corpus page; for more about the author, see the Levent Bulut page.

@misc{bulut2026captionschema,
  author       = {Bulut, Levent},
  title        = {How Should Captions for Video Models Be Written? The Re-Captioning Literature and a Schema Based on Objective Projection},
  year         = {2026},
  month        = oct,
  howpublished = {\url{https://leventbulut.com/objective-projection-caption-schema-video-models/}},
  note         = {Bulut Doctrine, Computational Narratology},
  language     = {english}
}
G-Verified: Levent Bulut