Showing Emotion in AI Video: Objective Projection

How do you show emotion in text-to-video tools? Prompt design through light, sound, motion and space with Objective Projection, and its limits.

Share
Showing Emotion in AI Video: Objective Projection
Cinematic Nostalgia: Emotional Storytelling in a Vintage Train Station

Showing Emotion When Making Film with AI: Objective Projection and Text-to-Video Prompt Design

Summary and Cinematographic Framework

A video model cannot put the sentence "she was terrified" on screen; it has to decide what will appear in the frame. Cinema is therefore by nature a "showing" medium, and Objective Projection's principle of building emotion through light, sound, motion and space translates almost directly into video prompts. This article shows that translation with a table and a sample prompt; but it also states its limits plainly: video models still struggle to produce physically consistent scenes, and it has never been measured whether prompts written this way evoke a stronger emotion in viewers.

A camera cannot say "she was afraid"

A novelist can write "The woman was terrified"; the reader reads the sentence and moves on. A director cannot film that sentence. There has to be a face, a room, a light and a sound in front of the camera. Text-to-video AI models are in the same position: when you write "a frightened woman" in a prompt, the model has to guess how to turn that abstract label into an image, and the result is often a stock one: widened eyes, an open mouth, an exaggerated expression.

Cinema has been thinking for a hundred years about how emotion arises from images. One of the most frequently cited examples of that thinking is the Kuleshov effect. According to the story, the Soviet director Lev Kuleshov edited an actor's expressionless face together with different images, such as a bowl of soup, a coffin or a child playing, and viewers said they saw a different emotion in the same face each time. Because the original film is lost, whether the effect really exists was debated for a long time. A replication attempt by Prince and Hensley in 1992 did not find it. A study by Barratt and colleagues in 2016, using a more careful design, found that viewers tended to attribute emotions congruent with the context to expressionless faces.

The shared lesson of these findings is this: in an image, emotion is often read not from the face itself but from the context around it. But the effect is not mechanical; it varies with design, context and viewer. This is also what Objective Projection proposes for writing.

Tools change, the method remains

This article is not about a particular tool. The most striking example of how fast this field changes is OpenAI's Sora: according to the company's help center, Sora's web and app experiences were discontinued on April 26, 2026, and the API was scheduled to be discontinued on September 24, 2026. Today tools such as Google's Veo, Runway and Kling stand out in this field; tomorrow the list may change again.

What does not change is the question of what a scene is built from on screen. The Sora 2 prompting guide that OpenAI published in October 2025 remains instructive in how it answers that question, even though the product has been shut down. The guide likens writing a prompt to briefing a cinematographer who has never seen your storyboard. It says that details you leave out (time of day, weather, wardrobe, camera angle) will be invented by the model, that specific visible descriptions should be used instead of vague adjectives, that repeating character descriptions with the same wording across shots matters for continuity, and that focusing on one camera move and one action per shot gives better results.

These recommendations are very close to the principle that Objective Projection, developed by Levent Bulut, proposes for writing: observable physical detail instead of an abstract label. What Objective Projection can add here is to turn this intuition into a systematic checklist rather than scattered tips.

From Objective Projection to a video prompt

Objective Projection rests on two prohibitions and four techniques. The prohibitions are that emotion is not named (Emotion Embargo) and that similes are not used (Exclusion of Similes). The techniques are Materialized Metaphor, Micro-Focus, Temporal Anchor and Atmosphere Contradiction. The method's dataset also codes each scene along six physical parameters: light, heat, sound, motion, atmosphere and space. The table below shows what each of these can correspond to in a video prompt.

Objective Projection principleCounterpart in a video prompt
Emotion EmbargoNo emotion words in the prompt. Instead of "a sad face," the actor's physical action: lifting a glass and setting it down without drinking.
Exclusion of SimilesNo similes of the "like a ghost" kind. A simile describes nothing the camera can record; the model may ignore it or try to place it on screen literally.
Materialized MetaphorA concrete object that carries the emotional weight: a coat still hanging on its hook, a second cup going cold on the table.
Micro-FocusAt the critical moment, a close-up (insert) on a single object. One action and one camera move per shot.
Temporal AnchorA clock on the wall, light that marks the time of day, a detail that marks the season.
Atmosphere ContradictionA background indifferent to the character's state: a television from the next room, laughing children passing the window. In models that generate sound, described as diegetic sound.
LightLight source, direction, color and change: a single fluorescent tube, a flickering bulb, a window going dark.
HeatVisible consequences of heat: a fogged window, sweat on a forehead, breath turning to vapor.
SoundSound design: a single constant hum, one sound that breaks the silence.
MotionActor and camera motion: a hand that stays still, a slow push-in.
AtmosphereThe visible quality of the air: dust, fog, smoke, particles hanging in a beam of light.
SpaceFraming and space: a narrow corridor, a single exit, the distance between two characters.

Example: The same scene, two prompts

Scenario: A man waits in a hospital corridor for news of his wife, who is in surgery.

Prompt written with labels:

A sad and anxious man waits in a hospital corridor, feeling
hopeless, like a ghost. Emotional, cinematic.

This prompt contains three emotion labels, a simile and two vague adjectives. The model has to guess the camera angle, the lighting, the time, the sound and what the actor is doing.

Prompt written with Objective Projection principles:

Scene: A hospital corridor at 3:40 a.m. A man in his fifties sits
on a plastic chair, still wearing a wet raincoat. One fluorescent
tube above him flickers; the rest of the corridor is dim.

Cinematography: Static wide shot, eye level, deep focus. Cut to a
close-up insert of his hand holding a paper cup of tea.

Action (one per shot):
- Shot 1: He does not move. A cleaner pushes a floor polisher slowly
  past him and disappears at the far end.
- Shot 2: The tea in the cup is still. A thin skin has formed on its
  surface. His thumb presses a small dent into the rim.

Sound: Constant hum of the fluorescent tube. Distant rhythmic beep
of a machine. The polisher's motor fades out.

No emotion is named in this prompt. Temporal Anchor (3:40), Atmosphere Contradiction (the cleaner passing by indifferently), Micro-Focus (the hand holding the cup), Materialized Metaphor (the untouched tea with a skin on it) and the light, sound and motion parameters are all described as things the camera can record. Building the emotion is left to the viewer.

Consistency: What can a prompt solve, and what can it not?

The most frequent problem when building a multi-shot scene with AI is consistency: the character's face, the costume or the light of the location change from shot to shot. A prompt can do something about this, but not everything. Repeating the character and location description with the same words in every shot, keeping the logic of the light fixed, and using reference images where the tool supports them can help improve consistency. OpenAI's guide also warns that small changes in wording can alter the character's identity or the focus of the scene.

Whether a character's identity is preserved across shots, however, depends largely on the model's own capability. A well-written prompt does not extend the limits of that capability; it only leaves the model less room for guessing.

A physical prompt is not a physical result

Objective Projection's reliance on physical detail runs into an important limit with video models: these models cannot always render the physical world correctly. In the VideoPhy study, Bansal and colleagues tested the leading models of the time with 688 prompts describing interactions between materials; according to human evaluation, the best model produced a video that followed both the prompt and physical commonsense in only 19.7 percent of cases. The PhyGenBench study by Meng and colleagues likewise reported that models struggle to produce videos consistent with physical commonsense, and that problems involving dynamics are not solved simply by scaling up the model or by prompt engineering.

Interpretation: These studies evaluate 2024 models, and later models have probably improved. But the lesson still holds: writing "a skin has formed on the tea" in a prompt does not mean the model will show it convincingly. Objective Projection tells the model what to show; whether the model can show it is a separate question.

What is not yet known

The central claim of the method proposed in this article has not been tested. There is no measurement showing that video prompts written with Objective Projection principles evoke the intended emotion in viewers more reliably than prompts written with emotion labels. The question is not hard to test: the same scenes are generated with both types of prompt, the videos are shown to viewers who do not know which prompt produced them, and the viewers are asked to identify the emotion in the scene. Until such a study is done, this article should be read as a proposed method.

The counterpart of the same problem in writing, the tendency of language models to state emotion rather than build it, was discussed in Beyond Hallucination. How stories are transformed in the move from book to film is examined in Why Are Book Adaptations So Addictive?. Criticisms of the Doctrine are collected on a separate page.

Limitations

This article is a proposed practice, not an empirical study. The proposed prompt structure has not been tested systematically on any video model. The capabilities, names and access conditions of tools change quickly; the tools mentioned reflect the situation on the publication date. The physical commonsense studies evaluated 2024 models. Findings on the Kuleshov effect are mixed; context has been shown to affect how a facial expression is read, but the size of the effect varies with design. Objective Projection was developed for writing; this adaptation to video is the author's proposal.

Frequently Asked Questions

How should emotion be written in AI video prompts?

Instead of naming the emotion, describing physical details that the camera can record gives more control: light, sound, the actor's action, objects and space. The model guesses the details you leave out; an abstract emotion label often turns into a stock expression.

Why is Objective Projection a suitable method for video?

Because a camera can only record what can be seen and heard. Objective Projection's principle of building emotion through light, heat, sound, motion and space matches what a video prompt has to describe anyway. Whether this method evokes a stronger emotion in viewers has not yet been measured.

Can Sora still be used?

No. According to OpenAI's help center, Sora's web and app experiences were discontinued on April 26, 2026, and the API was scheduled to be discontinued on September 24, 2026. The principles in this article are not tied to any particular tool.

Does a good prompt keep a character the same across shots?

Not on its own. Repeating the character description with the same words in every shot and using reference images where supported can help, but preserving identity depends largely on the model's own capability.

References

How to cite this article

You can use the BibTeX record below to cite this article. The author's other registered works are listed on the Levent Bulut corpus page; for more about the author, see the Levent Bulut page.

@misc{bulut2026aivideo,
  author       = {Bulut, Levent},
  title        = {Showing Emotion When Making Film with AI: Objective Projection and Text-to-Video Prompt Design},
  year         = {2026},
  month        = sep,
  howpublished = {\url{https://leventbulut.com/objective-projection-ai-video-prompting/}},
  note         = {Bulut Doctrine, Computational Narratology},
  language     = {english}
}
G-Verified: Levent Bulut