Skip to content

Dialogue Scene

Style example
Best for:
Short film scenesExplainer hostsCharacter vignettesAd spokespeoplePreviz table reads

Character-driven scenes with synchronized, lip-synced speech and natural performance. Only Veo 3.1 and Sora 2 generate native dialogue — write the exact lines in quotation marks, keep them short, and direct the delivery like a script.

Veo 3.1 handles multi-person conversations with distinct voices. Attribute each line clearly and specify tone.

Two-Person Exchange Drama
Cinematic medium two-shot in [ENVIRONMENT — e.g., a dim diner booth at night], warm practical light overhead, rain on the window behind. [CHARACTER A — e.g., an older detective] leans forward and says wearily, "[LINE A — keep under 10 words]." [CHARACTER B — e.g., a nervous young man] looks away and replies quietly, "[LINE B]." Shallow depth of field, subtle handheld drift. Ambient noise: rain, distant dishes clinking, low diner hum.
Direct-to-Camera Host Explainer
Medium shot of [HOST — describe age, style, energy] standing in [ENVIRONMENT — e.g., a bright modern studio with soft gradient backdrop], speaking directly to camera with warm confident energy: "[SCRIPT LINE — 15-20 words max for 8 seconds]." Natural hand gestures synchronized to the emphasis. Soft key light, gentle background bokeh, 16:9. Ambient noise: clean room tone, subtle upbeat music bed underneath.
Emotional Close-Up Performance
Slow push-in to a tight close-up of [CHARACTER] in [ENVIRONMENT], eyes glistening, jaw tightening. After a long beat they whisper, "[LINE — under 8 words]," and exhale. Motivated window light on one side of the face, the other falling into shadow. Filmic grain, muted palette. SFX: room tone, a faint clock ticking, no music.
  • Budget ~2.5 words per second: an 8s Veo clip fits roughly 15–20 spoken words total. Overstuffed scripts get rushed or cut off.
  • Direct the read: “wearily,” “briskly,” “whispers” — delivery adjectives steer the voice performance.
  • Attribute every line: “A says… B replies…” prevents voice-swapping between characters.
  • Silence is a tool: A beat before or after a line (“after a long pause”) makes performances feel human.
  • No-audio platforms: Generate clean mouth movement and dub in post; keep the shot medium or closer so lip sync tools can track.