Skip to content
Updated for 2026

The reliable way is one speaker per clip: multi-person dialogue attribution is fragile because the model binds a quoted line to whichever character description sits nearest it, not to the person you meant. When you must keep two people in frame, name the speaker by visible appearance immediately before the quote and state that the other person stays silent.

Make the Right Character Speak: Fixing Multi-Person Dialogue in AI Video

You write a line for the man in the grey shirt and the woman delivers it, lips moving in sync with words you meant for someone else. Or the line splits down the middle — each person mouths half of it, both perfectly synced to audio neither of them should be saying. This is one of the most common failures in multi-character AI video, and it happens because of how the model reads a prompt: not as a screenplay with character columns, but as prose. It binds spoken words to whichever character description sits closest to the quoted line. With one person in frame that's rarely ambiguous. With two or more, that binding gets fragile fast. The fix is mostly about how you structure the prompt, but it has a real limit — know that limit before you burn a render on it.

Make a video from one prompt. $2.99 a video · Ready in about a minute · No subscription.

Key facts

Root cause Proximity binding The model attaches a quote to whichever character description sits closest to it, not the one you intended
Strongest fix One speaker per clip Two people talking is two clips — this is the single biggest reliability gain available
Attribution cue Visible appearance "The man in the grey t-shirt says" beats a name, a role, or a position in frame
Format to avoid Screenplay style NAME: line and parentheticals read as prose to the model — the colon doesn't bind the way it does on a page
Reliability ceiling Not guaranteed Even a correct prompt can still misattribute — regenerating is a legitimate second try

How to Use MakeThisVid

From prompt to downloadable MP4, ready to deploy.

  1. Give each speaker their own clip

    This is by far the strongest lever, and the one people resist because it feels like more work. Two people talking is two clips, not one. Attribution reliability drops sharply the moment a second mouth is in the frame — if the dialogue is the point, isolating the speaker removes the ambiguity entirely instead of fighting it.

  2. Attribute by what they look like, not who they are

    Identify the speaker by visible appearance immediately before the quote: "The man in the grey t-shirt says, '...'". Appearance outperforms names, seating position, and role labels like "the manager" — it's the description the model has the easiest time matching back to a specific person already rendered in the scene.

  3. Keep the attribution and the quote adjacent

    Don't describe the room, the lighting, or the background between naming the speaker and giving the line. Every sentence of scene description you insert between the two is a chance for the binding to slip onto the wrong person. Name the speaker, then quote them, with nothing in between.

  4. Say the other person stays silent

    State it directly: "the woman beside him listens without speaking." Leaving it implied gives the model room to assign her a line of her own, or to split your intended line between the two of them.

  5. Skip screenplay formatting

    NAME: line, with parentheticals for tone or action, reads as prose to the model, not as a script cue. The colon doesn't carry the binding people expect it to. Write it as a described action instead: someone says something, in a full sentence.

  6. Know when to stop prompting and regenerate

    Even done exactly right, multi-character attribution is not fully reliable — it varies between runs of the identical prompt. If a clean render comes back with the line on the wrong face, that's not a prompt you failed to write well enough. Regenerate before you spend more time rewriting a prompt that was already correct.

Who Uses MakeThisVid for This

Two-person testimonial or interview

Split the question and the answer into separate clips rather than trying to hold both people in one shot with alternating lines. Stitch them together after.

Product demo with a narrator on screen

If a second person appears in the background, describe them as present and not speaking — don't leave their silence to be assumed.

Scripted exchange between two characters

Generate each side of the exchange as its own clip, attributing the speaker by appearance each time, then edit the two together for the back-and-forth.

Frequently Asked Questions

The model binds a quoted line to whichever character description sits closest to it in the prompt, not necessarily the person you intended. With more than one person in frame, that binding is fragile and can land on the wrong face.
Technically yes, but reliability drops sharply once a second person is in frame. If the dialogue matters, generate each speaker in their own clip instead — it is the single most effective fix available.
Describe how they look, immediately before the quote: "the woman with short red hair says, '...'". Visible appearance binds more reliably than a name, a role, or where someone is standing.
Appearance works better than a name for binding a line to the right person. And if the name belongs to a real person, that introduces a separate problem — see the page on using real names in AI video prompts.
No. The model reads screenplay formatting as prose, not as a script cue — the colon doesn't carry the binding a reader expects. Write the attribution and quote as a normal descriptive sentence instead.
Yes. State it explicitly, such as "the man beside her listens without speaking." Leaving it unstated leaves room for the model to give the second person a line of their own, or to split the intended line between both people.
That happens. Even a correctly structured prompt does not guarantee correct attribution — results vary between runs of the same prompt. Regenerate before assuming the prompt itself needs more rewriting.

Generate a clean single-speaker clip

$2.99 makes one clean video, ready in about a minute — or add 3 more takes for $6.99 — any photo, any look, same clean 1080p with commercial use. Pay once. No subscription required.

Make a video — $2.99

Broken renders are remade or refunded.

or 3 more takes — $6.99 Any photo, any look · About $2.33 a take · No subscription required.