Skip to main content

Multiple Consistent Characters in One AI Scene

Keep multiple consistent characters in one AI scene: per-character reference blocks, a scene assembly order, tool limits and one-character-at-a-time repairs.

Founder of IImagined.ai

Published
Oct 11, 2026
Reading time
10 min read
Quick answer

To keep multiple consistent characters in one AI scene, give each character their own reference image and fixed description block, tell the model which image is which person and where each stands, and repair one character at a time instead of regenerating the scene. Google documents character consistency for up to four characters on Nano Banana 2.1 and five on Nano Banana Pro; FLUX 3 Image takes up to ten references and Midjourney's Edit Model four. With character LoRAs, inpaint each face separately, because chained LoRAs blend.

To keep multiple consistent characters in one AI scene, give each character their own reference image and their own fixed description block, tell the model which image is which person and where each one stands, and repair one character at a time instead of regenerating the whole scene. Stay inside what your model documents: character consistency for up to four characters on Nano Banana 2.1 and five on Nano Banana Pro, up to ten references on FLUX 3 Image, and four on Midjourney's Edit Model.

Limits checked in October 2026 against Google's Gemini image generation docs, Black Forest Labs' FLUX.2 overview, character consistency guide and FLUX 3 Image editing and bounding box docs, Midjourney's Edit Model docs, OpenAI's image generation guide and ComfyUI's multiple LoRAs and inpainting tutorials. The method below is assembled from those vendors' prompting guidance. It is not a scored test of ours, and the example characters are invented.

This page assumes you can already hold one face. If not, start with the consistent character AI guide, which covers the single-character methods this one builds on.

What you will have at the end

A prompt structure that keeps two to four characters separate, an order for assembling a shared scene, and a repair routine for the frame where one face slipped. It works in any tool that reads reference images, and it adapts to LoRAs.

Why two characters drift when one does not

With one character, every trait in the prompt belongs to the same person. With two, the model has to decide who owns the copper hair and who owns the beard, and it gets no help from a sentence that lists both. The usual failures are swapped features, two faces averaged into siblings, and a wrong head count. Midjourney's docs acknowledge the first one directly, with a tip for when features get mixed up between characters.

The fix is ownership. Each trait, outfit and reference image is assigned to a named character, and each character is assigned a place in the frame. That is the whole method; the rest is doing it in the right order.

AI consistent multiple characters: what each tool documents

Check the limit before you plan a cast. A scene with five recurring characters is outside what most models document.

ToolDocumented reference limitWhat the docs say about several people
Nano Banana 2.1 (Gemini)Up to 4 character images, inside 14 referencesGoogle's docs include a group photo built from five portraits; multi-turn editing is the recommended way to iterate
Nano Banana ProUp to 5 character images, plus up to 3 style referencesDescribed as the premium choice for the most complex visual tasks
FLUX.2 [pro]Up to 8 references by API, 10 in the playgroundBlack Forest Labs shows a couple from one image placed into the street of another
FLUX 3 ImageUp to 10 referencesRefer to images by number; box rows fix where each element goes, and layout examples seat up to four people
Midjourney Edit Model (V8.2)Up to 4 reference imagesIf features mix between characters, the docs suggest combining the characters into one reference image
GPT Image 2.5 (ChatGPT)Several input images; the docs example uses fourOpenAI says it may occasionally struggle to keep recurring characters consistent
Character LoRAs in ComfyUIOne LoRA per characterChained LoRAs blend, so repair each face separately with inpainting

Two notes on that table. Google's limit is about character images, not total references, so on Nano Banana 2.1 four characters with one image each uses the whole documented allowance. And FLUX 3 Image is the only one here with documented placement control: a prompt can end with rows that give each element an id, a box and a description, with box coordinates from 0 to 1000. Black Forest Labs says boxes guide placement and are not clipping masks.

Pre-flight for a multi-character scene
  • Each character already holds up alone: one approved hero image per character
  • The hero images share a similar light and framing, so one does not dominate
  • Each character has a written block of fixed traits, saved somewhere you can paste from
  • The two blocks differ in at least three visible traits: hair, build, age, clothing colour
  • The cast fits your model's documented character limit
  • You know who stands where before you write the prompt
  • Every face is an original persona, or a person who has agreed to it

The character-blocking method

Blocking is the theatre word for deciding where each actor stands. Here it means writing the prompt in owned blocks instead of one flowing sentence.

Character blocking, end to end
  1. 01
    Lock each character alone

    One approved hero image each

  2. 02
    Write one block per character

    Fixed traits, pasted word for word

  3. 03
    Block the scene

    Who stands where, doing what

  4. 04
    Assemble with every reference

    Image 1 is A, image 2 is B

  5. 05
    Repair one face at a time

    Edit, never regenerate

A full prompt has five parts, in this order: the image roles, one block per character, the blocking, the scene, and a keep line. This is an example with two invented characters:

Image 1 is MARA. Image 2 is JONAH.

MARA: woman in her late 20s, shoulder-length copper hair with a centre parting,
freckles across the nose, slim build. Wearing a grey wool coat.

JONAH: man in his early 30s, short black curly hair, trimmed beard,
broad shoulders. Wearing a navy bomber jacket.

Blocking: MARA sits on the left of a cafe table, facing the camera.
JONAH sits on the right, turned toward MARA, mid-laugh. Two people only.

Scene: a small cafe by a window, soft morning daylight, eye level, 4:5 vertical.

Keep MARA's face and hair exactly as in image 1, and JONAH's exactly as in image 2.
Do not blend their features.

Every part is there for a documented reason. Naming images by number is how Black Forest Labs' docs address references. The separate blocks follow Google's advice to be hyper-specific and to break complex scenes into steps. The keep line borrows the wording of Google's own edit example, which starts with "keep everything the same". For ready-made variations, our Gemini face consistency prompt pack has a group-scene set.

Consistent characters in the same scene: the assembly order

Order matters more with two characters than with one, because every regeneration puts both faces at risk again.

Assemble the scene in eight steps
  1. 1
    Approve each character alone

    Generate or pick one hero image per character. Result: a face you would accept in a final post, for each.

  2. 2
    Write the two blocks

    Hair, age range, build, one distinguishing mark, and the outfit for this scene. Save them as text.

  3. 3
    Attach the references in a fixed order

    Character A first, B second, every time. Result: image 1 always means the same person.

  4. 4
    Paste the full prompt

    Image roles, both blocks, blocking, scene and keep line. Generate a few candidates.

  5. 5
    Check each face against its hero

    Compare one character at a time. Result: you know exactly which face failed, if any.

  6. 6
    Repair by editing

    Feed the frame back with the failed character's reference: keep everything the same, change only that face to match the image.

  7. 7
    Lock the passing frame

    Save it as the scene reference. Result: the next shot starts from two approved faces already together.

  8. 8
    Change one thing per new shot

    New pose or new place, not both. Reuse the blocks unchanged.

Step six is where most of the saving is. Google calls multi-turn conversation the recommended way to iterate on images, and its consistency example says to include previously generated images in later prompts. An edit that names one person leaves the other face alone; a fresh generation rolls the dice on both.

Two characters, AI image consistency: copy the features word for word

The second half of the method is boring on purpose. Once a block is written, it is pasted, never retyped. A model has no way to know that "the redhead" in one prompt and "the woman with copper hair" in the next are the same person.

Paraphrased or pasted
Paraphrased each time
  • Traits reworded from memory in every prompt
  • Both characters described in one sentence
  • Outfits listed in the scene description
  • Positions left for the model to choose
  • A failed face means a full regeneration
Blocked and pasted
  • The same block, word for word, in every prompt
  • One labelled block per character
  • Each outfit inside its owner's block
  • Left, right and eyeline stated in a blocking line
  • A failed face is repaired with a one-person edit

This is a writing convention, not a measured result. It follows from how these models read prompts: the more consistently a trait is worded and the more clearly it is attached to one person, the less the model has to guess. The same idea sits behind the description blocks in our ChatGPT consistent character guide and the identity block in Flux consistent character.

Two character LoRAs in one scene

If your characters live in LoRAs, the blocking idea still holds, but the assembly changes. Loading two face LoRAs together does not give you two faces.

  1. Generate the scene with neither LoRA loaded, or with only one. Get the composition, poses and light right first.
  2. Mask the first character's face in the mask editor and inpaint it with only that character's LoRA loaded. Result: one correct face, and the other person untouched.
  3. Mask the second character and repeat with the second LoRA.
  4. Check both against their heroes, then save the frame as a reference for the next shot.

ComfyUI's inpainting tutorial covers the masking steps, and its grow-mask setting softens the edge of the repainted area. If a LoRA puts its face on every person in the frame even when used alone, it has absorbed more than the identity; our guide to face LoRA overfitting covers that diagnosis.

Group photo, AI, same characters: going past two

Three and four characters work the same way with more blocks, as long as the tool documents that many. Past the limit, change the approach instead of adding references and hoping.

  • Stay inside the documented number. Four character images on Nano Banana 2.1, five on Pro. Google's own example asks for an office group photo of the people in five attached portraits.
  • Combine characters into one reference. Midjourney's docs suggest putting the characters into a single reference image when their features mix. A clean side-by-side sheet also frees up reference slots.
  • Build in passes. Compose two or three people, lock the frame, then add the next person with an edit that keeps everyone already in the frame unchanged.
  • Keep faces large enough to matter. In a wide group shot each face is a small part of the frame. Black Forest Labs notes that elements given very small boxes often fail to appear at all, so frame tighter, or fix faces afterwards with edits.

For moving scenes, the same cast needs a different binding step; AI video character consistency covers that. And if references stop holding once the cast is on screen for weeks, that is the point to train a LoRA per character; the AI Influencers program takes you through reference tools first, then LoRA training for consistent faces and ControlNet for poses.

Troubleshooting a multi-character scene

SymptomFix
Features swapped between charactersRewrite the prompt as separate labelled blocks, one per character, each tied to its image number. State left and right explicitly.
Both faces look like siblingsThe model averaged them. Make the two blocks differ in at least three visible traits, and repair one face with an edit that names only that person.
One face is right, the other is genericDo not regenerate. Feed the image back with that character's reference and ask for a change to that face only.
An extra person appears, or someone is missingState the count in the scene line and name each person once. On FLUX 3 Image, give each person a box row.
Outfits traded placesPut each outfit inside its owner's block, not in the scene line.
A third or fourth character breaks the first twoBuild in passes: lock the pair, then add the next person with an edit that keeps everyone already in frame unchanged.

Whose faces you can put in a scene

Two characters means two sets of rights. Google's docs carry a plain reminder to make sure you have the necessary rights to any images you upload, and every Nano Banana image includes a SynthID watermark. Build scenes from original personas. If a real person appears, including you, that person needs to have agreed to the use, and a second real person needs their own consent, not yours. The likeness section of consistent characters without a LoRA lists the rules by tool.

Multiple consistent characters: FAQ

How do I keep multiple characters consistent in one AI image?

Give each character their own reference image and their own fixed description, tell the model which image is which person and where each one stands, and end with a line saying whose face must match which image. Then repair one character at a time with an edit, instead of regenerating the whole scene. Stay inside the number of character references your model documents.

Which AI image generator handles multiple consistent characters?

Several document it. Google lists character consistency for up to four character images on Nano Banana 2.1 and five on Nano Banana Pro. FLUX 3 Image combines up to ten references and can place elements with bounding boxes. Midjourney's Edit Model takes up to four reference images. We have not scored them against each other, so this is a list of documented limits. Checked October 2026.

How many characters can Nano Banana keep consistent?

Google's Gemini API docs say Nano Banana 2.1 and Nano Banana 2 support character resemblance for up to four characters in a single workflow, inside a total of 14 reference images. Nano Banana Pro takes up to five character images and up to three style references. Nano Banana 2 Lite is described as not optimized for multiple reference inputs, so skip it for this job.

Why do my two AI characters swap features or look like siblings?

Because the model was given features without owners. A prompt that lists red hair, a beard and a grey coat in one sentence leaves the model to decide who gets what. Put each character in a separate labelled block, tie each block to one reference image, and state positions. Midjourney's docs suggest combining the characters into a single reference image when features still mix.

Can I use two character LoRAs in one image?

Not by simply stacking them. ComfyUI's own tutorial describes chaining two LoRAs as producing a blended result, which is what you want for styles and the opposite of what you want for two faces. Generate the scene first, then inpaint each character's face separately with only that character's LoRA loaded, masking one person at a time.

How do I make a group photo with the same AI characters?

Stay within the documented limit first: four character images on Nano Banana 2.1, five on Pro. Google's own example builds an office group photo from five portraits in one request. Beyond that, build the group in passes: compose two or three people, lock that image, then add the next person with an edit that says to keep everyone already in the frame unchanged.

All Access · all four programs · $99/mo

Two characters hold. Now give them something to do.

AI Influencers, included in All Access, covers creating a persona, reference tools and LoRA training for consistent faces, ControlNet for poses and the video lessons that put characters in motion, with the other three programs, live coaching and the private community in one subscription.

Start All Access — $99/mo →30-day money-back guarantee
Free

Get the prompt patterns that are holding up this month

Join the free Telegram channel for reference-image and multi-character workflows as the models change.