To keep multiple consistent characters in one AI scene, give each character their own reference image and fixed description block, tell the model which image is which person and where each stands, and repair one character at a time instead of regenerating the scene. Google documents character consistency for up to four characters on Nano Banana 2.1 and five on Nano Banana Pro; FLUX 3 Image takes up to ten references and Midjourney's Edit Model four. With character LoRAs, inpaint each face separately, because chained LoRAs blend.
To keep multiple consistent characters in one AI scene, give each character their own reference image and their own fixed description block, tell the model which image is which person and where each one stands, and repair one character at a time instead of regenerating the whole scene. Stay inside what your model documents: character consistency for up to four characters on Nano Banana 2.1 and five on Nano Banana Pro, up to ten references on FLUX 3 Image, and four on Midjourney's Edit Model.
Limits checked in October 2026 against Google's Gemini image generation docs, Black Forest Labs' FLUX.2 overview, character consistency guide and FLUX 3 Image editing and bounding box docs, Midjourney's Edit Model docs, OpenAI's image generation guide and ComfyUI's multiple LoRAs and inpainting tutorials. The method below is assembled from those vendors' prompting guidance. It is not a scored test of ours, and the example characters are invented.
This page assumes you can already hold one face. If not, start with the consistent character AI guide, which covers the single-character methods this one builds on.
What you will have at the end
A prompt structure that keeps two to four characters separate, an order for assembling a shared scene, and a repair routine for the frame where one face slipped. It works in any tool that reads reference images, and it adapts to LoRAs.
Why two characters drift when one does not
With one character, every trait in the prompt belongs to the same person. With two, the model has to decide who owns the copper hair and who owns the beard, and it gets no help from a sentence that lists both. The usual failures are swapped features, two faces averaged into siblings, and a wrong head count. Midjourney's docs acknowledge the first one directly, with a tip for when features get mixed up between characters.
The fix is ownership. Each trait, outfit and reference image is assigned to a named character, and each character is assigned a place in the frame. That is the whole method; the rest is doing it in the right order.
AI consistent multiple characters: what each tool documents
Check the limit before you plan a cast. A scene with five recurring characters is outside what most models document.
| Tool | Documented reference limit | What the docs say about several people |
|---|---|---|
| Nano Banana 2.1 (Gemini) | Up to 4 character images, inside 14 references | Google's docs include a group photo built from five portraits; multi-turn editing is the recommended way to iterate |
| Nano Banana Pro | Up to 5 character images, plus up to 3 style references | Described as the premium choice for the most complex visual tasks |
| FLUX.2 [pro] | Up to 8 references by API, 10 in the playground | Black Forest Labs shows a couple from one image placed into the street of another |
| FLUX 3 Image | Up to 10 references | Refer to images by number; box rows fix where each element goes, and layout examples seat up to four people |
| Midjourney Edit Model (V8.2) | Up to 4 reference images | If features mix between characters, the docs suggest combining the characters into one reference image |
| GPT Image 2.5 (ChatGPT) | Several input images; the docs example uses four | OpenAI says it may occasionally struggle to keep recurring characters consistent |
| Character LoRAs in ComfyUI | One LoRA per character | Chained LoRAs blend, so repair each face separately with inpainting |
Two notes on that table. Google's limit is about character images, not total references, so on Nano Banana 2.1 four characters with one image each uses the whole documented allowance. And FLUX 3 Image is the only one here with documented placement control: a prompt can end with rows that give each element an id, a box and a description, with box coordinates from 0 to 1000. Black Forest Labs says boxes guide placement and are not clipping masks.
- Each character already holds up alone: one approved hero image per character
- The hero images share a similar light and framing, so one does not dominate
- Each character has a written block of fixed traits, saved somewhere you can paste from
- The two blocks differ in at least three visible traits: hair, build, age, clothing colour
- The cast fits your model's documented character limit
- You know who stands where before you write the prompt
- Every face is an original persona, or a person who has agreed to it
The character-blocking method
Blocking is the theatre word for deciding where each actor stands. Here it means writing the prompt in owned blocks instead of one flowing sentence.
- 01Lock each character alone
One approved hero image each
- 02Write one block per character
Fixed traits, pasted word for word
- 03Block the scene
Who stands where, doing what
- 04Assemble with every reference
Image 1 is A, image 2 is B
- 05Repair one face at a time
Edit, never regenerate
A full prompt has five parts, in this order: the image roles, one block per character, the blocking, the scene, and a keep line. This is an example with two invented characters:
Image 1 is MARA. Image 2 is JONAH.
MARA: woman in her late 20s, shoulder-length copper hair with a centre parting,
freckles across the nose, slim build. Wearing a grey wool coat.
JONAH: man in his early 30s, short black curly hair, trimmed beard,
broad shoulders. Wearing a navy bomber jacket.
Blocking: MARA sits on the left of a cafe table, facing the camera.
JONAH sits on the right, turned toward MARA, mid-laugh. Two people only.
Scene: a small cafe by a window, soft morning daylight, eye level, 4:5 vertical.
Keep MARA's face and hair exactly as in image 1, and JONAH's exactly as in image 2.
Do not blend their features.Every part is there for a documented reason. Naming images by number is how Black Forest Labs' docs address references. The separate blocks follow Google's advice to be hyper-specific and to break complex scenes into steps. The keep line borrows the wording of Google's own edit example, which starts with "keep everything the same". For ready-made variations, our Gemini face consistency prompt pack has a group-scene set.
Consistent characters in the same scene: the assembly order
Order matters more with two characters than with one, because every regeneration puts both faces at risk again.
- 1Approve each character alone
Generate or pick one hero image per character. Result: a face you would accept in a final post, for each.
- 2Write the two blocks
Hair, age range, build, one distinguishing mark, and the outfit for this scene. Save them as text.
- 3Attach the references in a fixed order
Character A first, B second, every time. Result: image 1 always means the same person.
- 4Paste the full prompt
Image roles, both blocks, blocking, scene and keep line. Generate a few candidates.
- 5Check each face against its hero
Compare one character at a time. Result: you know exactly which face failed, if any.
- 6Repair by editing
Feed the frame back with the failed character's reference: keep everything the same, change only that face to match the image.
- 7Lock the passing frame
Save it as the scene reference. Result: the next shot starts from two approved faces already together.
- 8Change one thing per new shot
New pose or new place, not both. Reuse the blocks unchanged.
Step six is where most of the saving is. Google calls multi-turn conversation the recommended way to iterate on images, and its consistency example says to include previously generated images in later prompts. An edit that names one person leaves the other face alone; a fresh generation rolls the dice on both.
Two characters, AI image consistency: copy the features word for word
The second half of the method is boring on purpose. Once a block is written, it is pasted, never retyped. A model has no way to know that "the redhead" in one prompt and "the woman with copper hair" in the next are the same person.
- Traits reworded from memory in every prompt
- Both characters described in one sentence
- Outfits listed in the scene description
- Positions left for the model to choose
- A failed face means a full regeneration
- The same block, word for word, in every prompt
- One labelled block per character
- Each outfit inside its owner's block
- Left, right and eyeline stated in a blocking line
- A failed face is repaired with a one-person edit
This is a writing convention, not a measured result. It follows from how these models read prompts: the more consistently a trait is worded and the more clearly it is attached to one person, the less the model has to guess. The same idea sits behind the description blocks in our ChatGPT consistent character guide and the identity block in Flux consistent character.
Two character LoRAs in one scene
If your characters live in LoRAs, the blocking idea still holds, but the assembly changes. Loading two face LoRAs together does not give you two faces.
- Generate the scene with neither LoRA loaded, or with only one. Get the composition, poses and light right first.
- Mask the first character's face in the mask editor and inpaint it with only that character's LoRA loaded. Result: one correct face, and the other person untouched.
- Mask the second character and repeat with the second LoRA.
- Check both against their heroes, then save the frame as a reference for the next shot.
ComfyUI's inpainting tutorial covers the masking steps, and its grow-mask setting softens the edge of the repainted area. If a LoRA puts its face on every person in the frame even when used alone, it has absorbed more than the identity; our guide to face LoRA overfitting covers that diagnosis.
Group photo, AI, same characters: going past two
Three and four characters work the same way with more blocks, as long as the tool documents that many. Past the limit, change the approach instead of adding references and hoping.
- Stay inside the documented number. Four character images on Nano Banana 2.1, five on Pro. Google's own example asks for an office group photo of the people in five attached portraits.
- Combine characters into one reference. Midjourney's docs suggest putting the characters into a single reference image when their features mix. A clean side-by-side sheet also frees up reference slots.
- Build in passes. Compose two or three people, lock the frame, then add the next person with an edit that keeps everyone already in the frame unchanged.
- Keep faces large enough to matter. In a wide group shot each face is a small part of the frame. Black Forest Labs notes that elements given very small boxes often fail to appear at all, so frame tighter, or fix faces afterwards with edits.
For moving scenes, the same cast needs a different binding step; AI video character consistency covers that. And if references stop holding once the cast is on screen for weeks, that is the point to train a LoRA per character; the AI Influencers program takes you through reference tools first, then LoRA training for consistent faces and ControlNet for poses.
Troubleshooting a multi-character scene
| Symptom | Fix |
|---|---|
| Features swapped between characters | Rewrite the prompt as separate labelled blocks, one per character, each tied to its image number. State left and right explicitly. |
| Both faces look like siblings | The model averaged them. Make the two blocks differ in at least three visible traits, and repair one face with an edit that names only that person. |
| One face is right, the other is generic | Do not regenerate. Feed the image back with that character's reference and ask for a change to that face only. |
| An extra person appears, or someone is missing | State the count in the scene line and name each person once. On FLUX 3 Image, give each person a box row. |
| Outfits traded places | Put each outfit inside its owner's block, not in the scene line. |
| A third or fourth character breaks the first two | Build in passes: lock the pair, then add the next person with an edit that keeps everyone already in frame unchanged. |
Whose faces you can put in a scene
Two characters means two sets of rights. Google's docs carry a plain reminder to make sure you have the necessary rights to any images you upload, and every Nano Banana image includes a SynthID watermark. Build scenes from original personas. If a real person appears, including you, that person needs to have agreed to the use, and a second real person needs their own consent, not yours. The likeness section of consistent characters without a LoRA lists the rules by tool.
Multiple consistent characters: FAQ
How do I keep multiple characters consistent in one AI image?
Give each character their own reference image and their own fixed description, tell the model which image is which person and where each one stands, and end with a line saying whose face must match which image. Then repair one character at a time with an edit, instead of regenerating the whole scene. Stay inside the number of character references your model documents.
Which AI image generator handles multiple consistent characters?
Several document it. Google lists character consistency for up to four character images on Nano Banana 2.1 and five on Nano Banana Pro. FLUX 3 Image combines up to ten references and can place elements with bounding boxes. Midjourney's Edit Model takes up to four reference images. We have not scored them against each other, so this is a list of documented limits. Checked October 2026.
How many characters can Nano Banana keep consistent?
Google's Gemini API docs say Nano Banana 2.1 and Nano Banana 2 support character resemblance for up to four characters in a single workflow, inside a total of 14 reference images. Nano Banana Pro takes up to five character images and up to three style references. Nano Banana 2 Lite is described as not optimized for multiple reference inputs, so skip it for this job.
Why do my two AI characters swap features or look like siblings?
Because the model was given features without owners. A prompt that lists red hair, a beard and a grey coat in one sentence leaves the model to decide who gets what. Put each character in a separate labelled block, tie each block to one reference image, and state positions. Midjourney's docs suggest combining the characters into a single reference image when features still mix.
Can I use two character LoRAs in one image?
Not by simply stacking them. ComfyUI's own tutorial describes chaining two LoRAs as producing a blended result, which is what you want for styles and the opposite of what you want for two faces. Generate the scene first, then inpaint each character's face separately with only that character's LoRA loaded, masking one person at a time.
How do I make a group photo with the same AI characters?
Stay within the documented limit first: four character images on Nano Banana 2.1, five on Pro. Google's own example builds an office group photo from five portraits in one request. Beyond that, build the group in passes: compose two or three people, lock that image, then add the next person with an edit that says to keep everyone already in the frame unchanged.
Two characters hold. Now give them something to do.
AI Influencers, included in All Access, covers creating a persona, reference tools and LoRA training for consistent faces, ControlNet for poses and the video lessons that put characters in motion, with the other three programs, live coaching and the private community in one subscription.
Get the prompt patterns that are holding up this month
Join the free Telegram channel for reference-image and multi-character workflows as the models change.