A Stable Diffusion consistent character comes from three decisions in order: one base model family, one approved hero image, and one identity method. A reference adapter works from a single image, a character LoRA holds across more angles, and a hybrid adds ControlNet for pose. Published VRAM figures run from 8GB for SDXL to 12GB for training a FLUX.2 klein LoRA.
A Stable Diffusion consistent character comes from three decisions made in order: one base model family, one approved hero image, and one identity method, which is a reference adapter working from an image, a trained character LoRA, or both with ControlNet holding the pose. This tutorial walks through all three workflows and gives the published VRAM figure for each, so you can pick the one your graphics card can run.
Hardware figures and licenses checked October 2026 against Stability AI's SDXL 1.0 and Stable Diffusion 3.5 announcements, the SDXL and SD 1.5 model cards, the original Stable Diffusion repository, ComfyUI's FLUX.2 klein, LoRA and ControlNet docs, Black Forest Labs' klein training guide, the sd-scripts FLUX.1 and SDXL training docs, and the PuLID and InstantID repositories. The VRAM numbers are the publishers' statements, not our measurements.
"Stable Diffusion" here means the open-weight workflow as people run it in 2026: Stability AI's own models plus the Flux models that load into the same tools. If you have not chosen between local and hosted yet, our consistent character AI guide compares every method first. This page assumes you have picked local.
What you will have at the end
One base model you will not switch away from, one hero image, an identity method that reproduces the face on demand, and a pose control you can add when a shot needs it.
- ComfyUI installed and updated, or another UI you already know
- Your GPU's VRAM written down: it decides the base model in step 1
- A fixed identity block: age range, hair, eyes, two distinguishing marks
- A fictional adult persona, or your own face, and nobody else's
- Disk space for one base model, its VAE and encoders, and one identity model
- For the LoRA route: a plan for 10 to 20 varied images of one face
- The license of every model you will download, read before paid work
Step 1: choose the base family by VRAM and tooling
Everything downstream is built for one family. A LoRA, ControlNet or face adapter made for SDXL does not load on SD 1.5 or Flux, so this choice comes first and stays fixed.
| Base | Published VRAM figure | Identity tools published for it | License note |
|---|---|---|---|
| SD 1.5 | The 2022 release notes name a GPU with at least 10GB; the smallest model here, with an 860M UNet | LoRA, IP-Adapter, ControlNet 1.1 | CreativeML OpenRAIL-M; trained at 512 x 512 |
| SDXL 1.0 and fine-tunes | 8GB consumer GPUs (Stability AI) | LoRA, IP-Adapter, InstantID, PuLID, pose ControlNets | OpenRAIL++-M for the base; each fine-tune sets its own |
| SD 3.5 Medium | 9.9GB excluding text encoders (Stability AI) | LoRA; the three face adapters list no SD 3.5 version | Community License: free under $1M annual revenue |
| FLUX.2 [klein] 4B | 8.4GB distilled, 9.2GB base (Comfy); about 13GB (BFL) | Built-in editing with up to 4 references; LoRA on the base variant | Apache 2.0 |
| FLUX.1 [dev] | PuLID-FLUX demo peaks under 15GB with fp8 and offloading | LoRA, PuLID-FLUX | FLUX.1 [dev] model license; read it before paid work |
- SDXL is the default for persona work on a mid-range card. Stability AI's announcement says SDXL 1.0 should work effectively on consumer GPUs with 8GB of VRAM, and the widest set of identity tools targets it.
- SD 1.5 is for old hardware and quick tests. It was trained at 512 x 512, so faces in full-body shots have few pixels to live in. SDXL's base model is 3.5B parameters against SD 1.5's 860M UNet.
- SD 3.5 is the newest Stable Diffusion, with the thinnest face tooling. Stability describes 3.5 Medium as needing 9.9GB of VRAM, excluding text encoders. None of the IP-Adapter, InstantID or PuLID repositories we checked lists an SD 3.5 version, so plan on a LoRA there.
- The Flux stack trades ecosystem for built-in references. FLUX.2 [klein] 4B edits from up to four reference images with no custom nodes and is Apache 2.0.
Expected result: one base checkpoint downloaded and rendering a plain portrait from your identity block. ComfyUI's README says its weight streaming can run large models on as little as 4GB of VRAM with 8GB of RAM, so a smaller card still works, slowly. Our best ComfyUI models guide lists specific checkpoints for realistic people and their licenses.
Step 2: write the identity block the way your base reads it
The identity block is the fixed description you paste into every prompt, whichever route you take next. The wording stays the same; where it goes depends on the family.
- SD 1.5 and SDXL. In ComfyUI the positive and negative prompts are two separate text nodes wired into the sampler. Put the identity block first in the positive prompt and keep the negative list short and specific; our negative prompt guide for realistic people explains what each term does.
- FLUX.2. Black Forest Labs' prompting guide says the model does not support negative prompts and pays more attention to what comes first. Lead with the identity block and describe what you want instead of what you do not.
- Any family. Leave clothing, mood and style words out of the block. They belong in the scene part of the prompt, where they can change.
Expected result: a text file holding one block you never retype. The free AI influencer prompt generator writes a first draft from a persona brief.
Route A: Stable Diffusion consistent character from image
The reference route needs no training. Generate hero candidates from your identity block, approve one, then condition every later generation on it.
- 1Generate and approve the hero
Front-facing, even light, sharp. Expected result: one image you never regenerate.
- 2Crop the face for the adapter
A tight, clean face crop. PuLID's ComfyUI notes say reference quality is very important.
- 3Load one identity adapter
IP-Adapter Plus Face on SDXL is the simplest. InstantID and PuLID are the single-image ID options.
- 4Start from the documented setting
IP-Adapter: weight about 0.8 and more steps. InstantID: raise the two scales for similarity, lower the adapter scale if it over-saturates.
- 5Fix the seed while you tune
Change one setting per run. Expected result: the face moves toward the hero and the scene stays put.
- 6Approve angle views
Three-quarter, profile, full body. These become your reference set and the start of a dataset.
On FLUX.2 [klein] the same route is the built-in image-edit template: load the hero in the first image node and prompt by position. The node chains, file names and failure points for each adapter are in our ComfyUI consistent character workflow templates, and the license position of each is covered in consistent characters without a LoRA.
Route B: Stable Diffusion consistent character LoRA
A LoRA is a small set of extra weights trained on your character. A trigger word then brings the face back in any prompt, at any angle the dataset covered. It costs a curated dataset and a training run.
- 1Build the dataset
Varied angles, expressions, light and outfits of one face, culled hard. Expected result: a folder you would be happy to see copied exactly.
- 2Caption with a trigger word
One consistent trigger. Black Forest Labs' guide asks for it in every caption and for images of 1024px or more.
- 3Train on your render base
SDXL LoRA on the SDXL family you render with; a klein LoRA on the klein Base variant, not the distilled one.
- 4Save several checkpoints
Keep intermediate saves so you can pick the one that holds the face without copying backgrounds.
- 5Compare with fixed prompts and seeds
Same prompts, same seeds, one checkpoint at a time. Expected result: one file that passes your hard angles.
- 6Load it and set strength
In ComfyUI the Load LoRA node exposes strength_model and strength_clip. Files go in ComfyUI/models/loras.
The numbers that matter are documented elsewhere on this site rather than repeated here: how many images a face LoRA needs compiles the trainer docs, the free LoRA training steps calculator turns your image count into steps and epochs, and the LoRA training guide for consistent faces covers trainers, captions and checkpoint testing. For klein, Black Forest Labs suggests 1,500 to 3,000 steps for a character LoRA.
Route C: Stable Diffusion consistent character with ControlNet (the hybrid)
ControlNet does not hold a face. It holds structure: an OpenPose skeleton fixes the body, hands and head position, and depth or canny fix shapes and edges. The hybrid workflow gives each tool one job.
- 01Base model
One family, fixed
- 02Character LoRA
Carries the face in the weights
- 03Reference adapter
Optional: pulls a specific look from one image
- 04ControlNet
Pose from an OpenPose skeleton, or depth
- 05Face detailer
Re-renders a small face with the same LoRA
- 06Review
Against the hero, at full size
- Render with the LoRA alone first. Expected result: the face holds in a medium shot. If it does not, fix the LoRA before adding anything.
- Add ControlNet for the pose. ComfyUI's Apply ControlNet node has a strength value and start and end percentages. Lower the strength or end it early if the pose is right but the image looks stiff.
- Add the reference adapter only if you need it. It helps when one shoot must match one specific image; it also brings that image's hair and light.
- Finish with a detailer. The SDXL model card lists "faces and people in general may not be generated properly" among its limitations, and a face that fills few pixels drifts first. A detailer crops the face, re-renders it larger with your LoRA loaded, and pastes it back.
Which control model fits which base, and a pose-battery recipe for shooting a full set, are in our ControlNet guide.
VRAM budgets by method
Rendering and training have different budgets. These are the figures each publisher states; real use shifts with resolution, batch size and offloading.
| Method | Stack | Published figure | Source |
|---|---|---|---|
| Reference, SDXL | SDXL plus IP-Adapter Plus Face | 8GB class: Stability's SDXL figure; the adapter and its image encoder load on top | Stability AI |
| Reference, FLUX.2 | [klein] 4B image-edit template | 8.4GB distilled or 9.2GB base on Comfy's test card | docs.comfy.org |
| Reference, FLUX.1 | FLUX.1 [dev] plus PuLID-FLUX | Under 15GB with fp8 and offloading; about 11GB with aggressive offloading, very slow | PuLID repository |
| LoRA training, SDXL | sd-scripts or kohya-colab | Runs on Colab's free T4 per kohya-colab; sd-scripts advises caching and mixed precision | kohya-colab, sd-scripts |
| LoRA training, FLUX.2 | [klein] 4B Base | 12GB VRAM and 32GB RAM minimum | Black Forest Labs |
| LoRA training, FLUX.1 | sd-scripts with block swapping | 24GB for basic settings, down to 8GB with 28 blocks swapped | sd-scripts |
If you would rather rent than buy, our free online LoRA training comparison lists the options and their documented limits.
SDXL or a Flux stack for a persona?
Both can carry a persona. The difference is where the identity tooling lives and what the license lets you do with the result.
- 8GB class for rendering, per Stability AI
- LoRA, IP-Adapter, InstantID and PuLID all target it
- Pose ControlNets are published for it
- Negative prompts work in the usual way
- Base is OpenRAIL++-M; check each fine-tune
- 8.4GB to 9.2GB reported by Comfy; BFL says about 13GB
- Up to four reference images, no custom nodes
- LoRA trains on the Base variant from 12GB
- One Apache 2.0 model for generation and editing
- No pose-control template in ComfyUI's klein guide
A reasonable path is SDXL if you already own SDXL checkpoints and LoRAs, and klein 4B if you are starting clean and want one license to read. Our Flux consistent character guide covers the Flux side in detail. The AI Influencers program goes deeper where this tutorial stops, with lessons on Flux and SDXL setup, hardware requirements, LoRA training for consistent faces and body consistency with ControlNet.
Troubleshooting
- The face holds in close-ups and drifts in full-body shots. Too few pixels on the face. Add a detailer pass that loads the same LoRA or adapter.
- Every image copies the reference photo's hair and lighting. Adapter weight too high. Step it down, or move from IP-Adapter to a LoRA.
- The LoRA copies backgrounds or one outfit. Overtrained or an unvaried dataset. Pick an earlier checkpoint and widen the dataset.
- The pose is right and the face is a stranger. ControlNet is running without an identity method. Add the LoRA or adapter; ControlNet carries no identity.
- Red nodes after an update. The IP-Adapter and PuLID ComfyUI node packs have been maintenance-only since April 2025. Keep a copy of the ComfyUI version your graph was built on.
- Drift you cannot explain. Work through the causes in our AI face consistency guide, one variable at a time.
Stable Diffusion consistent character: FAQ
How do I make a consistent character in Stable Diffusion?
Pick one base model family and stay on it, approve one hero image, then hold the face with an identity method. A reference adapter such as IP-Adapter Plus Face works from a single image, a character LoRA trained on a curated set holds across more angles, and ControlNet fixes the pose so the identity method only has to carry the face. Finish small faces with a detailer pass.
Do I need a LoRA for a consistent character in Stable Diffusion?
Not to start. A reference adapter conditions each generation on one image with no training, which is enough for a new persona and for building a dataset. A LoRA becomes worth it when you need the face across profile views, full-body shots and hundreds of images, because it stores the identity in weights instead of rereading one picture.
Can I create a consistent character from one image in Stable Diffusion?
Yes, on SDXL. IP-Adapter Plus Face reads a face crop as an image prompt, and InstantID and PuLID turn one face image into an identity embedding. One image only shows one angle, so expect drift on views it does not cover. InstantID's checkpoints and the InsightFace models both ID adapters use are licensed for non-commercial research only.
How much VRAM do I need for a consistent character workflow?
Stability AI says SDXL 1.0 should work on consumer GPUs with 8GB of VRAM and puts SD 3.5 Medium at 9.9GB excluding text encoders. Comfy reports 8.4GB to 9.2GB for FLUX.2 klein 4B. For LoRA training, Black Forest Labs lists 12GB as the minimum for klein 4B Base, and sd-scripts documents FLUX.1 training from 24GB down to 8GB with block swapping.
Is SDXL or Flux better for a consistent character?
They suit different builders, and we have not scored one against the other. SDXL has the widest set of published identity tools: LoRAs, IP-Adapter, InstantID, PuLID and pose ControlNets. FLUX.2 klein 4B reads up to four reference images natively, trains LoRAs on its base variant and is Apache 2.0. Choose by the tools you need and the license you can work under.
Why does ControlNet not keep my character's face the same?
Because that is not its job. ControlNet steers structure: an OpenPose skeleton fixes where the head, limbs and hands go, and depth or canny fix shapes and edges. None of them carries identity. Pair ControlNet with a LoRA or a reference adapter for the face, and use it to stop the pose changing while you test the identity method.
You have the stack. Now build the persona on it.
AI Influencers, included in All Access, covers the build end to end: Flux and SDXL setup, the GPU guide, LoRA training for consistent faces, body consistency with ControlNet and photorealism, with the other three programs, weekly coaching and the private community in one subscription.
Plan the training run free
Turn your image count into steps and epochs with the free calculator, then post your hero image and first LoRA tests in the free Discord for feedback.