Face LoRA overfitting means the LoRA memorised its training photos instead of learning the face, so every output repeats one pose, crop, expression or background. Fix it cheapest-first: use an earlier checkpoint, then repair the dataset's variety and captions, and only then lower steps, learning rate or rank on the next run. Six symptoms map to six different fixes, and most of them are dataset fixes, not settings.
Face LoRA overfitting means the LoRA has memorised its training photos instead of learning the face, so every output repeats the same pose, crop, expression or background. Fix it in this order: switch to an earlier checkpoint, then repair the dataset's variety and captions, and only then lower the steps, learning rate or rank for the next run.
Checked in October 2026 against Scenario's character training guide, Black Forest Labs' klein LoRA guide and training example, Civitai's developer docs for SDXL, Flux 1 and Flux 2 Klein, Hugging Face's SDXL LoRA write-up, the sd-scripts advanced training docs and AI Toolkit's example config. The fixes are quoted from those documents. Matching each symptom to a cause is our reading of them; we ran no training benchmark for this page.
This page is for someone who already trained a LoRA and does not like what came out. If you have not trained yet, the consistent character AI guide explains where a LoRA fits, and the LoRA training guide covers the first run.
What face LoRA overfitting is, and what it is not
A face LoRA has two jobs that pull against each other: hold the likeness, and stay out of the way of the prompt. Overfitting is when the first job wins completely. Scenario's docs describe the result as a character that appears stuck in the scenarios you trained on, with outputs that reproduce specific training poses or backgrounds.
It is easy to confuse with two other failures. An under-trained LoRA follows prompts but the face is generic. A LoRA trained on a messy dataset does neither job well. They need opposite fixes, so place your LoRA on this grid before you touch a setting.
The diagnosis chart: six symptoms and the parameter to change
Overfitting is not one problem with one dial. Find the row that matches your samples, try the free fix, and change the listed parameter only if you retrain.
| What you see | What it usually means | Free fix, today | Change on the next run |
|---|---|---|---|
| 1. The same pose, crop or composition every time | The LoRA memorised compositions: too many steps for this set, or one framing repeated in it | Go back to an earlier checkpoint | Steps or epochs down; near-duplicate framings out |
| 2. The training room or backdrop keeps appearing | One background dominates the dataset, so it became part of the character | Prompt a specific new setting on an earlier checkpoint | Dataset: mix the backgrounds and caption each one |
| 3. A training outfit shows up unprompted | The outfit repeats and the captions never mention it, so it attached to the trigger word | Name a different outfit explicitly in the prompt | Captions: describe what should vary; vary the outfits |
| 4. One expression; new angles and expressions are ignored | Narrow coverage of angles and expressions, hardened by a long run | Earlier checkpoint, then test again | Dataset: at least three distinct poses or angles; fewer repeats |
| 5. Waxy, airbrushed or over-sharpened skin | Too much capacity, or too high a learning rate | Earlier checkpoint | Rank down; learning rate down |
| 6. It only behaves at reduced strength | The weights are overcooked and overpower the base model at full strength | Run it at lower strength as a stopgap | Steps down, rank down, alpha to half of rank |
Count the last column: three of the six rows are fixed in the dataset or its captions, not in the trainer. That matches what the vendor docs keep saying. Scenario calls repetitive framing the most common cause of weak character recall, and Hugging Face's write-up says fewer, better-curated images beat more images of middling quality.
How to confirm it: the fixed-prompt test
Do not diagnose from one unlucky image. Run a small, fixed set of prompts against every checkpoint you saved, changing nothing else. Black Forest Labs' guide puts the principle in six words: watch the samples, not the loss.
- 1Fix everything except the checkpoint
Same base model, sampler, steps, size and seed for every image, with the LoRA at full strength.
- 2Prompt 1: a neutral portrait
Trigger word plus a plain head-and-shoulders shot. This is the likeness check.
- 3Prompt 2: a profile
Or any angle that is rare or missing in the dataset.
- 4Prompt 3: a new expression
Laughing, surprised or eyes closed: whichever your dataset has least of.
- 5Prompt 4: full body, new outfit, new place
An outfit and a location that appear nowhere in the training set.
- 6Run all four on every checkpoint
Lay the results out as a grid, prompts across and checkpoints down.
- 7Read down each column
Likeness arriving while prompts still work marks the keeper. Prompts failing as you go down is overfitting.
- 8Repeat the last checkpoint at lower strength
If it only obeys prompts at reduced strength, the run went past the point you wanted.
The grid answers two questions at once: whether the LoRA is overfit, and which earlier checkpoint is not. If you train on Civitai, Training Studio's Compare epochs view builds the same grid from your sample prompts; our Civitai LoRA training walkthrough covers it.
LoRA overfitting fix: the order to try things
Work from the cheapest fix to the most expensive. The first two cost nothing, the middle two cost an hour of dataset work, and only the last two need new settings.
- 01Earlier checkpoint
Free, and often all you need
- 02Lower strength
A stopgap, not a cure
- 03Dataset variety
Cut duplicates, add what is missing
- 04Captions
Describe what should vary
- 05Shorter run
Fewer steps or epochs
- 06Rate and rank
Lower learning rate, rank or alpha
Start with the checkpoint you already have. Scenario notes that the last epoch is set as the default but earlier epochs often produce more flexible characters. Black Forest Labs says the same about its own runs: pick the checkpoint whose samples look best, not necessarily the last one.
Lower strength only buys time. Turning the LoRA down in your generator weakens everything it learned, the memorised compositions and the likeness together. It can rescue a batch of images today. It does not repair the file.
LoRA only one pose: fix the dataset first
When every output has the same pose or crop, the pose was part of what repeated. Scenario lists it as a named pitfall: same pose in every image, and the model bakes the pose into the identity. No setting separates a face from a pose the dataset never separates.
- Cut near-duplicates. Scenario says too many similar images push the model toward overfitting, and that a curated set of 5 to 15 images often beats a set of 20 or more where images become redundant.
- Cover at least three poses or angles, even in a set of five to eight images: front, profile, three-quarter, and the back when possible.
- Mix the backgrounds. In Scenario's words, when all images are shot in the same room, the character wants to be in that room.
- Decide what the outfit is. Either it is part of the identity, or it is a variable. If it is a variable, vary it and caption it.
Captions do the same job from the other side. Scenario's rule is to caption the elements that are meant to vary in the generated outputs. Black Forest Labs explains the mechanism for styles: leave the style out of the captions and the model bakes it into the weights. Whatever your captions never mention gets absorbed into the trigger word. For a face that is what you want for the bone structure, and the opposite of what you want for a hoodie.
The character LoRA dataset checklist gives a shot list built to make the face the only constant, and how many images to train a face LoRA covers why adding more of the same makes this worse.
Face LoRA too stiff: steps, epochs and when to stop
If early checkpoints are looser than late ones, the run simply went on too long for that dataset. The loss curve will not warn you. Black Forest Labs writes that loss keeps dropping well past the point where the images start to overfit.
Both are style LoRAs, not faces, and the Hugging Face runs changed batch size and repeats as well. The pattern is the point: the keeper sat well before the end. Source: Black Forest Labs klein training example and Hugging Face SDXL LoRA write-up, checked October 2026
So the practical rule is to make stopping cheap. Save a checkpoint every 250 steps or every epoch, render sample prompts as you go, and treat the step count as a ceiling you expect to stop short of. Black Forest Labs puts the visual peak of most of its style LoRAs between steps 750 and 1,500.
Small datasets overfit fastest, because each image is seen more often. For a set of 5 to 8 images, Scenario suggests dropping the learning rate to 5e-5 so the model does not memorise too aggressively. The step arithmetic and the documented targets per trainer are in face LoRA training steps and learning rate, and the free LoRA Training Steps Calculator shows how many times each image is seen for a given set of repeats and epochs.
Dim, alpha and learning rate: the parameters to change
When the dataset is sound and an earlier checkpoint is not enough, these are the documented dials. Change one per run, or you will not know which one worked.
| Parameter | Documented change | Where it comes from |
|---|---|---|
| Steps or epochs | Lower them; keep the checkpoint where prompts still worked | Civitai: lower steps for an overfit SDXL or Flux LoRA. Scenario: earlier epochs often give more flexible characters |
| Rank (network dim) | 16 to 24 on SDXL, 8 to 12 on Flux 1 | Civitai developer docs |
| Alpha | Half of the rank on SDXL | Civitai; sd-scripts also suggests about half of network_dim |
| Learning rate | 1e-4 to 2e-4 for an overcooked Klein LoRA; 5e-5 for a set of 5 to 8 images | Civitai for Klein; Scenario for small sets |
| Text encoder learning rate | Lower than the model's rate, such as 5e-5 beside 1e-4 | Hugging Face: the text encoder tends to overfit faster |
| network_dropout | A dropout rate inside the LoRA modules | sd-scripts: can be effective in suppressing overfitting |
| scale_weight_norms | 1.0 as a starting point | sd-scripts: can help prevent overfitting by limiting weight size |
- Rank is capacity. A higher rank can store more, including things you did not want stored. Hugging Face compared ranks 4, 16, 32 and 64 on SDXL and found rank 64 tended toward a more airbrushed look with less realistic skin texture.
- Alpha scales the learning rate. Civitai notes that alpha scales the effective rate by alpha divided by rank, and lists alpha equal to rank with a high learning rate among the causes of an SDXL LoRA that memorises.
- Learning rate burns first. Civitai's fix for overcooked or broken Klein samples is to lower the rate to 1e-4 to 2e-4, and it says to keep Flux 1 at or below 5e-4.
- The text encoder overfits faster. If you train it, Hugging Face recommends a lower rate for it than for the model.
Different trainers expose different dials. The sd-scripts options above exist in Kohya-based tools; Civitai's trainer exposes rank, alpha and learning rate; the built-in ComfyUI node exposes rank and learning rate only, as our ComfyUI training walkthrough explains.
Overfit, or just a narrow dataset?
These two look alike in a single image and need opposite fixes. The checkpoint grid tells them apart: overfitting gets worse as steps rise, while a gap in the dataset is there from the first checkpoint to the last.
- Earlier checkpoints follow prompts better than later ones
- Outputs reproduce training poses or backgrounds
- Skin and contrast get harsher as steps rise
- Lower strength brings prompt control back
- Fix: an earlier checkpoint, then a shorter or gentler run
- Every checkpoint fails on the same angle or expression
- Profiles look like a different person at every epoch
- One background or outfit follows the trigger from the start
- Lower strength loses the face without adding variety
- Fix: add the missing images and caption what varies
This split is our reading of the documented symptoms, and real LoRAs often show both. If you cannot tell, fix the dataset first. A wider dataset tolerates more steps, so it helps either way.
Prevent it on the next run
- No two images share the same angle, expression, light and background
- At least three distinct poses or angles, even in a set of five to eight images
- Backgrounds are mixed; no single room appears in most of the set
- Outfits vary, or the repeated one is captioned so it stays optional
- Captions name the trigger word and describe what should vary
- Checkpoints save every 250 steps or every epoch, and none are deleted automatically
- Sample prompts render during training, including an angle and an outfit not in the set
- Learning rate and rank start at the trainer's documented defaults, not above them
- If the text encoder is trained, its learning rate is lower than the model's
- You can name the symptom the last run showed, and you changed one thing to address it
Where you train changes how much of this is automatic. Civitai saves ten epochs and renders samples by default, FluxGym ships with samples switched off, and the fast trainers on fal and Replicate hand back one final file. Our Flux LoRA walkthrough lists each trainer's defaults, including the ones that sit on the hot side.
When to stop fixing and retrain
Retrain when no saved checkpoint passes the fixed-prompt test, or when the only checkpoint that follows prompts has a weak likeness. Keep the old run's grid: it tells you where the keeper would have been, which is your new step ceiling. Then change the one thing your row in the chart points to.
A LoRA that passes the test is still only a face. Turning it into a persona that holds up across outfits, scenes and video is a longer build, and the AI Influencers program goes through it in order: creating the persona, LoRA training for consistent faces with a checkpoint comparison matrix, and ControlNet for body consistency.
Face LoRA overfitting: FAQ
What is face LoRA overfitting?
It is a LoRA that has memorised its training photos instead of learning the face. The likeness is exact, but outputs repeat the poses, crops, expressions and backgrounds of the dataset and ignore prompts that ask for anything else. Scenario's training docs describe it as a character that appears stuck in the scenarios you trained on and cannot break out of training compositions.
How do I fix an overfit LoRA without retraining?
Use an earlier checkpoint. Scenario's docs say earlier epochs often produce more flexible characters, and Black Forest Labs says to pick the checkpoint whose samples look best, not necessarily the last one. If you only saved the final file, lower the LoRA strength in your generator as a stopgap: it weakens the memorised compositions, and the likeness with them.
Why does my LoRA only generate one pose?
Because the pose was part of what repeated. Scenario lists the same pose in every image as a pitfall: the model bakes the pose into the identity. It also says too many similar images push a model toward overfitting. Remove near-duplicate framings, include at least three distinct poses or angles, and try an earlier checkpoint before you change any training setting.
Why is my face LoRA too stiff, with the same expression every time?
Either the dataset shows one expression, or the run went on long enough to harden it. Check the first by testing an early checkpoint: if it is just as stiff, the cause is the dataset, so add expressions and angles. If early checkpoints are looser, the cause is exposure, so keep an earlier checkpoint or shorten the next run.
Does lowering network dim fix LoRA overfitting?
It is one documented fix. Civitai's developer docs say to drop the rank to 16 to 24 when an SDXL LoRA overfits and to 8 to 12 on Flux 1, or to set alpha to half of the rank on SDXL. Hugging Face found rank 64 gave SDXL faces a more airbrushed look than lower ranks. Try an earlier checkpoint first; it costs nothing.
How many steps is too many for a face LoRA?
There is no fixed number, because it depends on the image count, learning rate and rank. Black Forest Labs warns that loss keeps dropping well past the point where images start to overfit, so the loss curve will not tell you. Save a checkpoint every 250 steps or every epoch, render the same prompts on each, and stop where prompts start losing their grip.
How do I tell overfitting from underfitting?
Look at what fails. Scenario describes an underfit character as one whose identity drifts, with face, hair or signature traits that look generic or inconsistent. An overfit one has the opposite problem: the identity is exact, but outputs reproduce training poses or backgrounds. Weak likeness with working prompts means train longer; exact likeness with dead prompts means you went too far.
You fixed the LoRA. Now build what it is for.
AI Influencers, included in All Access, covers creating the persona, LoRA training with checkpoint comparison, ControlNet for body consistency and the video and content lessons after it, with the other three programs, live coaching and the private community in one subscription.
Stuck on a symptom that is not in the chart?
Post your sample grid in the free Discord and compare notes with other persona builders, and check how often each image is seen with the free steps calculator before the next run.