Skip to main content

Face LoRA Overfitting: Why Your LoRA Only Makes One Photo

Face LoRA overfitting diagnosed: six symptoms, from one repeated pose to waxy skin, each mapped to the dataset, steps, rank or learning rate fix to make.

Founder of IImagined.ai

Published
Oct 11, 2026
Reading time
13 min read
Quick answer

Face LoRA overfitting means the LoRA memorised its training photos instead of learning the face, so every output repeats one pose, crop, expression or background. Fix it cheapest-first: use an earlier checkpoint, then repair the dataset's variety and captions, and only then lower steps, learning rate or rank on the next run. Six symptoms map to six different fixes, and most of them are dataset fixes, not settings.

Face LoRA overfitting means the LoRA has memorised its training photos instead of learning the face, so every output repeats the same pose, crop, expression or background. Fix it in this order: switch to an earlier checkpoint, then repair the dataset's variety and captions, and only then lower the steps, learning rate or rank for the next run.

Checked in October 2026 against Scenario's character training guide, Black Forest Labs' klein LoRA guide and training example, Civitai's developer docs for SDXL, Flux 1 and Flux 2 Klein, Hugging Face's SDXL LoRA write-up, the sd-scripts advanced training docs and AI Toolkit's example config. The fixes are quoted from those documents. Matching each symptom to a cause is our reading of them; we ran no training benchmark for this page.

This page is for someone who already trained a LoRA and does not like what came out. If you have not trained yet, the consistent character AI guide explains where a LoRA fits, and the LoRA training guide covers the first run.

What face LoRA overfitting is, and what it is not

A face LoRA has two jobs that pull against each other: hold the likeness, and stay out of the way of the prompt. Overfitting is when the first job wins completely. Scenario's docs describe the result as a character that appears stuck in the scenarios you trained on, with outputs that reproduce specific training poses or backgrounds.

It is easy to confuse with two other failures. An under-trained LoRA follows prompts but the face is generic. A LoRA trained on a messy dataset does neither job well. They need opposite fixes, so place your LoRA on this grid before you touch a setting.

Likeness against flexibility: where a checkpoint can land
Strong likeness
Overfit. The face is exact, but pose, expression and background come from the training photos.
The keeper. The face is recognisable, and a new prompt still changes everything else.
Weak likeness
Dataset problem. Conflicting or repetitive images: neither a clear face nor prompt control.
Under-trained. Prompts work, but the face is generic. Use a later checkpoint or train longer.
Ignores prompts
Follows prompts

The diagnosis chart: six symptoms and the parameter to change

Overfitting is not one problem with one dial. Find the row that matches your samples, try the free fix, and change the listed parameter only if you retrain.

What you seeWhat it usually meansFree fix, todayChange on the next run
1. The same pose, crop or composition every timeThe LoRA memorised compositions: too many steps for this set, or one framing repeated in itGo back to an earlier checkpointSteps or epochs down; near-duplicate framings out
2. The training room or backdrop keeps appearingOne background dominates the dataset, so it became part of the characterPrompt a specific new setting on an earlier checkpointDataset: mix the backgrounds and caption each one
3. A training outfit shows up unpromptedThe outfit repeats and the captions never mention it, so it attached to the trigger wordName a different outfit explicitly in the promptCaptions: describe what should vary; vary the outfits
4. One expression; new angles and expressions are ignoredNarrow coverage of angles and expressions, hardened by a long runEarlier checkpoint, then test againDataset: at least three distinct poses or angles; fewer repeats
5. Waxy, airbrushed or over-sharpened skinToo much capacity, or too high a learning rateEarlier checkpointRank down; learning rate down
6. It only behaves at reduced strengthThe weights are overcooked and overpower the base model at full strengthRun it at lower strength as a stopgapSteps down, rank down, alpha to half of rank

Count the last column: three of the six rows are fixed in the dataset or its captions, not in the trainer. That matches what the vendor docs keep saying. Scenario calls repetitive framing the most common cause of weak character recall, and Hugging Face's write-up says fewer, better-curated images beat more images of middling quality.

How to confirm it: the fixed-prompt test

Do not diagnose from one unlucky image. Run a small, fixed set of prompts against every checkpoint you saved, changing nothing else. Black Forest Labs' guide puts the principle in six words: watch the samples, not the loss.

The test that separates overfit from unlucky
  1. 1
    Fix everything except the checkpoint

    Same base model, sampler, steps, size and seed for every image, with the LoRA at full strength.

  2. 2
    Prompt 1: a neutral portrait

    Trigger word plus a plain head-and-shoulders shot. This is the likeness check.

  3. 3
    Prompt 2: a profile

    Or any angle that is rare or missing in the dataset.

  4. 4
    Prompt 3: a new expression

    Laughing, surprised or eyes closed: whichever your dataset has least of.

  5. 5
    Prompt 4: full body, new outfit, new place

    An outfit and a location that appear nowhere in the training set.

  6. 6
    Run all four on every checkpoint

    Lay the results out as a grid, prompts across and checkpoints down.

  7. 7
    Read down each column

    Likeness arriving while prompts still work marks the keeper. Prompts failing as you go down is overfitting.

  8. 8
    Repeat the last checkpoint at lower strength

    If it only obeys prompts at reduced strength, the run went past the point you wanted.

The grid answers two questions at once: whether the LoRA is overfit, and which earlier checkpoint is not. If you train on Civitai, Training Studio's Compare epochs view builds the same grid from your sample prompts; our Civitai LoRA training walkthrough covers it.

LoRA overfitting fix: the order to try things

Work from the cheapest fix to the most expensive. The first two cost nothing, the middle two cost an hour of dataset work, and only the last two need new settings.

Cheapest fix first
  1. 01
    Earlier checkpoint

    Free, and often all you need

  2. 02
    Lower strength

    A stopgap, not a cure

  3. 03
    Dataset variety

    Cut duplicates, add what is missing

  4. 04
    Captions

    Describe what should vary

  5. 05
    Shorter run

    Fewer steps or epochs

  6. 06
    Rate and rank

    Lower learning rate, rank or alpha

Start with the checkpoint you already have. Scenario notes that the last epoch is set as the default but earlier epochs often produce more flexible characters. Black Forest Labs says the same about its own runs: pick the checkpoint whose samples look best, not necessarily the last one.

Lower strength only buys time. Turning the LoRA down in your generator weakens everything it learned, the memorised compositions and the likeness together. It can rescue a batch of images today. It does not repair the file.

LoRA only one pose: fix the dataset first

When every output has the same pose or crop, the pose was part of what repeated. Scenario lists it as a named pitfall: same pose in every image, and the model bakes the pose into the identity. No setting separates a face from a pose the dataset never separates.

  • Cut near-duplicates. Scenario says too many similar images push the model toward overfitting, and that a curated set of 5 to 15 images often beats a set of 20 or more where images become redundant.
  • Cover at least three poses or angles, even in a set of five to eight images: front, profile, three-quarter, and the back when possible.
  • Mix the backgrounds. In Scenario's words, when all images are shot in the same room, the character wants to be in that room.
  • Decide what the outfit is. Either it is part of the identity, or it is a variable. If it is a variable, vary it and caption it.

Captions do the same job from the other side. Scenario's rule is to caption the elements that are meant to vary in the generated outputs. Black Forest Labs explains the mechanism for styles: leave the style out of the captions and the model bakes it into the weights. Whatever your captions never mention gets absorbed into the trigger word. For a face that is what you want for the bone structure, and the opposite of what you want for a hoodie.

The character LoRA dataset checklist gives a shot list built to make the face the only constant, and how many images to train a face LoRA covers why adding more of the same makes this worse.

Face LoRA too stiff: steps, epochs and when to stop

If early checkpoints are looser than late ones, the run simply went on too long for that dataset. The loss curve will not warn you. Black Forest Labs writes that loss keeps dropping well past the point where the images start to overfit.

Steps trained, and the checkpoint the authors kept
BFL klein example: steps configured
3,000
BFL klein example: checkpoint kept
1,500
Hugging Face SDXL: first run, overfit
1,500
Hugging Face SDXL: runs kept
1,000

Both are style LoRAs, not faces, and the Hugging Face runs changed batch size and repeats as well. The pattern is the point: the keeper sat well before the end. Source: Black Forest Labs klein training example and Hugging Face SDXL LoRA write-up, checked October 2026

So the practical rule is to make stopping cheap. Save a checkpoint every 250 steps or every epoch, render sample prompts as you go, and treat the step count as a ceiling you expect to stop short of. Black Forest Labs puts the visual peak of most of its style LoRAs between steps 750 and 1,500.

Small datasets overfit fastest, because each image is seen more often. For a set of 5 to 8 images, Scenario suggests dropping the learning rate to 5e-5 so the model does not memorise too aggressively. The step arithmetic and the documented targets per trainer are in face LoRA training steps and learning rate, and the free LoRA Training Steps Calculator shows how many times each image is seen for a given set of repeats and epochs.

Dim, alpha and learning rate: the parameters to change

When the dataset is sound and an earlier checkpoint is not enough, these are the documented dials. Change one per run, or you will not know which one worked.

ParameterDocumented changeWhere it comes from
Steps or epochsLower them; keep the checkpoint where prompts still workedCivitai: lower steps for an overfit SDXL or Flux LoRA. Scenario: earlier epochs often give more flexible characters
Rank (network dim)16 to 24 on SDXL, 8 to 12 on Flux 1Civitai developer docs
AlphaHalf of the rank on SDXLCivitai; sd-scripts also suggests about half of network_dim
Learning rate1e-4 to 2e-4 for an overcooked Klein LoRA; 5e-5 for a set of 5 to 8 imagesCivitai for Klein; Scenario for small sets
Text encoder learning rateLower than the model's rate, such as 5e-5 beside 1e-4Hugging Face: the text encoder tends to overfit faster
network_dropoutA dropout rate inside the LoRA modulessd-scripts: can be effective in suppressing overfitting
scale_weight_norms1.0 as a starting pointsd-scripts: can help prevent overfitting by limiting weight size
  • Rank is capacity. A higher rank can store more, including things you did not want stored. Hugging Face compared ranks 4, 16, 32 and 64 on SDXL and found rank 64 tended toward a more airbrushed look with less realistic skin texture.
  • Alpha scales the learning rate. Civitai notes that alpha scales the effective rate by alpha divided by rank, and lists alpha equal to rank with a high learning rate among the causes of an SDXL LoRA that memorises.
  • Learning rate burns first. Civitai's fix for overcooked or broken Klein samples is to lower the rate to 1e-4 to 2e-4, and it says to keep Flux 1 at or below 5e-4.
  • The text encoder overfits faster. If you train it, Hugging Face recommends a lower rate for it than for the model.

Different trainers expose different dials. The sd-scripts options above exist in Kohya-based tools; Civitai's trainer exposes rank, alpha and learning rate; the built-in ComfyUI node exposes rank and learning rate only, as our ComfyUI training walkthrough explains.

Overfit, or just a narrow dataset?

These two look alike in a single image and need opposite fixes. The checkpoint grid tells them apart: overfitting gets worse as steps rise, while a gap in the dataset is there from the first checkpoint to the last.

Two causes of a rigid LoRA
Overfitting: too much exposure
  • Earlier checkpoints follow prompts better than later ones
  • Outputs reproduce training poses or backgrounds
  • Skin and contrast get harsher as steps rise
  • Lower strength brings prompt control back
  • Fix: an earlier checkpoint, then a shorter or gentler run
Narrow dataset: too little variety
  • Every checkpoint fails on the same angle or expression
  • Profiles look like a different person at every epoch
  • One background or outfit follows the trigger from the start
  • Lower strength loses the face without adding variety
  • Fix: add the missing images and caption what varies

This split is our reading of the documented symptoms, and real LoRAs often show both. If you cannot tell, fix the dataset first. A wider dataset tolerates more steps, so it helps either way.

Prevent it on the next run

Before you press train again
  • No two images share the same angle, expression, light and background
  • At least three distinct poses or angles, even in a set of five to eight images
  • Backgrounds are mixed; no single room appears in most of the set
  • Outfits vary, or the repeated one is captioned so it stays optional
  • Captions name the trigger word and describe what should vary
  • Checkpoints save every 250 steps or every epoch, and none are deleted automatically
  • Sample prompts render during training, including an angle and an outfit not in the set
  • Learning rate and rank start at the trainer's documented defaults, not above them
  • If the text encoder is trained, its learning rate is lower than the model's
  • You can name the symptom the last run showed, and you changed one thing to address it

Where you train changes how much of this is automatic. Civitai saves ten epochs and renders samples by default, FluxGym ships with samples switched off, and the fast trainers on fal and Replicate hand back one final file. Our Flux LoRA walkthrough lists each trainer's defaults, including the ones that sit on the hot side.

When to stop fixing and retrain

Retrain when no saved checkpoint passes the fixed-prompt test, or when the only checkpoint that follows prompts has a weak likeness. Keep the old run's grid: it tells you where the keeper would have been, which is your new step ceiling. Then change the one thing your row in the chart points to.

A LoRA that passes the test is still only a face. Turning it into a persona that holds up across outfits, scenes and video is a longer build, and the AI Influencers program goes through it in order: creating the persona, LoRA training for consistent faces with a checkpoint comparison matrix, and ControlNet for body consistency.

Face LoRA overfitting: FAQ

What is face LoRA overfitting?

It is a LoRA that has memorised its training photos instead of learning the face. The likeness is exact, but outputs repeat the poses, crops, expressions and backgrounds of the dataset and ignore prompts that ask for anything else. Scenario's training docs describe it as a character that appears stuck in the scenarios you trained on and cannot break out of training compositions.

How do I fix an overfit LoRA without retraining?

Use an earlier checkpoint. Scenario's docs say earlier epochs often produce more flexible characters, and Black Forest Labs says to pick the checkpoint whose samples look best, not necessarily the last one. If you only saved the final file, lower the LoRA strength in your generator as a stopgap: it weakens the memorised compositions, and the likeness with them.

Why does my LoRA only generate one pose?

Because the pose was part of what repeated. Scenario lists the same pose in every image as a pitfall: the model bakes the pose into the identity. It also says too many similar images push a model toward overfitting. Remove near-duplicate framings, include at least three distinct poses or angles, and try an earlier checkpoint before you change any training setting.

Why is my face LoRA too stiff, with the same expression every time?

Either the dataset shows one expression, or the run went on long enough to harden it. Check the first by testing an early checkpoint: if it is just as stiff, the cause is the dataset, so add expressions and angles. If early checkpoints are looser, the cause is exposure, so keep an earlier checkpoint or shorten the next run.

Does lowering network dim fix LoRA overfitting?

It is one documented fix. Civitai's developer docs say to drop the rank to 16 to 24 when an SDXL LoRA overfits and to 8 to 12 on Flux 1, or to set alpha to half of the rank on SDXL. Hugging Face found rank 64 gave SDXL faces a more airbrushed look than lower ranks. Try an earlier checkpoint first; it costs nothing.

How many steps is too many for a face LoRA?

There is no fixed number, because it depends on the image count, learning rate and rank. Black Forest Labs warns that loss keeps dropping well past the point where images start to overfit, so the loss curve will not tell you. Save a checkpoint every 250 steps or every epoch, render the same prompts on each, and stop where prompts start losing their grip.

How do I tell overfitting from underfitting?

Look at what fails. Scenario describes an underfit character as one whose identity drifts, with face, hair or signature traits that look generic or inconsistent. An overfit one has the opposite problem: the identity is exact, but outputs reproduce training poses or backgrounds. Weak likeness with working prompts means train longer; exact likeness with dead prompts means you went too far.

All Access · all four programs · $99/mo

You fixed the LoRA. Now build what it is for.

AI Influencers, included in All Access, covers creating the persona, LoRA training with checkpoint comparison, ControlNet for body consistency and the video and content lessons after it, with the other three programs, live coaching and the private community in one subscription.

Start All Access — $99/mo →30-day money-back guarantee
Free

Stuck on a symptom that is not in the chart?

Post your sample grid in the free Discord and compare notes with other persona builders, and check how often each image is seen with the free steps calculator before the next run.