Skip to main content

Face LoRA Training Steps and Learning Rate: Settings That Work

Face LoRA training steps explained: the images x repeats x epochs formula, documented step targets, learning rate by optimizer and the signs of overtraining.

Founder of IImagined.ai

Published
Oct 11, 2026
Reading time
8 min read
Quick answer

Face LoRA training steps equal images times repeats times epochs, divided by batch size. Documented targets for a face or character run from about 1,000 to 2,500 steps at a learning rate of 1e-4 with AdamW8bit, or 1.0 with Prodigy. Save checkpoints every 250 steps and keep the one that holds the face while still following new prompts.

Face LoRA training steps come from one formula: images × repeats × epochs ÷ batch size. The documented targets for a face or character sit between about 1,000 and 2,500 steps at a learning rate of 1e-4 with AdamW8bit, and the right place to stop is whichever saved checkpoint holds the face while still following new prompts.

Run your own numbers in the free LoRA Training Steps Calculator: enter images, repeats, epochs and batch size, and it returns the step total kohya will show, plus training time and GPU cost. This page is the explanation behind it: the math, what the trainers document for steps and learning rate, and how to tell when a run has gone too far.

Settings checked in October 2026 against the sd-scripts training and FLUX.1 docs, AI Toolkit's FLUX example config, FluxGym's source, Black Forest Labs' klein training guide and example, Civitai's developer docs, Replicate's guide and the Prodigy README. These are the trainers' documented defaults and ranges, not the result of a benchmark we ran.

Steps and learning rate are the last two dials in a longer workflow. If you are still choosing between a LoRA and a reference-image method, start with the consistent character AI guide; for the full training walkthrough, see the LoRA training guide.

Face LoRA training steps: the formula

Total steps equal images times repeats, divided by batch size and rounded up, then multiplied by epochs.

How kohya-based trainers count steps
  1. 01
    Images × repeats

    Images shown per epoch

  2. 02
    ÷ batch size

    Rounded up

  3. 03
    Steps per epoch

    Summed across folders

  4. 04
    × epochs

    Full passes over the set

  5. 05
    Total steps

    The number on the progress bar

Kohya-based trainers (sd-scripts, the Kohya SS interface and FluxGym) never ask for a step count. You set repeats and epochs, and the trainer derives the total. Kohya's own explanation is that steps per epoch equal images times repeats divided by batch size, rounded up, and total steps equal that figure times the epochs. FluxGym shows the result as "Expected training steps" and ships with 10 repeats and 16 epochs.

AI Toolkit and the hosted trainers on Replicate, fal and Civitai work the other way round: you type a step count, and each step is one batch. Either way the question is the same, which is how many times the trainer sees each image. With gradient accumulation, kohya counts optimizer updates, so an accumulation of 4 shows about a quarter of the batch steps.

How to use the LoRA training steps calculator

The calculator takes the same inputs as the formula and shows the total before you commit GPU time to it.

Six inputs, one step total
  1. 1
    Enter images and repeats per folder

    One row per dataset folder. Add a row if you train a second folder with its own repeats.

  2. 2
    Set epochs

    Each epoch is one full pass and, on kohya-based trainers, one chance to save a checkpoint.

  3. 3
    Set batch size

    Use the value your VRAM allows. Doubling it halves the steps for the same images seen.

  4. 4
    Add gradient accumulation if you use it

    Leave it at 1 otherwise.

  5. 5
    Optional: seconds per step and GPU price

    Read the speed from the progress bar of a run that has settled.

  6. 6
    Compare the total with the documented range

    Then change repeats or epochs until it lands where you want it.

Worked example: 20 face images

Take the calculator's starting values: one folder of 20 images at 10 repeats, 10 epochs, batch size 2. Each epoch shows 200 images in 100 batches, so the run is 1,000 steps. At 1.5 seconds per step on a GPU rented for $0.50 an hour, both illustrative inputs, that is 25 minutes and about $0.21.

Now change one thing at a time. Batch size 1 doubles the run to 2,000 steps. FluxGym's defaults of 10 repeats and 16 epochs turn the same 20 images into 3,200 steps, close to the top of the range the AI Toolkit config calls good. At 26 images those defaults pass 4,000 steps, so on a bigger set lower the epochs or repeats instead of accepting them.

Step totals from documented defaults
Replicate, FLUX.1 fast trainer
1,000
fal, FLUX.1 fast training
1,000
AI Toolkit, FLUX.1 example
2,000
Civitai, Flux and SDXL
2,000
FluxGym defaults, 20 images
3,200

FluxGym computes epochs × images × repeats; the others take a step count directly. Source: each trainer's docs, config or source code, checked October 2026

How many steps for a face LoRA?

Between about 1,000 and 2,500 on current FLUX and SDXL trainers, with the smaller figures for smaller sets. Here is what each trainer documents.

SourceModelDocumented stepsNotes
BFL klein training exampleFLUX.2 [klein]800 to 1,200For a character LoRA of 10 to 15 images; 1,200 to 2,000 for 20 to 40 images
BFL klein training guideFLUX.2 [klein]1,500 to 3,000Character LoRAs in general; style LoRAs 1,500 to 2,500
Civitai developer docsFlux 1, Flux 2 Klein, SDXL2,000 defaultSuggests 1,500 to 2,500 when a Flux character LoRA is under-trained; hard cap 10,000
AI Toolkit example configFLUX.1 [dev]2,000Comment in the file: 500 to 4,000 is a good range; saves every 250 steps
Replicate docsFLUX.1 fast trainer1,000Says to leave it there: fewer under-learns, more adds cost without much gain
fal model pageFLUX.1 fast training1,000 defaultPrice scales linearly with steps
Astria docsFLUX fine-tunes27 per imagePortrait preset, 300-step minimum; 100 per image on the High preset
Hugging Face, Jan 2024SDXL75 to 120 per imageTested on a six-photo face set

Sources: Black Forest Labs example and guide, Civitai recipes for Flux 1, Flux 2 Klein and SDXL, AI Toolkit, Replicate, fal, Astria and Hugging Face.

Black Forest Labs' two pages disagree on characters: 800 to 1,200 steps in one, 1,500 to 3,000 in the other. The lower figure is tied to a dataset of 10 to 15 images, so use it for a small face set. The spread is the lesson. Do not train to a number; save checkpoints through the range and pick. The step target also follows the image count, which is why how many images to train a face LoRA is the question to settle first.

Face LoRA learning rate, by optimizer

Use 1e-4 with AdamW8bit unless your trainer's docs say otherwise. Adaptive optimizers take 1.0 instead, because they estimate the rate themselves.

OptimizerDocumented rateWhere it comes from
AdamW or AdamW8bit1e-4Used in the sd-scripts SDXL and FLUX.1 examples, the AI Toolkit example and Civitai's Flux defaults. BFL gives 8e-5 to 1e-4 for klein LoRAs.
Adafactor, fixed rate1e-4 to 5e-4Civitai's SDXL default optimizer: 1e-4 by default, with 5e-4 called typical for character and style LoRAs. sd-scripts documents it for low-VRAM FLUX.1 runs.
Prodigy1.0The README recommends lr=1.0 for all networks, plus safeguard_warmup, use_bias_correction and weight_decay=0.01 for diffusion models.
D-Adaptation1.0The README says to set LR to 1.0 and to try Prodigy if the estimated rate comes out too low.

Sources: sd-scripts SDXL and FLUX.1 docs, Civitai developer docs, the Prodigy and D-Adaptation READMEs.

  • There is a ceiling. Civitai's docs call Flux 1 sensitive to high rates and say to keep it at or below 5e-4, and to keep klein between 1e-4 and 5e-4. FluxGym ships with 8e-4, above that ceiling, so if a FluxGym run burns early, the learning rate is the first thing we would lower.
  • Alpha changes the effective rate. Civitai notes that alpha scales the learning rate by alpha divided by rank. The sd-scripts range of 1e-4 to 1e-3 is given for an alpha of 1; with alpha equal to rank, stay at the low end.
  • The text encoder gets less. sd-scripts recommends a smaller rate for the text encoder than for the U-Net and uses 1e-5 in its SDXL example; Civitai defaults to 5e-5. The AI Toolkit FLUX example does not train the text encoder at all.
  • Read the loss before you move it. Black Forest Labs starts at 1e-4, raises to 2e-4 if the loss falls slowly and drops to 5e-5 if it oscillates. It also suggests higher rates for characters than for styles.

Overtraining symptoms to watch

Steps and learning rate multiply. Too much of either and the LoRA stops learning the face and starts memorising the photos. Black Forest Labs' guide says to monitor sample outputs to avoid overfitting, and samples show the problem before the loss does.

What the checkpoints look like as steps rise
  1. Early
    A generic face

    The trigger word changes the output, but the person is only loosely similar. Under-trained.

  2. Middle
    Likeness arrives

    The face is recognisable at normal strength and prompts still change pose, outfit and light. Keep this one.

  3. Late
    Prompts lose their grip

    The face is exact, but expressions, angles and backgrounds drift toward the training photos.

  4. Too late
    Copies

    Outputs reproduce the crop and lighting of a training image, and skin turns waxy or over-sharpened.

Signs of an overtrained face LoRA
  • The same pose, crop or head tilt keeps coming back across different prompts
  • A training outfit or background appears when you did not ask for it
  • Prompts for a new expression or angle are ignored
  • Skin looks waxy, burned or over-sharpened at normal strength
  • The LoRA only behaves when you lower its strength well below 1
  • An earlier checkpoint follows the same prompt better than the final one

Fix it in this order, cheapest first:

  1. Use an earlier checkpoint. It costs nothing, which is why you save one every 250 steps or every epoch.
  2. Lower the steps on the next run to where the good checkpoint was.
  3. Lower the learning rate. Civitai's advice for an over-trained klein LoRA is to drop to 1e-4 to 2e-4.
  4. Lower the rank. For Flux 1, Civitai suggests a network dimension of 8 to 12 when a LoRA overfits.
  5. Fix the dataset. Near-duplicate images multiply exposure in a way no setting undoes. The character LoRA dataset checklist covers culling.

The opposite problem, a face that never quite arrives, gets the opposite treatment: Civitai suggests raising Flux steps to 1,500 to 2,500 and keeping the rate between 1e-4 and 3e-4. Settings are one part of a persona build. The AI Influencers program covers the rest in order: designing the character, building the dataset, testing checkpoints against a fixed prompt set and turning the LoRA into a content system.

Face LoRA steps and learning rate: FAQ

How many steps should I train a face LoRA for?

Most documented starting points sit between 1,000 and 2,500 steps. Replicate and fal default to 1,000 for FLUX.1, the AI Toolkit example config and Civitai default to 2,000, and Black Forest Labs gives 800 to 1,200 steps for a FLUX.2 klein character LoRA of 10 to 15 images. Save a checkpoint every 250 steps and keep the best one instead of trusting the last. Checked October 2026.

What is the best learning rate for a face LoRA?

Start at 1e-4 with AdamW8bit. That value appears in the sd-scripts SDXL and FLUX.1 examples, the AI Toolkit example config and Civitai's defaults, and Black Forest Labs gives 8e-5 to 1e-4 for FLUX.2 klein LoRAs. Civitai advises keeping Flux 1 at or below 5e-4. With Prodigy or D-Adaptation you set the learning rate to 1.0 and the optimizer finds the rate itself.

How do I calculate LoRA training steps?

Multiply images by repeats, divide by batch size and round up to get steps per epoch, then multiply by epochs. Twenty images at 10 repeats and batch size 2 is 100 steps per epoch, so 10 epochs is 1,000 steps. Trainers such as AI Toolkit, Replicate and fal skip the formula and ask for a step count directly, where one step is one batch.

Is 2,000 steps too many for a face LoRA?

It depends on the dataset. Two thousand is the default in the AI Toolkit FLUX.1 example and on Civitai, so it is a normal figure, but Black Forest Labs suggests 800 to 1,200 steps for a character set of only 10 to 15 images. With a small set, save checkpoints from about 750 steps onward and compare them; the best one is often earlier than the final step.

What learning rate do I use with Prodigy?

Set it to 1.0. The Prodigy README recommends lr=1.0 for all networks because the optimizer estimates the real rate during training. For diffusion models it also recommends safeguard_warmup=True, use_bias_correction=True and weight_decay=0.01, with either no scheduler or cosine annealing without restarts. Typing 1e-4 into Prodigy by habit gives a run that barely learns.

How do I know my face LoRA is overtrained?

Look at samples, not the loss. An overtrained face LoRA repeats the poses, crops and backgrounds of its training photos, ignores prompts for a new expression or outfit, and often shows waxy or over-sharpened skin. If the face only behaves when you lower the LoRA strength, you have gone too far. Go back to an earlier checkpoint before changing any setting.

Do repeats or epochs matter more for step count?

Neither. They multiply, so 10 repeats over 10 epochs shows each image the same 100 times as 1 repeat over 100 epochs. The practical difference is that kohya-based trainers can save a checkpoint every epoch, so more epochs with fewer repeats gives you more checkpoints to compare, and repeats let you weight one folder of images above another.

All Access · all four programs · $99/mo

The settings take an afternoon. The persona is the real project.

AI Influencers, included in All Access, walks through the full build: the character, the dataset, LoRA training and checkpoint testing, then content and monetisation, with the other three programs, live coaching and the private community in one subscription.

Start All Access — $99/mo →30-day money-back guarantee
Free · no signup

Run the numbers before you train

Get steps per epoch, total steps, training time and GPU cost from your own images, repeats, epochs and batch size, and join the free Telegram channel for persona workflows that are working now.