Face LoRA training steps equal images times repeats times epochs, divided by batch size. Documented targets for a face or character run from about 1,000 to 2,500 steps at a learning rate of 1e-4 with AdamW8bit, or 1.0 with Prodigy. Save checkpoints every 250 steps and keep the one that holds the face while still following new prompts.
Face LoRA training steps come from one formula: images × repeats × epochs ÷ batch size. The documented targets for a face or character sit between about 1,000 and 2,500 steps at a learning rate of 1e-4 with AdamW8bit, and the right place to stop is whichever saved checkpoint holds the face while still following new prompts.
Run your own numbers in the free LoRA Training Steps Calculator: enter images, repeats, epochs and batch size, and it returns the step total kohya will show, plus training time and GPU cost. This page is the explanation behind it: the math, what the trainers document for steps and learning rate, and how to tell when a run has gone too far.
Settings checked in October 2026 against the sd-scripts training and FLUX.1 docs, AI Toolkit's FLUX example config, FluxGym's source, Black Forest Labs' klein training guide and example, Civitai's developer docs, Replicate's guide and the Prodigy README. These are the trainers' documented defaults and ranges, not the result of a benchmark we ran.
Steps and learning rate are the last two dials in a longer workflow. If you are still choosing between a LoRA and a reference-image method, start with the consistent character AI guide; for the full training walkthrough, see the LoRA training guide.
Face LoRA training steps: the formula
Total steps equal images times repeats, divided by batch size and rounded up, then multiplied by epochs.
- 01Images × repeats
Images shown per epoch
- 02÷ batch size
Rounded up
- 03Steps per epoch
Summed across folders
- 04× epochs
Full passes over the set
- 05Total steps
The number on the progress bar
Kohya-based trainers (sd-scripts, the Kohya SS interface and FluxGym) never ask for a step count. You set repeats and epochs, and the trainer derives the total. Kohya's own explanation is that steps per epoch equal images times repeats divided by batch size, rounded up, and total steps equal that figure times the epochs. FluxGym shows the result as "Expected training steps" and ships with 10 repeats and 16 epochs.
AI Toolkit and the hosted trainers on Replicate, fal and Civitai work the other way round: you type a step count, and each step is one batch. Either way the question is the same, which is how many times the trainer sees each image. With gradient accumulation, kohya counts optimizer updates, so an accumulation of 4 shows about a quarter of the batch steps.
How to use the LoRA training steps calculator
The calculator takes the same inputs as the formula and shows the total before you commit GPU time to it.
- 1Enter images and repeats per folder
One row per dataset folder. Add a row if you train a second folder with its own repeats.
- 2Set epochs
Each epoch is one full pass and, on kohya-based trainers, one chance to save a checkpoint.
- 3Set batch size
Use the value your VRAM allows. Doubling it halves the steps for the same images seen.
- 4Add gradient accumulation if you use it
Leave it at 1 otherwise.
- 5Optional: seconds per step and GPU price
Read the speed from the progress bar of a run that has settled.
- 6Compare the total with the documented range
Then change repeats or epochs until it lands where you want it.
Worked example: 20 face images
Take the calculator's starting values: one folder of 20 images at 10 repeats, 10 epochs, batch size 2. Each epoch shows 200 images in 100 batches, so the run is 1,000 steps. At 1.5 seconds per step on a GPU rented for $0.50 an hour, both illustrative inputs, that is 25 minutes and about $0.21.
Now change one thing at a time. Batch size 1 doubles the run to 2,000 steps. FluxGym's defaults of 10 repeats and 16 epochs turn the same 20 images into 3,200 steps, close to the top of the range the AI Toolkit config calls good. At 26 images those defaults pass 4,000 steps, so on a bigger set lower the epochs or repeats instead of accepting them.
FluxGym computes epochs × images × repeats; the others take a step count directly. Source: each trainer's docs, config or source code, checked October 2026
How many steps for a face LoRA?
Between about 1,000 and 2,500 on current FLUX and SDXL trainers, with the smaller figures for smaller sets. Here is what each trainer documents.
| Source | Model | Documented steps | Notes |
|---|---|---|---|
| BFL klein training example | FLUX.2 [klein] | 800 to 1,200 | For a character LoRA of 10 to 15 images; 1,200 to 2,000 for 20 to 40 images |
| BFL klein training guide | FLUX.2 [klein] | 1,500 to 3,000 | Character LoRAs in general; style LoRAs 1,500 to 2,500 |
| Civitai developer docs | Flux 1, Flux 2 Klein, SDXL | 2,000 default | Suggests 1,500 to 2,500 when a Flux character LoRA is under-trained; hard cap 10,000 |
| AI Toolkit example config | FLUX.1 [dev] | 2,000 | Comment in the file: 500 to 4,000 is a good range; saves every 250 steps |
| Replicate docs | FLUX.1 fast trainer | 1,000 | Says to leave it there: fewer under-learns, more adds cost without much gain |
| fal model page | FLUX.1 fast training | 1,000 default | Price scales linearly with steps |
| Astria docs | FLUX fine-tunes | 27 per image | Portrait preset, 300-step minimum; 100 per image on the High preset |
| Hugging Face, Jan 2024 | SDXL | 75 to 120 per image | Tested on a six-photo face set |
Sources: Black Forest Labs example and guide, Civitai recipes for Flux 1, Flux 2 Klein and SDXL, AI Toolkit, Replicate, fal, Astria and Hugging Face.
Black Forest Labs' two pages disagree on characters: 800 to 1,200 steps in one, 1,500 to 3,000 in the other. The lower figure is tied to a dataset of 10 to 15 images, so use it for a small face set. The spread is the lesson. Do not train to a number; save checkpoints through the range and pick. The step target also follows the image count, which is why how many images to train a face LoRA is the question to settle first.
Face LoRA learning rate, by optimizer
Use 1e-4 with AdamW8bit unless your trainer's docs say otherwise. Adaptive optimizers take 1.0 instead, because they estimate the rate themselves.
| Optimizer | Documented rate | Where it comes from |
|---|---|---|
| AdamW or AdamW8bit | 1e-4 | Used in the sd-scripts SDXL and FLUX.1 examples, the AI Toolkit example and Civitai's Flux defaults. BFL gives 8e-5 to 1e-4 for klein LoRAs. |
| Adafactor, fixed rate | 1e-4 to 5e-4 | Civitai's SDXL default optimizer: 1e-4 by default, with 5e-4 called typical for character and style LoRAs. sd-scripts documents it for low-VRAM FLUX.1 runs. |
| Prodigy | 1.0 | The README recommends lr=1.0 for all networks, plus safeguard_warmup, use_bias_correction and weight_decay=0.01 for diffusion models. |
| D-Adaptation | 1.0 | The README says to set LR to 1.0 and to try Prodigy if the estimated rate comes out too low. |
Sources: sd-scripts SDXL and FLUX.1 docs, Civitai developer docs, the Prodigy and D-Adaptation READMEs.
- There is a ceiling. Civitai's docs call Flux 1 sensitive to high rates and say to keep it at or below 5e-4, and to keep klein between 1e-4 and 5e-4. FluxGym ships with 8e-4, above that ceiling, so if a FluxGym run burns early, the learning rate is the first thing we would lower.
- Alpha changes the effective rate. Civitai notes that alpha scales the learning rate by alpha divided by rank. The sd-scripts range of 1e-4 to 1e-3 is given for an alpha of 1; with alpha equal to rank, stay at the low end.
- The text encoder gets less. sd-scripts recommends a smaller rate for the text encoder than for the U-Net and uses 1e-5 in its SDXL example; Civitai defaults to 5e-5. The AI Toolkit FLUX example does not train the text encoder at all.
- Read the loss before you move it. Black Forest Labs starts at 1e-4, raises to 2e-4 if the loss falls slowly and drops to 5e-5 if it oscillates. It also suggests higher rates for characters than for styles.
Overtraining symptoms to watch
Steps and learning rate multiply. Too much of either and the LoRA stops learning the face and starts memorising the photos. Black Forest Labs' guide says to monitor sample outputs to avoid overfitting, and samples show the problem before the loss does.
- EarlyA generic face
The trigger word changes the output, but the person is only loosely similar. Under-trained.
- MiddleLikeness arrives
The face is recognisable at normal strength and prompts still change pose, outfit and light. Keep this one.
- LatePrompts lose their grip
The face is exact, but expressions, angles and backgrounds drift toward the training photos.
- Too lateCopies
Outputs reproduce the crop and lighting of a training image, and skin turns waxy or over-sharpened.
- The same pose, crop or head tilt keeps coming back across different prompts
- A training outfit or background appears when you did not ask for it
- Prompts for a new expression or angle are ignored
- Skin looks waxy, burned or over-sharpened at normal strength
- The LoRA only behaves when you lower its strength well below 1
- An earlier checkpoint follows the same prompt better than the final one
Fix it in this order, cheapest first:
- Use an earlier checkpoint. It costs nothing, which is why you save one every 250 steps or every epoch.
- Lower the steps on the next run to where the good checkpoint was.
- Lower the learning rate. Civitai's advice for an over-trained klein LoRA is to drop to 1e-4 to 2e-4.
- Lower the rank. For Flux 1, Civitai suggests a network dimension of 8 to 12 when a LoRA overfits.
- Fix the dataset. Near-duplicate images multiply exposure in a way no setting undoes. The character LoRA dataset checklist covers culling.
The opposite problem, a face that never quite arrives, gets the opposite treatment: Civitai suggests raising Flux steps to 1,500 to 2,500 and keeping the rate between 1e-4 and 3e-4. Settings are one part of a persona build. The AI Influencers program covers the rest in order: designing the character, building the dataset, testing checkpoints against a fixed prompt set and turning the LoRA into a content system.
Face LoRA steps and learning rate: FAQ
How many steps should I train a face LoRA for?
Most documented starting points sit between 1,000 and 2,500 steps. Replicate and fal default to 1,000 for FLUX.1, the AI Toolkit example config and Civitai default to 2,000, and Black Forest Labs gives 800 to 1,200 steps for a FLUX.2 klein character LoRA of 10 to 15 images. Save a checkpoint every 250 steps and keep the best one instead of trusting the last. Checked October 2026.
What is the best learning rate for a face LoRA?
Start at 1e-4 with AdamW8bit. That value appears in the sd-scripts SDXL and FLUX.1 examples, the AI Toolkit example config and Civitai's defaults, and Black Forest Labs gives 8e-5 to 1e-4 for FLUX.2 klein LoRAs. Civitai advises keeping Flux 1 at or below 5e-4. With Prodigy or D-Adaptation you set the learning rate to 1.0 and the optimizer finds the rate itself.
How do I calculate LoRA training steps?
Multiply images by repeats, divide by batch size and round up to get steps per epoch, then multiply by epochs. Twenty images at 10 repeats and batch size 2 is 100 steps per epoch, so 10 epochs is 1,000 steps. Trainers such as AI Toolkit, Replicate and fal skip the formula and ask for a step count directly, where one step is one batch.
Is 2,000 steps too many for a face LoRA?
It depends on the dataset. Two thousand is the default in the AI Toolkit FLUX.1 example and on Civitai, so it is a normal figure, but Black Forest Labs suggests 800 to 1,200 steps for a character set of only 10 to 15 images. With a small set, save checkpoints from about 750 steps onward and compare them; the best one is often earlier than the final step.
What learning rate do I use with Prodigy?
Set it to 1.0. The Prodigy README recommends lr=1.0 for all networks because the optimizer estimates the real rate during training. For diffusion models it also recommends safeguard_warmup=True, use_bias_correction=True and weight_decay=0.01, with either no scheduler or cosine annealing without restarts. Typing 1e-4 into Prodigy by habit gives a run that barely learns.
How do I know my face LoRA is overtrained?
Look at samples, not the loss. An overtrained face LoRA repeats the poses, crops and backgrounds of its training photos, ignores prompts for a new expression or outfit, and often shows waxy or over-sharpened skin. If the face only behaves when you lower the LoRA strength, you have gone too far. Go back to an earlier checkpoint before changing any setting.
Do repeats or epochs matter more for step count?
Neither. They multiply, so 10 repeats over 10 epochs shows each image the same 100 times as 1 repeat over 100 epochs. The practical difference is that kohya-based trainers can save a checkpoint every epoch, so more epochs with fewer repeats gives you more checkpoints to compare, and repeats let you weight one folder of images above another.
The settings take an afternoon. The persona is the real project.
AI Influencers, included in All Access, walks through the full build: the character, the dataset, LoRA training and checkpoint testing, then content and monetisation, with the other three programs, live coaching and the private community in one subscription.
Run the numbers before you train
Get steps per epoch, total steps, training time and GPU cost from your own images, repeats, epochs and batch size, and join the free Telegram channel for persona workflows that are working now.