Skip to main content

LoRA Training Guide for Consistent AI Influencer Faces

Train a character LoRA for more consistent AI influencer faces. Learn dataset curation, Kohya SS settings, checkpoint testing, and troubleshooting.

Founder of IImagined.ai

Published
Jan 22, 2026
Updated
Oct 1, 2026
Reading time
14 min read
Quick answer

Pick the base model you will generate with first, because a LoRA only works on its own architecture. Collect consented images that clearly show one identity, caption the things that should stay changeable, train with a matching trainer (sd-scripts or Kohya SS for SDXL and FLUX.1; AI Toolkit, OneTrainer or musubi-tuner for FLUX.2, Qwen-Image and Z-Image), save checkpoints, and keep the one that holds the face while still following new prompts. There is no universal recipe, so compare checkpoints instead of trusting loss.

LoRA can improve character consistency, but no training recipe guarantees the same face in every generation. Results depend on the base model, dataset, captions, training exposure, checkpoint, and inference settings. The reliable workflow is to control those variables, save intermediate checkpoints, and compare them with a repeatable test set.

This guide uses the sd-scripts project (version 0.12.0, released September 24, 2026) as the command-line reference and Kohya SS as its graphical interface. It also covers AI Toolkit, OneTrainer and musubi-tuner, which train the newer FLUX.2, Qwen-Image and Z-Image models that sd-scripts does not. These tools change often, so match every command to your installed version and your base model.

Checked against the trainers' READMEs, the model cards and the hosted trainers' pricing pages on October 1, 2026. This update adds FLUX.2, Qwen-Image and Z-Image options, current trainer versions and hosted-trainer prices to the July 2026 version.

What a character LoRA actually trains

Low-Rank Adaptation freezes a pretrained model and learns smaller injected matrices instead of updating the full set of model weights. The original LoRA paper introduced the method for language models; diffusion tooling later adapted the approach to image generation.

For a character, the adapter learns patterns associated with the training images and captions. It does not store a formal identity model, and it cannot repair conflicting references. A clean adapter can still drift when you change the base checkpoint, sampler, resolution, prompt, seed, VAE, or LoRA strength.

Choose the base model before the trainer

Pick the model you will use for production before preparing a run. An adapter is architecture-specific: an SDXL LoRA will not work on FLUX, and a FLUX.1 LoRA is not a FLUX.2 LoRA. Check the license of the weights as well, because it decides whether you may use the model commercially.

Base modelWeights licenseTrainers with LoRA supportNotes
SDXL 1.0 and its fine-tunesCreativeML Open RAIL++-M (fine-tunes add their own terms)sd-scripts, Kohya SS, OneTrainer, AI ToolkitSupported by every trainer in this guide except FluxGym and musubi-tuner.
FLUX.1 [dev] (12B)FLUX.1 [dev] Non-Commercial License; the model card says outputs can be used commerciallysd-scripts, Kohya SS, FluxGym, AI Toolkit, OneTrainersd-scripts documents training down to 8 GB of VRAM with block swapping.
FLUX.2 [dev] (32B)FLUX Non-Commercial License; the model card says outputs can be used commerciallyAI Toolkit, OneTrainer, musubi-tunerSupports multi-reference editing; the model card's consumer-GPU example runs a 4-bit quantized model on an RTX 4090.
FLUX.2 [klein] Base 4BApache 2.0AI Toolkit, OneTrainer, musubi-tunerBFL calls the undistilled base models ideal for LoRA training; its guide says a 4B LoRA run fits in 24 GB of VRAM.
FLUX.2 [klein] Base 9BFLUX Non-Commercial LicenseAI Toolkit, OneTrainer, musubi-tunerThe model card says it fits in about 29 GB of VRAM.
Qwen-Image and Qwen-Image-2512 (20B)Apache 2.0AI Toolkit, OneTrainer, musubi-tunerThe 2512 release (December 2025) targets more realistic people.
Qwen-Image-2.1 (September 2026)Qwen Research License: non-commercial research or evaluation onlyAI ToolkitTakes up to 10 reference images and aims to preserve identity, so test it before training.
Z-Image (6B) and Z-Image-TurboApache 2.0AI Toolkit, OneTrainer, musubi-tunerThe model card calls the non-distilled base a good base for LoRA training; Turbo is the faster variant.
Chroma1-BaseApache 2.0sd-scripts, AI Toolkit, OneTrainerTrains through the sd-scripts FLUX.1 script with --model_type=chroma.

Sources: model cards for SDXL, FLUX.1 [dev], FLUX.2 [dev], klein Base 4B, klein Base 9B, Qwen-Image-2512, Qwen-Image-2.1, Z-Image and Chroma1-Base, plus each trainer's README.

In sd-scripts, the entry point follows the architecture: train_network.py for SD 1.x and 2.x, sdxl_train_network.py, sd3_train_network.py, and flux_train_network.py for FLUX.1 and Chroma, plus separate scripts for Lumina, HunyuanImage-2.1 and Anima. The repository lists its currently supported model families. Read the model-specific document shipped with your version before copying a command from another architecture.

Pick a trainer

TrainerInterfaceImage models it listsNotes
sd-scripts 0.12.0Command lineSD 1.x/2.x, SDXL, SD3/3.5, FLUX.1, Lumina, HunyuanImage-2.1, AnimaRequires PyTorch 2.6.0 or later. The commands in this guide follow it.
Kohya SS v26.0.0Browser GUI over sd-scriptsSD 1.5/2.x, SDXL, SD3, FLUX.1, Lumina Image 2.0, Anima, HunyuanImage-2.1Builds the sd-scripts command for you. Released July 9, 2026 with sd-scripts 0.11.1.
AI ToolkitWeb UI and command lineFLUX.1, FLUX.2 [dev] and [klein], Qwen-Image (including 2512 and 2.1), Z-Image, Chroma, SDXL and moreUsed in BFL's own klein LoRA guide. Needs an NVIDIA GPU.
OneTrainerDesktop GUI and command lineZ-Image, Qwen Image, FLUX.1, FLUX.2 [dev] and [klein], Chroma, SD 1.5 to 3.5, SDXL and moreBuilt-in BLIP, BLIP-2 and WD-1.4 captioning and masked training. Python 3.10 to 3.13.
musubi-tuner 0.3.6Command lineFLUX.1 Kontext, FLUX.2 [dev] and [klein], Qwen-Image, Z-Image, plus video modelsRecommends 12 GB or more of VRAM for image training and 64 GB of system RAM.
FluxGymSimple web UI running sd-scriptsFLUX.1 onlyBuilt for 12, 16 and 20 GB cards.

If you already work in Kohya SS or train SDXL, stay with sd-scripts. For FLUX.2, Qwen-Image or Z-Image, use AI Toolkit, OneTrainer or musubi-tuner, because sd-scripts has no entry point for them.

Hosted trainers if you have no suitable GPU

Prices as listed on each service's page on October 1, 2026, in US dollars:

ServiceModelListed price
falFLUX.1 fast training$2 per run at the default 1,000 steps; cost scales with steps
falFLUX.2 [dev]$6.40 per 1,000 steps
falQwen-Image-2512$1.50 per 1,000 steps, 500-step minimum
falZ-Image Turbo$2.26 per 1,000 steps
ReplicateFLUX.1 fast trainerAbout $1.46 per run in its fine-tuning guide; GPU time is billed per second
CivitaiSD 1.5, SDXL, FLUX.1, FLUX.2, Chroma, Qwen-Image, Z-ImageFrom 500 Buzz for SD 1.5 and SDXL; FLUX and video models cost more

Hosted trainers choose most settings for you and bill per run or per step, so a failed experiment still costs money. Start small, keep every checkpoint the service lets you download, and run the same comparison described below. fal lists commercial use for these trainers, but the base model's license still applies, and Civitai has not allowed models of real people since May 22, 2025 (its rules cover celebrities, influencers and private individuals alike).

Build an identity-consistent dataset

Dataset quality matters more than hitting a universal image count. Replicate's FLUX fine-tuning guide says training can work with as few as two images but recommends at least 10 for best results, and BFL's klein training example, written for a style LoRA, gives 20 to 40 images at 1,024 px or higher as the optimal dataset size (checked October 2026). For a face, coverage of angles, expressions and light matters more than the total. Start with images you are permitted to use, then keep only references that look like the same subject and are clear enough to teach the intended facial details.

Keep references that add useful coverage

  • Sharp facial features without heavy compression, motion blur, or masking.
  • Front, three-quarter, and profile views represented naturally.
  • Useful variation in expression, lighting, framing, clothing, and background.
  • A coherent visual domain for the production style you intend to generate.
  • Captions that distinguish identity from changeable attributes.

Reject references that teach the wrong identity

  • Another person, blended faces, or images where the subject is ambiguous.
  • Distorted eyes, teeth, hands near the face, strong beauty filters, or generated artifacts.
  • Near-duplicates that increase exposure without adding a new view.
  • A permanent outfit, location, or color grade repeated across most references.
  • Images without clear consent, provenance, or sufficient usage rights.

Use a side-by-side contact sheet for the final review. Ask one question: would a neutral reviewer identify every reference as the same subject? Remove outliers before training. Keep a separate evaluation set out of training so the adapter is tested on views it did not memorize.

Building references for a character who does not exist

A fully synthetic character has no photos, so you make the reference set first. Write a fixed description of the face, generate many candidate portraits on your base model, and choose one hero image. Then create the other views from that hero with a reference-based editor, such as Qwen-Image-Edit or FLUX.2's multi-reference editing: front, three-quarter and profile, several expressions, different light and outfits. Curate hard, because small drifts in this set become the LoRA's idea of the face. Log the prompts, seeds and models you used so you can show where the character came from.

Caption identity and variables separately

Give the character a distinctive trigger token that is unlikely to collide with ordinary prompt words. BFL suggests invented tokens that are not real words, and Replicate warns against existing words such as "dog". Put the token consistently in each caption, then describe variable attributes such as angle, expression, clothing, framing, and setting. When captions omit a repeated attribute, the adapter is more likely to bind it to the character. BFL's guide explains the same mechanism from the other side: with a small dataset, whatever the captions leave out is learned from the images.

Match the caption style to the model. FLUX.1 reads captions through a T5-XXL encoder alongside CLIP-L, and FLUX.2 [dev] uses a single Mistral Small 3.1 text encoder, according to Hugging Face's FLUX.2 announcement, so plain descriptive sentences suit them. Tag-style captions suit checkpoints that were trained on tags; check the model card. The WD14 tagger ships with sd-scripts, and OneTrainer includes BLIP, BLIP-2 and WD-1.4 captioning.

Caption shuffling and keep_tokens can preserve a trigger at the front while varying the remaining tags. Inspect the final text files rather than trusting an automatic tagger blindly; a wrong eye color or hairstyle label creates conflicting supervision.

Use an explicit TOML dataset configuration

Current sd-scripts accepts a TOML file through --dataset_config. Images should sit directly inside image_dir. Set repeats explicitly in the subset instead of relying on directory naming conventions.

[general]
caption_extension = ".txt"
shuffle_caption = true
keep_tokens = 1

[[datasets]]
resolution = 1024
batch_size = 1
enable_bucket = true
bucket_no_upscale = true

  [[datasets.subsets]]
  image_dir = "C:/datasets/creator_subject/images"
  num_repeats = 1

This is an SDXL-oriented example, not a cross-model prescription. Change resolution, batch size, repeats, captions, and augmentation only after checking the base model requirements and your available memory. The official dataset configuration reference explains option precedence and the supported TOML scopes. musubi-tuner uses a similar TOML dataset file, while AI Toolkit uses YAML job configs and a web UI for setting up jobs.

class_tokens is only used as a fallback for an image that has no matching caption file. Because this workflow supplies a .txt caption for each image, the trainer reads those captions instead.

Start with a controlled pilot run

The command below shows the shape of an SDXL LoRA pilot. The rank, alpha, learning rate, epoch count, precision, optimizer, and caching choices are experiment inputs. They are not promises about likeness quality.

accelerate launch sdxl_train_network.py ^
  --pretrained_model_name_or_path="<sdxl-base-model>" ^
  --dataset_config="dataset.toml" ^
  --output_dir="output" ^
  --output_name="creator_subject" ^
  --save_model_as=safetensors ^
  --network_module=networks.lora ^
  --network_dim=16 ^
  --network_alpha=8 ^
  --learning_rate=1e-4 ^
  --optimizer_type=AdamW8bit ^
  --max_train_epochs=10 ^
  --save_every_n_epochs=1 ^
  --mixed_precision=bf16 ^
  --gradient_checkpointing

Windows uses ^ for line continuation in Command Prompt. PowerShell and other shells use different syntax. The sd-scripts training overview documents the common arguments, while the model-specific files define the options that apply to each architecture.

FLUX.1 on a smaller card

The sd-scripts FLUX.1 guide needs four files (the FLUX.1 [dev] model, CLIP-L, T5-XXL and the autoencoder) and the networks.lora_flux module. With --fp8_base, it recommends these settings by VRAM:

VRAMDocumented settings
24 GBBasic settings, batch size 2
16 GBBatch size 1 plus --blocks_to_swap
12 GB--blocks_to_swap 16 and 8-bit AdamW
10 GB--blocks_to_swap 22 and an fp8 T5-XXL file
8 GB--blocks_to_swap 28 and an fp8 T5-XXL file

Higher block-swap values save memory but slow training. For FLUX.2, Qwen-Image and Z-Image, start from AI Toolkit's example configs, OneTrainer's presets or musubi-tuner's per-model documents, and keep the same discipline: one dataset definition, explicit repeats and saved checkpoints.

Treat parameters as hypotheses

ControlWhat it changesHow to evaluate it
Rank and alphaAdapter capacity and scaling.Compare a small set of runs; more capacity can learn unwanted detail as well as identity.
Learning rateThe size of each optimizer update.Change it independently and compare saved checkpoints, not just final loss.
Repeats, epochs, and stepsTotal exposure of each reference.Watch for identity learning followed by prompt rigidity or copied compositions.
Resolution and bucketsTraining detail, crops, aspect-ratio handling, and memory use.Match the architecture and inspect whether faces are cropped or upscaled poorly.
Text-encoder trainingHow strongly the trigger and captions are adapted.Check prompt responsiveness and read the cache restrictions before enabling it.

The official examples are anchors, not recipes. The sd-scripts SDXL guide uses a network dimension of 32, alpha 16 and a 1e-4 learning rate; its FLUX.1 guide uses dimension 16, alpha 1 and 1e-4; BFL's klein guide trains at 1e-4 in AI Toolkit and saves a checkpoint every 250 steps. Change one variable per run and write down what changed.

Compare saved checkpoints under identical conditions

Save intermediate adapters with --save_every_n_epochs or --save_every_n_steps. Then run a fixed prompt matrix against each checkpoint. Hold these inference variables constant:

  • Base model, VAE, sampler, scheduler, steps, and image dimensions.
  • Positive and negative prompts, including the trigger token.
  • Seed and batch size.
  • LoRA strength and any other adapters in the graph.

Include neutral portraits, profiles, full-body compositions, varied lighting, expressions, and clothing not present in the training set. Score identity separately from prompt adherence and artifact rate. A useful checkpoint preserves identity while still responding to those changes.

Training and validation loss can expose a bad run, but the lowest number is not automatically the best image checkpoint. BFL's klein training guide tells you to monitor sample outputs to avoid overfitting, so judge checkpoints by the images they produce, since loss can keep falling after the images start to overfit. Review the images and held-out prompts together. The Diffusers LoRA guide also shows how adapter weights are saved and loaded in a reproducible training pipeline.

Troubleshoot the failure you can observe

SymptomLikely checksNext experiment
Identity changes between promptsOutlier references, contradictory captions, weak trigger use, or an incompatible base model.Re-curate the contact sheet, correct captions, and compare checkpoints on one base model.
Output copies training compositionsExcessive exposure, duplicate images, narrow framing, or too much adapter capacity.Test an earlier checkpoint and reduce one exposure variable at a time.
Profiles or expressions failThe dataset does not represent that view clearly enough.Add a small number of licensed, identity-consistent references for the missing condition.
Outfit or background follows the triggerA repeated attribute is missing from captions or dominates the dataset.Caption the attribute and add controlled variation without changing identity.
Out-of-memory errorBatch size, resolution, optimizer state, precision, and caching choices.Reduce batch size first, then evaluate mixed precision, gradient checkpointing, SDPA, block swapping, or documented caching options.
Adapter is ignored or corrupts outputArchitecture mismatch, wrong loader, unsupported format, or excessive inference strength.Verify the adapter metadata and load it with the matching model family and node or UI version.

The advanced sd-scripts guide documents memory controls and their tradeoffs. For example, gradient checkpointing reduces memory at a speed cost. Latent caching disables image augmentation, while text-encoder output caching disables caption augmentation and prevents text-encoder LoRA training in the same run, so it needs --network_train_unet_only.

Load the adapter in a compatible inference UI

Load the adapter with the same base model family it was trained on. The Automatic1111 feature guide documents placing compatible adapters under models/Lora and applying them with <lora:filename:multiplier>, and it points users to Kohya SS for training. Its latest release is still v1.10.1, so newer architectures usually need another UI.

In ComfyUI, put the file in ComfyUI/models/loras and use the Load LoRA node, whose strength_model and strength_clip inputs set how strongly the adapter changes the model and the text encoder (ComfyUI LoRA tutorial). Connect it to the same model and text-encoding path used by the workflow. If you are new to the graph, start with the ComfyUI overview and installation guide.

When reference images may be enough

Training is not the only route. These tools keep a character recognizable from reference images without a training run:

  • Qwen-Image-2.1 accepts up to 10 reference images and aims to preserve identity for people and products, under a non-commercial research license.
  • FLUX.2 [dev] and [klein] support multi-reference editing, per their model cards.
  • Midjourney's Edit model for V8.1 and V8.2 takes up to 4 reference images and replaces Omni Reference and Character Reference (Midjourney docs). See our Midjourney prompts guide.
  • Kling Elements bind a reference-built character to video shots; our Kling tutorial covers them.

Test references first. Train a LoRA when they cannot hold the identity across your shot list, or when you need the character inside your own workflow with weights that do not change when a hosted model updates.

Consent, licensing, and safe sharing

  • Use a real person's likeness only with informed permission that covers training and intended distribution. Many services refuse likeness models even with permission; Civitai bans them outright.
  • Verify the licenses for the base model, dataset images, trainer, and hosting destination. FLUX.1 [dev], FLUX.2 [dev] and klein 9B weights are non-commercial, and Qwen-Image-2.1 is research-only.
  • Do not train on minors or on private, intimate, stolen, or deceptive source material.
  • Document the character's provenance and disclose synthetic media where the platform, campaign, or audience context requires it. Our face swap guide summarizes the likeness laws and platform rules.
  • Before publishing adapter weights, assume recipients can generate outputs outside your planned use case.

A repeatable character LoRA workflow

  1. Choose the production model family, check its license, and pick a trainer that supports it.
  2. Collect consented, licensed references and remove identity outliers.
  3. Caption the trigger token, identity, and changeable attributes deliberately.
  4. Declare the dataset explicitly in TOML (or the trainer's own config).
  5. Run a small pilot and save intermediate checkpoints.
  6. Compare checkpoints with the same held-out prompts and inference settings.
  7. Change one training variable, record the result, and repeat only when the evidence supports it.

LoRA training FAQ

How many images do I need for a character LoRA?

There is no universal number. Replicate's FLUX guide says as few as two images can work but recommends at least 10, and BFL's klein example gives 20 to 40 as optimal for a style LoRA. For a face, coverage matters more than count: sharp front, three-quarter and profile views in varied light, with near-duplicates removed.

Which base model should I train on in 2026?

The one you will generate with, because a LoRA only works on its own architecture. For commercial work, check the weights license first: FLUX.2 [klein] 4B, Qwen-Image, Qwen-Image-2512, Z-Image and Chroma are Apache 2.0, SDXL uses CreativeML Open RAIL++-M, and FLUX.1 [dev], FLUX.2 [dev] and klein 9B use non-commercial licenses, although the two [dev] model cards say generated outputs can be used commercially. Qwen-Image-2.1 is limited to research and evaluation.

Can I train a LoRA without my own GPU?

Yes. fal, Replicate and Civitai run the job for you and bill per run, per step or in Buzz, as listed in the table above. You still need to test the checkpoints yourself.

What rank and learning rate should I start with?

Start from the trainer's documented example for your model, then change one variable at a time. sd-scripts' SDXL example uses dimension 32 and alpha 16, its FLUX.1 example uses dimension 16 and alpha 1, and both use a 1e-4 learning rate.

Can I train a LoRA of a real person?

Only with that person's informed permission for training and every intended use, and many platforms refuse such models anyway. Sexual or deceptive use of someone's likeness is illegal in many places regardless of how the model was made.

Primary sources reviewed

Technical instructions, licenses, prices and source links were reviewed on October 1, 2026. Recheck the installed tool version and the pricing page before starting a training run.

Continue the AI influencer workflow

Building the image set first? The 40-shot character LoRA dataset checklist covers angles, lighting, outfits and culling before you train.

Browse the AI Influencers topic hub, compare current ComfyUI model families, try the Z-Image Turbo tutorial, estimate the cost to create an AI influencer, and plan the next stage with the AI influencer monetization guide. The AI Influencers course covers character design, LoRA-based face consistency in ComfyUI, and the disclosure rules for publishing.

Operator program · recommended for this article

Want the full AI Influencers playbook?

The complete pipeline for building virtual brands at scale — identity engineering, ComfyUI production, IP governance, and the distribution flywheel.

9 modules · one-time purchase · 30-day money-back guaranteeiimagined.ai by Anyro
All-Access subscription

Every program. Member benefits.
One subscription.

Use all four premium programs with weekly live coaching, a private community, and the resource vault.

Confirm current lessons, downloadable resources and member-benefit arrangements before purchasing.

  • All 4 premium programs plus free Futures Trading
  • Weekly live coaching calls
  • Private community access
  • Resource vault and templates
  • 30-day money-back guarantee, cancel anytime
$99/ month
$99 for the first month · $702 to buy all four standalone
Start All-AccessOr browse standalone programs
30-day money-back guarantee · $99/month · cancel anytime