Skip to main content
← Journal·AI InfluencersOct 7, 2026·8 min read

Character LoRA Dataset Checklist: 40 Shots That Train Clean Faces

A printable 40-shot character LoRA dataset checklist: angles, expressions, lighting, outfits and backgrounds, plus captions, culling and file prep.

A

Founder of IImagined.ai

Quick answer

A clean character LoRA comes from a dataset where only the character repeats. This checklist plans 40 shots: 19 face angles and expressions, 6 lighting setups, 6 half-body outfit and background combos, 5 full-body poses and 4 variation and context shots. Capture all 40, cull to the 20 to 40 that are sharp and on-model, caption what changes, and export as JPG or PNG.

A character LoRA learns whatever repeats across its dataset. If the face is the only thing that repeats, you get a LoRA that puts that face into any scene. If the same hoodie, wall and window light repeat too, you get a LoRA that keeps drawing the hoodie, the wall and the light. The checklist below is built to make the face the only constant.

For every way to keep one face consistent, compared side by side, start with our consistent character AI guide.

It is a capture plan of 40 shots: generate or photograph all of them, then cull. This page is the dataset step of our LoRA training guide for consistent AI influencer faces, which covers the training settings. Once the set is ready, the free LoRA Training Steps Calculator works out steps and epochs for its size.

The 40-shot character LoRA checklist

The list leans on the face, because the face is what people notice when a character drifts, and adds body and context shots so the model also learns proportions and how the character sits in a scene.

ShotsGroupWhat to varyFraming
1-6Face, frontNeutral, soft smile, big smile, serious, eyes closed, talkingClose-up, shoulders up
7-12Face, three-quarterLeft and right, three expressions eachClose-up
13-16Face, profileLeft and right, neutral and smilingClose-up, ears and jaw visible
17-19Face, high and low angleLooking down at camera, looking up, over the shoulderClose-up
20-25Lighting setWindow daylight, golden hour, overcast, studio softbox, indoor lamp, night street lightHead and shoulders
26-31Half bodyThree outfits x two backgroundsWaist up
32-36Full bodyStanding, walking, sitting, leaning, one action poseHead to toe
37-38Hair variationHair up and hair down (only if the persona changes hair)Head and shoulders
39-40Context shotsHolding an object, mid-activity in a real settingHalf or full body
How the 40 shots split
40 shots
  • Face angles and expressions19
  • Lighting set6
  • Half body6
  • Full body5
  • Hair and context4

Our template, not a measured optimum. Shift the weights toward body shots for a full character, toward the face for a face-only LoRA.

Printable version

Print this or keep it next to your generator. Tick a box only when the shot is sharp, on-model and different from the ones you already have.

Character LoRA dataset: tick as you go
  • Front face: neutral, soft smile, big smile, serious, eyes closed, mid-speech (6)
  • Three-quarter left and right, three expressions each (6)
  • Profile left and right, jaw and ear visible (4)
  • High angle, low angle, over-the-shoulder (3)
  • Six lighting setups: window, golden hour, overcast, softbox, lamp, night (6)
  • Half body: three outfits on two backgrounds (6)
  • Full body: stand, walk, sit, lean, action pose (5)
  • Hair up and down, only if the persona changes hair (2)
  • Two context shots: holding an object, mid-activity (2)

What to vary, and what to keep fixed

Vary everything you want to control later with a prompt: angle, expression, lighting, outfit, background, framing. Keep fixed everything that defines the character: face shape, eye color, skin tone and texture, distinguishing marks, and hair if the persona always wears it the same way.

The test for each image is simple. Cover the face with your thumb: is the rest of the image different from the other images? Then uncover it: is the face unmistakably the same person? A yes to both means the image belongs in the set.

Where the 40 shots come from

For a fictional AI persona, you generate the set. Lock the face first with a reference method such as PuLID, InstantID or a strong character sheet, then work through the checklist one group at a time. Expect to throw away a good share of generations because the face drifted; keep only images where the face clearly matches your best reference.

For a real person, such as your own AI twin, you photograph the set. Use a phone in good light, change rooms and clothes between groups, and avoid heavy beauty filters, which the LoRA will learn as skin texture. Only train on a real person's face with their written consent.

Which reference method you use matters less than using one consistently for the whole set. Our Flux Kontext guide covers editing one locked face into new scenes, which is the fastest way to fill the outfit and background groups without the face drifting. If you are still designing the persona itself, start with how to create an AI influencer and come back here once the face is settled.

Cull before you train

Forty captured shots rarely means forty training images. As a rule of thumb, most face and character LoRAs on Flux and SDXL train well on roughly 20 to 40 strong images, and a smaller clean set beats a bigger repetitive one.

Culling order
  1. 1
    Remove soft and broken images

    Blur, motion smear, extra fingers, warped teeth or ears

  2. 2
    Remove off-model faces

    Compare each against your one best reference; when in doubt, cut it

  3. 3
    Remove near-duplicates

    Two shots with the same angle, light and expression count once

  4. 4
    Check the balance

    Every checklist group should still have at least one or two images

  5. 5
    Crop out distractions

    Watermarks, text, other faces and logos go before training

Captions and the trigger word

Captions tell the trainer which parts of the image are variable. A good caption names the trigger word, then describes what changes in that image. Do not caption the traits that should always come with the character: if you write "green eyes" on every image, green eyes become something you have to prompt.

Shot typeExample caption
Close-up, three-quarterphoto of [trigger], three-quarter view facing left, slight smile, soft window light, plain grey background
Half body, outfitphoto of [trigger], waist up, wearing a denim jacket, standing in a cafe, warm indoor light
Full bodyphoto of [trigger], full body, walking on a city sidewalk, overcast daylight, jeans and white sneakers

In AI Toolkit you can write [trigger] in the caption file and it is replaced with the trigger word from your config. Use a rare made-up token, not a real name or common word, so it does not collide with something the base model already knows.

Reading the results: what a bad dataset looks like after training

Most LoRA problems blamed on settings are dataset problems. Before you change the learning rate, match what you see in the samples against this table and fix the set first; retraining on the same images with new numbers usually moves the problem rather than removing it.

What you seeLikely dataset causeFix in the set
Same expression in every outputMost images share one expressionAdd the expression group back; cut duplicates of the dominant one
Face only works in one lightOne lighting setup across the setAdd hard light, backlight and outdoor shots
Outfit appears even when not promptedSame outfit in most images, never captionedVary clothes and caption them in every image
Background bleeds into new scenesOne room or wall dominatesAdd at least four distinct backgrounds
Side profiles look like a different personFew or no profile shotsAdd three-quarter and full profile shots
Skin looks plastic or filteredBeauty filter or heavy upscaler on the sourcesReplace with unfiltered or less processed images

Keep notes on each training run: dataset version, image count, steps and what changed. The free steps calculator helps keep the step count proportional when the image count changes between runs, so you are only testing one variable at a time.

File prep before upload

Checked October 2026 against the AI Toolkit README: supported formats are JPG, JPEG and PNG (WebP is noted as having issues), and each caption is a .txt file with the same name as its image. Images are never upscaled; they are downscaled and sorted into aspect-ratio buckets, so you do not need to crop everything square. Kohya-based trainers work the same way with bucketing switched on.

Before you hit train
  • Every image is JPG or PNG, at or above the training resolution
  • Every image has a matching .txt caption with the trigger word
  • No two images share the same angle, light, outfit and background
  • No watermarks, text overlays or other people in frame
  • You kept one untouched reference image outside the set to judge results

Then train, and check sample images at intervals. If samples start copying one photo from your set, you are into overfitting; lower the steps or the learning rate, or add more varied images. The AI Influencers program walks through the full persona build, from the first reference face to a trained LoRA and a posting schedule.

Character LoRA dataset: FAQ

How many images do I need for a character LoRA?

A common rule of thumb is 20 to 40 good images for a face or character LoRA on modern base models like Flux and SDXL, and fewer clean images beat many repetitive ones. The 40-shot list on this page is a capture plan: shoot or generate all 40, then cull the weak and near-duplicate ones before training.

What should a character LoRA dataset include?

Mostly the face from several angles (front, three-quarter, profile), a range of expressions, at least three lighting setups, a few outfits and a few backgrounds, plus some half-body and full-body shots so the model learns proportions. Anything that repeats in every image, such as one hoodie or one wall, tends to get baked into the character.

Should I crop my LoRA training images to 1024x1024?

Not necessarily. Trainers like AI Toolkit and Kohya use aspect ratio bucketing, which downscales and groups images by shape, so you can keep portrait and landscape crops. AI Toolkit says images are never upscaled, so start from images at or above your training resolution and crop out distractions rather than forcing squares.

Do I need captions for a face LoRA?

Captions help. A short caption with your trigger word plus what changes in the image (angle, expression, outfit, lighting, background) tells the trainer which parts are variable, so they stay promptable later. Leave out traits that should always be part of the character, such as eye color, because captioning them makes them optional.

Can I train a character LoRA from AI-generated images?

Yes, and it is the usual route for a fictional AI persona: generate a consistent set with a reference tool, keep only the images where the face matches, and train on those. The risk is that generator artifacts like waxy skin or the same lighting in every image get learned too, so mix setups deliberately and cull hard.

What file format should LoRA training images be?

Use JPG or PNG. The AI Toolkit README (checked October 2026) lists jpg, jpeg and png as the supported formats and notes that WebP has issues. Caption files are plain text with the same name as the image and a .txt extension.

All Access · all four programs · $99/mo

The dataset is the hard part. Learn the whole persona pipeline.

AI Influencers, included in All Access, covers building a consistent persona, training LoRAs and turning the character into content and income streams, with the other three programs, live coaching and the private community in one subscription.

Start All Access — $99/mo →30-day money-back guarantee
Free · no signup

Plan your dataset, free

Build a shot list sized to your base model and image budget, and join the free Telegram channel for persona workflows that are working right now.

About the author

Written by Anyro, Founder of IImagined.ai. IImagined.ai is a founder-led education platform teaching Instagram growth, AI influencers, digital products, and AI automation.

Results vary; no income is guaranteed.

All-Access subscription

Every program. Member benefits.
One subscription.

Use all four premium programs with weekly live coaching, a private community, and the resource vault.

Confirm current lessons, downloadable resources and member-benefit arrangements before purchasing.

  • All 4 premium programs plus free Futures Trading
  • Weekly live coaching calls
  • Private community access
  • Resource vault and templates
  • 30-day money-back guarantee, cancel anytime
$99/ month
$99 for the first month · $702 to buy all four standalone
Start All-AccessOr browse standalone programs
30-day money-back guarantee · $99/month · cancel anytime