Skip to main content
← Journal·AI InfluencersOct 7, 2026·10 min read

PhotoMaker: Photo-Realistic Personas From Reference Stacks

How PhotoMaker AI stacks several reference photos into one identity for SDXL personas: ComfyUI setup, V1 vs V2, settings, and when it beats InstantID.

A

Founder of IImagined.ai

Quick answer

PhotoMaker is TencentARC's open-source SDXL adapter that turns a stack of reference photos into one identity you call with the word 'img' in a prompt, with no LoRA training. It suits recurring personas when you have several good images of the same face and want pose and style to stay editable. Use V1 through ComfyUI's native nodes or V2 through a community node pack, keep style strength around the 20% default, and move to a character LoRA once the persona needs the same face across hundreds of posts.

PhotoMaker is a free, open-source SDXL adapter from TencentARC that merges several reference photos of one face into a single identity, so you can generate that persona in new scenes in seconds without training a LoRA. You write the class word followed by the trigger word, such as "woman img", and PhotoMaker injects the stacked identity at that spot in the prompt. Its edge over single-image methods is the stack: the more clean, varied references you give it, the steadier the face.

For every way to keep one face consistent, compared side by side, start with our consistent character AI guide.

Checked October 2026 against the PhotoMaker GitHub repository and its V2 notes, the ComfyUI PhotoMaker nodes, ComfyUI-PhotoMaker-Plus, the InstantID repository and the InsightFace license. Nothing below is our own benchmark; settings are the developers' defaults plus reasoning about how the method works.

PhotoMaker is one of several ways to keep a face consistent. If you are still choosing a method, start with our LoRA training guide for consistent AI influencer faces, which covers the heavier option PhotoMaker is often compared with, and the AI image generation for influencers guide for the wider photo workflow.

What PhotoMaker actually does

The paper title explains the idea: "Customizing Realistic Human Photos via Stacked ID Embedding". PhotoMaker encodes each reference image into an identity embedding, stacks those embeddings, and fuses them into the text embedding of the class word you marked with img. The rest of the prompt (clothes, location, lighting, camera) stays ordinary text, which is why PhotoMaker keeps strong prompt control while it carries the face.

Two versions matter. V1 was released on January 15, 2024, and was presented at CVPR 2024. V2 followed on July 22, 2024 with what TencentARC calls improved ID fidelity, especially from a single input image, thanks to new training strategies, more portrait data and a stronger ID encoder that uses InsightFace. Both run on SDXL, and the README describes PhotoMaker as an adapter you can pair with other SDXL base models and community LoRAs. The minimum GPU memory the developers list is 11 GB.

How a PhotoMaker generation is built
  1. 01
    Reference stack

    Several photos of the same face, face filling most of each frame

  2. 02
    ID encoder

    Each photo becomes an identity embedding (V2 adds InsightFace features)

  3. 03
    Stack and fuse

    Embeddings merge into the class word marked with "img"

  4. 04
    SDXL sampling

    Your checkpoint renders the scene; the ID merges in after the first steps

  5. 05
    Persona image

    Same face, new outfit, pose and setting from the text prompt

When ID stacking beats single-image methods

Single-image methods such as InstantID and IP-Adapter FaceID read one reference. If that photo has odd lighting, a three-quarter angle or a slight expression, the quirk tends to follow the face into every output. PhotoMaker averages across the stack, so one awkward reference has less pull. For a recurring persona, that is the property you want: the face should be the persona's face, not the face in one particular selfie.

Use something else when
  • You only have one reference: InstantID or V2 alone
  • You need a pixel-tight face match every time
  • You generate in Flux or SD 1.5
  • The persona will post hundreds of times: train a LoRA
  • You need exact pose control: add ControlNet
PhotoMaker fits when
  • You already have 4+ good images of the persona
  • You want outfits, poses and styles to change freely
  • You work in SDXL and want no training step
  • You need a fast way to make a LoRA dataset later
  • You want text prompts to stay in control
MethodReferencesHow identity is appliedBaseTrainingLicense notes
PhotoMaker V11+ images (more is better)Prompt-level ID mergeSDXLNoApache 2.0
PhotoMaker V21+ images (more is better)Prompt-level ID merge + InsightFaceSDXLNoApache 2.0; InsightFace models non-commercial
InstantID1 imageIP-Adapter + landmark ControlNetSDXLNoCode Apache 2.0; checkpoints and InsightFace models research-only
Character LoRA15-40 images typicalWeights trained on your setAny familyYesDepends on base model

TencentARC's V2 notes include side-by-side images against PhotoMaker V1, IP-Adapter-FaceID-Plus-V2 and InstantID on the same RealVisXL V4.0 base, and conclude that V2 holds identity and quality better. Treat that as the developer's own comparison with hand-picked outputs (best of four per method), not an independent benchmark. The same notes say V2 can be combined with IP-Adapter-FaceID, InstantID or a character LoRA for even tighter likeness, which is a useful hint: these methods stack with each other too.

Build the reference stack

For an AI influencer, the stack is usually not real photos at all. You design the persona first, generate a few dozen candidates, and keep the four to eight images where the face is unmistakably the same person. Those become the stack. If you do use photos of a real person, you need that person's explicit consent to generate new images of them, and most platforms require AI disclosure on top of that; read our Instagram AI label rules before posting.

Reference stack pre-flight
  • Same identity in every image: drop any shot where the face reads as a near-lookalike
  • Face fills most of the frame; crop tightly before loading
  • At least one frontal, one three-quarter left and one three-quarter right view
  • Neutral or soft expressions; avoid big laughs or squints in most of the stack
  • Even, soft light in most shots; one harsh-shadow photo is fine, five are not
  • No sunglasses, heavy hats, hands on the face or hair across the eyes
  • Similar apparent age and makeup across the set
  • Sharp, uncompressed images; no visible watermarks or text
  • Consent documented if any reference shows a real person

The order of the stack does not matter, but its makeup does. If six of eight photos are frontal, outputs will lean frontal and the persona can look stiff in profile. If the stack mixes two slightly different faces, PhotoMaker will happily average them into a third face you did not design. Curate harder than feels necessary.

PhotoMaker in ComfyUI, step by step

ComfyUI is the most practical place to run PhotoMaker for persona work because you can save the graph and rerun it for every content batch. If ComfyUI is not installed yet, follow our ComfyUI installation guide first.

From reference stack to persona image
  1. 1
    Pick the version

    V1 runs on ComfyUI's native PhotoMakerLoader and PhotoMakerEncode nodes. V2 needs ComfyUI-PhotoMaker-Plus plus insightface and onnxruntime.

  2. 2
    Download the weights

    Get photomaker-v1.bin or the V2 file from TencentARC on Hugging Face and put it in ComfyUI/models/photomaker.

  3. 3
    Load an SDXL checkpoint

    Use a photoreal SDXL model. PhotoMaker will not load on Flux or SD 1.5 graphs.

  4. 4
    Feed the stack

    Batch your cropped reference images into one image input for the PhotoMaker encode node.

  5. 5
    Write the prompt with the trigger

    Put "img" right after the class word: "photo of a woman img sitting in a cafe, natural window light".

  6. 6
    Set style strength and steps

    Start near the official defaults: 20% style strength, 50 steps, CFG 5. Lower steps only once results look right.

  7. 7
    Generate four and review

    Compare the face to the stack at thumbnail size and at full size. Keep only outputs that pass both.

  8. 8
    Save the graph

    Store the workflow with the stack so every future batch starts from identical settings.

Settings that matter

PhotoMaker has few knobs, and one does most of the work. In the official V2 demo, the style strength percentage decides at which sampling step the identity is merged in: the code computes the merge step as style strength divided by 100 times the step count, capped at step 30. A higher percentage leaves more early steps to the base model and your style prompt, which helps stylized looks but loosens the likeness. The README suggests 30 to 50 when you want a stylized result and the face still looks too photographic.

SettingStarting valueWhat it changes
Style strength (%)20 (demo default), range 15-50Sets the step where the identity merges in; higher = more stylization, weaker likeness
Sampling steps50 (demo default)Fewer steps are faster but the README warns ID fidelity may drop
Guidance scale (CFG)5 (demo default)Keep it moderate on photoreal SDXL checkpoints
Trigger wordimg, after the class word"woman img" or "man img"; it is the slot where the identity goes
Base modelAny SDXL checkpointPhotoreal SDXL checkpoints suit persona work best

The README adds two practical tips. First, upload more photos of the person to improve fidelity. Second, ethnicity words before the class word can help when the base model drifts, for example writing "Asian woman img". Keep the rest of the prompt descriptive about scene, wardrobe and camera rather than the face, because the face is the job of the stack. Our free AI Influencer Prompt Generator builds scene prompts in that shape; add the "img" trigger after the class word before you paste one in.

Adding pose control and LoRAs

PhotoMaker carries the face, not the body position. For specific poses, combine it with ControlNet: TencentARC ships example scripts for V2 with ControlNet, T2I-Adapter and IP-Adapter, and in ComfyUI you simply apply a ControlNet to the conditioning after the PhotoMaker encode. Our ControlNet guide covers OpenPose, depth and canny for persona shots. Style and clothing LoRAs load as usual on the SDXL model; keep their weights moderate so they do not fight the identity.

Choice of base checkpoint changes skin, light and realism more than any PhotoMaker setting. Our list of the best ComfyUI models for realistic people covers SDXL options that suit persona work.

From PhotoMaker to a character LoRA

Many creators use PhotoMaker as a bridge. It gives you a consistent persona on day one, and its best outputs, varied in angle, outfit and light, make a clean training set for a character LoRA later. A LoRA is slower to set up but carries the identity on any checkpoint of the family it was trained for, without the adapter, and tends to hold small details such as moles or tooth shape better over hundreds of posts. When you get there, the LoRA Training Steps Calculator helps you size the run.

Building the persona system around these tools, from the first reference set to a weekly content batch and the accounts it posts to, is what our AI Influencers program walks through.

Licensing and consent

PhotoMaker's repository states the project is licensed under Apache 2.0, apart from listed third-party components. V2 depends on InsightFace for face analysis, and InsightFace says its pretrained models are for non-commercial research only; InstantID carries the same InsightFace caveat and calls its own checkpoints research-only. If you run a commercial persona, read those licenses yourself or ask a lawyer. And whatever the license says, generating a recognizable real person without their consent can break platform rules and, in many places, the law. This is not legal advice.

PhotoMaker AI: FAQ

What is PhotoMaker AI?

PhotoMaker is an open-source image model from TencentARC that customizes realistic human photos without training a LoRA. You give it one or more reference photos of a face and a prompt in which the word "img" follows a class word such as "woman img", and it merges the identity from all the references into that part of the prompt. It works with SDXL base models and was presented at CVPR 2024.

How many reference images does PhotoMaker need?

One image works, but the official demo and README both say more photos of the same person improve identity fidelity. The point of the method is that it stacks the embeddings from several images, so a set of four to eight clear, varied shots of the same face is a sensible starting stack. The face should fill most of each image, because the official demo notes it does not crop to the face for you.

Is PhotoMaker better than InstantID?

They solve the problem differently. InstantID is built around a single reference image plus a face-landmark ControlNet, so it locks the face and its geometry tightly. PhotoMaker merges several references into the prompt, which tends to leave pose, outfit and style more editable. TencentARC published side-by-side comparisons favoring PhotoMaker V2, but those are the developer's own selections, so test both on your persona before choosing.

Does PhotoMaker work in ComfyUI?

Yes. ComfyUI has had native PhotoMakerLoader and PhotoMakerEncode nodes since January 2024, which run the original V1 model. PhotoMaker V2 needs a community node pack such as ComfyUI-PhotoMaker-Plus, which adds an InsightFace loader because V2 uses InsightFace to extract the face identity.

Can I use PhotoMaker with a custom SDXL checkpoint?

Yes. The README describes PhotoMaker as an adapter that works with other SDXL base models and community LoRAs; its own comparison used RealVisXL V4.0. It does not work with SD 1.5 or Flux checkpoints, because the adapter was trained for the SDXL text encoders.

Can I use PhotoMaker commercially?

The PhotoMaker code and weights are under the Apache 2.0 license, with some third-party components listed separately. V2 also relies on InsightFace face models, and InsightFace states its pretrained models are for non-commercial research only. Read both licenses, and get permission before using any real person's photos as references. This is not legal advice.

All Access · all four programs · $99/mo

The face is settled. Now build the persona around it.

AI Influencers, included in All Access, covers persona design, consistent faces, content batches and monetization, alongside the other three programs, live coaching and the private community in one subscription.

Start All Access — $99/mo →30-day money-back guarantee
Free · no signup

Write better persona prompts, free

Generate scene prompts for your PhotoMaker graph, and join the free Telegram channel for what is working in AI persona content right now.

About the author

Written by Anyro, Founder of IImagined.ai. IImagined.ai is a founder-led education platform teaching Instagram growth, AI influencers, digital products, and AI automation.

Results vary; no income is guaranteed.

All-Access subscription

Every program. Member benefits.
One subscription.

Use all four premium programs with weekly live coaching, a private community, and the resource vault.

Confirm current lessons, downloadable resources and member-benefit arrangements before purchasing.

  • All 4 premium programs plus free Futures Trading
  • Weekly live coaching calls
  • Private community access
  • Resource vault and templates
  • 30-day money-back guarantee, cancel anytime
$99/ month
$99 for the first month · $702 to buy all four standalone
Start All-AccessOr browse standalone programs
30-day money-back guarantee · $99/month · cancel anytime