PhotoMaker is TencentARC's open-source SDXL adapter that turns a stack of reference photos into one identity you call with the word 'img' in a prompt, with no LoRA training. It suits recurring personas when you have several good images of the same face and want pose and style to stay editable. Use V1 through ComfyUI's native nodes or V2 through a community node pack, keep style strength around the 20% default, and move to a character LoRA once the persona needs the same face across hundreds of posts.
PhotoMaker is a free, open-source SDXL adapter from TencentARC that merges several reference photos of one face into a single identity, so you can generate that persona in new scenes in seconds without training a LoRA. You write the class word followed by the trigger word, such as "woman img", and PhotoMaker injects the stacked identity at that spot in the prompt. Its edge over single-image methods is the stack: the more clean, varied references you give it, the steadier the face.
For every way to keep one face consistent, compared side by side, start with our consistent character AI guide.
Checked October 2026 against the PhotoMaker GitHub repository and its V2 notes, the ComfyUI PhotoMaker nodes, ComfyUI-PhotoMaker-Plus, the InstantID repository and the InsightFace license. Nothing below is our own benchmark; settings are the developers' defaults plus reasoning about how the method works.
PhotoMaker is one of several ways to keep a face consistent. If you are still choosing a method, start with our LoRA training guide for consistent AI influencer faces, which covers the heavier option PhotoMaker is often compared with, and the AI image generation for influencers guide for the wider photo workflow.
What PhotoMaker actually does
The paper title explains the idea: "Customizing Realistic Human Photos via Stacked ID Embedding". PhotoMaker encodes each reference image into an identity embedding, stacks those embeddings, and fuses them into the text embedding of the class word you marked with img. The rest of the prompt (clothes, location, lighting, camera) stays ordinary text, which is why PhotoMaker keeps strong prompt control while it carries the face.
Two versions matter. V1 was released on January 15, 2024, and was presented at CVPR 2024. V2 followed on July 22, 2024 with what TencentARC calls improved ID fidelity, especially from a single input image, thanks to new training strategies, more portrait data and a stronger ID encoder that uses InsightFace. Both run on SDXL, and the README describes PhotoMaker as an adapter you can pair with other SDXL base models and community LoRAs. The minimum GPU memory the developers list is 11 GB.
- 01Reference stack
Several photos of the same face, face filling most of each frame
- 02ID encoder
Each photo becomes an identity embedding (V2 adds InsightFace features)
- 03Stack and fuse
Embeddings merge into the class word marked with "img"
- 04SDXL sampling
Your checkpoint renders the scene; the ID merges in after the first steps
- 05Persona image
Same face, new outfit, pose and setting from the text prompt
When ID stacking beats single-image methods
Single-image methods such as InstantID and IP-Adapter FaceID read one reference. If that photo has odd lighting, a three-quarter angle or a slight expression, the quirk tends to follow the face into every output. PhotoMaker averages across the stack, so one awkward reference has less pull. For a recurring persona, that is the property you want: the face should be the persona's face, not the face in one particular selfie.
- You only have one reference: InstantID or V2 alone
- You need a pixel-tight face match every time
- You generate in Flux or SD 1.5
- The persona will post hundreds of times: train a LoRA
- You need exact pose control: add ControlNet
- You already have 4+ good images of the persona
- You want outfits, poses and styles to change freely
- You work in SDXL and want no training step
- You need a fast way to make a LoRA dataset later
- You want text prompts to stay in control
| Method | References | How identity is applied | Base | Training | License notes |
|---|---|---|---|---|---|
| PhotoMaker V1 | 1+ images (more is better) | Prompt-level ID merge | SDXL | No | Apache 2.0 |
| PhotoMaker V2 | 1+ images (more is better) | Prompt-level ID merge + InsightFace | SDXL | No | Apache 2.0; InsightFace models non-commercial |
| InstantID | 1 image | IP-Adapter + landmark ControlNet | SDXL | No | Code Apache 2.0; checkpoints and InsightFace models research-only |
| Character LoRA | 15-40 images typical | Weights trained on your set | Any family | Yes | Depends on base model |
TencentARC's V2 notes include side-by-side images against PhotoMaker V1, IP-Adapter-FaceID-Plus-V2 and InstantID on the same RealVisXL V4.0 base, and conclude that V2 holds identity and quality better. Treat that as the developer's own comparison with hand-picked outputs (best of four per method), not an independent benchmark. The same notes say V2 can be combined with IP-Adapter-FaceID, InstantID or a character LoRA for even tighter likeness, which is a useful hint: these methods stack with each other too.
Build the reference stack
For an AI influencer, the stack is usually not real photos at all. You design the persona first, generate a few dozen candidates, and keep the four to eight images where the face is unmistakably the same person. Those become the stack. If you do use photos of a real person, you need that person's explicit consent to generate new images of them, and most platforms require AI disclosure on top of that; read our Instagram AI label rules before posting.
- Same identity in every image: drop any shot where the face reads as a near-lookalike
- Face fills most of the frame; crop tightly before loading
- At least one frontal, one three-quarter left and one three-quarter right view
- Neutral or soft expressions; avoid big laughs or squints in most of the stack
- Even, soft light in most shots; one harsh-shadow photo is fine, five are not
- No sunglasses, heavy hats, hands on the face or hair across the eyes
- Similar apparent age and makeup across the set
- Sharp, uncompressed images; no visible watermarks or text
- Consent documented if any reference shows a real person
The order of the stack does not matter, but its makeup does. If six of eight photos are frontal, outputs will lean frontal and the persona can look stiff in profile. If the stack mixes two slightly different faces, PhotoMaker will happily average them into a third face you did not design. Curate harder than feels necessary.
PhotoMaker in ComfyUI, step by step
ComfyUI is the most practical place to run PhotoMaker for persona work because you can save the graph and rerun it for every content batch. If ComfyUI is not installed yet, follow our ComfyUI installation guide first.
- 1Pick the version
V1 runs on ComfyUI's native PhotoMakerLoader and PhotoMakerEncode nodes. V2 needs ComfyUI-PhotoMaker-Plus plus insightface and onnxruntime.
- 2Download the weights
Get photomaker-v1.bin or the V2 file from TencentARC on Hugging Face and put it in ComfyUI/models/photomaker.
- 3Load an SDXL checkpoint
Use a photoreal SDXL model. PhotoMaker will not load on Flux or SD 1.5 graphs.
- 4Feed the stack
Batch your cropped reference images into one image input for the PhotoMaker encode node.
- 5Write the prompt with the trigger
Put "img" right after the class word: "photo of a woman img sitting in a cafe, natural window light".
- 6Set style strength and steps
Start near the official defaults: 20% style strength, 50 steps, CFG 5. Lower steps only once results look right.
- 7Generate four and review
Compare the face to the stack at thumbnail size and at full size. Keep only outputs that pass both.
- 8Save the graph
Store the workflow with the stack so every future batch starts from identical settings.
Settings that matter
PhotoMaker has few knobs, and one does most of the work. In the official V2 demo, the style strength percentage decides at which sampling step the identity is merged in: the code computes the merge step as style strength divided by 100 times the step count, capped at step 30. A higher percentage leaves more early steps to the base model and your style prompt, which helps stylized looks but loosens the likeness. The README suggests 30 to 50 when you want a stylized result and the face still looks too photographic.
| Setting | Starting value | What it changes |
|---|---|---|
| Style strength (%) | 20 (demo default), range 15-50 | Sets the step where the identity merges in; higher = more stylization, weaker likeness |
| Sampling steps | 50 (demo default) | Fewer steps are faster but the README warns ID fidelity may drop |
| Guidance scale (CFG) | 5 (demo default) | Keep it moderate on photoreal SDXL checkpoints |
| Trigger word | img, after the class word | "woman img" or "man img"; it is the slot where the identity goes |
| Base model | Any SDXL checkpoint | Photoreal SDXL checkpoints suit persona work best |
The README adds two practical tips. First, upload more photos of the person to improve fidelity. Second, ethnicity words before the class word can help when the base model drifts, for example writing "Asian woman img". Keep the rest of the prompt descriptive about scene, wardrobe and camera rather than the face, because the face is the job of the stack. Our free AI Influencer Prompt Generator builds scene prompts in that shape; add the "img" trigger after the class word before you paste one in.
Adding pose control and LoRAs
PhotoMaker carries the face, not the body position. For specific poses, combine it with ControlNet: TencentARC ships example scripts for V2 with ControlNet, T2I-Adapter and IP-Adapter, and in ComfyUI you simply apply a ControlNet to the conditioning after the PhotoMaker encode. Our ControlNet guide covers OpenPose, depth and canny for persona shots. Style and clothing LoRAs load as usual on the SDXL model; keep their weights moderate so they do not fight the identity.
Choice of base checkpoint changes skin, light and realism more than any PhotoMaker setting. Our list of the best ComfyUI models for realistic people covers SDXL options that suit persona work.
From PhotoMaker to a character LoRA
Many creators use PhotoMaker as a bridge. It gives you a consistent persona on day one, and its best outputs, varied in angle, outfit and light, make a clean training set for a character LoRA later. A LoRA is slower to set up but carries the identity on any checkpoint of the family it was trained for, without the adapter, and tends to hold small details such as moles or tooth shape better over hundreds of posts. When you get there, the LoRA Training Steps Calculator helps you size the run.
Building the persona system around these tools, from the first reference set to a weekly content batch and the accounts it posts to, is what our AI Influencers program walks through.
Licensing and consent
PhotoMaker's repository states the project is licensed under Apache 2.0, apart from listed third-party components. V2 depends on InsightFace for face analysis, and InsightFace says its pretrained models are for non-commercial research only; InstantID carries the same InsightFace caveat and calls its own checkpoints research-only. If you run a commercial persona, read those licenses yourself or ask a lawyer. And whatever the license says, generating a recognizable real person without their consent can break platform rules and, in many places, the law. This is not legal advice.
PhotoMaker AI: FAQ
What is PhotoMaker AI?
PhotoMaker is an open-source image model from TencentARC that customizes realistic human photos without training a LoRA. You give it one or more reference photos of a face and a prompt in which the word "img" follows a class word such as "woman img", and it merges the identity from all the references into that part of the prompt. It works with SDXL base models and was presented at CVPR 2024.
How many reference images does PhotoMaker need?
One image works, but the official demo and README both say more photos of the same person improve identity fidelity. The point of the method is that it stacks the embeddings from several images, so a set of four to eight clear, varied shots of the same face is a sensible starting stack. The face should fill most of each image, because the official demo notes it does not crop to the face for you.
Is PhotoMaker better than InstantID?
They solve the problem differently. InstantID is built around a single reference image plus a face-landmark ControlNet, so it locks the face and its geometry tightly. PhotoMaker merges several references into the prompt, which tends to leave pose, outfit and style more editable. TencentARC published side-by-side comparisons favoring PhotoMaker V2, but those are the developer's own selections, so test both on your persona before choosing.
Does PhotoMaker work in ComfyUI?
Yes. ComfyUI has had native PhotoMakerLoader and PhotoMakerEncode nodes since January 2024, which run the original V1 model. PhotoMaker V2 needs a community node pack such as ComfyUI-PhotoMaker-Plus, which adds an InsightFace loader because V2 uses InsightFace to extract the face identity.
Can I use PhotoMaker with a custom SDXL checkpoint?
Yes. The README describes PhotoMaker as an adapter that works with other SDXL base models and community LoRAs; its own comparison used RealVisXL V4.0. It does not work with SD 1.5 or Flux checkpoints, because the adapter was trained for the SDXL text encoders.
Can I use PhotoMaker commercially?
The PhotoMaker code and weights are under the Apache 2.0 license, with some third-party components listed separately. V2 also relies on InsightFace face models, and InsightFace states its pretrained models are for non-commercial research only. Read both licenses, and get permission before using any real person's photos as references. This is not legal advice.
The face is settled. Now build the persona around it.
AI Influencers, included in All Access, covers persona design, consistent faces, content batches and monetization, alongside the other three programs, live coaching and the private community in one subscription.
Write better persona prompts, free
Generate scene prompts for your PhotoMaker graph, and join the free Telegram channel for what is working in AI persona content right now.