Advanced ComfyUI work comes down to a handful of techniques: ControlNet to copy pose, edges or depth; area and mask conditioning to prompt parts of the image separately; reference-image models such as Qwen-Image-Edit and FLUX.2 where IP-Adapter used to be the default; model, two-pass and tiled upscaling; few-step models for fast drafts; subgraphs to keep large graphs manageable; and the /prompt API for batches. Add them one at a time to an official template and compare each change at a fixed seed.
Checked against ComfyUI's official documentation, its example workflows and the source code of the nodes named here on October 1, 2026, when v0.38.0 was the latest ComfyUI release. This guide explains what each technique does and when to use it; for complete step-by-step builds, see our advanced ComfyUI workflows guide.
How hard is ComfyUI to learn?
ComfyUI is harder to pick up than a prompt box because you build the pipeline yourself: each node does one job, and links pass models, conditioning, latents and images between them. The hard part is usually matching, not the interface: a checkpoint, its text encoder, VAE, LoRAs and ControlNets must belong to the same model family, and shared workflows often need custom nodes you have not installed.
The tooling now does more of that matching. Built-in templates check for missing model files and link to them, ComfyUI-Manager offers to install missing nodes, and App mode turns a finished workflow into a simple form. There is no reliable universal timeline, so track milestones:
- Run an official template unchanged.
- Change one thing, such as the checkpoint, a LoRA or the resolution, and explain the difference.
- Fix a missing node or missing model error on your own.
- Add one technique from this guide to a working graph.
- Run the graph from a script through the API.
Still at step 1? Start with our ComfyUI installation and first-image guide and ComfyUI-Manager guide.
Which technique solves which problem
| Problem | Technique | Key nodes |
|---|---|---|
| Match a pose, outline or depth from a reference | ControlNet | Load ControlNet Model, Apply ControlNet |
| Different prompts for different parts of the frame | Area or mask conditioning | Conditioning (Set Area), Conditioning (Set Mask), Conditioning (Combine) |
| Keep a face, product or style from a reference image | Reference-image editing models, or IP-Adapter on older models | Model-specific templates |
| Repaint one region | Inpainting | Mask Editor, VAE Encode (for Inpainting) |
| Bigger, sharper output | Model upscaling, two-pass or tiled refinement, SeedVR2 | Upscale Image (using Model), Ultimate SD Upscale |
| Faster drafts | Few-step models and LoRAs | Model-specific sampler settings |
| A graph too big to manage | Subgraphs, blueprints, partial execution | Selection toolbox |
| Many variations or scheduled jobs | Seeds, list outputs and the HTTP API | KSampler seed control, /prompt |
| A step no node covers | A custom node | INPUT_TYPES or the V3 schema |
ControlNet: copy structure from a reference
ControlNet, introduced by Lvmin Zhang and Maneesh Agrawala in 2023, adds a condition image such as an edge map, a depth map or pose keypoints to guide generation. Start with the downloadable workflow in the official ControlNet tutorial and use its specified checkpoint, control model and input image before swapping in your own.
How the data flows
- Encode the prompts into positive and negative conditioning.
- Load a ControlNet that matches your base model family. Comfy's examples pair SD1.5 ControlNets with SD1.5 checkpoints.
- Preprocess the reference into the image type the model expects: a scribble, Canny edges, a depth map or an OpenPose skeleton. Core ComfyUI does not include every preprocessor, so the docs point to comfyui_controlnet_aux and ComfyUI-Advanced-ControlNet.
- Apply ControlNet takes the positive and negative conditioning, the ControlNet, the preprocessed image and optionally a VAE, and outputs new positive and negative conditioning.
- Connect that conditioning, the model and a latent to the KSampler, then decode with the matching VAE.
Settings that matter
- strength: how much the ControlNet influences the result. Higher values increase its influence; they do not guarantee an exact pose or identity.
- start_percent and end_percent: when guidance begins and ends. In Comfy's example, 0.2 starts it a fifth of the way through sampling and 0.8 stops it four fifths of the way through.
- denoise (on the KSampler): the KSampler reference describes how much of the initial latent is changed: 1.0 for text-to-image, lower values to keep more of an encoded input image. It is a separate control from ControlNet strength.
Older shared workflows may show Apply ControlNet (Old), which Comfy has deprecated in favor of the current node.
Stacking several ControlNets
Chain Apply ControlNet nodes, passing each node's conditioning to the next. For different regions, such as a pose on the left and a scribble on the right, Comfy's mixing guide suggests similar strengths, for example 1.0 each, so one region does not suppress the other. For several controls on one subject, such as pose plus depth, every reference image must line up with that subject. Union ControlNets bundle several control types into one model; the Set Union ControlNet Type node picks the type or leaves it on auto.
Newer model families ship their own control models. Z-Image-Turbo's Fun Union ControlNet, for example, is a model patch covering Canny, HED, depth, pose and MLSD (Z-Image-Turbo guide). Our ControlNet guide goes deeper on control types.
Regional prompting with area and mask conditioning
A single prompt applies to the whole image, so attributes can bleed between subjects. Regional prompting gives each part of the frame its own conditioning. The core nodes:
- Conditioning (Set Area) applies a prompt to a rectangle given in pixels (x, y, width, height) with a strength; Conditioning (Set Area with Percentage) takes the same values as fractions of the image.
- Conditioning (Set Mask) applies a prompt to a painted mask, with
set_cond_areaset to default or mask bounds. - Conditioning (Combine) merges the regional prompts and your whole-frame prompt into one conditioning for the sampler.
The pattern: encode one prompt for the whole scene and one per region, run each regional prompt through Set Area or Set Mask, combine them all, and send the result to the KSampler's positive input. ComfyUI's area composition example adds two cautions: Stable Diffusion is most consistent near its training resolution, so a subject in its own square area helps in wide images, and a second pass without the area prompts can blend attributes between subjects, such as hair colors. For a single fix, an instruction-editing model that changes only the object you name is often simpler.
Reference images: IP-Adapter and what replaced it
SD1.5 and SDXL workflows kept a face, product or style from a reference image with IP-Adapter, through the ComfyUI_IPAdapter_plus node pack. Its developer put the repository in maintenance-only mode on April 14, 2025, so treat it as a legacy option for those model families. If you still use it, its README lists ip-adapter-plus-face_sdxl_vit-h.safetensors as the SDXL face model: put it in ComfyUI/models/ipadapter, with the CLIP-ViT-H-14-laion2B-s32B-b79K.safetensors image encoder in models/clip_vision. Newer models accept reference images natively, and each has an official template:
| Model | Reference support | Model license |
|---|---|---|
| Qwen-Image-Edit-2511 | Instruction editing with reference images; better character and multi-person consistency than 2509 | Apache 2.0 |
| FLUX.2 [klein] 4B | Single- and multi-reference editing; Comfy lists 8.4 GB of VRAM for the distilled model | Apache 2.0 |
| FLUX.2 [dev] | Up to 10 reference images | FLUX [dev] Non-Commercial License |
| Qwen-Image 2.1 | Text encoder node accepts up to 16 image slots | Qwen Research License (research or evaluation only) |
Licenses come from each model's Hugging Face page. Black Forest Labs' non-commercial license lets you use outputs commercially but limits the model itself to non-commercial use unless you buy a commercial license. Give each reference image one job and refer to it by its slot in the prompt; Qwen-Image 2.1's template treats image_1 as the image being edited. For a recurring character, a trained character LoRA is still the most direct route.
Inpainting: repaint one region
When most of an image works, repaint only the broken part: draw a mask in the Mask Editor, encode with VAE Encode (for Inpainting), whose grow_mask_by setting widens the mask so the edit blends in, and sample. Comfy's inpainting tutorial shows that a model trained for inpainting blends better than a general checkpoint.
Upscaling: four methods that behave differently
Comfy's upscaling guide separates conservative upscalers, which preserve the original, from creative ones, which add detail but can invent it. It also warns against relying on upscaling to fix AI artifacts such as extra fingers; fix those first.
| Method | How it works | Use it for |
|---|---|---|
| Upscale model (ESRGAN family) | A trained upscaler enlarges the image; Comfy suggests RealESRGAN for general use, BSRGAN for text, SwinIR for natural textures | Fast enlargement, no diffusion pass |
| Two-pass (hires fix) | Generate small, upscale, then sample again at a lower denoise | Adding detail at a higher resolution |
| Tiled refinement (Ultimate SD Upscale) | Upscales, then runs image-to-image tile by tile | Large outputs on modest hardware |
| SeedVR2 | One-step diffusion restoration, native in ComfyUI | Faithful image and video upscales |
Sources: basic upscale tutorial, two-pass example, Ultimate SD Upscale README and SeedVR2 guide.
Ultimate SD Upscale inputs (INPUT_TYPES)
Ultimate SD Upscale is a GPL-3.0 custom node pack (registry ID comfyui_ultimatesdupscale) that appears under image/upscaling. The main node's INPUT_TYPES, from its source file:
- Sampling:
image,model,positive,negative,vae,seed,steps(default 20),cfg(8.0),sampler_name,scheduleranddenoise(0.2). - Upscale:
upscale_model,upscale_by(default 2, range 0.05 to 4),mode_type(Linear, Chess or None),tile_widthandtile_height(512),mask_blur(8) andtile_padding(32). - Seam fix:
seam_fix_mode(None, Band Pass, Half Tile, or Half Tile + Intersections),seam_fix_denoise(1.0),seam_fix_width(64),seam_fix_mask_blur(8) andseam_fix_padding(16). - Other:
force_uniform_tiles(on),tiled_decode(off) andbatch_size(1; values above 1 needforce_uniform_tiles).
Variants: UltimateSDUpscaleNoUpscale refines an image you have already upscaled, UltimateSDUpscaleCustomSample takes a custom sampler and sigmas, and UltimateSDUpscaleGuider works with a guider. A running server lists any installed node's inputs at /object_info/{node_class}, for example /object_info/UltimateSDUpscale (server routes).
SeedVR2
ByteDance's SeedVR2 is natively supported (PR #14424) in 3B and 7B sizes under Apache 2.0, with FP16, FP8 and INT8 files. Comfy's tip is to downscale an image to 0.35 megapixels with ImageScaleToTotalPixels before upscaling it with SeedVR2.
Faster generation: few-step models and LoRAs
Distilled models and step-reduction LoRAs cut sampling to a handful of steps. Each needs its own sampler settings, so start from its official example rather than lowering the step count of a normal workflow.
| Option | Documented settings | Source |
|---|---|---|
| SDXL Turbo | 1 to 4 steps, guidance disabled, trained at 512x512; ComfyUI uses the SDTurboScheduler node | ComfyUI Turbo example, Diffusers docs |
| LCM LoRA | 4 to 8 steps, guidance 1.0 to 2.0; in ComfyUI the lcm sampler with the sgm_uniform or simple scheduler | Diffusers LCM guide, ComfyUI LCM example |
| Z-Image-Turbo | 8 function evaluations; fits 16 GB consumer GPUs, per Comfy | Z-Image-Turbo guide |
| FLUX.2 [klein] 4B distilled | 4 steps; about 1.2 seconds on an RTX 5090 with 8.4 GB of VRAM, per Comfy | Klein guide |
| Lightning and LightX2V LoRAs | 4-step LoRAs for Qwen-Image-Edit-2511 and Wan 14B video | Qwen Edit 2511 guide |
- Use one acceleration method at a time. SDXL Turbo and an LCM LoRA are separate approaches, so do not combine them into one speed recipe.
- Check what cfg does. Many few-step setups run at cfg 1, and Comfy's Qwen-Image 2.1 guide notes that at cfg 1 ComfyUI skips the negative conditioning pass, so a negative prompt has no effect.
- Check the license. Stability lists SDXL Turbo under its Community License, free for commercial use below $1 million in annual revenue; Z-Image-Turbo and FLUX.2 [klein] 4B are Apache 2.0.
Measure the whole workflow, not just the sampler
We have not benchmarked these options, and speed figures rarely transfer between machines. Record your GPU, precision, resolution, sampler and steps; separate the first model load from warm runs; keep prompts and seeds fixed; count retries, upscaling and review time; and change one setting at a time.
Subgraphs, blueprints and partial execution
- Subgraphs (frontend 1.24.3 or later): select nodes and click the subgraph icon to package them as one node with exposed inputs and outputs. Double-click to edit it, or unpack it back into nodes.
- Subgraph blueprints (frontend 1.27.7 or later): publish a subgraph to the node library and reuse it across workflows.
- Partial execution: select an output node and run only the branch that feeds it.
- App mode (frontend 1.41.13 or later): expose chosen inputs and outputs as a simple form. Share links work on Comfy Cloud only.
Sources: Comfy's subgraph guide, partial execution and App mode pages.
Batch generation and the API
Inside the app
- batch_size on the latent node renders several images in one pass and uses more VRAM.
- Seed control: the KSampler's seed widget can change the seed after each run, so queueing the workflow several times gives you variations.
- List outputs: when a node outputs a list, ComfyUI runs the downstream nodes once per item (data lists). The custom node below uses this to turn one run into many prompts.
From a script
- Export the graph with File, then Export Workflow (API), which keys nodes by ID and drops layout data (API format).
- POST it to
/prompt; the server returns aprompt_id, ornode_errorsif validation fails. - Poll
/history/{prompt_id}or listen on the/wsWebSocket, then fetch files from/view.
A minimal loop using Python's standard library, adapted from ComfyUI's basic_api_example.py. Node IDs 3 (KSampler) and 6 (positive prompt) come from that example, so check the IDs in your own export:
import json, time, urllib.request
SERVER = "http://127.0.0.1:8188"
workflow = json.load(open("workflow_api.json")) # File > Export Workflow (API)
def post(path, data):
req = urllib.request.Request(SERVER + path, data=json.dumps(data).encode("utf-8"),
headers={"Content-Type": "application/json"})
return json.loads(urllib.request.urlopen(req).read())
prompts = ["portrait in a sunlit cafe", "portrait on a rainy street at night"]
for i, text in enumerate(prompts):
workflow["6"]["inputs"]["text"] = text
workflow["3"]["inputs"]["seed"] = 1000 + i
prompt_id = post("/prompt", {"prompt": workflow})["prompt_id"]
while True: # history appears once the run has finished
history = json.loads(urllib.request.urlopen(f"{SERVER}/history/{prompt_id}").read())
if prompt_id in history:
break
time.sleep(1)
print(text, "->", list(history[prompt_id]["outputs"]))These routes have no login. ComfyUI binds to 127.0.0.1 by default, and its security policy leaves securing anything you expose with --listen to you. For new integrations, Comfy points to its versioned Comfy API v2, in beta, which runs on Comfy Cloud with an API key and on self-hosted ComfyUI through a small proxy.
Custom nodes: when to write your own
Write a custom node only when built-in nodes, a subgraph or an existing registry pack cannot handle a step you repeat. A node is a Python class: INPUT_TYPES declares its inputs, RETURN_TYPES its outputs and FUNCTION the method ComfyUI calls, and NODE_CLASS_MAPPINGS registers it (node properties, loading lifecycle). That V1 schema is still fully supported; new nodes can also use the V3 schema.
This example turns a character trigger and a theme into a list of prompt variations. It only writes text; it does not render images or check identity. OUTPUT_IS_LIST makes ComfyUI run the text encoder and sampler once per prompt, and the seed input keeps runs reproducible: ComfyUI reruns a node only when its inputs change, and its docs recommend a seed input for any node that uses randomness.
# ComfyUI/custom_nodes/batch_scenarios/__init__.py
import random
SCENES = {
"Fashion": {
"actions": ["walking toward the camera", "adjusting a jacket sleeve", "sitting at a cafe table"],
"places": ["boutique", "rooftop at dusk", "city street"],
"outfits": ["tailored blazer", "linen summer dress", "streetwear set"],
"light": ["golden hour light", "overcast daylight", "neon night light"],
},
"Fitness": {
"actions": ["tying running shoes", "stretching after a run", "lifting a kettlebell"],
"places": ["outdoor track", "home gym", "park path"],
"outfits": ["running jacket and leggings", "tank top and shorts", "matching training set"],
"light": ["early morning light", "bright midday sun", "soft window light"],
},
}
class BatchScenarioGenerator:
"""Builds prompt variations. OUTPUT_IS_LIST makes ComfyUI run the
downstream nodes once per prompt; the seed makes each run reproducible."""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"character": ("STRING", {"default": "ava_character, photo of a woman"}),
"theme": (list(SCENES.keys()),),
"count": ("INT", {"default": 8, "min": 1, "max": 64}),
"seed": ("INT", {"default": 0, "min": 0, "max": 0xFFFFFFFF, "control_after_generate": True}),
}
}
RETURN_TYPES = ("STRING", "INT")
RETURN_NAMES = ("prompts", "count")
OUTPUT_IS_LIST = (True, False)
FUNCTION = "build"
CATEGORY = "examples/prompts"
def build(self, character, theme, count, seed):
rng = random.Random(seed)
scene = SCENES[theme]
prompts = [
f"{character}, {rng.choice(scene['actions'])}, {rng.choice(scene['places'])}, "
f"wearing a {rng.choice(scene['outfits'])}, {rng.choice(scene['light'])}"
for _ in range(count)
]
return (prompts, count)
NODE_CLASS_MAPPINGS = {"BatchScenarioGenerator": BatchScenarioGenerator}
NODE_DISPLAY_NAME_MAPPINGS = {"BatchScenarioGenerator": "Batch Scenario Generator"}Restart ComfyUI, find the node under examples/prompts, and connect prompts to the text input of a CLIP Text Encode node. Replace ava_character with your LoRA's trigger word; the word alone does not load a LoRA. Rerun a small test workflow after every ComfyUI update.
Custom nodes are ordinary Python with your user's permissions, which ComfyUI's security policy trusts like any other software you install. Install them from the Comfy Registry through ComfyUI-Manager, and read our guide to finding and vetting ComfyUI workflows before loading a stranger's graph.
Advanced ComfyUI FAQ
Is ComfyUI hard to learn?
It is harder than a prompt-only tool because you build the pipeline from nodes, and the models, text encoders, VAEs, LoRAs and ControlNets in a graph must belong to the same model family. Built-in templates, missing-node installs, subgraphs and App mode now remove much of that setup. There is no reliable universal timeline, so track milestones: run a template, change one node, fix a missing model yourself, then add one technique at a time.
What is the difference between ControlNet strength and denoise?
ControlNet strength sets how much the control image steers the conditioning. Denoise, on the KSampler, sets how much of the initial latent is changed: 1.0 for text-to-image, lower values to keep more of an encoded input image. They are separate controls, and raising either one does not guarantee an exact pose or face.
How do I do regional prompting in ComfyUI?
Encode one prompt for the whole image and one for each region, apply Conditioning (Set Area) or Conditioning (Set Mask) to each regional prompt, merge them with Conditioning (Combine), and send the result to the KSampler's positive input. Mixing ControlNets by region is another option.
What replaced IP-Adapter in ComfyUI?
ComfyUI_IPAdapter_plus has been in maintenance-only mode since April 14, 2025 and still suits SD1.5 and SDXL. Newer models accept reference images natively, with official templates for Qwen-Image-Edit-2511 and FLUX.2 [klein] 4B (both Apache 2.0) and FLUX.2 [dev] and Qwen-Image 2.1 (non-commercial model licenses).
What are the inputs of Ultimate SD Upscale?
The UltimateSDUpscale node takes an image, model, positive and negative conditioning, VAE and upscale model, plus upscale_by (default 2), seed, steps (20), cfg (8.0), sampler, scheduler, denoise (0.2), tile size (512 by 512), mask_blur (8), tile_padding (32), seam-fix settings, force_uniform_tiles, tiled_decode and batch_size. A running server lists them at /object_info/UltimateSDUpscale.
Can I use these techniques for commercial work?
The techniques are free to use, but each model has its own license. Qwen-Image-Edit-2511, Z-Image-Turbo, FLUX.2 [klein] 4B and SeedVR2 are Apache 2.0. FLUX.2 [dev] and Qwen-Image 2.1 restrict how you use the model itself, and SDXL Turbo is free for commercial use only below $1 million in annual revenue. Read the license on each model page before client work.
Related Reads
Want the full AI Influencers playbook?
The complete pipeline for building virtual brands at scale — identity engineering, ComfyUI production, IP governance, and the distribution flywheel.