Skip to main content

n8n Queue Mode: When You Need It and How to Set It Up

n8n queue mode explained: when one instance stops being enough, a worker-count formula with RAM maths, and a compose file for main, workers and Redis.

Founder of IImagined.ai

Published
Oct 11, 2026
Reading time
12 min read
Quick answer

n8n queue mode splits one n8n process into a main instance that receives triggers and webhooks and workers that run the executions, with Redis as the queue and Postgres as the store. You need it when peak concurrent executions, roughly executions per second times average run time, outgrow one process, or when restarts must not interrupt work. This guide gives the sizing formulas, a compose file with two workers, and the checks that prove a worker is doing the job.

n8n queue mode splits one n8n process into a main instance that receives triggers and webhooks and worker processes that run the executions, with Redis passing jobs between them and Postgres storing the data. You need it when your peak number of simultaneous executions is more than a single process can run without the editor and webhooks slowing down. Below that point, a concurrency limit on one instance is simpler and usually enough.

Checked October 2026 against n8n's docs on queue mode, concurrency, performance, task runners, binary data and the queue mode variables, n8n's official withPostgresAndWorker compose example, and the source of n8n 2.42.6, the stable release on 9 October 2026. The compose file below is built from that example and was not run for this page, so test it on a spare server first.

This guide belongs to our n8n hub and has one job: help you decide whether you need queue mode, size it, and switch an existing instance over. For a full production stack with HTTPS and backups, use the n8n Docker Compose setup guide. For the other limits of running n8n yourself, see n8n self-hosted limitations.

How n8n queue mode works

In regular mode one Node.js process serves the editor, receives webhooks, fires schedules and runs every execution. Queue mode keeps the first three jobs on the main instance and moves the fourth to workers. This is the process flow from n8n's docs:

One execution in queue mode
  1. 01
    Main creates the execution

    It handles timers and webhook calls, and generates the execution without running it.

  2. 02
    Redis queues the ID

    The execution ID waits in Redis until a worker is free.

  3. 03
    A worker picks it up

    The next available worker in the pool takes the message.

  4. 04
    The worker loads the workflow

    It reads the workflow and credentials from Postgres and runs the nodes.

  5. 05
    Results go to Postgres

    The worker writes the result and tells Redis the execution has finished.

  6. 06
    Redis notifies main

    The main instance learns the run is done and answers a waiting webhook caller.

Source: n8n queue mode docs, checked October 2026

Three facts about availability. Queue mode is included in the free Community Edition. Multi-main mode, which runs a second main instance for high availability, and the Workers view under Settings are Enterprise features. On n8n Cloud, queue mode exists on Enterprise plans only.

When do you need n8n queue mode? The maths

Later than most guides suggest. Raw throughput is rarely the reason, as n8n's own benchmark shows:

What one n8n instance can do
220/s
Workflow executions per second n8n says a single instance can handle
  • Example single-instance benchmark: a Webhook Trigger and an Edit Fields node
  • Hardware in that example: one c5a.large instance with 4 GB RAM and Postgres
  • Real workflows do far more per run, so measure your own

Source: n8n docs, Measure performance, checked October 2026

What breaks first is concurrency. In regular mode n8n "doesn't limit how many production executions may run at the same time", and the docs warn that too many can "thrash the event loop, causing performance degradation and unresponsiveness". So the number to estimate is how many executions overlap at your busiest moment:

peak concurrent executions = executions per second at peak x average run time in seconds
workers                    = round up (peak concurrent / worker concurrency) + 1 spare
Postgres connections       = about (main + workers + webhook processors) x DB_POSTGRESDB_POOL_SIZE
RAM                        = main + (workers x measured worker memory) + runners + Redis + Postgres

# Worked example with illustrative inputs
2 executions per second x 12 seconds     = 24 running at once
24 / 10 (default concurrency) = 2.4      -> 3 workers, plus 1 spare = 4
(1 main + 4 workers) x 2 (default pool)  = about 10 Postgres connections

The first line is the one that decides. Two executions a second that each take 12 seconds means 24 running at once, on the same process that serves your editor. The fixes come in order of effort:

What you seeDo thisWhy
The editor freezes or webhooks time out during burstsSet N8N_CONCURRENCY_PRODUCTION_LIMIT on the single instanceExtra runs wait in a queue. Move to queue mode when that wait becomes the problem.
Runs sit behind the concurrency limit at normal loadQueue modeWorkers add capacity. A limit can only ration what one process has.
The whole instance restarts on memory spikesFix binary data mode and batch sizes, then queue modeA worker that crashes leaves the editor and the webhook receiver running.
Deploys interrupt running executionsQueue modeWorkers get a shutdown window, 30 seconds by default, and can be restarted one at a time.
You need two main instances for high availabilityQueue mode with an Enterprise licenseMulti-main is not part of the Community Edition.

Queue mode also costs something. The docs note that for webhooks the request is received by the main process and the execution is passed to a worker, "which can add some overhead and latency". You add Redis, at least two more containers and a second place for things to go wrong. If a concurrency limit solves the problem, stop there.

How many n8n queue mode workers do you need?

Enough to cover your peak concurrent executions at the concurrency you set, plus one spare so a restart or a crash does not drop you below the peak. Four numbers go into the plan:

  • Concurrency per worker. The default is 10 jobs in parallel. n8n recommends 5 or higher, because "low concurrency values with a large numbers of workers can exhaust your database's connection pool".
  • Memory per worker. n8n publishes no per-worker figure, and it cannot: memory depends on the size of your JSON and binary data, the number of nodes and how much runs at once. Run a realistic load and read docker stats for the worker containers. That measured number is the one that goes in the RAM line.
  • Postgres connections. Each n8n process keeps a pool, 2 connections by default (DB_POSTGRESDB_POOL_SIZE). Keep the total well under your database's connection limit.
  • Redis memory. Small for job IDs. The exception is large webhook responses: the docs say to budget about 1.5 times N8N_WEBHOOK_RESPONSE_RELAY_SIZE_MAX, 64 MiB by default, for each response in flight.
Add workers or raise concurrency?
Memory has headroom
Raise --concurrency on the workers you have. It is the cheapest capacity.
Add a worker. Each one is its own Node.js process, so it brings another core into play.
Memory is tight
Lower concurrency or shrink batch sizes before you add anything.
Add workers on a larger or a second server.
CPU has headroom
CPU is saturated

Our sizing rule, read from docker stats under peak load. It is not an n8n recommendation.

Prerequisites

Pre-flight checklist
  • A working single-instance n8n on Docker Compose, and a fresh backup of its database and .n8n folder
  • PostgreSQL 16, 17 or 18 as the database: queue mode over SQLite is not supported
  • The encryption key your instance already uses, copied from its config or environment
  • Enough memory for Redis, two workers and three task runner containers on top of what you run today
  • No workflows that rely on files written to the n8n container's local disk
  • The same n8n version on the main instance, every worker and every runner image
  • A quiet hour: switching modes restarts the instance

If you are still on SQLite, move to Postgres first. Our Docker Compose guide covers that stack, and the self-hosted cost breakdown helps with sizing the server.

Set up n8n queue mode with Docker Compose: the steps

The procedure at a glance
  1. 1
    Back up

    Dump Postgres and copy the .n8n folder. Expected result: a restore you have tested, or at least files off the server.

  2. 2
    Write .env

    Pin the version and reuse the existing encryption key. Expected result: six values in a file only you can read.

  3. 3
    Add the queue services

    One shared block for main and workers, so they cannot drift apart. Expected result: eight services defined.

  4. 4
    Start the stack

    docker compose up -d. Expected result: postgres and redis healthy, the other six running.

  5. 5
    Check each worker

    Read the worker log and call its health endpoint. Expected result: "n8n worker is now ready" and a status of ok.

  6. 6
    Prove a run used a worker

    Trigger a published workflow. Expected result: the main log says Enqueued execution and a worker log says Worker started execution.

  7. 7
    Scale

    Add a worker block with its runner, or raise --concurrency, then run up -d again.

Step 1: back up, then write the .env file

Take the backup before anything else. Then put the settings every container shares in one place. Reuse the encryption key and the database credentials your instance already has: a new key makes every saved credential unreadable.

# .env  (chmod 600 .env, and never commit it)
N8N_VERSION=2.42.6
POSTGRES_USER=n8n
POSTGRES_PASSWORD=paste-secret-1
POSTGRES_DB=n8n
ENCRYPTION_KEY=the-key-your-instance-already-uses
RUNNERS_AUTH_TOKEN=paste-secret-2

Step 2: add the queue services to compose.yaml

This is n8n's official withPostgresAndWorker example with a second worker, an explicit concurrency, database binary storage and one database user to keep it short. The x-shared block is the point of the design: the main instance and every worker inherit the same settings, so a mismatched key or Redis host cannot creep in.

Starting fresh, the file works as it is. Converting an existing stack, keep your own postgres service, volume names and database credentials, and take from this file the x-shared settings, redis, the workers and the runners. A compose file that points at differently named volumes starts n8n on an empty database, which looks exactly like data loss until you switch the names back.

volumes:
  db_storage:
  n8n_storage:
  redis_storage:

x-shared: &shared
  restart: always
  image: docker.n8n.io/n8nio/n8n:${N8N_VERSION}
  environment:
    - DB_TYPE=postgresdb
    - DB_POSTGRESDB_HOST=postgres
    - DB_POSTGRESDB_PORT=5432
    - DB_POSTGRESDB_DATABASE=${POSTGRES_DB}
    - DB_POSTGRESDB_USER=${POSTGRES_USER}
    - DB_POSTGRESDB_PASSWORD=${POSTGRES_PASSWORD}
    - EXECUTIONS_MODE=queue
    - QUEUE_BULL_REDIS_HOST=redis
    - QUEUE_HEALTH_CHECK_ACTIVE=true
    - N8N_ENCRYPTION_KEY=${ENCRYPTION_KEY}
    - N8N_DEFAULT_BINARY_DATA_MODE=database
    - OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS=true
    - N8N_RUNNERS_MODE=external
    - N8N_RUNNERS_AUTH_TOKEN=${RUNNERS_AUTH_TOKEN}
    - N8N_RUNNERS_BROKER_LISTEN_ADDRESS=0.0.0.0
  volumes:
    - n8n_storage:/home/node/.n8n
  depends_on:
    redis:
      condition: service_healthy
    postgres:
      condition: service_healthy

x-runner: &runner
  restart: always
  image: n8nio/runners:${N8N_VERSION}

services:
  postgres:
    image: postgres:18
    restart: always
    environment:
      - POSTGRES_USER
      - POSTGRES_PASSWORD
      - POSTGRES_DB
      - PGDATA=/var/lib/postgresql/data
    volumes:
      - db_storage:/var/lib/postgresql/data
    healthcheck:
      test: ['CMD-SHELL', 'pg_isready -h localhost -U ${POSTGRES_USER} -d ${POSTGRES_DB}']
      interval: 5s
      timeout: 5s
      retries: 10

  redis:
    image: redis:7-alpine
    restart: always
    volumes:
      - redis_storage:/data
    healthcheck:
      test: ['CMD', 'redis-cli', 'ping']
      interval: 5s
      timeout: 5s
      retries: 10

  n8n:
    <<: *shared
    ports:
      - 5678:5678

  n8n-runner:
    <<: *runner
    environment:
      - N8N_RUNNERS_AUTH_TOKEN=${RUNNERS_AUTH_TOKEN}
      - N8N_RUNNERS_TASK_BROKER_URI=http://n8n:5679
    depends_on:
      - n8n

  n8n-worker-1:
    <<: *shared
    command: worker --concurrency=10
    depends_on:
      - n8n

  n8n-worker-1-runner:
    <<: *runner
    environment:
      - N8N_RUNNERS_AUTH_TOKEN=${RUNNERS_AUTH_TOKEN}
      - N8N_RUNNERS_TASK_BROKER_URI=http://n8n-worker-1:5679
    depends_on:
      - n8n-worker-1

  n8n-worker-2:
    <<: *shared
    command: worker --concurrency=10
    depends_on:
      - n8n

  n8n-worker-2-runner:
    <<: *runner
    environment:
      - N8N_RUNNERS_AUTH_TOKEN=${RUNNERS_AUTH_TOKEN}
      - N8N_RUNNERS_TASK_BROKER_URI=http://n8n-worker-2:5679
    depends_on:
      - n8n-worker-2

Step 3: start it and check each worker

docker compose up -d
docker compose ps
docker compose logs n8n-worker-1 | tail -20
docker compose exec n8n-worker-1 wget -qO- http://localhost:5678/healthz/readiness

Expected result: the worker log ends with n8n worker is now ready, followed by the version and Concurrency: 10. The readiness call prints {"status":"ok"} once the worker's database and Redis connections are up. Repeat for the second worker.

Step 4: prove a production run went through a worker

Publish a small workflow with a Schedule Trigger set to every minute and a Code node, wait for it to fire, then compare docker compose logs n8n with the worker logs. The main instance logs Enqueued execution with an ID, and one worker logs Worker started execution and Worker finished execution for the same ID. The Code node matters: it only succeeds if that worker's task runner is connected. Because the file sets OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS=true, test runs from the editor take the same route.

Step 5: scale

To add capacity, copy the n8n-worker-2 and n8n-worker-2-runner blocks, rename them to 3, point the new runner at http://n8n-worker-3:5679 and run docker compose up -d. To change concurrency, edit the command line. Workers on a second server work the same way, as long as they can reach the same Postgres and Redis and carry the same key.

The queue mode settings that matter

SettingWhereWhat it does
EXECUTIONS_MODE=queueMain and every workerSwitches n8n from running executions itself to queueing them
QUEUE_BULL_REDIS_HOSTMain and every workerWhere Redis is. Port, password, TLS and cluster nodes have their own QUEUE_BULL_REDIS_ variables
N8N_ENCRYPTION_KEYMain and every worker, identicalWorkers decrypt saved credentials with it
N8N_DEFAULT_BINARY_DATA_MODE=databaseMain and every workerFilesystem storage is not supported in queue mode. S3 is the paid alternative
worker --concurrency=10Each worker's commandJobs one worker runs in parallel. The default is 10 and n8n recommends 5 or more
N8N_CONCURRENCY_PRODUCTION_LIMITWorkers, optionalWhen set to anything other than -1, it overrides the --concurrency flag
QUEUE_HEALTH_CHECK_ACTIVE=trueWorkersTurns on /healthz and /healthz/readiness on each worker, port 5678 by default
N8N_GRACEFUL_SHUTDOWN_TIMEOUTWorkersSeconds a worker gets to finish its jobs on shutdown. The default is 30
OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS=trueMain, n8n 2.x onlySends editor test runs to workers. Removed in n8n 3.0, where that is always the behaviour

The concurrency override is easy to miss. In worker.ts, n8n takes the limit from N8N_CONCURRENCY_PRODUCTION_LIMIT when that is set to anything other than -1, and only then falls back to the flag. If you set that variable for a single instance last year, it now quietly sets your worker concurrency.

Scale n8n self-hosted further: webhooks, health checks and metrics

Workers are the first layer. When the main instance itself becomes the bottleneck for incoming requests, n8n has three more tools:

  • Webhook processors. Extra n8n processes started with the webhook command that only receive production webhooks. Put a load balancer in front, send /webhook/* and /webhook-waiting/* to them and everything else, including /webhook-test/*, to the main instance. The docs advise against adding the main process to that pool.
  • Health checks. With QUEUE_HEALTH_CHECK_ACTIVE=true every worker serves /healthz and /healthz/readiness, which is what a container orchestrator needs to restart a stuck worker.
  • Metrics. N8N_METRICS=true exposes /metrics for Prometheus, and N8N_METRICS_INCLUDE_QUEUE_METRICS=true adds job counts for queue mode. A queue that keeps growing means you are short of workers.

Workflows that stop without an error are a separate problem from capacity; our checklist of five ways n8n workflows fail silently covers those.

What changes for queue mode in n8n 3.0

Three items on n8n's 3.0 breaking-changes page touch this setup. Manual executions always run on workers and OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS is removed, so delete that line when you upgrade. The in-memory binary data mode is removed. And the internal task runner mode is deprecated, which the external runners in this file already avoid. That page says 3.0 "has been released", but on 11 October 2026 the latest stable release on GitHub was 2.42.6, so treat 3.0 as not generally available yet. Our n8n 3.0 upgrade guide tracks the rest.

Troubleshooting

What you seeCauseFix
Worker log: "Concurrency is set to less than 5. THIS CAN LEAD TO AN UNSTABLE ENVIRONMENT."Concurrency is below n8n's recommended floorRaise --concurrency to 5 or more and run fewer workers
"Credentials could not be decrypted"A worker has a different N8N_ENCRYPTION_KEY from the main instanceSet the same key on every instance and restart
Code node fails with "Task request timed out"That worker has no reachable task runnerGive each worker its own runner sidecar and point the runner at that worker's port 5679
Files are missing in later nodesBinary data is kept in memory or on one container's diskSet N8N_DEFAULT_BINARY_DATA_MODE=database on every instance
"The response is too large to be sent back from the worker"A webhook response exceeded N8N_WEBHOOK_RESPONSE_RELAY_SIZE_MAX, 64 MiB by defaultReturn less, raise the limit, or from n8n 2.34.0 turn on N8N_WEBHOOK_RESPONSE_RELAY_OFFLOAD_ENABLED
Executions are created but never startNo worker is connected to the same RedisCheck the worker log for "n8n worker is now ready" and compare the QUEUE_BULL_REDIS_ settings

Running infrastructure like this is its own skill, and it carries over. Our AI SaaS Builder program has no n8n lessons, but it does cover the same questions for an app you build yourself: deployment, caching and performance, and infrastructure for scale.

n8n queue mode: FAQ

What is queue mode in n8n?

Queue mode splits n8n into a main instance and worker processes. The main instance handles the editor, timers and incoming webhooks and creates each execution without running it. It passes the execution ID to Redis, a worker picks it up, runs the workflow and writes the result to the database. You scale by adding or removing workers. It is set with EXECUTIONS_MODE=queue.

When do I need n8n queue mode?

When your peak number of simultaneous executions is more than one n8n process can run without the editor and webhooks slowing down, or when a restart must not interrupt running work. Estimate the peak as executions per second multiplied by average run time. Below that point, setting N8N_CONCURRENCY_PRODUCTION_LIMIT on a single instance is simpler: extra runs wait in a queue.

How many workers do I need for n8n queue mode?

Divide your peak concurrent executions by the worker concurrency, round up and add one spare. A worker runs 10 jobs in parallel by default, so 24 concurrent executions need three workers plus a spare. n8n recommends a concurrency of 5 or higher, because many workers with low concurrency can exhaust the database connection pool. Measure memory per worker under load before adding more.

Is n8n queue mode free?

On self-hosted n8n, yes. The n8n edition comparison says queue mode is included in the free Community Edition. Multi-main mode, which runs more than one main instance for high availability, and the Workers view in Settings need an Enterprise license. On n8n Cloud, queue mode is available on Enterprise plans only. Checked October 2026 against the n8n docs.

Does n8n queue mode need Redis and Postgres?

Yes. Redis is the message broker that holds the queue of pending executions, and the database stores workflows, credentials and results, so the main instance and every worker need access to both. The n8n docs say running a distributed setup over SQLite is not supported. Use PostgreSQL 16, 17 or 18, and give every instance the same N8N_ENCRYPTION_KEY.

What is the default worker concurrency in n8n?

Ten. A worker started with n8n worker runs up to 10 jobs in parallel, and the --concurrency flag changes that, for example n8n worker --concurrency=5. If N8N_CONCURRENCY_PRODUCTION_LIMIT is set to anything other than -1, it overrides the flag. Below 5, n8n 2.42.6 logs a warning that the setup can become unstable. Checked October 2026.

Does queue mode change in n8n 3.0?

In one way. The 3.0 breaking-changes page says manual executions always run on workers in queue mode and removes OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS, so workers need memory for editor test runs too. On n8n 2.x you get the same behaviour by setting that variable to true. On 11 October 2026 the latest stable release on GitHub was still 2.42.6.

All Access · all four programs · $99/mo

Your automations can take the load. Now build the product that sends it.

AI SaaS Builder, included in All Access, covers what sits around a stack like this: a Next.js and Supabase app, Claude API features, Claude Code and MCP, deployment and scaling on Vercel, and Stripe billing. It has no n8n lessons. All Access adds the other three programs, live coaching and the private community.

Start All Access — $99/mo →30-day money-back guarantee
Free

Sizing a queue mode setup?

Share your peak numbers and compose file in the free Discord and compare notes with other people running n8n on their own servers.