n8n queue mode splits one n8n process into a main instance that receives triggers and webhooks and workers that run the executions, with Redis as the queue and Postgres as the store. You need it when peak concurrent executions, roughly executions per second times average run time, outgrow one process, or when restarts must not interrupt work. This guide gives the sizing formulas, a compose file with two workers, and the checks that prove a worker is doing the job.
n8n queue mode splits one n8n process into a main instance that receives triggers and webhooks and worker processes that run the executions, with Redis passing jobs between them and Postgres storing the data. You need it when your peak number of simultaneous executions is more than a single process can run without the editor and webhooks slowing down. Below that point, a concurrency limit on one instance is simpler and usually enough.
Checked October 2026 against n8n's docs on queue mode, concurrency, performance, task runners, binary data and the queue mode variables, n8n's official withPostgresAndWorker compose example, and the source of n8n 2.42.6, the stable release on 9 October 2026. The compose file below is built from that example and was not run for this page, so test it on a spare server first.
This guide belongs to our n8n hub and has one job: help you decide whether you need queue mode, size it, and switch an existing instance over. For a full production stack with HTTPS and backups, use the n8n Docker Compose setup guide. For the other limits of running n8n yourself, see n8n self-hosted limitations.
How n8n queue mode works
In regular mode one Node.js process serves the editor, receives webhooks, fires schedules and runs every execution. Queue mode keeps the first three jobs on the main instance and moves the fourth to workers. This is the process flow from n8n's docs:
- 01Main creates the execution
It handles timers and webhook calls, and generates the execution without running it.
- 02Redis queues the ID
The execution ID waits in Redis until a worker is free.
- 03A worker picks it up
The next available worker in the pool takes the message.
- 04The worker loads the workflow
It reads the workflow and credentials from Postgres and runs the nodes.
- 05Results go to Postgres
The worker writes the result and tells Redis the execution has finished.
- 06Redis notifies main
The main instance learns the run is done and answers a waiting webhook caller.
Source: n8n queue mode docs, checked October 2026
Three facts about availability. Queue mode is included in the free Community Edition. Multi-main mode, which runs a second main instance for high availability, and the Workers view under Settings are Enterprise features. On n8n Cloud, queue mode exists on Enterprise plans only.
When do you need n8n queue mode? The maths
Later than most guides suggest. Raw throughput is rarely the reason, as n8n's own benchmark shows:
- Example single-instance benchmark: a Webhook Trigger and an Edit Fields node
- Hardware in that example: one c5a.large instance with 4 GB RAM and Postgres
- Real workflows do far more per run, so measure your own
Source: n8n docs, Measure performance, checked October 2026
What breaks first is concurrency. In regular mode n8n "doesn't limit how many production executions may run at the same time", and the docs warn that too many can "thrash the event loop, causing performance degradation and unresponsiveness". So the number to estimate is how many executions overlap at your busiest moment:
peak concurrent executions = executions per second at peak x average run time in seconds
workers = round up (peak concurrent / worker concurrency) + 1 spare
Postgres connections = about (main + workers + webhook processors) x DB_POSTGRESDB_POOL_SIZE
RAM = main + (workers x measured worker memory) + runners + Redis + Postgres
# Worked example with illustrative inputs
2 executions per second x 12 seconds = 24 running at once
24 / 10 (default concurrency) = 2.4 -> 3 workers, plus 1 spare = 4
(1 main + 4 workers) x 2 (default pool) = about 10 Postgres connectionsThe first line is the one that decides. Two executions a second that each take 12 seconds means 24 running at once, on the same process that serves your editor. The fixes come in order of effort:
| What you see | Do this | Why |
|---|---|---|
| The editor freezes or webhooks time out during bursts | Set N8N_CONCURRENCY_PRODUCTION_LIMIT on the single instance | Extra runs wait in a queue. Move to queue mode when that wait becomes the problem. |
| Runs sit behind the concurrency limit at normal load | Queue mode | Workers add capacity. A limit can only ration what one process has. |
| The whole instance restarts on memory spikes | Fix binary data mode and batch sizes, then queue mode | A worker that crashes leaves the editor and the webhook receiver running. |
| Deploys interrupt running executions | Queue mode | Workers get a shutdown window, 30 seconds by default, and can be restarted one at a time. |
| You need two main instances for high availability | Queue mode with an Enterprise license | Multi-main is not part of the Community Edition. |
Queue mode also costs something. The docs note that for webhooks the request is received by the main process and the execution is passed to a worker, "which can add some overhead and latency". You add Redis, at least two more containers and a second place for things to go wrong. If a concurrency limit solves the problem, stop there.
How many n8n queue mode workers do you need?
Enough to cover your peak concurrent executions at the concurrency you set, plus one spare so a restart or a crash does not drop you below the peak. Four numbers go into the plan:
- Concurrency per worker. The default is 10 jobs in parallel. n8n recommends 5 or higher, because "low concurrency values with a large numbers of workers can exhaust your database's connection pool".
- Memory per worker. n8n publishes no per-worker figure, and it cannot: memory depends on the size of your JSON and binary data, the number of nodes and how much runs at once. Run a realistic load and read
docker statsfor the worker containers. That measured number is the one that goes in the RAM line. - Postgres connections. Each n8n process keeps a pool, 2 connections by default (
DB_POSTGRESDB_POOL_SIZE). Keep the total well under your database's connection limit. - Redis memory. Small for job IDs. The exception is large webhook responses: the docs say to budget about 1.5 times
N8N_WEBHOOK_RESPONSE_RELAY_SIZE_MAX, 64 MiB by default, for each response in flight.
Our sizing rule, read from docker stats under peak load. It is not an n8n recommendation.
Prerequisites
- A working single-instance n8n on Docker Compose, and a fresh backup of its database and .n8n folder
- PostgreSQL 16, 17 or 18 as the database: queue mode over SQLite is not supported
- The encryption key your instance already uses, copied from its config or environment
- Enough memory for Redis, two workers and three task runner containers on top of what you run today
- No workflows that rely on files written to the n8n container's local disk
- The same n8n version on the main instance, every worker and every runner image
- A quiet hour: switching modes restarts the instance
If you are still on SQLite, move to Postgres first. Our Docker Compose guide covers that stack, and the self-hosted cost breakdown helps with sizing the server.
Set up n8n queue mode with Docker Compose: the steps
- 1Back up
Dump Postgres and copy the .n8n folder. Expected result: a restore you have tested, or at least files off the server.
- 2Write .env
Pin the version and reuse the existing encryption key. Expected result: six values in a file only you can read.
- 3Add the queue services
One shared block for main and workers, so they cannot drift apart. Expected result: eight services defined.
- 4Start the stack
docker compose up -d. Expected result: postgres and redis healthy, the other six running.
- 5Check each worker
Read the worker log and call its health endpoint. Expected result: "n8n worker is now ready" and a status of ok.
- 6Prove a run used a worker
Trigger a published workflow. Expected result: the main log says Enqueued execution and a worker log says Worker started execution.
- 7Scale
Add a worker block with its runner, or raise --concurrency, then run up -d again.
Step 1: back up, then write the .env file
Take the backup before anything else. Then put the settings every container shares in one place. Reuse the encryption key and the database credentials your instance already has: a new key makes every saved credential unreadable.
# .env (chmod 600 .env, and never commit it)
N8N_VERSION=2.42.6
POSTGRES_USER=n8n
POSTGRES_PASSWORD=paste-secret-1
POSTGRES_DB=n8n
ENCRYPTION_KEY=the-key-your-instance-already-uses
RUNNERS_AUTH_TOKEN=paste-secret-2Step 2: add the queue services to compose.yaml
This is n8n's official withPostgresAndWorker example with a second worker, an explicit concurrency, database binary storage and one database user to keep it short. The x-shared block is the point of the design: the main instance and every worker inherit the same settings, so a mismatched key or Redis host cannot creep in.
Starting fresh, the file works as it is. Converting an existing stack, keep your own postgres service, volume names and database credentials, and take from this file the x-shared settings, redis, the workers and the runners. A compose file that points at differently named volumes starts n8n on an empty database, which looks exactly like data loss until you switch the names back.
volumes:
db_storage:
n8n_storage:
redis_storage:
x-shared: &shared
restart: always
image: docker.n8n.io/n8nio/n8n:${N8N_VERSION}
environment:
- DB_TYPE=postgresdb
- DB_POSTGRESDB_HOST=postgres
- DB_POSTGRESDB_PORT=5432
- DB_POSTGRESDB_DATABASE=${POSTGRES_DB}
- DB_POSTGRESDB_USER=${POSTGRES_USER}
- DB_POSTGRESDB_PASSWORD=${POSTGRES_PASSWORD}
- EXECUTIONS_MODE=queue
- QUEUE_BULL_REDIS_HOST=redis
- QUEUE_HEALTH_CHECK_ACTIVE=true
- N8N_ENCRYPTION_KEY=${ENCRYPTION_KEY}
- N8N_DEFAULT_BINARY_DATA_MODE=database
- OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS=true
- N8N_RUNNERS_MODE=external
- N8N_RUNNERS_AUTH_TOKEN=${RUNNERS_AUTH_TOKEN}
- N8N_RUNNERS_BROKER_LISTEN_ADDRESS=0.0.0.0
volumes:
- n8n_storage:/home/node/.n8n
depends_on:
redis:
condition: service_healthy
postgres:
condition: service_healthy
x-runner: &runner
restart: always
image: n8nio/runners:${N8N_VERSION}
services:
postgres:
image: postgres:18
restart: always
environment:
- POSTGRES_USER
- POSTGRES_PASSWORD
- POSTGRES_DB
- PGDATA=/var/lib/postgresql/data
volumes:
- db_storage:/var/lib/postgresql/data
healthcheck:
test: ['CMD-SHELL', 'pg_isready -h localhost -U ${POSTGRES_USER} -d ${POSTGRES_DB}']
interval: 5s
timeout: 5s
retries: 10
redis:
image: redis:7-alpine
restart: always
volumes:
- redis_storage:/data
healthcheck:
test: ['CMD', 'redis-cli', 'ping']
interval: 5s
timeout: 5s
retries: 10
n8n:
<<: *shared
ports:
- 5678:5678
n8n-runner:
<<: *runner
environment:
- N8N_RUNNERS_AUTH_TOKEN=${RUNNERS_AUTH_TOKEN}
- N8N_RUNNERS_TASK_BROKER_URI=http://n8n:5679
depends_on:
- n8n
n8n-worker-1:
<<: *shared
command: worker --concurrency=10
depends_on:
- n8n
n8n-worker-1-runner:
<<: *runner
environment:
- N8N_RUNNERS_AUTH_TOKEN=${RUNNERS_AUTH_TOKEN}
- N8N_RUNNERS_TASK_BROKER_URI=http://n8n-worker-1:5679
depends_on:
- n8n-worker-1
n8n-worker-2:
<<: *shared
command: worker --concurrency=10
depends_on:
- n8n
n8n-worker-2-runner:
<<: *runner
environment:
- N8N_RUNNERS_AUTH_TOKEN=${RUNNERS_AUTH_TOKEN}
- N8N_RUNNERS_TASK_BROKER_URI=http://n8n-worker-2:5679
depends_on:
- n8n-worker-2Step 3: start it and check each worker
docker compose up -d
docker compose ps
docker compose logs n8n-worker-1 | tail -20
docker compose exec n8n-worker-1 wget -qO- http://localhost:5678/healthz/readinessExpected result: the worker log ends with n8n worker is now ready, followed by the version and Concurrency: 10. The readiness call prints {"status":"ok"} once the worker's database and Redis connections are up. Repeat for the second worker.
Step 4: prove a production run went through a worker
Publish a small workflow with a Schedule Trigger set to every minute and a Code node, wait for it to fire, then compare docker compose logs n8n with the worker logs. The main instance logs Enqueued execution with an ID, and one worker logs Worker started execution and Worker finished execution for the same ID. The Code node matters: it only succeeds if that worker's task runner is connected. Because the file sets OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS=true, test runs from the editor take the same route.
Step 5: scale
To add capacity, copy the n8n-worker-2 and n8n-worker-2-runner blocks, rename them to 3, point the new runner at http://n8n-worker-3:5679 and run docker compose up -d. To change concurrency, edit the command line. Workers on a second server work the same way, as long as they can reach the same Postgres and Redis and carry the same key.
The queue mode settings that matter
| Setting | Where | What it does |
|---|---|---|
| EXECUTIONS_MODE=queue | Main and every worker | Switches n8n from running executions itself to queueing them |
| QUEUE_BULL_REDIS_HOST | Main and every worker | Where Redis is. Port, password, TLS and cluster nodes have their own QUEUE_BULL_REDIS_ variables |
| N8N_ENCRYPTION_KEY | Main and every worker, identical | Workers decrypt saved credentials with it |
| N8N_DEFAULT_BINARY_DATA_MODE=database | Main and every worker | Filesystem storage is not supported in queue mode. S3 is the paid alternative |
| worker --concurrency=10 | Each worker's command | Jobs one worker runs in parallel. The default is 10 and n8n recommends 5 or more |
| N8N_CONCURRENCY_PRODUCTION_LIMIT | Workers, optional | When set to anything other than -1, it overrides the --concurrency flag |
| QUEUE_HEALTH_CHECK_ACTIVE=true | Workers | Turns on /healthz and /healthz/readiness on each worker, port 5678 by default |
| N8N_GRACEFUL_SHUTDOWN_TIMEOUT | Workers | Seconds a worker gets to finish its jobs on shutdown. The default is 30 |
| OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS=true | Main, n8n 2.x only | Sends editor test runs to workers. Removed in n8n 3.0, where that is always the behaviour |
The concurrency override is easy to miss. In worker.ts, n8n takes the limit from N8N_CONCURRENCY_PRODUCTION_LIMIT when that is set to anything other than -1, and only then falls back to the flag. If you set that variable for a single instance last year, it now quietly sets your worker concurrency.
Scale n8n self-hosted further: webhooks, health checks and metrics
Workers are the first layer. When the main instance itself becomes the bottleneck for incoming requests, n8n has three more tools:
- Webhook processors. Extra n8n processes started with the
webhookcommand that only receive production webhooks. Put a load balancer in front, send/webhook/*and/webhook-waiting/*to them and everything else, including/webhook-test/*, to the main instance. The docs advise against adding the main process to that pool. - Health checks. With
QUEUE_HEALTH_CHECK_ACTIVE=trueevery worker serves/healthzand/healthz/readiness, which is what a container orchestrator needs to restart a stuck worker. - Metrics.
N8N_METRICS=trueexposes/metricsfor Prometheus, andN8N_METRICS_INCLUDE_QUEUE_METRICS=trueadds job counts for queue mode. A queue that keeps growing means you are short of workers.
Workflows that stop without an error are a separate problem from capacity; our checklist of five ways n8n workflows fail silently covers those.
What changes for queue mode in n8n 3.0
Three items on n8n's 3.0 breaking-changes page touch this setup. Manual executions always run on workers and OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS is removed, so delete that line when you upgrade. The in-memory binary data mode is removed. And the internal task runner mode is deprecated, which the external runners in this file already avoid. That page says 3.0 "has been released", but on 11 October 2026 the latest stable release on GitHub was 2.42.6, so treat 3.0 as not generally available yet. Our n8n 3.0 upgrade guide tracks the rest.
Troubleshooting
| What you see | Cause | Fix |
|---|---|---|
| Worker log: "Concurrency is set to less than 5. THIS CAN LEAD TO AN UNSTABLE ENVIRONMENT." | Concurrency is below n8n's recommended floor | Raise --concurrency to 5 or more and run fewer workers |
| "Credentials could not be decrypted" | A worker has a different N8N_ENCRYPTION_KEY from the main instance | Set the same key on every instance and restart |
| Code node fails with "Task request timed out" | That worker has no reachable task runner | Give each worker its own runner sidecar and point the runner at that worker's port 5679 |
| Files are missing in later nodes | Binary data is kept in memory or on one container's disk | Set N8N_DEFAULT_BINARY_DATA_MODE=database on every instance |
| "The response is too large to be sent back from the worker" | A webhook response exceeded N8N_WEBHOOK_RESPONSE_RELAY_SIZE_MAX, 64 MiB by default | Return less, raise the limit, or from n8n 2.34.0 turn on N8N_WEBHOOK_RESPONSE_RELAY_OFFLOAD_ENABLED |
| Executions are created but never start | No worker is connected to the same Redis | Check the worker log for "n8n worker is now ready" and compare the QUEUE_BULL_REDIS_ settings |
Running infrastructure like this is its own skill, and it carries over. Our AI SaaS Builder program has no n8n lessons, but it does cover the same questions for an app you build yourself: deployment, caching and performance, and infrastructure for scale.
n8n queue mode: FAQ
What is queue mode in n8n?
Queue mode splits n8n into a main instance and worker processes. The main instance handles the editor, timers and incoming webhooks and creates each execution without running it. It passes the execution ID to Redis, a worker picks it up, runs the workflow and writes the result to the database. You scale by adding or removing workers. It is set with EXECUTIONS_MODE=queue.
When do I need n8n queue mode?
When your peak number of simultaneous executions is more than one n8n process can run without the editor and webhooks slowing down, or when a restart must not interrupt running work. Estimate the peak as executions per second multiplied by average run time. Below that point, setting N8N_CONCURRENCY_PRODUCTION_LIMIT on a single instance is simpler: extra runs wait in a queue.
How many workers do I need for n8n queue mode?
Divide your peak concurrent executions by the worker concurrency, round up and add one spare. A worker runs 10 jobs in parallel by default, so 24 concurrent executions need three workers plus a spare. n8n recommends a concurrency of 5 or higher, because many workers with low concurrency can exhaust the database connection pool. Measure memory per worker under load before adding more.
Is n8n queue mode free?
On self-hosted n8n, yes. The n8n edition comparison says queue mode is included in the free Community Edition. Multi-main mode, which runs more than one main instance for high availability, and the Workers view in Settings need an Enterprise license. On n8n Cloud, queue mode is available on Enterprise plans only. Checked October 2026 against the n8n docs.
Does n8n queue mode need Redis and Postgres?
Yes. Redis is the message broker that holds the queue of pending executions, and the database stores workflows, credentials and results, so the main instance and every worker need access to both. The n8n docs say running a distributed setup over SQLite is not supported. Use PostgreSQL 16, 17 or 18, and give every instance the same N8N_ENCRYPTION_KEY.
What is the default worker concurrency in n8n?
Ten. A worker started with n8n worker runs up to 10 jobs in parallel, and the --concurrency flag changes that, for example n8n worker --concurrency=5. If N8N_CONCURRENCY_PRODUCTION_LIMIT is set to anything other than -1, it overrides the flag. Below 5, n8n 2.42.6 logs a warning that the setup can become unstable. Checked October 2026.
Does queue mode change in n8n 3.0?
In one way. The 3.0 breaking-changes page says manual executions always run on workers in queue mode and removes OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS, so workers need memory for editor test runs too. On n8n 2.x you get the same behaviour by setting that variable to true. On 11 October 2026 the latest stable release on GitHub was still 2.42.6.
Your automations can take the load. Now build the product that sends it.
AI SaaS Builder, included in All Access, covers what sits around a stack like this: a Next.js and Supabase app, Claude API features, Claude Code and MCP, deployment and scaling on Vercel, and Stripe billing. It has no n8n lessons. All Access adds the other three programs, live coaching and the private community.
Sizing a queue mode setup?
Share your peak numbers and compose file in the free Discord and compare notes with other people running n8n on their own servers.