Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To control how many n8n workflows run at once, use the setting that matches your deployment: self-hosted regular mode uses N8N_CONCURRENCY_PRODUCTION_LIMIT, queue mode uses each worker’s --concurrency flag, and n8n Cloud concurrency is set by plan. To slow requests sent to an external API, configure batching on the workflow’s HTTP Request node. These controls solve related but different problems: execution concurrency limits parallel work in n8n; request pacing helps manage traffic to a particular service.

Choose the control for your n8n deployment

Option What it controls Where to configure it Important distinction
Self-hosted regular mode Production executions running simultaneously on the instance N8N_CONCURRENCY_PRODUCTION_LIMIT Applies to production executions started by a webhook or trigger node
Queue mode Jobs a worker runs in parallel n8n worker --concurrency=N Setting is per worker; capacity also depends on worker count and database capacity
n8n Cloud Production executions allowed by the plan Plan-defined Check the current quota for your plan
HTTP Request batching Items grouped into requests and the delay between batches from that node HTTP Request node → Batching Paces requests from a node; it does not set instance-wide execution concurrency

Limit production executions in self-hosted regular mode

Regular mode has no production concurrency cap by default. Set N8N_CONCURRENCY_PRODUCTION_LIMIT to a positive integer to cap simultaneous production executions. For example, this shell setting configures a limit of 20:

export N8N_CONCURRENCY_PRODUCTION_LIMIT=20

n8n queues excess production executions FIFO until a slot becomes available. This limit covers production executions started by a webhook or trigger node; it does not cover manual runs, sub-workflow runs, error executions, or CLI-started executions. The environment-variable reference lists -1 as the default, which disables the regular-mode limit. See n8n’s concurrency control documentation and its execution environment variables reference.

Apply the variable through the environment used by your n8n process, then restart or redeploy n8n as appropriate for your process manager or hosting platform. The exact deployment procedure differs by platform. Queued executions cannot be retried while they remain queued; cancelling or deleting one removes it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Set per-worker concurrency in queue mode

Queue mode is a scaling architecture: a main n8n instance places work in Redis, worker processes execute jobs, and a database stores workflow and execution data. Set parallelism for a worker with the --concurrency flag. n8n documents a default of 10 and recommends a value of 5 or higher; its example starts a worker with concurrency set to 5:

n8n worker --concurrency=5

The flag sets jobs handled in parallel by that worker, not a single global limit for the entire deployment. Adding workers increases available processing capacity, but database connection capacity matters too: n8n warns that many workers with low concurrency can exhaust the connection pool, causing delays and failures. Read n8n’s queue mode documentation before adopting this architecture.

Queue mode uses Redis and a database, and it has storage constraints: filesystem binary-data storage is unsupported. n8n advises against SQLite for queue execution mode; a distributed queue setup over SQLite is unsupported. The documentation describes queue mode as offering the best scalability, but that does not remove the need to size and operate its supporting infrastructure.

Check concurrency limits on n8n Cloud

Cloud concurrency is determined by the plan, and executions above the plan’s capacity wait FIFO. Because quotas can vary by plan and may change, check the current n8n pricing page for the number that applies to your account rather than relying on a fixed quota quoted elsewhere. Queue mode is listed as available for n8n Cloud Enterprise by contacting n8n.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pace requests to an external API with HTTP Request batching

An instance concurrency limit is not a way to enforce a third-party API’s exact quota. If one workflow sends too many requests to a service, configure the HTTP Request node’s Batching option. Set Items per Batch to control how many items are grouped together and Batch Interval to set the delay between batches, in milliseconds. See the HTTP Request node documentation.

  1. Open the workflow and select the HTTP Request node that calls the API.

  2. Open the node’s Batching option.

  3. Set Items per Batch and Batch Interval in milliseconds to pace that node’s requests.

  4. Compare the resulting request pattern with the API provider’s documented quota, time window, and error behavior, then adjust the settings as needed.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no single safe batch size or interval for every API: providers define different quotas and windows. Confirm the service’s current rules and inspect its error responses. n8n’s documentation lists a separate Handle rate limits guide, but the specific retry and backoff steps are not established here; do not assume batching alone handles every provider’s rate-limit response.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Which setting should you change?

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.