Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When Redis cannot accept a BullMQ job, that job has not been queued—and may have no durable record anywhere. In an incident described by Matheus Morett, production work arrived faster than it could be processed; Redis filled, enqueue attempts failed, and negotiation messages were lost. His response was bullmq-outbox: save failed enqueue attempts in a separate store, then replay them through BullMQ when the queue can accept work again. It changes an immediate failure into delayed processing only when the fallback store and recovery path remain available.

What happens when Redis fills up and BullMQ cannot add a job?

BullMQ uses Redis to store queue state. In Morett’s account of an April 2025 production incident, incoming messages outpaced the workload’s processing capacity. Waiting work accumulated, Redis reached capacity, and attempts to enqueue more jobs failed. Because the failed jobs had not made it into the queue, the application had no queue-side record to process later; Morett says negotiation messages were lost. This is his reported experience, not an independently audited incident analysis. Morett’s account of building bullmq-outbox.

The important distinction is between a job that is queued but waiting and an enqueue that Redis rejects. A backed-up queue can still hold work for consumers to process. A rejected enqueue needs another durable record if the application is expected to recover it.

Why Redis Cluster does not split one hot queue across nodes

Redis Cluster distributes keys among 16,384 hash slots, with each node responsible for a subset. That can increase the cluster’s overall capacity and spread different queues across nodes. But Redis requires keys touched together by multi-key commands, transactions, and Lua scripts to share a slot. Hash tags—text inside curly braces in a key—can force keys with the same tag into one slot. Redis Cluster specification.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

BullMQ scripts operate on related keys belonging to a queue. Those keys need to be colocated for the relevant multi-key operations, so one queue’s operations cannot be spread across multiple slots. Adding nodes can help distribute separate queues; it cannot divide a single slot among nodes. That is the sense in which Morett says, “Redis is the ceiling.” It is a limit on the capacity available to a hot queue or slot, not a claim that Redis Cluster provides no scaling benefit.

Before treating a full Redis node as proof that one queue has reached its limit, distinguish among the actual bottlenecks:

  • Slow consumers: jobs arrive faster than workers complete them, so waiting work grows.
  • One hot queue or slot: one queue’s colocated keys are under pressure even if other cluster nodes have room.
  • Retained completed or failed jobs: finalized jobs can continue occupying Redis memory.
  • Cluster-wide pressure: the overall Redis deployment may be short of capacity.

BullMQ’s current guide says completed and failed jobs are retained in dedicated sets by default. It documents count- and age-based automatic removal, which is lazy: removal happens as jobs are processed rather than as a guarantee that all old records vanish at a precise time. Unique job IDs can help with idempotent adds, but an ID stops preventing a later add once its job has been removed. Retention settings may reduce memory used by finalized jobs; they cannot recover an enqueue that Redis rejected before storing it. BullMQ guide to automatic job removal.

How the outbox turns a rejected enqueue into delayed work

Morett’s first implementation wrapped calls to queue.add(). If an add failed, it saved the queue name, job name, payload, options, and a PENDING status in DynamoDB. A cron job ran every 15 minutes, loaded pending rows, and tried to enqueue them through the real BullMQ queue. The failed add’s original error was still rethrown to the caller, leaving the application responsible for deciding what to tell the user. Morett’s implementation and package description.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The later bullmq-outbox package exposes createOutbox({ store }), wrapQueue(queue), and flush(limit). The adopter supplies the storage functions; the described interface includes save, loadPending, markProcessed, and markFailed. Morett says the package ships no storage adapters and provides example implementations to copy for Postgres, Redis, MongoDB, and DynamoDB. These are the author’s descriptions; package maintenance, current API details, and compatibility were not independently verified.

The author describes wrapQueue as a Proxy and the package as structurally typed, without a direct BullMQ dependency, to pass through methods and support BullMQ v5, v6, and Pro. Treat that as a design claim, not a guarantee about compatibility with a current release. The recovery sequence is conceptually simple:

  1. Attempt the normal enqueue. Call the wrapped queue’s add method with the intended job and options.
  2. Persist a rejected add elsewhere. If the add fails, save enough information in the independent store to reconstruct it.
  3. Keep the original error visible. Do not silently tell the caller that work is queued when only a fallback record exists.
  4. Replay pending records. A recovery process loads them and adds them to the actual BullMQ queue once it can accept work.
  5. Record the replay outcome. Mark records processed or failed according to the store’s policy, and retain enough state to handle retries and investigation.

What reliable replay requires

Use a separate failure domain

The fallback store must still accept writes when queue Redis is unavailable or full. The recovery scheduler also needs to run independently of the failure it is meant to repair. In Morett’s original example, the scheduler used a dedicated Redis; if it depended on the same impaired Redis as the queue, it might not run. Separate storage alone is not enough if the process responsible for replaying it shares the outage.

Preserve identity and retry behavior

Morett specifically calls out preserving the job’s jobId, attempts, and backoff settings. Retaining the ID can support idempotent re-enqueue, and restoring the retry options avoids changing the job’s intended failure behavior. But queue-level identity does not establish exactly-once execution of downstream side effects: a worker can perform an external action and fail before the queue records completion. Consumers still need suitable idempotency or deduplication for effects such as sending a payment or updating another system.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make delay an explicit product trade-off

An outbox preserves work; it does not make an unavailable queue realtime. Morett says he excludes real-time conversation queues because waiting as long as the original 15-minute replay interval could be worse than dropping a stale interaction. Decide whether delayed delivery is acceptable for each job type instead of buffering every failure automatically.

Plan for the fallback store’s own failures

If the outbox store cannot save a failed enqueue, the original durability gap remains. The application needs a defined response for that case—such as surfacing an error for retry or escalation—rather than treating an unsaved job as recoverable. Store retention, backup, access controls, and capacity also matter: replay only helps while the pending record still exists and can be read.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Monitor recovery delay, not just replay volume

A replay count tells operators how many jobs moved; it does not tell them how long users waited. Morett identifies ageMs from onJobRequeued as the more useful recovery measure: the age of a job when it is re-enqueued. Alerting on the oldest pending record or replay age makes a stalled scheduler visible even when it processes a small number of jobs. In the author’s example output, a flush reports 12 requeued, 0 failed, and 3 skipped; those are example results, not a performance benchmark or expected rate.

Morett also recommends reserving memory for the Redis service so exhaustion produces a catchable error rather than a stalled connection, and says the package’s integration tests use a real Redis configured near its memory limit. These are the author’s operational advice and test description, not independently reproduced results. Monitoring a BullMQ dashboard can help expose queue state, but recovery-age tracking should include the outbox records and scheduler, not only the live queue.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When this pattern is a good fit

An outbox is most useful when losing a rejected enqueue is unacceptable and processing it later remains valuable. It adds storage, replay logic, and another operational path, so compare it with the actual failure mode rather than treating it as a general queue-capacity fix.

Situation What to address Would an outbox help?
Consumers cannot keep up Worker throughput and backlog growth It can preserve rejected adds, but it does not make consumers faster.
One queue’s slot is hot while the cluster has spare capacity Queue distribution and the limits of colocated keys It can buffer failures; it cannot split that queue’s slot across nodes.
Finalized jobs consume memory BullMQ retention and automatic-removal settings Not as a substitute for retention policy.
Jobs are rejected and remain useful after a delay Durable fallback storage and independent replay Yes, if both are available and replay is safe.
Work is time-sensitive and stale after a delay Product behavior for late or dropped work Possibly not; delayed replay may be worse than an explicit failure.

The central design choice is not whether to retry forever. It is whether a rejected job should be preserved, where it can be stored during Redis pressure, and how the system will safely and visibly bring it back to the queue. Morett’s outbox is one implementation of that choice, with its value depending on an independent store, an independent recovery path, preserved job semantics, and a delay the application can tolerate.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.