iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To track BullMQ failures safely, record processor exceptions separately from worker or connection errors, classify retryable and permanent failures, and make application-side effects safe to repeat. BullMQ’s events, stalled-job recovery, and optional PostgreSQL backend each address part of the problem; none automatically makes a Postgres write, an external API call, and a queue acknowledgement one atomic transaction.
What “rollback-safe” can—and cannot—mean
A job can fail before a database transaction commits, after it commits but before the worker acknowledges the job, or while the worker is unable to renew its lock. In the latter cases, a job may run again. That means a successful Postgres write does not, by itself, prove that BullMQ recorded the job as complete.
Keep the transaction boundary explicit in your design. A Postgres transaction can group the application-table changes that are issued through that transaction. It does not automatically include BullMQ queue state when the queue uses Redis, nor does it roll back an external service call. BullMQ’s optional PostgreSQL backend documents transactions for its own queue-state transitions, not a general distributed transaction across queue state, arbitrary application writes, and external effects.
Consequently, “rollback-safe” should mean that your application has defined what commits together and has a deliberate response to retries or duplicate delivery—not that choosing Postgres for one part of the system guarantees end-to-end atomicity.
#1 Best Overall
Separate job failures from worker and infrastructure errors
Use distinct signals for different failure paths. A processor exception belongs to the job’s failure and retry path. A BullMQ error event can indicate a connection problem or another operational issue. A stalled job is a lock-renewal problem and can be recovered separately from ordinary processor retries.
| Failure or signal | What it tells you | What to do |
|---|---|---|
| Processor throws an ordinary error | The job’s processing attempt failed; configured retry behavior may apply. | Log the exception with job context, then allow the error to follow the configured retry policy. |
Processor throws UnrecoverableError |
The failure should not use the normal retry path. | Use it for a permanent failure that should move the job to the failed set without the configured retries. |
Worker or queue emits error |
An operational signal, including possible connection issues; it is not by itself a classification of a particular job’s processor failure. | Attach handlers and forward the error to structured logs or monitoring. |
| Job becomes stalled | The worker did not renew the active lock in time; the job can return to waiting or eventually reach the failed set after its allowed stalls. | Investigate event-loop blocking, worker health, and shutdown behavior. |
BullMQ’s production guidance recommends handling worker and queue error events; an attached handler can also help prevent an unhandled error. These handlers do not replace processor error logging or retention of failed-job state.
Log enough context to investigate without exposing payloads
For processor exceptions, a useful structured record includes the queue name, job name, stable job identifier, attempt information, error class, message and stack, timestamp, and a correlation identifier connecting the job to the request or domain record that initiated it. This is an implementation recommendation, not a prescribed BullMQ logging schema.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #2
Do not log every job payload by default. BullMQ’s production guidance says queue payload data is stored in clear text. Keep sensitive data out of payloads where possible; if a sensitive field must be present, encrypt it before enqueueing. Retain failed-job records according to a policy that leaves enough information for debugging while meeting your data-retention requirements.
Example: route operational events and processor errors separately
This pattern illustrates the separation. Adapt initialization, logger fields, and the installed BullMQ API to your application; it deliberately omits payload contents.
const worker = new Worker(queueName, async (job) => {
try {
return await processJob(job);
} catch (error) {
logger.error({
event: "job_processor_failed",
queue: queueName,
jobName: job.name,
jobId: job.id,
attemptsMade: job.attemptsMade,
correlationId: job.data.correlationId,
err: error,
});
throw error; // Preserve BullMQ's configured failure/retry behavior.
}
});
worker.on("error", (error) => {
logger.error({ event: "bullmq_worker_error", err: error });
});
queue.on("error", (error) => {
logger.error({ event: "bullmq_queue_error", err: error });
});
Only include a correlation value in the example if your payload schema has one; avoid copying arbitrary job data into logs. A processor wrapper that logs and rethrows preserves the failure for BullMQ’s handling rather than turning it into a successful return.
Rank #3
Make retries safe at the application boundary
Before enabling or increasing retries, map what the processor changes and when each change commits. The question is whether rerunning the same logical work after a partial failure could apply a database mutation twice, issue a duplicate external request, or leave related records inconsistent.
- Identify the application Postgres transaction, if any, and list the writes it commits together.
- Identify effects outside that transaction, such as calls to another service; a database rollback cannot undo an already completed external call.
- Decide how the application recognizes repeated work and avoids repeating a side effect. The precise mechanism depends on the domain and architecture; BullMQ’s cited guidance does not prescribe one universal idempotency recipe.
- Test failure points around both the write and job completion, including a worker stop after a database commit. Confirm the outcome against the actual queue backend and application transaction boundary.
When BullMQ uses Redis and the application stores business data in Postgres, treat queue-state recovery and Postgres transaction behavior as separate mechanisms. Do not infer that one database commit controls both. If your worker’s transaction and acknowledgement sequence has not been verified, describe the boundary as unresolved rather than promising rollback safety.
Prevent stalled jobs and shut workers down cleanly
BullMQ locks a job while a worker processes it and expects the worker to renew that lock periodically. CPU-heavy synchronous work can block Node.js’s event loop, interrupt renewal, and cause BullMQ to treat the job as stalled. A stalled job can be returned to waiting and run again, or reach the failed set after its allowed stalls.
Rank #4
BullMQ’s stalled-job documentation emphasizes returning control to the Node.js event loop often enough to avoid this problem. Keep processors from monopolizing the event loop; for CPU-heavy work, use an appropriate process or thread-isolation design and verify the API and behavior against the BullMQ version you deploy.
Graceful shutdown reduces avoidable stalls during deployments and service termination. BullMQ’s worker.close() stops the worker from picking up new jobs and waits for active work to finish or fail. The method has no built-in timeout, so account for job duration and the deployment platform’s termination grace period.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Close the worker when the process receives a termination signal
let closing = false;
async function shutdown() {
if (closing) return;
closing = true;
await worker.close();
await queue.close();
}
process.once("SIGTERM", shutdown);
process.once("SIGINT", shutdown);
Integrate this with the service’s existing cleanup and error handling. A platform may stop the process if its grace period expires before active jobs finish; BullMQ’s stalled-job mechanism exists to recover work after an ungraceful shutdown, but that recovery is another reason your application effects must tolerate a rerun.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choose Redis or the optional PostgreSQL queue backend for the workload
BullMQ’s official documentation describes Redis as the default and most battle-tested backend. Its optional PostgreSQL backend can suit teams that want to avoid operating a separate Redis instance or keep job state alongside relational data. The PostgreSQL backend requires PostgreSQL 13 or newer, with 14 or newer recommended, and the pg package.
The backend choice does not remove the need to benchmark your own workload. BullMQ publishes the following rough illustrative measurements from an Apple Silicon laptop using local PostgreSQL, trivial no-op jobs, and default durable settings. The documentation warns that results depend on hardware, PostgreSQL configuration, and network placement between workers and the database. These are publisher-reported figures from the current documentation page accessed in 2026, not independently verified production promises.
| Operation and conditions | PostgreSQL | Redis |
|---|---|---|
Sequential add() |
Around 7,000 jobs/s | Around 7,500 jobs/s |
Concurrent add() |
Around 15,000 jobs/s | Around 38,000 jobs/s |
Concurrent bulk addBulk() |
Around 45,000 jobs/s | Around 52,000 jobs/s |
| Processing, one worker at concurrency 1 | Around 2,300 jobs/s | Around 6,000 jobs/s |
| Processing, concurrency 8–32 | Around 11,000 jobs/s | Around 18,000 jobs/s |
Use workload-specific tests and capacity planning rather than treating these rates as a guarantee. Consider operational footprint, the backend your team already runs, database version, durability expectations, and the actual mix of enqueueing and processing in your application.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsQuick Recap
A practical failure-handling checklist
- Attach worker and queue
errorhandlers and route their events to operational logging or monitoring. - Log processor exceptions with stable job and correlation context, while excluding sensitive payload data.
- Use ordinary processor errors for failures eligible for configured retries; use
UnrecoverableErrorwhen a failure must bypass those retries and go to the failed set. - Keep failed-job retention sufficient for debugging, subject to your data policy.
- Check for event-loop starvation and test stalled-job recovery under realistic workload.
- Await
worker.close()during shutdown and align deployment grace periods with expected job duration. - Document exactly which application writes share a Postgres transaction and how repeated processing avoids duplicate effects.
- Benchmark the selected queue backend with your own workload before relying on capacity estimates.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

