Background Jobs and Queues in 2026: A Reliability Guide for Web Apps

Background Jobs and Queues in 2026: A Reliability Guide for Web Apps — cover image

Every web application eventually hits the same wall. A user clicks a button, and the page hangs for eight seconds while the server generates a PDF, resizes images, calls three third party APIs and sends a confirmation email. Some requests time out. Others succeed but leave the user wondering whether to click again. The fix is almost always the same: stop doing slow work inside the request. Accept the request, put the work on a queue, and let background workers handle it.

This guide explains how background jobs and queues work, how to choose between popular tools, and how to make them reliable. It is aimed at startup and SME teams that want a dependable foundation without building a complex platform.

What Belongs in a Background Job

A useful rule is that anything the user does not need to see the result of immediately should not block the response. Typical candidates include:

Keep inside the request only what the user must see to proceed, such as validating input and writing the core record. Everything else can follow asynchronously, with the interface showing a status like "processing" and updating when the work completes.

How a Queue System Works

The pattern has three parts. A producer, usually your web server, creates a job with a payload and places it on a queue. A broker stores the job safely. A worker pulls jobs off the queue, runs the handler and marks the job done or failed.

Good systems add more behaviour on top: delayed jobs that run in the future, priorities so urgent work goes first, concurrency limits so one tenant cannot flood the system, rate limits to respect external APIs, and scheduled recurring jobs that replace fragile cron scripts.

Comparing BullMQ, SQS and Temporal

BullMQ and Redis based queues

BullMQ is popular in Node.js projects because it is simple, fast and feature rich. It supports delays, priorities, retries with backoff, rate limiting and repeatable jobs. It needs Redis, so you must run and monitor Redis with persistence configured correctly. It is an excellent choice when your team already uses Redis and wants to ship quickly. Similar libraries exist for Python, Ruby and other languages.

Amazon SQS and managed cloud queues

SQS is a managed queue with very little operational overhead. You do not patch servers or worry about memory limits. It offers visibility timeouts, dead letter queues and strong durability. The tradeoff is fewer built in features, so you write more of the logic for scheduling, rate limits and per tenant fairness yourself. It suits teams already on AWS who value reliability over convenience. Google Cloud Tasks and Azure Service Bus fill similar roles.

Temporal and workflow engines

Temporal and similar engines solve a different problem: long running, multi step processes that must survive crashes. Think of an onboarding flow that sends an email, waits three days, checks whether the user activated, then branches. The engine records each step so a failed worker resumes where it stopped. It has a learning curve and operational cost, so reserve it for workflows where recovery logic would otherwise become tangled code.

A simple way to choose

If you need to send emails and process uploads, start with BullMQ or SQS. If you have multi step business processes with waits and compensations, such as order fulfilment or payment retries, evaluate a workflow engine. Resist adopting the heaviest tool for the simplest problem. You can migrate later if your handlers are written cleanly.

A Real World Example: An Order Confirmation That Kept Timing Out

Consider an illustrative scenario. An online store handles checkout in a single request. After payment, the server updates inventory, creates an invoice PDF, emails the customer, pushes the order to a courier API and posts to the accounting system. During a festive sale, the courier API slows down, checkout requests start timing out, and customers are charged without seeing a confirmation. Some retry the payment.

The team restructures the flow. The request now records the payment and the order, enqueues five jobs and immediately shows the confirmation page. The courier job retries with exponential backoff when the API is slow. If it still fails, it lands in a dead letter queue and an alert tells operations to follow up manually. Customers no longer wait on third party systems, and a slow partner stops being a checkout outage.

Step by Step: Adding Reliable Background Jobs

  1. List the slow or unreliable work. Review your slowest endpoints and every call to an external service. These are your first job candidates.
  2. Pick a broker that fits your stack. Prefer something your team can operate. A managed service beats a self hosted one if nobody wants to be on call for Redis.
  3. Define small, focused job types. One job should do one thing, such as "send receipt email." Smaller jobs retry more cleanly than one giant job.
  4. Pass identifiers, not large payloads. Store an order ID in the job and load fresh data when it runs. This avoids stale data and keeps the queue light.
  5. Make every handler idempotent. Jobs can run twice. Use unique keys for side effects so a repeat execution does not double charge or double send. Our article on idempotency keys for preventing double charges explains the pattern in detail.
  6. Configure retries with backoff and jitter. Retry transient failures several times with growing delays. Do not retry errors that can never succeed, such as validation failures.
  7. Add a dead letter queue. Send exhausted jobs somewhere visible, with the error message and payload, so someone can fix and replay them.
  8. Set timeouts and concurrency limits. A hung job should not block a worker forever, and a spike should not overwhelm your database or a partner API.
  9. Expose job status to users. For exports and imports, store a status record and show progress or completion in the interface.
  10. Monitor and alert. Track queue depth, job age, failure rate and processing time. Alert on growing backlogs and on dead letter queue items.

Reliability Patterns Worth Adopting Early

The outbox pattern. A classic bug appears when you save a record and then enqueue a job: if the process crashes between the two steps, the record exists but the job never runs. The outbox pattern writes the job into a database table in the same transaction as the record, and a small relay publishes it to the queue. This guarantees that either both happen or neither does.

Fair scheduling across tenants. In multi tenant products, one customer importing a huge file can starve everyone else. Use separate queues, per tenant concurrency limits or round robin scheduling so that heavy users do not degrade the experience of light ones.

Graceful shutdown and deploys. Workers should finish or safely release their current job when they receive a shutdown signal. Otherwise every deploy creates a burst of duplicate or lost work. If your deployment process is still manual and risky, consider the practices in our guide to zero downtime database migrations, since job payloads and schemas must stay compatible during rollouts.

Versioned payloads. Jobs queued before a deploy may be processed by code released after it. Include a version field and keep handlers backward compatible for at least one release cycle.

Key Benefits of Moving Work Off the Request Path

Mistakes That Cause Late Night Pages

The most common mistake is assuming exactly once delivery. Almost every queue offers at least once delivery, which means duplicates will happen. Design for them from the start.

Another is infinite retries. A poison message that always fails can loop forever, consuming workers and hiding real problems. Cap attempts and move failures aside.

A third is ignoring queue depth until customers complain. A backlog of ten thousand jobs does not announce itself. Alert on the age of the oldest job, not just the count.

Finally, teams often forget that jobs touch real data. Log carefully, redact personal information from payloads and make sure job dashboards are protected by proper access control.

If a task can fail, retry, or take longer than a blink, it belongs in a queue, not in a request.

Getting Expert Help

Designing queues, retries and monitoring correctly takes experience, especially when payments, inventory and partner APIs are involved. If you are planning a new product or stabilising an existing one, our team can help. See how our web development services approach reliability, performance and maintainable architecture from day one.

Conclusion

Background jobs are one of the highest leverage upgrades a growing web application can make. They protect user experience, isolate failures and give you room to scale. Start by moving your slowest and least reliable work behind a queue, choose the simplest broker your team can operate, and invest early in idempotency, retries, dead letter queues and monitoring.

Do those few things well and your application will keep working even when partners slow down, traffic spikes or a deploy goes slightly wrong. That quiet reliability is exactly what customers pay for, even though they never see the queue behind it.

Frequently Asked Questions

What is a background job?
A background job is work your application performs outside the user's web request, such as sending email, generating a report or processing an upload. The request returns quickly while a separate worker completes the task.
When do we need a queue instead of a cron job?
Use cron for tasks that run on a fixed schedule with no per user trigger. Use a queue when work is triggered by events, needs retries, must scale with load or should be processed in parallel by several workers.
Which queue tool should a small team choose?
If you already run Redis, BullMQ is a quick start for Node.js apps. If you are on AWS and want minimal operations, SQS is dependable. Choose a workflow engine like Temporal when jobs involve many steps, long waits or complex recovery.
What is a dead letter queue?
A dead letter queue holds jobs that failed after all retries. It stops bad jobs from blocking healthy ones and gives your team a place to inspect, fix and replay failures.
How do we avoid running a job twice?
Assume every job can run more than once and make handlers idempotent. Use unique keys for side effects like payments or emails so a repeat execution produces the same result.