Every web application eventually hits the same wall. A user clicks a button, and the page hangs for eight seconds while the server generates a PDF, resizes images, calls three third party APIs and sends a confirmation email. Some requests time out. Others succeed but leave the user wondering whether to click again. The fix is almost always the same: stop doing slow work inside the request. Accept the request, put the work on a queue, and let background workers handle it.
This guide explains how background jobs and queues work, how to choose between popular tools, and how to make them reliable. It is aimed at startup and SME teams that want a dependable foundation without building a complex platform.
A useful rule is that anything the user does not need to see the result of immediately should not block the response. Typical candidates include:
Keep inside the request only what the user must see to proceed, such as validating input and writing the core record. Everything else can follow asynchronously, with the interface showing a status like "processing" and updating when the work completes.
The pattern has three parts. A producer, usually your web server, creates a job with a payload and places it on a queue. A broker stores the job safely. A worker pulls jobs off the queue, runs the handler and marks the job done or failed.
Good systems add more behaviour on top: delayed jobs that run in the future, priorities so urgent work goes first, concurrency limits so one tenant cannot flood the system, rate limits to respect external APIs, and scheduled recurring jobs that replace fragile cron scripts.
BullMQ is popular in Node.js projects because it is simple, fast and feature rich. It supports delays, priorities, retries with backoff, rate limiting and repeatable jobs. It needs Redis, so you must run and monitor Redis with persistence configured correctly. It is an excellent choice when your team already uses Redis and wants to ship quickly. Similar libraries exist for Python, Ruby and other languages.
SQS is a managed queue with very little operational overhead. You do not patch servers or worry about memory limits. It offers visibility timeouts, dead letter queues and strong durability. The tradeoff is fewer built in features, so you write more of the logic for scheduling, rate limits and per tenant fairness yourself. It suits teams already on AWS who value reliability over convenience. Google Cloud Tasks and Azure Service Bus fill similar roles.
Temporal and similar engines solve a different problem: long running, multi step processes that must survive crashes. Think of an onboarding flow that sends an email, waits three days, checks whether the user activated, then branches. The engine records each step so a failed worker resumes where it stopped. It has a learning curve and operational cost, so reserve it for workflows where recovery logic would otherwise become tangled code.
If you need to send emails and process uploads, start with BullMQ or SQS. If you have multi step business processes with waits and compensations, such as order fulfilment or payment retries, evaluate a workflow engine. Resist adopting the heaviest tool for the simplest problem. You can migrate later if your handlers are written cleanly.
Consider an illustrative scenario. An online store handles checkout in a single request. After payment, the server updates inventory, creates an invoice PDF, emails the customer, pushes the order to a courier API and posts to the accounting system. During a festive sale, the courier API slows down, checkout requests start timing out, and customers are charged without seeing a confirmation. Some retry the payment.
The team restructures the flow. The request now records the payment and the order, enqueues five jobs and immediately shows the confirmation page. The courier job retries with exponential backoff when the API is slow. If it still fails, it lands in a dead letter queue and an alert tells operations to follow up manually. Customers no longer wait on third party systems, and a slow partner stops being a checkout outage.
The outbox pattern. A classic bug appears when you save a record and then enqueue a job: if the process crashes between the two steps, the record exists but the job never runs. The outbox pattern writes the job into a database table in the same transaction as the record, and a small relay publishes it to the queue. This guarantees that either both happen or neither does.
Fair scheduling across tenants. In multi tenant products, one customer importing a huge file can starve everyone else. Use separate queues, per tenant concurrency limits or round robin scheduling so that heavy users do not degrade the experience of light ones.
Graceful shutdown and deploys. Workers should finish or safely release their current job when they receive a shutdown signal. Otherwise every deploy creates a burst of duplicate or lost work. If your deployment process is still manual and risky, consider the practices in our guide to zero downtime database migrations, since job payloads and schemas must stay compatible during rollouts.
Versioned payloads. Jobs queued before a deploy may be processed by code released after it. Include a version field and keep handlers backward compatible for at least one release cycle.
The most common mistake is assuming exactly once delivery. Almost every queue offers at least once delivery, which means duplicates will happen. Design for them from the start.
Another is infinite retries. A poison message that always fails can loop forever, consuming workers and hiding real problems. Cap attempts and move failures aside.
A third is ignoring queue depth until customers complain. A backlog of ten thousand jobs does not announce itself. Alert on the age of the oldest job, not just the count.
Finally, teams often forget that jobs touch real data. Log carefully, redact personal information from payloads and make sure job dashboards are protected by proper access control.
If a task can fail, retry, or take longer than a blink, it belongs in a queue, not in a request.
Designing queues, retries and monitoring correctly takes experience, especially when payments, inventory and partner APIs are involved. If you are planning a new product or stabilising an existing one, our team can help. See how our web development services approach reliability, performance and maintainable architecture from day one.
Background jobs are one of the highest leverage upgrades a growing web application can make. They protect user experience, isolate failures and give you room to scale. Start by moving your slowest and least reliable work behind a queue, choose the simplest broker your team can operate, and invest early in idempotency, retries, dead letter queues and monitoring.
Do those few things well and your application will keep working even when partners slow down, traffic spikes or a deploy goes slightly wrong. That quiet reliability is exactly what customers pay for, even though they never see the queue behind it.