Webhook Design for SaaS: Reliable Delivery, Retries and Security

Every SaaS product eventually gets the same request from a customer: "Can you tell our system when something changes?" Polling your API every minute is wasteful for them and expensive for you. Webhooks solve this, but a webhook system that silently drops events or leaks data does more damage than having none at all. This guide covers how to design webhooks that customers actually trust.

The advice below comes from the kind of integration work we do while building SaaS platforms, and it applies whether you are a two person startup or a growing product team. Where we use numbers, they are illustrative unless we name a source.

Why Webhooks Are Harder Than They Look

A webhook is just an HTTP POST. That simplicity is misleading, because the receiver is a server you do not control. It may be down for maintenance, slow, behind a firewall, or badly written. Your job is to deliver reliably to an unreliable destination, and to do it without letting one slow customer degrade everyone else.

The core problems fall into four groups:

Ignore any one of these and your support inbox will tell you within a month.

Start With a Clear Event Contract

Before writing delivery code, define what an event looks like. A stable contract is the difference between a webhook customers build on and one they fear.

Envelope and payload

Wrap every event in a consistent envelope with a unique event ID, an event type such as invoice.paid, a creation timestamp, an API version and a data object. Consumers will write code against this shape, so treat changes like public API changes.

Thin versus fat payloads

A fat payload contains the full resource. A thin payload contains only the ID and type, and the consumer fetches the rest through your API. Fat payloads are convenient but can leak stale or sensitive data. Thin payloads are safer and always current, at the cost of an extra API call. Many teams send a moderate payload with the fields most consumers need and the resource ID for anything else.

Versioning

Pin each endpoint to an API version at creation time. When you change a field, the customer upgrades on their schedule instead of breaking on yours.

Real-World Example: A Billing Events Integration

Consider a subscription billing product with a customer whose accounting system needs to know when invoices are paid. In a typical illustrative scenario, the first version simply posts to a URL inside the request that marks the invoice paid. Then the customer's server takes eight seconds to respond and the payment request times out for the end user.

The fix follows a pattern we recommend across builds: write the event to a database table in the same transaction as the business change, then let a separate worker deliver it. The payment flow stays fast, and delivery problems never touch the user. This is the same outbox idea that underpins the approach in our guide to event-driven architecture and queues for startups.

The consumer, meanwhile, receives an event ID with each delivery. If a retry brings the same ID twice, their code notices it has already processed that ID and returns success without acting again. Neither side loses money or trust.

Step-by-Step: Building a Reliable Webhook System

  1. Persist events first. Save each event to an outbox table inside the same database transaction as the change that caused it. This guarantees you never emit an event for a rolled back change, or lose one for a committed change.
  2. Deliver from a worker queue. A background worker picks up pending events and sends them. Give each endpoint its own concurrency limit so one slow customer cannot starve the rest.
  3. Set strict timeouts. A 5 to 10 second timeout is a common choice. Treat any 2xx as success and everything else, including timeouts, as a failure to retry.
  4. Retry with exponential backoff and jitter. For example, retry after 1 minute, 5 minutes, 30 minutes, 2 hours and then every few hours for up to three days. Jitter prevents a thundering herd when a customer's server comes back online.
  5. Sign every request. Compute an HMAC of the timestamp and raw body with a per-endpoint secret and send it in a header. Document exactly how to verify it, with code samples in the languages your customers use.
  6. Disable failing endpoints gracefully. After a long run of failures, pause the endpoint, email the account owner, and keep the undelivered events so they can be replayed later.
  7. Ship a delivery log and replay button. Show each attempt, response code and response body snippet. Let customers resend any event with one click.
  8. Add a test event tool. Let developers fire a sample event at their endpoint during setup, before they go live.

Security Details That Matter

Webhook security is mostly about proving authenticity and limiting damage.

Signatures and replay protection

Include the timestamp inside the signed content and tell consumers to reject any request older than a few minutes. Without this, an attacker who captures one valid request can resend it forever.

Protecting your own network

Because customers supply the URL, your delivery worker can be tricked into calling internal addresses. This is a server side request forgery risk. Block private IP ranges, resolve DNS yourself and check the result before connecting, and disallow redirects to internal hosts. Send webhooks from a fixed set of IPs and publish that list so customers can allow it through firewalls.

Secret rotation

Let customers rotate secrets without downtime by accepting two active secrets for a short overlap window. If you are also tightening broader access controls, our piece on zero trust security architecture for startups covers the wider mindset.

Handling Duplicates and Ordering

Webhooks are delivered at least once. That is a promise you should state in your docs, because exactly once delivery across a network is not something any provider can honestly guarantee.

Give consumers what they need to cope:

If your events for a single resource must arrive in order, partition your queue by resource ID and deliver those serially. Just be honest in the docs about which event types carry ordering guarantees.

Observability for Your Own Team

Customers are not the only ones who need visibility. Track delivery success rate, median and 95th percentile delivery latency, retry counts, and the number of endpoints currently failing. Alert when success rate drops below your target, and review the slowest endpoints monthly. Treat a failing endpoint as a customer success signal: the account is probably having an integration problem you can help fix.

Key Benefits of Doing Webhooks Well

Common Mistakes to Avoid

A webhook is a promise. The payload is the easy part, and the promise about what happens when things go wrong is the product.

Build or Buy?

A minimal in-house version, with an outbox table, a worker and signature headers, can often be built in a couple of weeks by a small team. That estimate is illustrative and depends on your stack. The long tail is where the effort hides: rate controls, dashboards, replay, alerting and documentation. If webhooks are a side feature, a managed delivery service may be cheaper than the engineering time. If they are central to your platform, owning the system pays off. Our SaaS development team helps founders make this call based on their product and volume, and we also build the customer facing dashboards that go with it.

Conclusion

Good webhooks feel boring, and that is the goal. Persist events transactionally, deliver from a queue, sign everything, retry with backoff, and give customers a log and a replay button. Do those five things and you avoid most of the pain that teams discover the hard way. Start with a clear event contract today, because it is the one part that is expensive to change later.

Frequently Asked Questions

What is a webhook in a SaaS product?
A webhook is an HTTP callback your product sends to a customer's server when something happens, such as a payment succeeding or a record being updated. It lets customers react to events without polling your API repeatedly.
How should webhook payloads be secured?
Sign every payload with an HMAC using a per-endpoint secret, include a timestamp in the signed content, and require HTTPS. Consumers verify the signature and reject old timestamps to block replay attacks.
How many times should a failed webhook be retried?
Most teams retry with exponential backoff and jitter over several hours to a few days. A common pattern is immediate retries for the first few minutes, then longer gaps, before marking the endpoint as failing and notifying the customer.
Why do webhook consumers need to be idempotent?
Delivery is at least once, so the same event can arrive twice after a timeout or retry. Storing each event ID and ignoring repeats keeps customer systems from double charging or double creating records.
Should we build webhooks ourselves or use a managed service?
Building a basic version is quick, but retries, logs, replay and rate controls take real engineering time. Early stage teams often start with a small in-house queue and move to a managed delivery service once volume or customer demands grow.