Every SaaS product eventually gets the same request from a customer: "Can you tell our system when something changes?" Polling your API every minute is wasteful for them and expensive for you. Webhooks solve this, but a webhook system that silently drops events or leaks data does more damage than having none at all. This guide covers how to design webhooks that customers actually trust.
The advice below comes from the kind of integration work we do while building SaaS platforms, and it applies whether you are a two person startup or a growing product team. Where we use numbers, they are illustrative unless we name a source.
A webhook is just an HTTP POST. That simplicity is misleading, because the receiver is a server you do not control. It may be down for maintenance, slow, behind a firewall, or badly written. Your job is to deliver reliably to an unreliable destination, and to do it without letting one slow customer degrade everyone else.
The core problems fall into four groups:
Ignore any one of these and your support inbox will tell you within a month.
Before writing delivery code, define what an event looks like. A stable contract is the difference between a webhook customers build on and one they fear.
Wrap every event in a consistent envelope with a unique event ID, an event type such as invoice.paid, a creation timestamp, an API version and a data object. Consumers will write code against this shape, so treat changes like public API changes.
A fat payload contains the full resource. A thin payload contains only the ID and type, and the consumer fetches the rest through your API. Fat payloads are convenient but can leak stale or sensitive data. Thin payloads are safer and always current, at the cost of an extra API call. Many teams send a moderate payload with the fields most consumers need and the resource ID for anything else.
Pin each endpoint to an API version at creation time. When you change a field, the customer upgrades on their schedule instead of breaking on yours.
Consider a subscription billing product with a customer whose accounting system needs to know when invoices are paid. In a typical illustrative scenario, the first version simply posts to a URL inside the request that marks the invoice paid. Then the customer's server takes eight seconds to respond and the payment request times out for the end user.
The fix follows a pattern we recommend across builds: write the event to a database table in the same transaction as the business change, then let a separate worker deliver it. The payment flow stays fast, and delivery problems never touch the user. This is the same outbox idea that underpins the approach in our guide to event-driven architecture and queues for startups.
The consumer, meanwhile, receives an event ID with each delivery. If a retry brings the same ID twice, their code notices it has already processed that ID and returns success without acting again. Neither side loses money or trust.
Webhook security is mostly about proving authenticity and limiting damage.
Include the timestamp inside the signed content and tell consumers to reject any request older than a few minutes. Without this, an attacker who captures one valid request can resend it forever.
Because customers supply the URL, your delivery worker can be tricked into calling internal addresses. This is a server side request forgery risk. Block private IP ranges, resolve DNS yourself and check the result before connecting, and disallow redirects to internal hosts. Send webhooks from a fixed set of IPs and publish that list so customers can allow it through firewalls.
Let customers rotate secrets without downtime by accepting two active secrets for a short overlap window. If you are also tightening broader access controls, our piece on zero trust security architecture for startups covers the wider mindset.
Webhooks are delivered at least once. That is a promise you should state in your docs, because exactly once delivery across a network is not something any provider can honestly guarantee.
Give consumers what they need to cope:
If your events for a single resource must arrive in order, partition your queue by resource ID and deliver those serially. Just be honest in the docs about which event types carry ordering guarantees.
Customers are not the only ones who need visibility. Track delivery success rate, median and 95th percentile delivery latency, retry counts, and the number of endpoints currently failing. Alert when success rate drops below your target, and review the slowest endpoints monthly. Treat a failing endpoint as a customer success signal: the account is probably having an integration problem you can help fix.
A webhook is a promise. The payload is the easy part, and the promise about what happens when things go wrong is the product.
A minimal in-house version, with an outbox table, a worker and signature headers, can often be built in a couple of weeks by a small team. That estimate is illustrative and depends on your stack. The long tail is where the effort hides: rate controls, dashboards, replay, alerting and documentation. If webhooks are a side feature, a managed delivery service may be cheaper than the engineering time. If they are central to your platform, owning the system pays off. Our SaaS development team helps founders make this call based on their product and volume, and we also build the customer facing dashboards that go with it.
Good webhooks feel boring, and that is the goal. Persist events transactionally, deliver from a queue, sign everything, retry with backoff, and give customers a log and a replay button. Do those five things and you avoid most of the pain that teams discover the hard way. Start with a clear event contract today, because it is the one part that is expensive to change later.