All articles
Backend 7 min readJun 2026

Self-healing webhooks for real-time voice AI

At 50,000 callbacks a day, idempotency and retry logic aren't nice-to-haves — they're the product. Here's the pattern.

A voice-AI coaching product I built the backend for processes 50,000+ webhook callbacks a day from a real-time voice provider — call started, transcript ready, analysis complete. At that volume, the webhook layer isn't plumbing behind the product; it is the product. If callbacks are dropped, duplicated, or processed out of order, the user gets broken feedback. Here's how I made it self-healing.

Assume every webhook arrives at least twice

Providers retry on any non-2xx, network blips duplicate deliveries, and your own retries add more. So idempotency isn't optional. Give every event a stable id, record processed ids, and make handlers safe to run twice. The cheapest reliable version is a unique constraint that turns a duplicate into a no-op.

ts
async function handleEvent(evt: WebhookEvent) {
  // Idempotency guard: insert-or-ignore on the event id
  const first = await db.processedEvents
    .insert({ id: evt.id })
    .onConflictDoNothing();

  if (!first.inserted) return; // already handled — safe no-op
  await process(evt);
}

Acknowledge fast, process async

Do the least possible work in the request handler: verify the signature, persist the raw event, return 200. Then let a worker do the real processing off a queue. This keeps you under the provider's timeout (which itself triggers retries), and it decouples spikes from your processing capacity.

At real-time voice scale, idempotency and retry logic aren't nice-to-haves — they're the product.

Retry with backoff, then dead-letter

Transient failures — a downstream timeout, a rate limit — should retry with exponential backoff and jitter. But retries can't be infinite. After a capped number of attempts, move the event to a dead-letter queue where it's visible and replayable, instead of silently lost or blocking the pipeline forever.

  • Exponential backoff with jitter to avoid thundering-herd retries.
  • A max-attempts cap, then dead-letter for human/automated replay.
  • Alert on dead-letter growth — it's your early warning system.

Handle out-of-order delivery

Webhooks are not ordered. "Analysis complete" can beat "call started." Design handlers to tolerate it: key state by the call id, upsert rather than assume a prior insert, and use event timestamps to ignore stale updates. Never assume the event you need has already arrived.

Verify signatures — always

A public webhook endpoint is an open door. Verify the provider's signature on every request and reject anything that doesn't match, before you persist or process. Reliability and security are the same discipline here.

The takeaway

A self-healing webhook layer comes down to a few boring, non-negotiable habits: idempotency by default, fast acknowledge with async processing, bounded retries into a dead-letter queue, order-independent handlers, and signature verification. Get those right and 50,000 events a day become a system you can sleep through.

Building something with AI?

I'm Muhammad Usman Haider, a full stack ai engineer. If you want a system like the ones in this article, let's talk.

Start a project