Retries and failures

Timeouts, backoff, what counts as permanent, and where a delivery goes when it runs out of attempts.

Return 2xx fast, work afterwards

Timeout

The platform abandons the request if your endpoint has not responded in time. The per-endpoint request timeout defaults to 30 seconds and is not settable through the API today. A timeout is recorded as a failed attempt and retried.

Treat 30 seconds as a ceiling, not a budget. An endpoint that regularly takes seconds to answer is an endpoint that will start timing out under load, and every timeout costs it an attempt from its retry budget.

Success and failure

Your responseTreated as
200–299Success. No further attempts.
408, 425, 429Failure, retried. These are about timing, not about the request being wrong.
Any other 4xx (e.g. 400, 401, 403, 404, 410)Failure, permanent. Not retried.
Any 5xxFailure, retried.
No response — timeout, DNS failure, TLS failure, connection resetFailure, retried.

Permanent versus retried

The distinction matters because a permanent failure stops immediately rather than spending the whole retry budget rediscovering the same answer.

Retried — the condition may clear on its own
  • 5xx from your endpoint.
  • 408 Request Timeout, 425 Too Early, 429 Too Many Requests.
  • No response at all: request timeout, DNS failure, TLS handshake failure, connection reset.

These are the server's or the network's problem, and the next attempt may well find it fixed.

Permanent — repeating the request unchanged cannot help
  • Any 4xx other than 408, 425 and 429. A 400 says the body is wrong for you; a 401/403 says you will not accept it; a 404/410 says the endpoint is gone. An endpoint being torn down is a normal lifecycle event, not an outage, and the platform treats it as such.
  • The endpoint URL no longer passes the address rules — it has stopped being https, or its hostname now resolves to a private, loopback, link-local or reserved address. This is checked on every delivery, so a hostname that was public when you registered it and points inward today is refused at send time.
  • No usable signing secret for the webhook. Signing is fail-closed: the delivery is refused rather than sent unsigned. Rotating the secret resolves it.
  • The webhook has been deleted.

A permanent failure is dead-lettered on the spot.

Skipped — not a failure at all

A disabled endpoint (enabled: false) has its deliveries recorded as skipped. They are not retried, they do not count as failures and they do not dead-letter. Disabling is a deliberate operator state, so it does not generate failures for anyone to chase.

Events that occur while an endpoint is disabled are not held and replayed when you re-enable it. Disabling an endpoint loses those events.

Retry schedule

Each endpoint has its own delivery queue and its own retry counter, so a failing endpoint burns its own budget and does not slow down any other subscriber.

12 attempts before dead-letter, counting the first delivery. After the twelfth failed attempt the delivery is dead-lettered.

Exponential backoff. The delay after the nth failure is min(30s × 2^(n−1), 1 hour), plus up to 10% random jitter so that many endpoints failing together do not retry in lockstep.

After failed attemptNext attempt in roughly
130 seconds
21 minute
32 minutes
44 minutes
58 minutes
616 minutes
732 minutes
8 and later1 hour (the cap)

A delivery that fails every attempt therefore spans roughly five hours from first attempt to dead-letter. An endpoint that is down for a short deploy will almost certainly catch up; one that is down for a working day will not.

Dead-lettering

When a delivery is dead-lettered — twelve failed attempts, or one permanent failure — it stops. The platform makes no further attempts for that event.

Idempotency

Delivery is at least once. The same event can arrive more than once — most obviously when your 2xx was sent but never reached the platform, which from the platform's side is indistinguishable from your endpoint failing.

  1. Deduplicate on executionId

    executionId is stable across every retry of the same event. Record it, and make a second delivery with the same value a no-op.

  2. Make the record before the side effect

    Insert executionId under a unique constraint, then act. If the insert conflicts, you have already handled this event — return 200 and do nothing.

  3. Watch data.attempt

    An attempt above 1 tells you the platform has tried before. It is a signal that your previous answer did not land, not a reason to skip your own deduplication check.

Diagnosing a failing endpoint

  1. Read the delivery log

    GET /v3/webhooks/{uuid}/executions?limit=50 returns the most recent attempts with the response status, the duration and the error message the platform recorded. Start here — it tells you whether the platform ever reached you.

  2. Check the endpoint's counters

    GET /v3/webhooks/{uuid} carries totalAttempts, totalSuccesses, totalFailures, lastSuccessAt and lastFailureAt. A lastFailureAt that keeps moving with no lastSuccessAt means nothing has ever got through.

  3. Fire a test event

    POST /v3/webhooks/{uuid}/test goes through the real path — same signing, same address checks — as a single attempt with no retry, and answers with the status, duration and error inline.

  4. Check the usual suspects

    • 400 on everything: almost always signature verification against a re-serialised body. See Verifying signatures.
    • statusCode: null with a timeout error: your endpoint is doing work before answering.
    • Permanent refusal with no HTTP status: the URL has stopped resolving to a public address, or the signing secret cannot be resolved.