Skip to main content
Introducing packages.sweber.dev
Documentation menuRetries and failures

Retries and failures

The retry schedule, what counts as a failure, automatic disabling of endpoints, and how to retry or resend.

The schedule

A delivery gets up to eight attempts: one right away and seven retries after

RetryWait
15 seconds
25 minutes
330 minutes
42 hours
55 hours
610 hours
710 hours

That is about 27 hours in total, enough to ride out a long outage on the receiving side. Each wait varies by ±10 % so that many failed deliveries do not all come back at the same moment. If the endpoint answers 429 or 503 with a Retry-After header, Vector waits at least that long (at most 24 hours).

Change the schedule with retrySchedule, a list of seconds; its length plus one is the number of attempts:

createVector({ retrySchedule: [10, 60, 600] }); // four attempts over about 11 minutes

What counts as a failure

Success is any 2xx status. Everything else is a failure:

  • other statuses, including 3xx: Vector does not follow redirects, because a redirect could point into your internal network
  • network errors and TLS errors
  • no answer within timeoutMs (15 seconds by default)
  • a URL that is no longer allowed, for example because its host name now resolves to a private address

Each attempt is stored with status code, error, duration and the first 2 KiB of the response body (maxResponseBytes).

When a delivery gives up

After the last attempt the delivery's status becomes failed and the delivery.failed event fires. The endpoint's failureStreak goes up by one; any successful delivery resets it to zero.

When failureStreak reaches disableEndpointAfter (default 10), Vector disables the endpoint, sets disabledReason and fires endpoint.disabled. An endpoint that answers 410 Gone is disabled at once (disableOnGone: false turns that off). Disabled endpoints receive no new deliveries, and their pending deliveries are cancelled. Re-enable with vector.endpoints.update(id, { enabled: true }), which also resets the streak.

Tell your customers when this happens; that is what the events are for. Vector Pro includes ready-made alerts for Slack, email and webhooks.

Retry and resend

await vector.retry(deliveryId);                      // fresh set of attempts, starting now
await vector.resend(messageId);                      // again to every endpoint it went to
await vector.resend(messageId, { endpointId });      // to one endpoint, also a new one

Both take { tenant } to scope them to a customer. Retried and resent requests carry the same webhook-id as before, so receivers that deduplicate handle them only once. To recover everything that failed during an outage in one go, use recover from Vector Pro.