---
title: "Supplier Timeout Incident Playbook"
description: "Handle travel-search supplier timeouts with containment, retry budgets, fallbacks, circuit breakers and controlled recovery."
slug: "supplier-timeout-playbook"
translationKey: "playbook-supplier-timeout"
locale: "en"
type: "playbook"
category: "operations"
tags: ["supplier","timeout","resilience","incident","playbook"]
publishedAt: "2026-09-26"
updatedAt: "2026-09-26"
reviewedAt: "2026-09-26"
---

The goal during a supplier timeout incident is not to rescue every request. It is to **protect the overall search experience while isolating a degraded upstream dependency**.

## Trigger / symptom

Provider p95/p99 latency rises, timeout rate crosses a threshold, partial results increase, or worker threads and connection pools begin to saturate.

## Objective and scope

This playbook focuses on search/read paths. Booking/create-order writes require a separate retry design because they may be non-idempotent.

## Evidence needed

Collect provider/endpoint, timeout rate, latency percentiles, concurrency, connection-pool usage, request volume, HTTP/network error mix, retry count and provider incident status.

## Ordered response

1. Define blast radius: one endpoint, region or whole provider.
2. Verify timeout budget relative to upstream behavior.
3. Check for retry amplification.
4. Limit concurrency with bulkheads/pools.
5. Open a circuit breaker when failure exceeds the defined threshold.
6. Select fallback: cached/indicative data, alternate provider or partial response.
7. Preserve the UI contract; one supplier should not fail the entire search.
8. Probe recovery with small half-open traffic.

## Retry decision

Retry only idempotent requests when sufficient latency budget remains. Use exponential backoff with jitter and avoid synchronized retries across instances.

## Stop conditions

Timeout rate returns near baseline, the circuit closes safely, and queues/pools recover.

## Escalation criteria

Escalate with evidence when multiple regions/endpoints fail, SLA breach persists or a booking-critical path is affected.

## Metrics proving resolution

Track provider success rate, p95 latency, partial-result rate, retry amplification, circuit-open duration and overall search completion.

## Prevention

Define supplier-specific timeout, concurrency, circuit-breaker and fallback policies. Avoid one global timeout value for every integration.
