Supplier Timeout Incident Playbook
Handle travel-search supplier timeouts with containment, retry budgets, fallbacks, circuit breakers and controlled recovery.
The goal during a supplier timeout incident is not to rescue every request. It is to protect the overall search experience while isolating a degraded upstream dependency.
Trigger / symptom
Provider p95/p99 latency rises, timeout rate crosses a threshold, partial results increase, or worker threads and connection pools begin to saturate.
Objective and scope
This playbook focuses on search/read paths. Booking/create-order writes require a separate retry design because they may be non-idempotent.
Evidence needed
Collect provider/endpoint, timeout rate, latency percentiles, concurrency, connection-pool usage, request volume, HTTP/network error mix, retry count and provider incident status.
Ordered response
- Define blast radius: one endpoint, region or whole provider.
- Verify timeout budget relative to upstream behavior.
- Check for retry amplification.
- Limit concurrency with bulkheads/pools.
- Open a circuit breaker when failure exceeds the defined threshold.
- Select fallback: cached/indicative data, alternate provider or partial response.
- Preserve the UI contract; one supplier should not fail the entire search.
- Probe recovery with small half-open traffic.
Retry decision
Retry only idempotent requests when sufficient latency budget remains. Use exponential backoff with jitter and avoid synchronized retries across instances.
Stop conditions
Timeout rate returns near baseline, the circuit closes safely, and queues/pools recover.
Escalation criteria
Escalate with evidence when multiple regions/endpoints fail, SLA breach persists or a booking-critical path is affected.
Metrics proving resolution
Track provider success rate, p95 latency, partial-result rate, retry amplification, circuit-open duration and overall search completion.
Prevention
Define supplier-specific timeout, concurrency, circuit-breaker and fallback policies. Avoid one global timeout value for every integration.
Are you facing this in production?
We can review the symptom, data flow and integration behavior technically.