Provider Partial Outage / Degraded Mode Troubleshooting
Diagnose travel-provider partial outages where latency, timeouts or specific capabilities degrade without a total service failure.
Provider incidents are not always total failures. The hardest cases are partial degradation, such as search working while reprice/booking fails or only certain markets/capabilities breaking.
Symptom
Latency percentiles rise, timeouts increase, some endpoints/markets fail selectively, search still returns results while booking fails, or stale-cache rates rise.
Possible causes
Regional provider incident, throttling, downstream dependency failure, capacity saturation, auth/token problems or capability-specific outage.
Diagnosis
Break errors down by operation and market, compare p50/p95/p99, separate rate-limit/auth failures, run read/write smoke checks and correlate with provider status communication.
Recovery
Remove unhealthy capabilities from eligibility, tighten timeout budgets, use circuit breakers/bulkheads, serve cached/partial search where safe, fail closed for booking/reprice and expose degraded availability to users.
Prevention
Use provider health scoring, operation-level circuit breakers, bounded concurrency, retry budgets, capability-specific health and graceful-degradation runbooks.
Observability
Track provider success, operation-level errors, latency percentiles, circuit-open duration, degraded-mode traffic share, partial-result rate and recovery time.
Are you facing this in production?
We can review the symptom, data flow and integration behavior technically.