Webhook vs Polling Reconciliation
Design webhook, polling and reconciliation roles for travel booking systems, including duplicate, out-of-order and missing-event recovery.
Webhook and polling are not substitutes. In a production travel system, webhook is fast evidence, while polling and reconciliation provide completeness and recovery.
Reference flow
Webhook and reconciliation architecture
Why webhook is not enough
Webhook delivery can be lost, duplicated, reordered, retried or rejected during temporary endpoint failures. Therefore "no webhook means no event" is not a safe assumption.
Why polling is not enough
Continuous polling consumes rate limits, creates unnecessary traffic and delays state propagation. Polling should be targeted and state-aware.
Reconciliation scheduling
Good candidates include long-lived UNKNOWN/PENDING bookings, expected-but-missing webhooks, booking/payment divergence, incomplete cancellation/refund and provider references without a local terminal state.
Use backoff:
1 min -> 5 min -> 15 min -> 1 h -> manual thresholdDeduplication and ordering
Do not let ingress update state immediately. First verify authenticity, derive provider event identity, deduplicate, record occurredAt/sequence, then execute transition rules.
An older event must not roll back newer state.
Source precedence
Define precedence for conflicting evidence. One example is:
authoritative status lookup
> signed provider webhook
> synchronous API response
> locally inferred timeout stateThe exact order can vary by provider, but it should be explicit.
Failure modes
Typical failures include duplicate webhook, replay attack, stale event rollback, polling storms, parallel reconciliation of the same booking and eventual-consistency lag in provider lookups.
Observability
Track webhook receive rate, signature failures, duplicate and out-of-order ratios, webhook-to-state latency, reconciliation queue depth, resolution rate and pending/unknown age.
Production checklist
Require signature/auth verification, idempotent event apply, event storage, ordering/version control, state-aware polling, exponential backoff, reconciliation ownership locking and manual escalation thresholds.
Let’s review your architecture.
We can assess your travel distribution and metasearch architecture for scalability, failure modes and operations.