Webhook Replay & Idempotency Troubleshooting
Handle duplicate, delayed and out-of-order webhook events safely with deduplication, replay rules and idempotent state transitions.
Duplicate webhook delivery should be considered normal. The real failure is applying the same business event more than once.
Symptom
A booking updates twice, a refund is applied twice, event order is reversed or replay restores an older state.
Diagnosis
Check stable provider event IDs, delivery count, occurred/received timestamps, prior application, replay safety and whether older events can overwrite newer state.
Idempotency
Use provider event IDs when available. Otherwise build a deterministic key from provider, entity, event type, source timestamp and/or payload hash.
Replay
Reprocessing the same evidence set should not change final state or duplicate side effects.
Possible causes
Expect duplicates, out-of-order delivery, partial transactions, dedup expiry, reused provider IDs and poison events.
Recovery
Replay the same event twice in a controlled test. Business state and external side effects should change once.
Observability
Track duplicate delivery, dedup hits, out-of-order events, replay failures, poison events and processing lag.
Prevention
Keep dedup retention at least as long as provider replay windows and protect state transitions with explicit versions or sequences.
Out-of-order event recovery
Do not let an older event roll back a newer state. Prefer monotonic provider sequences when available; otherwise combine provider timestamps, local versions and allowed-transition rules.
For example, a delayed CONFIRMED event arriving after CANCELLED should be preserved as evidence but must not restore the materialized booking state.
Validation
Replay CONFIRMED → CANCELLED → delayed CONFIRMED for one booking. Final state should remain CANCELLED while the stale event remains visible in observability.
Are you facing this in production?
We can review the symptom, data flow and integration behavior technically.