Quick Answer
Authenticate each delivery, commit a durable inbox record with discoverable work, then acknowledge under the provider’s response contract. Workers recover unfinished effects using separate intake and business identities. Resolve uncertain external calls before replacement and reconcile gaps beyond provider retry windows.
Key Takeaways
- Duplicate and out-of-order deliveries are expected; durable receipt and recoverable progress make their consequences manageable.
- Verify the selected signature scheme and commit durable receipt with discoverable work before acknowledging.
- Apply one decision rule everywhere: if receipt cannot be persisted safely do not acknowledge success, and if receipt is stored then acknowledge and recover internally when later processing fails.
- Deduplicate intake and financial effects separately, commit effects with markers, and preserve distinct later refunds and returns.
- Build a provider delivery contract matrix before writing handlers and treat unknown ack, retry, disablement, or replay behavior as a launch blocker, while classifying worker failures into backoff-with-jitter retries versus a dead letter queue.
Why webhook retries fail in production even when the code looks fine#
Webhook failures can arise at the acknowledgment boundary or during downstream processing. A valid event may arrive twice, late or out of order. If receipt, progress and financial effects are not stored durably, a retry can lose work or apply it twice.
A clean happy path does not exercise a crash after receipt, a queue outage or concurrent workers. Decide what counts as received, how unfinished work is found and which identity protects each effect before relying on provider redelivery.
Why "working code" still breaks#
At-least-once delivery means duplicates are normal. If your handler performs business side effects before a dedupe boundary, a valid retry can apply those effects again.
| Signal | Meaning |
|---|---|
| Slow acknowledgment | Compare elapsed time with the selected provider’s documented deadline. |
| Redirect response | Stripe treats 3xx as failed delivery; configure the final endpoint URL. |
| 4xx or 5xx | Inspect the response and logs; these do not prove no downstream effect occurred. |
| Fast success with growing backlog | Delivery succeeded, but internal processing may still be stuck. |
Timeouts can cause redelivery while internal work remains unfinished. Adyen’s handling guide specifies a ten-second acknowledgment deadline; PayMongo documents thirty seconds. These are different provider contracts, not a general five-second failure threshold.
Inspect the provider delivery record and your durable intake record together. For Stripe, failed delivery status and HTTP response details diagnose the edge; they do not establish whether a worker completed a financial effect. Match the account, event identifier and internal progress before replaying.
What keeps failures recoverable#
Recoverability depends on three things working together: a provider-aware contract, replay-safe architecture, and operator-grade diagnostics. Keep the webhook endpoint thin. Verify authenticity, persist receipt, and return quickly. Run business processing behind that boundary, where your own retries, exponential backoff with jitter, and dead letter queue can handle unresolved failures.
Use one decision rule everywhere: if receipt cannot be persisted safely, do not acknowledge success. If receipt is persisted, acknowledge and recover internally if later processing fails. That separates delivery reliability from downstream recovery and makes replay predictable instead of guesswork.
Store the authenticated event payload, provider/account/environment identity, received time, payload fingerprint and processing state with restricted access and an appropriate retention policy. Never log signing secrets. Keep provider delivery evidence separate from internal completion evidence.
What this article is and is not about#
This article covers payment notification delivery and processing reliability: webhook endpoint behavior, retries, duplicates, timeout handling, replay safety, and recovery workflows. It does not cover card decline recovery or dunning strategy.
From here, the path is straightforward: verify the provider delivery contract, define the webhook boundary, design idempotency, route failures intentionally, and verify the flow before go-live. If you want a deeper dive, read How to Handle Failed Payments Across Multiple Payment Methods and Regions.
Webhook retry logic in payments and what it does not cover#
In payments, webhook retry logic means the provider redelivers the same event when your endpoint does not acknowledge it successfully. That is about delivery reliability, not completion of your downstream business work.
Separate provider redelivery from your own retry path#
Treat these as separate responsibilities. The provider manages delivery attempts to your endpoint. Your system manages retries for business steps after receipt is stored durably.
Before acknowledging, commit trusted receipt and a recoverable work item. Use a transactional inbox/outbox or let workers discover unprocessed inbox rows. A database insert followed by an unreliable queue push leaves a gap: the provider can stop retrying even though no worker will run.
What duplicate protection actually depends on#
Use the provider’s actual duplicate identity, scoped by account and environment, for intake. Use a separate business-effect key at the write boundary. Event delivery, financial posting and later refund or return are different identities; one payment-wide key must not suppress legitimate later movements.
Group duplicate intake by the provider-specific identity, but preserve each authenticated payload version with its fingerprint and receipt time. Load existing processing progress on conflict; a received row is not proof of completion. For Adyen Standard notifications, the same eventCode and pspReference can carry changed details: evaluate the latest provider-supported version instead of discarding it as already seen. Deduplicate the financial effect separately, and recover unfinished downstream work.
What a success response should mean#
For Stripe, a successful 2xx acknowledges delivery. For other providers, use their documented status and body contract. In this design, success means trusted receipt and work are durably recoverable, not that fulfillment or every financial effect has completed.
Automatic retries help with transient delivery failures. They do not fix persistent signature, payload validation, or authentication errors. For card decline recovery and sequencing, see Smart Dunning Strategies: How to Sequence Retry Logic for Maximum Recovery.
Compare provider delivery contracts before writing application code#
Build and approve a provider delivery contract matrix before you write handlers, and block launch if key fields are still unknown. Once you separate provider redelivery from your internal retry path, you need each provider's actual contract. Retry behavior is not uniform across services, and retries do not fix persistent payload or authentication errors.
Build the contract table first#
Treat the contract table as the first deliverable. Include the fields that affect incident response, recovery design, and alerting, even when the current answer is "not confirmed."
| Provider | Acknowledgment | Automatic redelivery | Gap recovery and operating caveat |
|---|---|---|---|
| Stripe | Successful 2xx; authenticate the raw request with the endpoint secret. | Live: up to 3 days with backoff; sandbox: 3 attempts over a few hours. | Dashboard resend up to 15 days; CLI up to 30 days. Manual resend does not cancel automatic retries. Disabled/deleted at retry time can stop further attempts. |
| PayMongo | HTTP 200–209 and JSON response within 30 seconds. | Up to 12 retries; the resource reference says 3 consecutive events exhausting retries auto-disable the endpoint. | Re-enable via Dashboard/API. Current resend guidance supports failed-delivery retry; re-enable alone does not replay missed events. Verify availability for the actual account and reconcile gaps from objects. |
| Adyen | Successful status such as 200/202 within 10 seconds; certain terminal events require 200. | A missing timely success response moves delivery into its retry queue. | Verify the webhook type’s retry/recovery controls. Standard duplicate identity includes eventCode and pspReference; retain changed payload versions and evaluate the latest details separately from effect deduplication. |
Record the source and check date for each row. These examples do not establish every provider’s policy. Keep account-specific status, restoration permissions and any unsupported recovery path visible in the runbook.
Use PayMongo as the incident-readiness baseline#
PayMongo’s Webhook Resource says three consecutive events that exhaust twelve retries can disable the endpoint. Alert on endpoint status as well as failures. The current manual retry guide supports resending failed deliveries after repair, while re-enabling alone does not replay missed events. Confirm which controls are available for the account and retain an object-based reconciliation path.
After restoring an endpoint, identify the missing period and affected resources. Resend supported failed deliveries and reconcile the remaining gap from authoritative objects. Endpoint health alone does not prove that internal payment state caught up.
Treat unknowns as launch blockers, not doc debt#
Keep a visible "known unknowns" column and resolve it before production. At minimum, confirm for each provider in scope:
- What counts as acknowledgment, whether HTTP
2xxonly, a body requirement, or structured JSON - What retry behavior is documented, and what is explicitly not guaranteed
- Whether endpoints can be disabled or paused, and how that is detected and reversed
- Whether missed deliveries can be replayed or resubmitted, or must be reconciled from authoritative provider objects
Resolve the recovery path for each provider in scope before relying on it. Where delivery cannot be replayed, define the authoritative object lookup and exception owner instead of assuming a resend button exists.
Capture verification evidence#
Once you resolve an unknown, keep the proof. Store the provider doc URL, check date, and exact excerpt or support response. During sandbox tests, keep delivery artifacts such as headers, payloads, timestamps, statuses, and error details so observed behavior can be compared with documented behavior.
Confirm account-, environment- and webhook-type-specific terms rather than implying a Gruv-specific integration contract. For the accounting side, see Xero Integration for Payout Platforms.
Decide what happens in the webhook endpoint and what moves to workers#
Set a strict boundary: the webhook endpoint should acknowledge receipt quickly, and workers should handle fulfillment. In practice, the endpoint should validate authenticity and basic schema, persist the receipt, return a fast acknowledgment, and hand off processing asynchronously.
This separation keeps delivery acknowledgment distinct from business execution. The provider needs confirmation that you received the event. Your system owns retries, fulfillment, and reconciliation after that point. Mixing those concerns in one path is what turns timeouts into duplicate deliveries and harder incident recovery.
Keep the endpoint narrow#
A practical endpoint should do only this:
| Task | Where it belongs |
|---|---|
| Verify the request is authentic and matches expected schema | Webhook endpoint |
| Persist an immutable receipt record before side effects | Webhook endpoint |
| Commit discoverable processing work with the receipt, or use a transactional outbox | Webhook endpoint |
| Return the provider-required success response after durable receipt/work commit | Webhook endpoint |
ledger journal writes | Workers |
payout batch changes | Workers |
| Notifications | Workers |
| Dependency-heavy lookups | Workers |
When endpoint work gets slow, timeout risk rises, and retry behavior is provider-specific enough that you should not depend on it to clean up design mistakes.
Keep financial mutations out of the acknowledgment path#
Do not perform money-state or payout-state mutations before acknowledgment. If a timeout hits around partial processing, you can end up unsure what committed, whether a retry will come, and whether a replay will apply the same side effect again.
Workers do not remove the need for idempotency, but they make the boundary clearer: receipt first, fulfillment second. That gives operators a cleaner recovery path when duplicates or failures happen.
Make failure behavior explicit#
Make the branches explicit:
- if receipt persistence fails, do not return success
- if receipt is stored but downstream worker processing fails, keep provider acknowledgment and recover internally
Before success, confirm authenticated receipt and recoverable work with a provider-specific intake identity. Preserve changed payload versions where the provider can update a duplicate notification, and evaluate them under that provider’s ordering rules. An existing receipt must still expose unfinished work. Never depend solely on redelivery or a volatile queue.
Design idempotency that survives retries and partial failures#
Design idempotency at the data boundary so retries and partial failures resolve as safe no-ops, not duplicate side effects. Use two dedupe controls for two different risks, and enforce both where writes happen.
Use two keys for two different failure modes#
A provider event ID is useful where supplied, but some webhook types use composite duplicate identities. Scope intake by provider, account and environment. Delivery deduplication does not solve out-of-order state changes or duplicate financial effects across different events.
Use:
- Provider-specific event or composite identity for duplicate intake, scoped by account and environment.
- Business-effect identity for each protected operation, including distinct later refunds or returns.
The event ID tells you whether this delivery is new. The business key tells you whether the protected effect already happened.
Put the stop at the persistence boundary#
Do duplicate protection before side effects, at the write boundary. Code-level checks alone can race under concurrency.
Commit each local effect and its applied marker in the same transaction, with uniqueness at the write boundary. A marker committed first can suppress work that never happened; an effect committed first can be repeated after a crash. For an external call, keep a durable attempt and resolve an unknown result under that API’s idempotency contract before replacement.
Track progress so retries can resume safely#
Persist processing state and use a lease or comparable concurrency control so workers can reclaim abandoned jobs. Treat progress as a recovery checkpoint, not an irreversible seen-event flag. If a financial effect completed but notification did not, resume the notification rather than repeating the posting.
A practical internal state machine is:
receivedvalidatedappliednotified
These are illustrative internal states, not provider states. Each marker needs evidence for its completed effect. Commit notification intent through an outbox; an external send can be uncertain after a crash, so use the notification service’s deduplication or reconciliation support where available.
Use a clear decision rule for replays#
When a new delivery maps to an already-applied effect, skip that effect and recover any unfinished downstream work. Do not skip an entire payment because one previous event was applied: a later refund or return has its own financial identity.
Build the payment event pipeline in the right order#
Once idempotency is in place at the write boundary, sequencing becomes the next reliability decision. Separate event receipt from business effects so retries stay operational, not financial. A practical internal order is receive, verify, persist raw input, enqueue, process, then audit. Use that as your design rule, not as a provider standard.
Keep receipt separate from business effects#
Delivery and processing fail for different reasons, and providers can retry when acknowledgment is not confirmed even if your server received the event. If acknowledgment is tied to ledger updates, fulfillment, or notifications, transport issues can trigger duplicate side effects.
After trusted durable intake, normalize the event and enforce provider-supported transition rules. Do not order Stripe events solely by arrival or event-created timestamp; timestamps can tie. Retrieve authoritative objects for current-state projections, while preserving distinct historical money movements for reconciliation.
A useful checkpoint is simple: every accepted event should have an immutable raw record plus internal processing state, so recovery does not depend on guesswork from logs.
Decide compliance gating explicitly#
If your flow includes KYC, KYB, or AML checks, treat gate placement as an explicit architecture choice and document it. The provider acknowledgment decision and the compliance decision do different jobs, so avoid blending them by accident.
Trusted receipt must remain separate from permission to initiate a new payment. A compliance condition may block new instructions where required, but cannot erase or suppress accounting for money that already moved. Record the actual financial event and route any control failure to review.
Make ownership and replay explicit#
In complex flows, missed events often become an ownership problem, not just a delivery problem. Before launch, define an ownership map per event type with:
- source event type
- owning service
- allowed state or money mutations
- downstream subscribers
- replay entry point
Replay points should match your internal states so recovery is controlled:
| Replay point | Recovery action |
|---|---|
| received | Recover unfinished work from inbox/outbox; do not create another receipt. |
| validated | Resume mapping with the original input and version. |
| applied | Confirm the financial effect and continue unfinished downstream tasks. |
| notified | Confirm communication evidence; resolve uncertain sends rather than assuming a marker proves delivery. |
This gives you a clean way to recover missed notifications without reapplying financial effects.
Set retry classes and backoff rules for internal workers#
Do not use one retry policy for every failure. Classify failures first: retry transient worker failures with exponential backoff plus jitter, and route non-retriable failures to a dead letter queue or equivalent review lane.
This keeps your internal recovery logic clear when provider redeliveries are also happening. If you mix those two retry clocks, duplicate inbound deliveries can look like internal progress, and stuck jobs can hide behind fresh traffic.
| Failure class | Typical signal | Internal handling | Stop or escalate rule |
|---|---|---|---|
| Transient processing or dependency failure | timeout, temporary unavailability, rate limiting | Retry with exponential backoff and jitter | Escalate if the same error repeats with no state change for the same event |
| Duplicate, replay, or idempotency conflict | same event reappears, idempotency conflict | Load completed and unfinished effects; skip only a verified completed effect. | Conflicting identity or amount requires review; an applied effect does not stop recovery of another unfinished step. |
| Non-retriable input or state failure | invalid payload or invalid state for your processor | Remove from the normal retry lane and send to dead letter queue for review | Escalate with payload reference and failure reason |
Make retry decisions visible on the job record, not only in logs. Per attempt, persist at least webhook event ID, idempotency key, failure class, retry count, next attempt time, and last processing state. That lets operators quickly tell whether the issue is provider delivery, worker processing, or a suppressed replay.
Keep provider redelivery and worker retries separate#
Provider retry policy is a different control surface from your internal worker retry policy. You may receive repeat deliveries at the edge while an accepted copy is already processing internally.
Track and alert on these windows separately: edge delivery behavior versus internal queue or job aging. If you only monitor end-to-end completion, one failure mode can hide the other.
Stop looping on the same identifiers#
Repeated failures on the same webhook event ID or idempotency key need a hard stop. If an event keeps failing without advancing state, stop auto-retrying at your class limit and escalate to incident or manual review.
Use the provider’s real signature scheme and endpoint secret. Stripe verifies the unmodified raw body and Stripe-Signature; apply the library’s timestamp tolerance and keep clocks accurate. A provider-generated redelivery receives a fresh signature, so distinguish signature age from event age. Adyen HMAC formats differ by webhook type; do not copy a generic X-Signature example into production.
Throttle bursty event classes before they stampede your stores#
High-volume bursts need rate limiting and buffering so retries do not become a correlated storm. Cap concurrency by event class and let queues absorb spikes to reduce overload, duplicate pressure, lock contention, and avoidable dead letter queue floods.
Need the full breakdown? Read Accounts Payable Aging Report for Platforms: How to Track Overdue Contractor Payments.
Prepare for missed events and disabled webhook states#
Use the selected provider’s retry and resend windows to recover undelivered notifications. Outside those windows, reconcile authoritative provider objects and your financial records. A healthy endpoint and a successful resend do not prove every earlier gap is closed.
Adyen’s handling guidance explicitly separates secure receipt, storage, acknowledgment and processing. Use the same separation in incident evidence: find the missing range, recover available messages and reconcile remaining state from authoritative provider records.
Before you close the incident, verify three things: delivery is healthy again, the backlog or gap has been backfilled, and your idempotency controls prevented duplicate application during catch-up. That is the difference between restoring the endpoint and restoring system state.
Frequently Asked Questions
What does webhook retry logic actually cover?
It covers provider redelivery when your endpoint does not acknowledge an event successfully. It does not guarantee that your downstream business processing completed correctly.
Why can working webhook code still fail in production?
Production adds duplicate delivery, timeouts, network jitter, slower commits, and partial failures between receipt and side effects. If the dedupe boundary is weak or the endpoint does too much synchronous work, retries can create duplicate outcomes and manual cleanup.
What should happen inside the webhook endpoint?
Authenticate using the selected webhook type’s scheme, then commit durable receipt and discoverable processing work before the required success response. Retain changed payload versions where provider duplicates can update details. Recover unfinished work and deduplicate financial effects separately.
How is a webhook event ID different from an idempotency key?
The provider-specific event identity detects duplicate intake; a business-effect key protects each financial or downstream operation. Scope identities by account and environment, commit local effects with their markers and give legitimate later refunds or returns separate identities.
How should internal worker retries differ from provider redelivery?
Provider retries are about delivering the event to your edge. Worker retries are your internal recovery mechanism after durable receipt, and they should use failure classes, backoff, jitter, and a dead-letter path for non-retriable cases.
What must the missed-event runbook confirm before closing an incident?
It should confirm that delivery health is restored, the backlog or gap has been backfilled, and idempotency controls prevented duplicate application during catch-up. That proves system state, not just endpoint availability, was restored.
Researched and edited by the Gruv editorial team. Gruv builds cross-border billing, payouts, and finance-operations software for global businesses.
Sources
Includes 3 external sources outside the trusted-domain allowlist.
Educational content only. Not legal, tax, or financial advice.
Related Posts

The Freelance Payment Penalty: A Modeled Audit of Platform Fees, FX Spreads, and Payout Delays
The money rarely disappears through a single, easy-to-spot fee. The real loss is stacked. A marketplace takes its commission, a processor adds a charge for international cards, a bank or payment company converts the currency at a spread, a platform holds the funds before release, and a wire sheds a little to intermediaries on the way in. Each layer looks defensible on its own, but the worker feels the combined result as a smaller deposit and a later payday.

How to Respond to a Subpoena for Business Records
Move fast, but do not produce records on instinct. If you need to **respond to a subpoena for business records**, your immediate job is to control deadlines, preserve records, and make any later production defensible.

A US Expat's Guide to Investing in UCITS ETFs to Avoid PFIC Issues
The real problem is a two-system conflict. U.S. tax treatment can punish the wrong fund choice, while local product-access constraints can block the funds you want to buy in the first place. For **us expat ucits etfs**, the practical question is not "Which product is best?" It is "What can I access, report, and keep doing every year without guessing?" Use this four-part filter before any trade:

