Quick Answer
One million payments per day averages 11.6 payments per second. Peak traffic, provider limits and recovery work determine the capacity you need; durable command identity prevents retries from turning that capacity into duplicate charges.
Key Takeaways
- Count business payments separately from API calls, provider attempts and webhook events.
- Persist request identity and the payment command atomically before external execution.
- Resolve an unknown provider outcome before rerouting or creating another attempt.
- Commit webhook processing markers with local state changes, and publish downstream work through an outbox.
- Reconcile provider settlement and bank cash separately from API acceptance.
Build for a million payments and their failure paths#
A million-payment day is not, by itself, a reason to split your system into dozens of services. It averages about 11.6 business payments per second. What matters is the busiest window, the work each payment creates, and what happens when a provider completes a charge just before your connection drops.
Start by defining a transaction for your capacity model. Here it means one business payment instruction. An authorization, capture, refund, status lookup and webhook are separate operations around that instruction. Counting them all as payments hides the load and makes success rates hard to interpret.
Turn daily volume into a peak capacity model#
| Illustrative input | Calculation | Capacity implication |
|---|---|---|
| 1,000,000 payments across 24 hours | 1,000,000 ÷ 86,400 = 11.57 payments/second | Average only; does not establish the peak |
| A ten-times-average busy window | 11.57 × 10 = 115.74 payments/second | A planning assumption to replace with observed or contracted demand |
| Six internal calls per payment | 115.74 × 6 = 694.44 calls/second | Measure each service separately; writes and reads have different costs |
| One provider call per payment, mean latency 0.8 seconds | 115.74 × 0.8 = 92.59 concurrent calls | About 93 in-flight calls at steady state before headroom |
| 20% additional provider attempts in that window | 115.74 × 1.20 = 138.89 calls/second | Illustrative retry amplification, subject to safe-retry rules and provider limits |
These are worked assumptions, not a benchmark or a guarantee. The concurrency calculation uses mean time in flight; use measured latency distributions and a bounded concurrency limit for operations. A spike lasting five minutes is a different queueing problem from a spike lasting an hour. Measure the largest tenant, synchronized renewals and batch imports as well as aggregate traffic.
Model authorization and capture separately if your flow calls the provider twice. Add status recovery, refunds, polling and webhook fetches to the provider budget. Check its account, endpoint and burst limits; additional API replicas cannot make an upstream limit disappear. Keep a separate budget for recovery traffic so a callback outage cannot consume every available provider request.
For example, if 116 jobs arrive each second while workers can complete only 100, backlog grows by about 16 jobs per second. Over ten minutes that adds about 9,600 jobs. When arrivals fall to 60 per second and service capacity stays at 100, the backlog takes roughly four minutes to drain. Those estimates assume constant rates; validate queue age and recovery time with your actual workload.
Define the API contract before the external call#
Return a durable payment ID when you accept a command. Say explicitly whether the response means accepted for processing, authorized, captured or settled. An HTTP success from your API must not make a client believe bank cash has arrived.
| Client situation | Contract behavior |
|---|---|
| New valid instruction | Create one durable command and return its payment ID and current state |
| Same tenant, operation and key; identical payload | Return the existing command and its current or saved response |
| Same key with a different amount, currency or recipient | Reject the mismatch; never reinterpret the original command |
| Concurrent duplicate while the first command is running | Return the existing in-progress state or a documented conflict; do not create a second command |
| Client timeout after acceptance | Lookup the payment or repeat the same command identity; do not invent a fresh key |
Scope the idempotency record to authenticated tenant, operation and client key. Store a canonical payload digest alongside it. Enforce uniqueness in the database, then create the command and any local reservation in the same transaction. A check followed by an unprotected insert lets concurrent requests both proceed.
Keep your internal record long enough to cover client retries, recovery and financial investigation. Provider retention is a separate limit. Stripe’s idempotency documentation says keys can be removed after they are at least 24 hours old; a reused key after pruning can initiate a new request. It also documents parameter checks and saved responses, including failures. Do not assume another provider uses the same rules.
Make crashes recoverable at the provider boundary#
Before sending an external instruction, persist an attempt ID, provider account, operation, amount, currency and the provider idempotency key. Use a conditional state update or transactional claim so two workers do not both start a new attempt. A worker lease helps schedule work, but expiry of the lease does not prove the previous external call failed.
| Failure point | Recovery action |
|---|---|
| Before the command transaction commits | No accepted command exists; the client may submit it again |
| After commit, before a provider call | A worker resumes the persisted command with the same attempt identity |
| Provider may have acted, but response is lost | Mark the attempt unknown; retrieve provider status or retry only under its documented idempotency contract |
| Provider success arrives before local state is committed | Recover the existing attempt through lookup, callbacks or reconciliation; commit its outcome once |
| Local outcome commits before downstream notification | The outbox publisher sends the committed notification; receivers tolerate duplicates |
Do not switch providers, change rails or issue a fresh key while the first attempt is unknown. A new destination is not a cancellation of the original request. Escalate unresolved attempts with their provider references and timestamps; create a replacement only after evidence establishes the original did not execute, or after a separately authorized reversal and new payment.
A local database transaction cannot atomically commit a remote provider charge. Keep that uncertainty visible instead of claiming end-to-end exactly-once delivery. Your practical goal is one business obligation, durable attempts and controlled recovery.
Keep payment state and accounting truth distinct#
Races show up when updates from different paths arrive close together or out of order. If multiple paths can write state without a guard, lifecycle integrity breaks.
For a card flow, requested, processing, authorized, captured and failed can describe execution; settlement, refunds and disputes need their own records. A captured payment can later be refunded or disputed. Do not treat a single terminal label as permission to discard every later event.
Validate a proposed transition against stored state in the same transaction that records it. Persist the triggering attempt or event ID and an append-only transition record. Use a unique business-effect identity for journal postings so a second event about the same capture cannot post the same financial effect twice.
A balanced double-entry ledger records your accounting entries; provider statements and bank records establish external settlement evidence. Reconciliation connects those records and explains timing differences, fees and mismatches. It should not silently overwrite the ledger to make a dashboard balance match.
Choose a database by the atomicity, query and recovery requirements of this write path. At this volume, measure a straightforward transactional design before introducing manual sharding. If you add derived stores, designate their data as projections and keep authoritative updates out of uncoordinated dual writes.
Receive webhooks durably, then process them once locally#
Verify the signature against the raw request body before trusting tenant routing or payload fields. Resolve the provider account to your own tenant mapping, validate the event, durably capture it and then acknowledge delivery. If capture fails, return a failure so the provider can retry. Keep expensive business processing out of the receiver.
Stripe’s webhook guidance documents duplicate deliveries, unordered events and asynchronous processing. Use the provider account and event ID to identify delivery duplicates. Distinct events can concern the same payment, so payment ID alone is not an event deduplication key.
In the worker transaction, check or insert the event-processing marker together with the guarded state change, journal effect and outgoing outbox record. If the worker crashes before commit, all those local changes roll back and the event can run again. If it crashes after commit, the marker prevents local effects from running twice. A marker written before the business transaction would instead risk losing an event.
Use the current provider resource where needed to resolve late or incomplete callbacks. Do not sort solely by receipt time or reject a legitimate refund because capture already completed. Separate delivery deduplication from the business-effect guard for two distinct events describing the same effect.
An outbox publisher may deliver a committed message more than once. Its downstream consumer needs its own atomic processing guard. The AWS transactional outbox pattern explains why the state update and outgoing event belong in one local transaction; it does not make the remote payment part of that transaction.
Apply bounded backoff with jitter to transient local processing failures. Send exhausted events to a visible dead-letter queue with event ID, payment ID, provider account, attempt count, last error and a named owner. A replay goes through the same guards as the first processing attempt.
Instrument operations for detection, triage and proof#
Track business acceptance, provider attempts and financially confirmed outcomes separately. Publish oldest unknown-attempt age, oldest unprocessed-event age, reconciliation exception count and amount, and queue drain time. Define the denominator of a duplicate-attempt rate; a safely replayed client request is different from an unintended second charge.
Carry tenant ID, payment ID, attempt ID, provider account and reference through requests, events and reconciliation rows. Store structured evidence with access controls and retention rules. Avoid putting raw payment credentials or unnecessary personal data into logs or idempotency keys.
Your operator view should answer which obligation is waiting, whether an external attempt might have executed, and what evidence permits the next action. API latency alone cannot answer those questions.
Related: How to Build a Deterministic Ledger for a Payment Platform.
Prove recovery before increasing traffic#
- Send simultaneous identical requests and confirm one durable command; send the same key with a changed payload and confirm rejection.
- Kill a worker before the provider call, after the call and after local commit. Confirm recovery preserves the attempt identity and exposes unknown outcomes.
- Deliver duplicate and out-of-order webhooks, including a later refund or dispute. Confirm each legitimate financial effect posts once.
- Stop outbox publishing, resume it and redeliver messages. Confirm committed events are recovered and downstream effects are guarded.
- Run the measured peak workload with realistic database writes, provider latency and rate limits. Include recovery traffic and the largest tenant burst.
- Compare command records, provider references, ledger entries, settlement reports and bank movements for the same sample. Explain every difference rather than relying on a matching total.
Assign thresholds to your own service objectives: maximum processing delay, unknown-attempt age, queue recovery time and unresolved reconciliation amount. Add orchestration when multiple durable steps need coordinated recovery; daily volume alone does not require it.
Frequently Asked Questions
How many transactions per second is one million per day?
Across 24 hours, 1,000,000 ÷ 86,400 is about 11.57 business payments per second. That average excludes additional API calls, provider attempts and webhooks. Size for the busiest measured window and each dependency’s limits.
Does one idempotency key protect the whole payment flow?
No. A client command key, provider attempt key, webhook event ID and local journal-effect identity protect different boundaries. Persist their relationship and commit each local processing guard with the effects it protects.
Can we retry with another provider after a timeout?
Resolve the original attempt first. A timeout means the outcome may be unknown, not that the payment failed. Use provider lookup, callbacks, reconciliation or its documented same-key retry behavior before authorizing a replacement.
When does a successful payment become settled cash?
API acceptance, authorization and capture do not by themselves establish settled bank cash. Match provider settlement records and bank movements to your ledger, recording fees and timing differences separately.
Do we need microservices to reach this volume?
The daily target alone does not decide that. Start with a transactional write path and bounded workers, measure peak load and recovery, and split components when independent capacity or operational ownership justifies the extra coordination.
Researched and edited by the Gruv editorial team. Gruv builds cross-border billing, payouts, and finance-operations software for global businesses.
Sources
Includes 1 external source outside the trusted-domain allowlist.
Educational content only. Not legal, tax, or financial advice.
Related Posts

The Freelance Payment Penalty: A Modeled Audit of Platform Fees, FX Spreads, and Payout Delays
The money rarely disappears through a single, easy-to-spot fee. The real loss is stacked. A marketplace takes its commission, a processor adds a charge for international cards, a bank or payment company converts the currency at a spread, a platform holds the funds before release, and a wire sheds a little to intermediaries on the way in. Each layer looks defensible on its own, but the worker feels the combined result as a smaller deposit and a later payday.

How to Respond to a Subpoena for Business Records
Move fast, but do not produce records on instinct. If you need to **respond to a subpoena for business records**, your immediate job is to control deadlines, preserve records, and make any later production defensible.

A US Expat's Guide to Investing in UCITS ETFs to Avoid PFIC Issues
The real problem is a two-system conflict. U.S. tax treatment can punish the wrong fund choice, while local product-access constraints can block the funds you want to buy in the first place. For **us expat ucits etfs**, the practical question is not "Which product is best?" It is "What can I access, report, and keep doing every year without guessing?" Use this four-part filter before any trade:

