Every e-commerce support team knows these tickets. “I was charged but never got a confirmation.” “My order says processing and has for three days.” “I got two charges for one order.” Behind each one is usually the same engineering story: an order that started its journey through payment, inventory, fraud checks and fulfilment, and stopped somewhere in the middle.
If you’re a backend or platform engineer at an online retailer, marketplace or D2C brand, this article is about why that happens and how to fix it structurally, so that “stuck order” stops being a weekly firefight.
Checkout is a distributed transaction
A modern checkout doesn’t happen in one database. A typical order touches:
- A payment provider to authorize, and later capture, the card
- An inventory service to reserve stock
- A fraud or risk service to score the order
- A pricing and promotions service to lock in discounts
- A fulfilment or warehouse system to pick and ship
- A notifications service to email the customer
Each runs as its own service, often owned by a different team, and some are external providers you don’t control. There’s no single transaction wrapping them all. If step four fails, steps one to three have already happened.
The three ways orders get stuck
Partial failure without compensation
Payment is authorized, then the inventory service times out. The code throws, the request fails, and nobody releases the payment authorization or tells the customer. The customer sees a pending charge and no order. Eventually the authorization expires, but the support ticket arrives first.
Retries that create duplicates
To fix the first problem, someone adds retries. Now the payment call times out on the client side but actually succeeded on the provider side. The retry charges the card again. Without idempotency keys and a durable record of what already succeeded, retries trade one problem for another.
Lost progress on restarts
A deployment rolls out at 2 p.m. on a busy day. Pods restart mid-checkout. Any order whose state was held in memory, or in a message that was acknowledged before it was processed, simply stops. There’s no error, just an order that never moves to the next status.
Why the common fixes don’t hold up
Teams usually patch this in layers. A nightly “stuck order” job finds orders in PROCESSING for more than an hour and tries to push them forward. A set of message queues with retry and dead-letter topics connects the services. A growing collection of if status == X and payment_state == Y conditions handles the edge cases.
It works, mostly. But the business logic of checkout ends up spread across consumers, cron jobs and status columns. Nobody can look at one place and say what happens when fraud scoring fails after inventory was reserved. And every new payment method or fulfilment option adds more branches.
The structural fix: an orchestrated saga on durable execution
The pattern that fixes this is well known: the saga. A long business transaction is broken into steps, and each step that changes state has a matching compensation step that undoes it. If a later step fails, the saga runs the compensations for everything that already succeeded, in reverse order.
Sagas can be choreographed (each service reacts to events) or orchestrated (one coordinator calls each step). For checkout, orchestration is usually easier to reason about, because the entire flow lives in one piece of code.
The missing ingredient in many saga implementations is durability. The orchestrator itself must survive crashes and deployments. If it dies after authorizing payment but before reserving stock, a new instance must pick up exactly where it left off, without re-authorizing.
That’s what durable execution engines provide. Each step is recorded in an event history before the orchestrator moves on. After a crash, the engine replays the history, skips completed steps using their recorded results, and continues. Here’s a simplified checkout saga using the open-source Dapr Workflow engine in Python:
import dapr.ext.workflow as wf
from datetime import timedelta
wfr = wf.WorkflowRuntime()
retry = wf.RetryPolicy(first_retry_interval=timedelta(seconds=2),
max_number_of_attempts=4, backoff_coefficient=2)
@wfr.workflow(name=”checkout”)
def checkout(ctx: wf.DaprWorkflowContext, order: dict):
done = []
try:
yield ctx.call_activity(authorize_payment, input=order, retry_policy=retry)
done.append(void_payment)
yield ctx.call_activity(reserve_stock, input=order, retry_policy=retry)
done.append(release_stock)
risk = yield ctx.call_activity(score_fraud, input=order, retry_policy=retry)
if risk[“decision”] == “reject”:
raise ValueError(“fraud check rejected order”)
yield ctx.call_activity(capture_payment, input=order, retry_policy=retry)
yield ctx.call_activity(create_shipment, input=order, retry_policy=retry)
yield ctx.call_activity(send_confirmation, input=order)
except Exception:
for undo in reversed(done):
yield ctx.call_activity(undo, input=order)
yield ctx.call_activity(notify_cancellation, input=order)
raise
A few things are worth pointing out:
- Every step is retried with backoff, so a two-second blip at the inventory service doesn’t fail the order.
- Completed steps are never repeated after a crash, because their results are replayed from history. That removes the “charged twice” class of bugs at the orchestration level.
- Compensation is explicit and ordered. Anyone reading the code can see exactly what happens when fraud rejects an order after stock was reserved.
- Deployments are safe. Restarting a pod mid-checkout doesn’t lose the order. Another instance picks it up.
You still need idempotency keys on calls to external providers, using the order ID as the key, because a step that crashed during a call may run again. Most payment providers support this, and it closes the remaining gap.
Handling the long tail: waits and timeouts
Real checkouts include steps that wait: 3-D Secure challenges, bank transfers that settle hours later, pre-orders that ship when stock arrives. Durable workflows handle these as durable timers or external events. The orchestrator can wait for a “payment settled” event with a 48-hour timeout, using no compute while it waits, and then either continue or compensate. This replaces a whole category of cron jobs that poll for orders in limbo.
Observability that support teams can use
A side effect of durable execution is that every order has a complete, ordered history of what happened. That’s useful for engineers, but it’s even more useful for support. Instead of “let me escalate to engineering,” an agent can look up the order’s workflow and see that payment authorized at 14:02, inventory timed out three times, and compensation voided the authorization at 14:03. Some teams expose a simplified version of this timeline directly in their support tooling.
Running it in production
You can run Dapr Workflow yourself on Kubernetes, with workflow state in a database you already operate. Retailers that would rather not run that infrastructure, or that need higher throughput and enterprise support, can use a managed option. Diagrid, the company founded by Dapr’s creators, offers one such platform, Catalyst, which runs the same workflow model either hosted or inside your own cloud account. Whichever you choose, peak trading events like Black Friday are the real test, so load-test the workflow engine with production-like order volumes and deliberate pod kills well before November.
A checklist for checkout reliability
- ☐ The checkout flow is defined in one orchestrator, not spread across consumers and cron jobs
- ☐ Each state-changing step has a documented compensation
- ☐ The orchestrator runs on a durable execution engine that survives restarts
- ☐ All external calls carry idempotency keys derived from the order ID
- ☐ Retry policies use backoff and limits, set per step
- ☐ Long waits (3-D Secure, bank transfers, pre-orders) are durable timers or events, not polling jobs
- ☐ Support can see an order’s step-by-step history without asking engineering
- ☐ The workflow has been load-tested at peak volume with pods being killed mid-run
Conclusion
Stuck orders aren’t caused by bad luck or one flaky service. They’re the predictable result of running a distributed transaction without a durable coordinator. Moving checkout onto an orchestrated saga with durable workflow execution gives every order a single, recoverable path from cart to doorstep, and gives your support team far fewer tickets that start with “I was charged but…”.

