Durable Execution: The Problem It Solves and How Temporal Works

Illustration representing reliable workflow execution

Learn how durable execution helps long-running processes survive failures, and how Temporal uses workflows, event history, activities, and replay to resume work.

Imagine an online order that needs to charge a card, reserve inventory, wait for a shipping update, and notify the customer. The process may take minutes or days. If a server restarts after charging the card but before recording the shipment, what happens next?

Without a durable process coordinator, application code must reconstruct what happened. It needs to avoid charging twice, decide which step to retry, track long waits, and handle messages that arrive late or more than once. Durable execution is a way to make that multi-step process recoverable when workers, servers, or networks fail.

What durable execution means

Durable execution means recording enough information about a running process that it can continue after a failure. Instead of keeping all progress only in a server's memory, the system persists the workflow's history and state transitions. When a worker disappears, another worker can use that recorded history to rebuild the workflow's state and continue.

The process can wait for a timer, an external event, or a human decision without requiring a server to remain occupied in memory the whole time. The workflow's progress is stored durably, and a worker can resume it when there is new work to do.

Durable execution addresses several familiar problems:

  • Process crashes: in-memory variables disappear when a process stops.
  • Partial progress: a multi-step operation can fail after some side effects have already happened.
  • Retries: a transient error may need another attempt, but careless retries can duplicate work.
  • Long waits: timers and human approvals are awkward to manage with a process that must stay alive.
  • Coordination: messages, timeouts, and service responses can arrive late, repeatedly, or out of order.

It does not make external systems infallible. A payment provider can still be unavailable, and a request can still have an ambiguous outcome. The application still needs idempotency, sensible retry policies, timeouts, and compensation where a completed side effect must be reversed.

A simple example

Consider a customer onboarding flow:

  1. Create the customer account.
  2. Send a verification email.
  3. Wait for the customer to verify the address.
  4. Provision the requested service.
  5. Send a welcome message.

If the process is implemented as one request handler, it may time out while waiting for verification. If it is spread across queues and scheduled jobs, the application needs to persist which step completed, correlate messages with the right customer, and decide what to do after each retry.

A durable workflow models those steps as one logical process. The process can sleep while awaiting verification, then resume when the verification event arrives. If the worker restarts during the wait, the workflow's recorded progress remains available.

How Temporal provides durable execution

Temporal is a platform for running durable workflows. A Temporal application has workflow code that describes the orchestration and activity code that performs external or otherwise non-deterministic work. A Temporal Service stores workflow event history and places tasks on task queues; worker processes poll those queues and execute workflow tasks and activities.

At a high level, the cycle looks like this:

  1. A client starts a workflow execution.
  2. A worker runs the workflow code and asks Temporal to schedule work, such as an activity or timer.
  3. The Temporal Service records workflow events in the execution's event history.
  4. Workers perform activities and report their results.
  5. The workflow continues from the recorded results and schedules the next work.

The event history is central. If a worker goes away, another worker can replay the workflow code against the recorded history. During replay, completed activities are represented by their recorded results rather than being run again. The workflow's state is reconstructed, and execution can proceed from the point where new work is needed.

Workflows and activities

A Workflow defines the process and coordinates its steps. Workflow code must be deterministic: given the same recorded history, it must make the same decisions and produce the same commands. This lets Temporal replay the workflow safely to rebuild its state.

An Activity performs work that interacts with the outside world, such as calling a payment API, sending an email, or writing to a database. Activities can fail for transient reasons, so they can use retry policies and timeouts. Because an activity may have completed an external side effect just before its worker lost contact, activity code should be designed to handle retries safely. Idempotency keys are one common technique when the external service supports them.

For example, a workflow might coordinate payment and fulfillment while activities call the payment and inventory services. If the workflow worker restarts after payment succeeds, replay uses the recorded workflow history to reconstruct progress. If an activity attempt itself has an uncertain outcome, the activity's retry and idempotency design determines how to avoid an unintended duplicate charge.

Timers, signals, and human input

Durable workflows can wait without keeping a worker busy. A workflow timer is recorded by the service and can trigger later progress. External events can be delivered to a workflow as signals, and the workflow can wait for a signal such as an approval or verification event.

This is useful for scheduled reminders, delayed retries, subscription renewals, order fulfillment, and other processes that span a long time. The application can express the wait as part of the workflow instead of relying on an in-memory sleep or building all timer recovery logic itself.

What durable execution does not remove

Temporal manages workflow progress, task delivery, event history, and supported retry behavior, but application design still matters:

  • Put external side effects in activities, not in replayed workflow logic.
  • Make activities safe to retry, especially when they can charge money or create records.
  • Set timeouts and retry policies based on the operation and its failure modes.
  • Keep workflow decisions deterministic, and follow Temporal's guidance when changing workflow code used by existing executions.
  • Use compensation or a saga-style process when earlier side effects need a business-level reversal.
  • Monitor workers and the Temporal Service, and plan how the service and its database are operated.

Durable execution changes the recovery model; it does not remove the need to understand failure semantics at system boundaries.

When to consider it

Durable execution is a strong fit when a business process has multiple steps, can take a long time, must survive deploys or crashes, or needs careful retries and timers. Examples include payments and fulfillment, onboarding, report generation, data pipelines, and approval flows.

For a short request that performs one quick operation, a regular service handler may be simpler. The value of a workflow platform grows when coordinating progress and recovery has become more complex than the business logic itself.

Takeaway

Durable execution makes process progress persistent so a workflow can recover after failures instead of starting from lost in-memory state. Temporal implements this with workflow event history, deterministic replay, workers, and activities. The model can simplify long-running coordination, while retries and external side effects still need deliberate, safe application design.

For implementation details, see Temporal's official documentation on what Temporal is and workflow tasks and replay.