Skip to content

Durable Execution Profile

The Durable Execution Profile adds persistence sufficient to suspend, resume, recover, reconcile, compensate, and cancel a run without changing its logical identity. It constrains durable boundaries and abstract history, not databases, queues, workers, internal serialization, retention periods, or service bindings. The independently claimed Run Records Capability defines an inspection projection.

Use of wf.wait_for, recovery, reconcile, undo, on_unknown: reconcile, or on_unknown: halt structurally activates this profile. Core on_unknown: abort and on_unknown: fail do not.

A claim to this profile MUST include a primary target under Profile Conformance; operational use requires Core Executor conformance. A valid Workfile requiring this profile is subject to the unsupported-workflow prohibition under Static Validation when the target does not claim it.

A durable boundary is a point from which the executor can continue without repeating a completed external effect or losing a binding, decision, timer, subscription, attempt result, or pending ambiguity. The executor MUST establish a durable boundary before it acknowledges an event, releases an earlier worker, or reports that a run is suspended or held.

A run is suspended while it is durably waiting for a timer, event, child run, recovery decision, or other profile-defined input. A run is held while an ambiguous action requires reconciliation or operator resolution.

Neither is a terminal outcome. Suspension releases no logical scope and does not change which step is active.

Resumption MUST restore the same run metadata, resolved definition, scopes, active step paths, attempt counts, fixed idempotency keys, timers, subscriptions, accepted events, and recovery generation. It MUST NOT re-evaluate an expression whose recorded result is itself a nondeterministic profile input or reissue a completed action; binding identity remains governed by Single Assignment.

A timer that elapsed while execution was unavailable becomes eligible on resumption but does not fire twice. The profile includes durable handling of wf.wait and state timers when the run executes on a Durable implementation.

Event suspension additionally requires the initiation capability corresponding to the resolved trigger kind. Child-run suspension requires the child and caller to remain linked until both reach the required boundary.

An unavailable prerequisite makes the resolved Workfile unsupported before execution.

wf.wait_for durably waits for one correlated event or a timeout. The table defines its literal trigger subject and required match and timeout configuration.

Key Role Required Type or domain Default
wf.wait_for Subject Yes Literal connector.trigger reference. —
match Configuration Yes Nonempty map supplying at least one declared correlation field and no other field. —
timeout Configuration Yes Nonnegative duration or timestamp, including a whole-value interpolation. —
configuration Configuration No Trigger configuration map; see Trigger Configuration. {}
connection Configuration No Nonempty string literal without interpolation; see Connection Selection. Unambiguous deployment default.
steps Configuration No Step list run after the subscription is active. Absent

The steps are ordinary nested steps of the construct and have paths such as approval.ask; they use no reserved segment.

Here connection is a construct key, governed by Connection Selection, rather than an action modifier. The permitted wf.wait_for modifiers are defined exclusively by Step Modifiers. Under on_fail: continue, a failed wait binds null; Static Nullability applies to every path in the wait binding.

recovery does not apply to wf.wait_for.

- approval:
wf.wait_for: messaging.approval
configuration: { channel: "{{ inputs.approval_channel }}" }
match: { request_id: "{{ inputs.request_id }}" }
timeout: 24h
steps:
- ask: { messaging.request_approval: { request_id: "{{ inputs.request_id }}" } }

The trigger MUST resolve to an external or poll source with nonempty correlate_on. configuration is governed by Trigger Configuration; the adjacent connection is governed by Connection Selection.

The configuration is evaluated once when the step begins, validated, and fixed through suspension. A configuration fault fails the step before it suspends.

The match expressions are evaluated in the enclosing scope when the step begins, validated against the corresponding payload fields, and fixed through suspension; the match expressions MUST NOT reference a binding created by steps.

timeout is evaluated once; a duration is measured from that evaluation. The step establishes its subscription and timer at one durable boundary, at which the subscription becomes active, then runs steps.

An event accepted before that boundary is not eligible. An event accepted afterward is eligible when each supplied match value equals the corresponding payload field and MUST be retained while steps is in progress; it does not interrupt the list.

The event or timeout is selected only after steps succeeds. An event accepted no later than the timeout deadline wins; otherwise the timeout wins.

Delivery identity prevents the same event from completing the step twice. Selection is recorded before the subscription and timer are released.

A failure or run termination in steps prevents selection and releases the timer and subscription after that disposition is durable. A step id in steps MUST NOT be status or event.

The step succeeds in both cases and binds a scope:

Path Type Meaning
status enum["received", "timeout"] Which condition completed the wait.
event Trigger output or null The delivered payload, or null on timeout.
each step id Declared step type Binding created by the optional steps list.

A timeout is not a failure. A payload MUST satisfy the resolved manifest and overlay before delivery.

A match fault fails the step before suspension; an invalid candidate event does not complete it. Cancellation releases the timer and subscription after the cancellation decision is durable.

recovery can appear at the Workfile root, on a state, or as the recovery modifier of a scope-producing construct. It contains redo, required stop_after, required then, and optional delay and notify.

redo is this_step, this_and_after, or whole_group and defaults to this_step.

When failure policy reaches a recovery scope, the executor selects the redo set as follows:

redo Redo set
this_step The failed step.
this_and_after The structural active suffix beginning with the failed step, as defined below.
whole_group All executed work in that scope.

For this_and_after, the redo set is determined from workflow structure, not from expression references or inferred data dependencies. It begins with the direct child of the recovery scope that contains the failed step and includes every later sibling in that scope’s ordered step list; that containing child is re-executed as a whole.

For a state, on_enter followed by steps forms one ordered list for this purpose. Within an active conditional branch or iteration, the same rule is applied to that branch or iteration before continuing with later siblings of its enclosing construct.

Completed independent branches of a wf.parallel are not in the suffix; the failed branch is rebuilt and the join is evaluated again. Before redo, the executor durably marks the selected active-history suffix superseded, performs required compensation, and begins a new recovery generation from the preceding boundary.

Superseded attempts and values remain history but no longer define active outcomes or bindings. The rebuilt generation creates each logical binding once; recovery MUST NOT mutate a binding in place.

stop_after is {failures, within}, where failures is a positive integer and within is a positive duration.

A recovery activation begins when its recovery decision is recorded. stop_after counts automatic recovery activations for that scope in the rolling interval ending at that recorded instant.

delay is a nonnegative duration literal and defaults to 0ms. It is measured from the recorded recovery decision to the start of compensation; during a positive delay the run is suspended on a durable timer, and compensation followed by re-execution begins automatically when the delay elapses.

When the positive activation limit is reached, the executor MUST NOT begin another recovery activation at that scope and MUST apply then. then is hold, fail, or escalate: hold suspends at that scope, fail fails the scope, and escalate applies recovery at the next enclosing recovery scope.

then: escalate is invalid with document.control when the recovery declaration has no next enclosing recovery scope.

While held by this rule, a deployment operator MAY authorize one recovery activation for the failure that caused the hold. The authorization and actor identity MUST be recorded before compensation; the activation follows the declared delay and does not count toward stop_after.

Absent such authorization or cancellation, the run remains held. notify is an opaque deployment notification target and has no workflow semantics.

Recovery is reached only after action retry and catch, under Failure Policy Precedence. on_fail: abort bypasses it; on_fail: continue, a containing on_iteration_fail: continue, or a containing on_branch_fail: continue consumes the failure before it reaches the recovery scope.

A recovery decision, its time input, selected redo set, and resulting generation MUST be recorded before compensation or re-execution begins.

A validator SHOULD warn when a Workfile with the Durable Execution Profile has no explicit on_unknown for a write action that declares idempotency: none. The warning SHOULD explain that unresolved ambiguity defaults to abort.

An ambiguous attempt, as defined by Connector Execution Contract, establishes neither success nor definite failure. The action step remains in progress and creates no binding.

This profile adds reconcile and halt to the Core on_unknown policies. The inherited policies and default are defined there.

halt durably holds the unresolved run for an operator decision. The decision can establish success with a result satisfying the output contract or establish definite failure with a stable code and details; authorization of a safe reissue remains governed by Core Ambiguity Disposition.

The executor MUST record the decision and actor identity supplied by the deployment before resuming. reconcile requires a sibling map with steps, outcome, and, when the action declares output, result.

The reconcile steps run in a nested scope after the ambiguous attempt. Each reconcile step has path <action-path>.(reconcile).<id>, with ordinary nesting and indexing continuing from that path.

outcome is a boolean expression evaluated after those steps:

  • true establishes that the action succeeded; result is evaluated, validated against the action output, and bound;
  • false establishes that the action had no effect; the ambiguous code becomes a definite retryable failure and ordinary remaining retry policy can reissue safely; and
  • a failed reconcile step or fault leaves the attempt ambiguous and applies halt.

A reconcile probe MUST NOT itself use the unresolved action’s binding. Its scope and results are visible only to outcome and result.

An action result already recorded as definite success MUST NOT be reconciled, retried, or replaced. A definite failure follows retry and catch, never on_unknown.

undo is an action call attached to an action step and describes compensation for that step’s successful active-generation effect. The compensation invocation has step path <action-path>.(undo).

The undo action runs only when Durable recovery is about to supersede that success; run failure, wf.fail, wf.stop, and cancellation do not implicitly compensate. Every successfully executed action in a redo set MUST be safely repeatable: its manifest declares inherent or keyed idempotency, or the step declares undo.

Static validation MUST reject a recovery scope for which a possible redo set violates this rule. Before re-executing a redo set, every successful action in it that declares undo is compensated in reverse completion order.

A successful action without undo is reissued under its inherent or keyed idempotency promise. Each undo’s arguments are evaluated from the still-available source generation, including the compensated step’s result.

The executor MUST durably record the intent to compensate before dispatch, and record its result before the next compensation. After interruption it resumes at the first compensation without a definite success.

Such a reissue is permitted only when the undo action is inherently idempotent or keyed with a durable fixed key; otherwise an ambiguous undo holds the run for operator resolution. A definite undo failure fails recovery and holds the run with the failed compensation identified; later redo work MUST NOT begin.

A successful compensation does not erase the original action or its result from history. Compensation is not a transaction, and no atomicity across actions is promised.

The Durable Execution Profile extends Core Cancellation with persistence and preserves its prescribed outcomes. The executor MUST record a cancellation decision durably before cancellation takes effect and before starting its effects.

The record contains any implementation-defined reason. A suspended or held run can be cancelled.

The executor durably releases pending subscriptions and timers and records propagated child-run cancellation requests. Results from work already dispatched remain subject to Late Results.

Cancellation does not run undo.

Durable execution requires a stable workflow identity across the lifetime of a run, and Durable history MUST record it. Durable history is an ordered semantic record sufficient to restore every nonterminal run at a durable boundary and to support the Replay Profile when claimed.

Durable history MUST identify the resolved definition and dependencies, run metadata, validated inputs and trigger data, scope and step paths, active recovery generation, and every logical decision that affects later work. For each operation it records evaluated arguments, connection identity, attempt number, timeout and idempotency context, dispatch status, result or ambiguity, stable code and details, retry delay and source, and resulting outcome or binding. Partial dry-run values retain the suppression information defined under Dry Runs.

Durable history also records guards and branch selections; iteration and state-entry positions; timers and their deadlines; subscriptions, correlation values, accepted events, and delivery identities; child-run links; handler, reconciliation, compensation, recovery, hold, operator, and cancellation decisions; and the terminal outcome, accepted stop path and reason when applicable, and successful workflow result when present. History MUST distinguish an invocation not dispatched, dispatched without a definite result, definitely failed, and definitely succeeded.

Durable history MUST distinguish active history from a superseded recovery suffix. A durable boundary is complete only when every value and decision needed after that boundary is recoverable or is explicitly marked unavailable in a way that prevents unsafe continuation.

Sensitive-field rules can require a value to be withheld. An implementation MUST NOT claim that history remains usable for resumption or replay after withholding a value required by that operation. The implementation MUST report the missing class of information instead of substituting null or reissuing work. History of a nonterminal run MUST NOT be reduced below what resumption requires.

Physical event schemas, database layout, transactional mechanism, compression, encryption, replication, and retention after terminal outcome are outside this specification; standardized export applies only where a claimed capability defines it. An implementation MUST document its retention and failure guarantees.

An implementation’s physical history choices MUST NOT weaken the logical requirements while the implementation claims a run is resumable or its history usable by a Replay operation.