Skip to content

Conformance Suite

Status: Draft · Workfile Standard revision 0.2.0

The suite has no version of its own. A claim names the suite it was made against by file set digest, so the digest is the suite’s identity.

This directory is the portable conformance suite for the Workfile Standard. Its cases cite the exact revision 0.2.0 requirement IDs they test. The specification remains authoritative, and passing applicable cases is evidence of conformance rather than a substitute for satisfying the prose requirements. Error codes come from the Error Code Registry.

The suite covers every validation code of that registry, every evaluation fault, and every core filter. The suite section also states the standard for completeness: a requirement that no case exercises is untested.

Directory Holds
validation/ One case per file. A document, and the codes an implementation MUST report for it.
digests/ One case per file. Content, and the digest that names it.
expressions/ Expression fixtures, as JSON Lines. One object per line.
connectors/ The test connector manifests that the conformance rules require, with their samples under fixtures/. Fixtures, not a catalog.
library/ Standard Library artifacts and their closed-format package fixtures. They are distributed beside the suite but are not portable case objects.
run-records/ Run-record discovery, inspection, and retention cases for wf.run-records.
traces/ Execution cases: a workflow, canned action results, and the expected outcome of every step.
triggers/ Trigger fixtures, as JSON Lines. The cron dialect of time.schedule is here.

requirements-index.json is the exact reverse index from every cited requirement ID to the cases that cite it.

Portable cases cite implementation requirements. IDs marked standard-maintenance or hosting in the requirement registry identify editorial or broad operational responsibilities outside the suite’s observations; the checker excludes them from case citations. Observable security requirements remain eligible.

run.mjs is the conformance runner. It drives any implementation that speaks the line protocol, so implementations that share no code produce comparable results.

cd conformance
npm install
node run.mjs -- <command> # e.g. node run.mjs -- wf conform
node run.mjs --verbose -- <command> # every applicable case, not only failures
node run.mjs --json -- <command> # the full result as JSON

The implementation writes its revision 0.2.0 claim first, then one response per request. The claim uses targets, capabilities, and profiles; the runner sends only cases within that claim. Each request includes the case’s requirements list and its capabilities deployment input when present. The implementation reports what it observed, and the runner decides whether that matches the case.

The runner and checks require Node.js 22 or later for source-aware JSON number parsing. canonical.mjs preserves exact integers and numeric kinds when reading JSON responses and YAML digest documents. check_canonical.mjs checks serialization and verifies that the runner rejects rounded integers, erased float kinds, and lost negative zero.

The runner reads the suite’s own fixtures to judge, which is why it lives here rather than in an implementation. Four details of that judging are worth naming.

  • An expression result keeps the int and float apart. {{ 6 / 3 }} is the float 2.0 and never the int 2, so the runner reads a number’s source text rather than trust a JSON parser that collapses the two.
  • A cron result compares as an instant, so an implementation may answer in any zone.
  • A trace case’s steps names steps the run never reached, and the record holds only the steps that did, so a missing entry reads as not_run.
  • A trace case’s bindings names a path into the run’s bindings, such as approval.event.status or scan.items[0].check. Execution step paths remain separate: that iteration step is scan[0].check.

check_schema.mjs checks the suite against itself: the structural layer, the fixture lints, and the digest cases. It judges no implementation, and the runner above is what does that. Run it after editing a fixture.

cd conformance
npm install
npm run check

It checks the structural schema boundary, current Error Code Registry membership, fixture format contracts, manifest samples, digest recomputation, published-schema compilation, and requirement traceability. The documentation check validates complete Workfile and connector-manifest YAML examples against their published schemas. Every case must contain a nonempty distinct requirements list of current IDs, and requirements-index.json must be its exact reverse index.

npm run check:docs first checks wiki links in all Git-tracked and new, non-ignored Markdown and MDX documents, including unpublished plan/ files. check_wiki_links.mjs rejects unknown names and registered targets whose document or heading no longer exists. It ignores fenced code, inline code, and literal link examples. This check runs in CI independently of which pages the site publishes.

check_docs.mjs then checks the selected documentation examples against the same schema. A fenced yaml block that states a top-level steps or states key is an example, and the schema checks its structure without a catalog, so an example that names a connector no catalog holds is still checked. An example that elides part of itself with ... is not a whole document, and the count of those is reported rather than hidden. These structural checks do not validate expressions or resolved connector contracts.

cd conformance
npm run check:docs
npm run check:links # repository-wide wiki links only
node check_docs.mjs --spec-only # only the normative specification examples

A case passes when the set of current registry codes an implementation reports equals expect, exactly. Message text is not compared. Without expect_results, an empty expect asserts the document is valid. With expect_results, an empty expect asserts only that no diagnostic code is reported; validity and support are checked separately.

expect_results is an optional nonempty map of result assertions. document and resolved accept valid, invalid, or incomplete; deployment accepts supported, unsupported, or incomplete. The implementation returns these observations in its response’s results map, alongside codes. The runner compares every asserted result exactly and fails a response that omits one. A case can therefore assert both incomplete resolved validity and a known unsupported deployment. Results not named by the case are not compared.

expect_warnings is an optional nonempty list of distinct warning assertions. Each is a closed record with requirement (a current, cited diagnostic.warning.* ID), path (a JSON Pointer into the primary workfile document), and explains (a nonempty list of distinct consequence observations). The defined observations are continuation-bypassed (the warning explains that unresolved ambiguity bypasses the expressed continuation policy), run-terminated (the warning explains that unresolved ambiguity terminates the run), ordinary-failure-remedy (the warning identifies on_unknown: fail as passing flow.action_ambiguous into catch, on_fail, and enclosing policy), external-write-uncertain (choosing fail does not establish whether the external write occurred), result-key-nullable (an omitted stop result makes otherwise non-null result keys nullable), and unmatched-route-empty-success (an unmatched open-string route succeeds with an empty branch). For raw source, the pointer addresses the parsed document. Assertions currently concern the primary Workfile; warnings from dependencies are outside this observation format.

The implementation reports observations in warnings, alongside codes and results, using the same fields. These records describe warnings actually emitted and what their messages explain; they are not predictions of what should be emitted. For unsafe continuation, an omitted stop result, or an implicit route fallback, the observation path identifies the affected step body, even if the native diagnostic points to a particular member of that body. A grouped diagnostic can be represented by one observation per affected path. The runner requires a record matching each asserted requirement and path whose explains includes every asserted consequence. Missing records or consequences fail the case; order, message wording, additional consequences, and additional warnings are not compared. Cases without expect_warnings impose no warning assertion. These are suite observations, not registered error codes or a production diagnostic wire format.

Every case resolves against the manifests in connectors/ and the canonical packages in library/. A case MAY add to that context with the keys that the suite section defines:

  • capability — the single capability claim required for the runner to select this case. Without this key, the case is not gated on an optional capability claim.
  • validator_claims — optional case-selection conditions on the implementation’s actual claim: capabilities and profiles list required claims; without_capabilities and without_profiles list claims that must be absent. Each list has distinct registered names. A name cannot be both required and excluded, including through capability. The runner skips a case when a condition is unmet; these conditions never change the implementation’s claim.
  • manifests — further connector or filter package manifests, inline.
  • overlays — schema overlays, inline.
  • connections — named connections, each with connector, granted_scopes, and parameters. Without this key, every connector resolves through one complete connection, named default, that grants every scope.
  • files — project-relative paths to file content, for call callees and templates.
  • capabilities — the capabilities available on the target deployment being assessed, independently of the validator’s own claim. Without this key, every capability is available.
  • profiles — the profiles available on that target deployment, independently of the validator’s own claim. Without this key, every profile is available. These support inputs assume compatible execution implementations for the listed features; they do not require the validator itself to execute them.
  • resolution — supplied catalog-resolution inputs, as described below.
  • previous_manifests — earlier connector manifests against which the selected manifests’ same-major compatibility is checked. These are comparison inputs, not additional selected versions. A contract violation uses dependency.invalid.

resolution is a list of closed records with required kind (connector or filter-package), namespace, and outcome. A kind/namespace pair occurs at most once. Each record replaces the default resolution input for that namespace, including any otherwise available suite artifact. These records supply resolver findings to the validator; they do not prescribe a production catalog API or version-constraint syntax. Unlisted namespaces use the ordinary fixture context. An unlisted namespace absent from that supplied context has no owner; it is not missing validation input.

Outcome Supplied input
input_missing No resolution input was supplied for this namespace.
no_owner Completed catalog lookup establishes no owner.
no_match Completed selection finds no artifact satisfying the supplied constraints.
version_conflict Supplied constraints require incompatible versions.
selected Required version selects the exact fixture manifest; optional digest supplies an expected document pin for verification.
content_unavailable An artifact with required exact version was selected or pinned, but its content is unavailable; optional digest identifies the pin.
owner_conflict Supplied catalog resolution establishes conflicting namespace owners.
identity_conflict Supplied catalog resolution establishes distinct contents with the same kind, namespace, and exact version.

version is a complete SemVer version, and digest is sha256: followed by 64 lowercase hexadecimal digits. Neither field is allowed on the other outcomes. Selection does not bypass manifest validation, member lookup, or argument checks. A selected fixture whose content fails pin verification is unavailable for dependent checks. The runner also supplies resolution and previous_manifests in the request when present; implementations can read the complete case at the request’s suite-relative path.

workfile is normally the parsed document. In a case that tests the serialization subset, workfile is a string that holds the raw document text, because the condition does not survive parsing.

The structural subset of these cases is also checkable with the JSON Schema. The schema cannot decide a case whose code needs a catalog, a connection, a file, or a scope — those cases are marked needs: resolution and a schema-only checker skips them.

A case states one digest an implementation must produce, and kind selects the rule: document, file, or file_set.

  • A document case holds parsed YAML under document, and the canonical form as text under canonical. The canonical form is checked too, so a failure names the rule that broke rather than a hash that differs.
  • A file case holds a string under file, and digests its UTF-8 bytes. The trailing newline is content, so it changes the digest.
  • A file_set case holds a map of relative path to content under files.

document_minimal and document_key_order hold one data model written two ways, so their digests are equal. The nested-map pair also reorders supplementary-plane keys and integer-like keys. List reordering changes the digest. The file_map_order_* pair shows that the same map reordering changes byte-sensitive file digests, and the file_set_unicode* pair covers UTF-16 ordering of file paths.

document_string_escapes and document_key_sort pin the two rules that implementations most often differ on — which code points a string escapes, and how keys sort. document_scalars and document_numeric_kinds preserve exact signed 64-bit integers, distinguish integral floats from integers, and preserve signed zero. Workfile borrows JCS UTF-16 property sorting with this numeric adaptation; it is not unmodified JCS.

One JSON object per line: expr, an optional bindings map, and exactly one of result and fault. A string binding in a position that expects a timestamp or a duration converts by the timestamp rule.

File Covers
core.jsonl Arithmetic, equality, ordering, membership, indexing, and conversion faults
maps.jsonl Unordered maps, UTF-16 key enumeration, recursive serialization, numeric round trips, and unchanged list ordering
operators.jsonl Precedence, short-circuit evaluation, the conditional, and literals
collections.jsonl The collection filters
selection.jsonl where, keep, and each, non-null tests, and key presence
strings.jsonl The string filters
numbers.jsonl The number and conversion filters
time.jsonl Timestamps, durations, and the time filters
escaping.jsonl The escaping and hashing filters

Where the core filter section defers to fixtures, the fixtures are normative. The escaping fixtures fix the exact character sets and quoting forms; fault.template has no expression fixture, because it occurs only in a template body — traces/template_render_fault.yaml covers it.

Every YAML case and every JSONL case object contains requirements, a nonempty list of distinct IDs from spec/requirements.json. A citation means the case directly observes some conformance behavior required by that ID; it does not imply that the case exhaustively proves the requirement.

requirements-index.json enumerates cases by cited ID. npm run check regenerates the same mapping in memory and fails if the committed index differs, an ID is unknown or retired, or any case lacks citations. The JSON runner report also groups results under byRequirement.

cron.jsonl holds the cron dialect of the canonical time.schedule trigger. One JSON object per line: cron, timezone, after, and then either next or invalid.

  • next is the first instant the expression matches after after. Timestamps compare as instants, so a case can name one in the zone that reads most clearly.
  • An absent next asserts that no instant matches, as 0 0 31 2 * does.
  • invalid: true asserts that the text is not a cron expression. These cases fix the grammar’s exclusions: no name forms, no @ keywords, no seconds field, and no step on a bare value.
  • why states what the case covers. It is documentation, and an implementation ignores it.

The zone cases are the reason this file exists. They fix the two transition rules: a local time that a gap skips fires at the first instant after the gap, and a local time that occurs twice fires on the first occurrence only.

A trace case supplies workfile, optional inputs, canned actions, optional agents, optional code, optional events, optional trigger, and expect.

  • actions entries are matched by step path and consumed in attempt order per path. A retried step consumes one entry per attempt. An entry’s outcome is success, failure, or unknown — an ambiguous attempt. A failure or unknown entry MAY carry details, the connector’s structured value that a handler reads as error.details, and retry_after_ms, the delay a Retry-After would have set.

  • code is a legacy fixture input for canned embedded-code results. No current portable case uses it: the language registry leaves wf.run reserved and invalid in this revision.

  • agents supplies deterministic agent scripts, each with step and a nonempty proposals list. Each proposal is either {tool, arguments} or {result}. A tool names a declared connector action, and arguments is its concrete argument map. The adapter presents proposals in order, delivering each next proposal only after the previous tool returns a result. A result proposal ends the script and supplies the candidate agent result. Tool dispatches consume ordinary actions entries. In these scripts, a tool proposal at zero-based position N uses fixture path <agent-step>[N]; retries consume further entries at that path. An adapter maps this fixture position to its stable logical tool-invocation position. When execution stops the agent, remaining proposals are not delivered; suppression supplies no fabricated tool response. These scripts replace an external agent runtime, not executor mediation or output validation.

  • Auxiliary work uses registered parenthesized segments: an action source is <step>.(source), a failure-handler step is <step>.(catch).<id>, a reconcile step is <step>.(reconcile).<id>, an undo invocation is <step>.(undo), and an exhaustion-handler step is <step>.(on_exhausted).<id>. wf.wait_for steps remain ordinary nested steps.

  • events supplies the outcome of each wf.wait_for, in order: { step, outcome: received, event: {...} } or { step, outcome: timeout }.

  • trigger states the trigger entry that started the run: { binding, payload }, where binding is the fired entry’s as name or manifest default. The run then binds the payload, the silent entries, and trigger.name exactly as a routed delivery would.

  • expect states the run’s terminal outcome and the outcome of every step. It MAY state bindings (path to value), outputs, and codes (the classified code per failed step).

  • A run that ends held is stated with held_at: the step that caused the hold and its classified code. steps then holds only the steps that reached an outcome, because a held step has none yet.

  • overlays resolves schema overlays for the run, as a validation case does. A validated_output action whose output holds an extension region is checked against them.

  • dry_run: true applies dry-run policy: writes without delegation and dependent work are suppressed, while independent work executes. A suppressed invocation consumes no canned action result. Affected output entries become null; successful calls preserve the suppression of individual child output members.

  • reduce names retained payload classes to remove after terminal outcome. It requires the observable history protocol below, which compares removed and retained data and each requested history operation.

  • cancel_after names a step path. When the run’s frontier command belongs to that path, the run is cancelled. The object form { at: <path>, only: callee } cancels the callee run that a call step at <path> started, on its own, and the calling run continues. cancel_ends_run, cancel_cascades_to_callee, and call_callee_cancelled_fails_the_step are the cases that use it.

  • A step outcome in expect.steps MAY be a list of admissible outcomes, for a step whose scheduling the standard leaves open. A sibling of a step that ended the run records not_run under a sequential implementation and cancelled under a concurrent one; parallel_fail_fast_cancels_siblings states both.

  • deadlines names the step paths whose timeout elapses. A bound a case does not name never fires, so every other case is unaffected. timeout_stops_a_loop is the case that uses it: the bound passes while the run waits out a repeat interval, so the next iteration never starts and the step fails flow.timeout.

A step that a trace does not reach still appears in expect, with the outcome not_run.

The recovery traces exercise selected on_unknown and undo outcomes: on_unknown_halt, ambiguity_safe_reissue, reconcile_effect, reconcile_no_effect, and undo_before_redo. They do not establish complete profile conformance.

Boundary timelines below can inject prior-attempt and superseded-generation completions, observe cancellation and recovery boundaries, resolve ambiguity, and request process loss and restoration. The history protocol below adds the three distinct Replay Profile operations and observable payload reduction.

Core Executor claims explicitly include wf.composition; this does not require any initiation capability. The time_now_core trace exercises canonical clock-action dispatch without scheduled initiation.

The dry_run_* traces distinguish suppression from ordinary null, exercise conservative reference dependencies and partial child results, and cover repeat, state, template, action-source, and agent boundaries. dry_run_states_* requires the Stateful Workflow Profile alone; cases with templates, action sources, composition, or agents state their required capability.

The stop_* cases exercise scoped stop results, default empty results, result faults, normal-completion joins, and competing stops. Trace expect.outputs observes the successful workflow result for either completion or a stop. The concurrent-stop case accepts either branch through a parent projection that checks result coherence and member count; it does not impose a branch order.

The call_handler_* and agent_handler_* cases distinguish terminal composition failures from local evaluation faults, cancellation, and ambiguity abort. They cover replacement contracts, child terminal fault codes, independent child cancellation, timeout recovery, and exhaustion-handler ordering. The catalog_error_namespace_* cases reserve each Standard namespace against exact and prefix declarations; vendor and executor-code cases check the boundaries. The ambiguity_attempts_* pairs preserve the shared attempt budget for inherent and keyed writes.

The for_each_nested_failure_origin and parallel_nested_failure_origin cases exercise the shared isolated-scope rules under Failure Policy Precedence: intervening construct outcomes, unreached siblings, and relative leaf error paths.

The timeline_recovery_*, timeline_resolution_*, and timeline_restart_* cases add focused recovery, resolution, and restart observations. These cases require the Durable Execution Profile; composition and event-initiation cases also select their required capability.

A trace can opt into workfile.timeline/1 with a timeline map. The request then includes observation_protocol: "workfile.timeline/1". Existing traces keep their partial assertions and canned-input interpretation. A timeline replaces actions, events, deadlines, cancel_after, agents, and code; combining them is invalid. reduce can accompany a timeline only through the history protocol below. Inputs, files, manifests, and ordinary final expect assertions still apply.

requires_profiles is an explicit list of Standard profile claims required to select a trace, including [] for Core alone. When absent, the existing trace filename selectors remain in use. capability retains its ordinary meaning. These selectors concern the implementation’s actual claim. They do not claim that the adapter can expose a particular observation boundary.

An adapter either returns timeline: {protocol: "workfile.timeline/1", observations: [...]} alongside its ordinary trace result, or responds with harness_unsupported: ["workfile.timeline/1"]. A missing protocol acknowledgement also produces harness-unsupported. This is an untested applicable case, distinct from outside the Standard claim and from an observed behavioral failure. It never counts as a pass, and makes the runner exit unsuccessfully. Once the protocol is acknowledged, malformed or missing observations fail the case. An adapter unable to expose any requested boundary reports harness support as missing; it does not omit that operation or synthesize a successful observation.

The adapter executes the Workfile under the fixture’s external inputs and reports actual logical effects, including rejected deliveries. The timeline is neither a predicted transcript nor an engine journal format. Native database keys, worker identities, physical timestamps, and persistence layouts are not compared.

timeline contains version: 1, handles, a nonempty operations list, and optional windows. Handle names are fixture-local aliases for distinct work identities. Each identity has kind and path:

Kind Additional identity fields Start boundary
action attempt, counting from 1 Actual connector dispatch, after the executor can no longer establish that dispatch did not occur
child None Child-run creation by the named call step
timer purpose, occurrence, and attempt for an attempt timeout Establishment of a logical timer or retry delay
subscription occurrence Subscription activation

A path is the suite’s step path, including auxiliary and indexed paths. Child paths use the call path as their prefix, such as call.fetch. The root run is root; a child run is identified by its call path. Timer purpose is attempt-timeout, construct-timeout, retry, wait, or subscription-timeout. occurrence counts from 1 for the path and purpose (or subscription path). It does not count deliveries. An attempt timeout also identifies its action attempt. These identities address the initial execution; version 2 adds recovery-generation identities below.

Every operation has a distinct id, an expect map (possibly empty), and exactly one control:

Control Meaning
await: handle Advance runnable work until that work has started. matched identifies its actual start in the cumulative event list. A start already observed can satisfy this operation.
deliver: {to, outcome, ...} Present a response to the identified work, including work that has already ended. Report whether it was applied or rejected as late.
expire: handle Deliver the named timer at or after its logical deadline, including a retired timer.
cancel: run-path Present an executor cancellation request, then observe its handling through the next scheduling boundary and settle runnable consequences.
settle: true Execute remaining runnable consequences until terminal or blocked on controlled external input. Report every intervening effect.

Action delivery uses success with result, or failure/unknown with code and optional details. Subscription delivery uses received with the event payload in result; each command is a distinct event delivery. Child delivery uses completed with result, failure with code, or cancelled. Here completed denotes a successful child notification from either a completed or stopped child; the observed child run retains its actual outcome. For a child, the adapter executes the supplied callee and buffers its terminal notification at the caller boundary. The delivery command’s payload checks that buffered notification; it does not replace child execution with an invented outcome. A mismatching child result fails the operation.

All external completions are withheld until their delivery commands. Timers fire only when explicitly expired; other timers can remain pending without advancing the fixture clock to their deadlines. Expiry advances logical time as needed, but does not automatically deliver other due timers. Late timer delivery is permitted by the scheduling rules. The adapter can control permitted jitter to reach a requested retry-delay boundary. An inability to expose that boundary is a harness limitation, not evidence that a particular jitter choice violates the Standard.

After a delivery, the executor processes its consequences to a scheduling boundary; it may also perform runnable work before the next observation. settle observes all runnable consequences. The final operation is always settle, including after terminal late deliveries, so appended work cannot hide outside the observed interval. Settling never fabricates a missing external result. A final run still blocked on such input cannot meet a terminal assertion.

Operations order only the specified external deliveries and observed starts. They impose no declaration order on independent branches and no physical-time or worker requirement. An atomic transition can report several effects in one operation. Cancellation taking effect and terminal transition can share that batch; the fixture does not require a nonterminal pause between them. Their logical order is still visible, so work started between those events is testable.

Each operation returns exactly one observation, in operation order, with id, status, events, and snapshot. Status is reached for await, applied or late for delivery/expiry, applied for a cancellation request presented to the executor, and settled for settle. Applied does not mean successful: a timeout or a fatal failure can be applied. A requested event is not evidence of its acceptance. Await also returns matched, the zero-based index of the actual start in the cumulative event list, including the current operation’s events.

events is an exhaustive ordered delta since the preceding observation, starting at run creation for the first operation. Report every logical event in these categories, including transient changes later reverted and work absent from the fixture’s handles. Do not coalesce state changes to their final value.

Event Fields and observation
start work: complete work identity; includes every dispatched attempt, child, timer, and subscription
release work: identity of a timer/subscription/work item made no longer pending
step-start path: every step beginning execution, including purely internal steps
cancel-requested, cancel-effective path: affected run; receipt and effectiveness are separate observations
state area, path, value: each logical state assignment or transition
output-start path: run whose workflow result evaluation begins
outputs path, value: produced workflow result
history path, value: retained information added or changed under an applicable history obligation

snapshot contains all five maps: runs, steps, bindings, attempts, and decisions. It is the complete logical state represented by the cumulative state events, not just the subset asserted by the case. The runner reconstructs these maps and requires exact agreement at every observation. Run and step maps use active before an outcome, with held for a held run. Binding entries use complete step paths and hold successful assigned values; absent assignments stay absent rather than being replaced by reference-resolution nulls. Attempt keys are <step-path>#<attempt> and their selected values are {outcome: success, result: ...}, {outcome: failure, code: ...}, or {outcome: unknown, code: ...}, with details when supplied. Pending attempts have no selected entry. Decisions use the existing trace decision path and branch value. The final run and reported step outcomes also agree with the ordinary final response.

Each operation’s expect can assert:

  • status: exact disposition;
  • snapshot: partial maps, with exact values for every named member, including explicit null; missing members fail;
  • absent: maps from snapshot area to lists of keys that have no assignment;
  • unchanged: true: the entire snapshot equals the preceding snapshot, and no state change, result evaluation, or result production occurred in the interval;
  • counts: records {match, exactly} or {match, at_least} counting matching events in the operation;
  • retained: partial observations of actual retained history, independent of active state, requiring requires_profiles: [durable-execution] (possibly with additional profiles).

An event matcher contains kind and optional event fields. Work fields match as a subset; value matches exactly, including nested objects. Counts can forbid an event with exactly: 0, or require an event without forbidding repetition using at_least: 1. Unknown fields and event kinds are rejected by the fixture checker. An unchanged snapshot alone cannot hide a temporary state change or forbid new work; unchanged checks mutations and counts check starts separately.

A window has from and through operation IDs, inclusive, and counts. Optional after identifies the first matching event within that interval and excludes that event and earlier events from the count. The marker must be observed. This can begin a prohibition at cancel-effective or terminal outcome even when later events share the same operation. It does not impose a global event order.

The Durable retention cases use retained entries keyed by attempt identity, with timeout, dispatched, outcome, code, and effect; terminal holds the terminal code and ambiguity disposition. releases maps fixture handles to true when their release is durably recorded. These observations read the implementation’s retained information; they are not copies of requested events. Additional retained details are allowed. Retention here is not evidence of survival across process loss; restart and replay need separate operations.

timeline_* cases cover timed-out attempts during retries, ended constructs and calls, cancellation cutoffs, terminal results, and retired timer/subscription deliveries. check_timeline.mjs feeds controlled transcripts through the actual runner to verify that stale acceptance, reverted mutations, missing observations, and forbidden extra work fail. Its passing transcripts validate the harness only. Execution evidence from an adapter supporting this protocol remains separate.

Recovery and restart observations (version 2)

Section titled “Recovery and restart observations (version 2)”

timeline.version: 2 negotiates workfile.timeline/2. It includes version 1’s controls and assertions and requires an explicit durable-execution profile selector. A version 1 acknowledgement cannot satisfy it. An adapter without recovery, operator, or actual process-restart support for a requested case returns harness_unsupported: ["workfile.timeline/2"]; it does not skip that control. Existing version 1 fixtures and responses keep their interpretation.

Every work identity in version 2 includes generation: 0 for initial execution, then the run’s recovery-generation ordinal. The number is a portable alias, not a native storage key. Attempts and timer occurrences count within the generation; this observation numbering does not reset an executor’s effective attempt budget. Preserved independent work retains its identity. A handle for old work continues to address that work after supersession or restoration. Child paths identify the owning run as in version 1. Recovery timers add purpose recovery.

Version 2 snapshots include two additional complete maps:

  • generations: run path to current generation ordinal;
  • suppression: binding or run-result path to the sorted list of suppressed member paths. Member paths use JSON Pointer escaping (~0, ~1); the empty string denotes the whole value. Unsuppressed values have no entry. Prefixes subsume descendants, so a whole-value marker is [""] alone.

The ordinary binding map still contains only assigned values. A partial value can contain explicit null members while the separate suppression map records why those members are unavailable. A superseded suppression decision is removed from active state; any suppression in the replacement generation comes from its own execution and restored inputs. Version 2 attempt keys are <step-path>@<generation>#<attempt>. Superseded attempts move out of active snapshots and remain observable in retained history.

An event {kind: remove, area, path} removes an existing active entry during supersession. Rebuilding a binding requires removal of its superseded assignment before its new assignment; restoration itself does not assign bindings again. All assignments and removals remain visible, including changes later reversed.

A recovery event is {kind: recovery, phase, scope, from, to, redo}. phase is decision, supersede, or begin, in that order for one transition. The first two events observe the durable decision and supersession, respectively; begin observes entry into the rebuilt generation after required compensation. scope is the recovery scope’s path (root for root recovery). redo is the exact structural redo set, represented by the shortest nonoverlapping subtree roots in structural order, including selected later siblings not yet executed. For a parallel suffix, these roots identify the failed branch and later selected work; completed independent branches are excluded. A this_step set names its failed step. The three events agree on scope, generations, and membership.

Compensation belongs to the source generation. Before its dispatch, a history event at <undo-attempt-key>.intent records value: {arguments: ...}; after a definite result, <undo-attempt-key>.result records value: {outcome, result} or {outcome, code}. These are normalized observations of the durable intent and result, not new history-storage fields. Other applicable retained details remain available in retained.operations. Cases observe reverse completion order, source-generation arguments, and completion before redo dispatch.

Windows additionally support before, an exclusive end marker, and ordered, a list of event matchers required in order within the selected interval. after is applied first, then before. Both markers must exist. ordered requires a subsequence, allowing other events and different operation batching; it does not impose an order on independent work. Existing counts still apply.

resolve: {to, actor, outcome, result?, code?, details?} presents a deployment- authorized ambiguity decision for an action handle. success supplies result; failure supplies a stable code and optional details. The adapter obtains fixture authorization from its test deployment and preserves the supplied actor identity. It observes {kind: resolution, to, actor, outcome, ...} when the decision is durably recorded, before resumption, with the same decision fields. An applied observation means the authorized decision was applied. It does not permit bypassing result validation or granting another attempt. Reconciliation uses ordinary Workfile execution and action delivery; it has no control that fabricates a reconciliation result. The resolution fixtures deliver an original completion only after resolution, when that completion is late.

interrupt: process-loss interrupts execution at the preceding observed durable boundary. It neither cancels the run nor advances work or creates a graceful shutdown checkpoint. The adapter terminates the execution process and discards its volatile run state. It returns status interrupted, an unchanged snapshot, no new events, complete retained observations, and lifecycle: {mode: process-loss, instance, boundary, stopped: true}. boundary equals the interruption operation ID; instance is an opaque identity for the actual execution process instance, independent of the adapter transport process. The acknowledgement certifies that the old execution process has stopped.

restore: {from: interruption-id} starts replacement execution from that retained boundary. Its status is restored; lifecycle contains mode: process-restart, the same boundary, a different actual process instance, and restored and retained maps read by the replacement executor before resumed work. restored is a complete snapshot. Both maps must exactly match the interrupted boundary’s logical state and retained information. Ordinary events and snapshot then report any resumed work. Logical timers, subscriptions, and child links remain the same work identities; reconnecting to them emits no new start. New logical work does emit start. Timers remain fixture-controlled, including after restart.

At interruption, retained contains complete maps metadata, resolution, operations, timers, subscriptions, events, children, terminal, recovery, and superseded; inapplicable maps are empty. Additional retained information is allowed and is also compared. The observation meanings are:

Map Retained information
metadata Run path to complete run metadata; includes root.
resolution Run path to resolved definition and dependency identities/pins; includes root.
operations Attempt key to evaluated arguments, connection, actual attempt count, fixed idempotency key, dispatch status, and known result or ambiguity.
timers Step path to logical timer identities, deadlines, and active/released state.
subscriptions Step path to identities, fixed configuration/correlation, and active/released state.
events Subscription path to retained accepted payloads and delivery identities; a single-event case uses {payload, delivery} with the accepting operation ID as the delivery alias.
children Call path to child-run identity/link and pending or accepted terminal notification.
terminal Run path to outcome, accepted stop path/reason when present, successful result and suppression when present.
recovery Current recovery decision, scope, source/resulting generation and redo membership; empty before recovery.
superseded Attempt key to retained source-generation arguments, results and outcomes.

Version 2’s other observations can make partial retained assertions using the same meanings. retained.resolution denotes resolved content; operator decisions are observed as retained.operator with to, actor, and decision fields. The adapter normalizes its actual retained information into these maps. It does not use the fixture, the runner’s snapshot, or a surviving in-memory copy of the executor state as a substitute for information restored by the replacement. A test deployment can retain external connector state and buffered deliveries; these do not replace the executor’s own durable state.

An optional restore dependencies list changes availability for that restore attempt only. Each entry names a previously pinned project file path and a source: retained hides the original source and permits only retained pinned bytes; missing makes all copies of those bytes unavailable; mismatch makes only the supplied UTF-8 content available under the original identity. Other dependencies stay unchanged. The adapter applies these controls at the resolver boundary, including caches; inability to control availability is unsupported harness functionality. A refusal has status refused, unchanged state, no execution events, the replacement-process acknowledgement, and diagnostic: {code: deployment.dependency, reason: content-unavailable} or reason: digest-mismatch. This is a restoration-operation refusal, not a terminal run outcome or a generic top-level adapter error. A later restore can use the original available dependencies and finish the case.

An actual process restart, in-process reconstruction, and a controlled transcript are different evidence. Only the first satisfies a restart case against an implementation. Passing runner results label this observation evidence: adapter-reported-process-restart; other timelines use evidence: logical-timeline. The label records the adapter’s evidence, not an independent proof of its process or storage claims. An execution report identifies the implementation, adapter, process-loss mechanism and affected execution state. An adapter that can only reconstruct in process reports restart support missing. check_recovery_timeline.mjs deliberately uses controlled transcripts to test comparison failures and transport; its passes establish runner correctness only.

Version 2 also supports the following focused controls:

  • recover: {scope, actor, authorized} supplies the deployment’s decision about one recovery activation after the automatic limit has held that scope. The observation includes recovery-authorization with those fields. Denial has status refused, reason unauthorized, and unchanged execution state. Authorization precedes the delayed recovery; retained observations expose automatic_count so operator activation cannot silently spend that counter.
  • resolve can include authorized: false to exercise rejection by deployment policy. A result outside the action contract is refused with reason invalid-result. Neither refusal creates an accepted resolution event or changes execution state. A later valid authorized decision can still resolve the hold. These reasons describe control refusals, not Standard error codes.
  • storage: {available: false} gates completion of durable commits; it does not inject a particular database error. A delivery awaiting that commit reports pending, which is not an acknowledgement of durable acceptance. Restoring availability completes pending commits. The response echoes each storage control. A new hold, suspension, terminal success/cancellation, or worker-loss acknowledgement cannot claim a durable boundary while the gate blocks it. This observes logical ordering, not hardware failure tolerance.
  • fences lists action handles to pause immediately before external dispatch. The adapter installs these acquisition-only fences before starting the run. await_fence: handle observes prepared work with status reached and the corresponding event index. release_fence: handle permits dispatch and has status applied. Preparation is not dispatch; no start can precede release. This permits capture after compensation is durable but before redo can create another effect. A history-operation copy does not inherit the acquisition fence. The control prescribes no native intent-record or journal format.

The recovery-delay timer belongs to the resulting generation and can start after the recovery decision, before that generation begins compensation and redo. Ordinary new-generation work still requires the generation’s begin boundary.

Coverage includes multiple recovery generations, structural suffixes, parallel siblings, compensation order and interruption, operator resolution and recovery authorization, reconciliation, suppression, commit ordering, and restoration of waits, child links, events, stop results, dependencies and subscription generations. The focused cases and their evidentiary limits are listed in PHASE5-COVERAGE.md. Execution evidence against a supporting Durable implementation remains pending.

history negotiates workfile.history/1. This protocol identifies each operation as reconstruct, reevaluate, or continue; it has no unqualified replay request. Cases explicitly require both durable-execution and replay. A history case can also supply a version 2 timeline for acquisition. Its single negotiation covers both the history controls and that timeline; the response includes both history and the ordinary workfile.timeline/2 observations. Existing timeline versions retain their meanings.

An adapter unable to perform any requested operation, control, or observation returns harness_unsupported: ["workfile.history/1"]. Missing acknowledgement is also harness-unsupported; acknowledgement followed by incomplete observations fails. The original run’s successful result cannot satisfy a history request. The machine-readable format and comparison rules are in history.mjs.

A history fixture contains version: 1, at, expect, and a nonempty operations list. at: terminal captures the original run’s terminal durable boundary. Otherwise at names a timeline operation and captures the boundary at its observation, before the next control. Capture retains an immutable test source. Each history operation receives an independent copy of that source; continuations run in isolated test deployments with the captured run identity and controlled external inputs. They cannot change the acquisition run or supply history to another operation. The ordinary top-level response still describes only the acquisition run. These copies test history operations, not process loss.

The response is history: {protocol, source, reduction?, observations}. source is read from actual retained execution state, before any requested reduction or controlled alteration. It contains:

Member Observation
boundary The fixture’s at alias.
run_id The actual source run identity; it agrees with recorded metadata.
metadata Complete recorded root run metadata as canonical Workfile JSON.
pins Complete resolved identity map. definition is the Workfile document digest; connector:name identifies a connector, and project files use their supplied paths. Other artifacts use stable kind-qualified identities. Values are sha256: digests.
facts Complete payload-independent logical facts, described below.
payloads Complete flat inventory of retained payload occurrences in the captured history, including duplicate representations of those occurrences.

The facts map contains runs, steps, codes, decisions, generations, suppression, dispatch, superseded, stops, timers, subscriptions, children, recovery, and contexts. All are maps except the list superseded; inapplicable collections are empty. Run paths, step paths, generations, decisions, and suppression follow timeline version 2. codes preserves stable failure and ambiguity codes, including superseded occurrences under their boundary identities. dispatch maps invocation boundaries to not-dispatched, ambiguous, failed, or succeeded. superseded lists the superseded boundaries in history order. stops maps run paths to accepted {path, reason} observations.

Timers, subscriptions, children and recovery retain their timeline version 2 identity and lifecycle information. A timer includes identity and active; its deadline remains recorded. contexts records each operation’s kind, resolved identity, connection, effective attempt/timeout context and fixed idempotency key where applicable. These are operational facts even if the same key also occurs in a removed argument payload. Arbitrary workflow values belong in payloads, not in facts. Native history layout is unconstrained.

An available payload entry is {class, value}. value is a string containing canonical Workfile JSON, so 1, 1.0, -0.0, and signed int64 values retain their runtime kinds and exact values over ordinary JSON transport. It uses the Standard’s canonical serialization, including its numeric adaptation, not plain JCS. The fixture and response checkers reject noncanonical encodings. Payload classes are suite controls over retained information, not additional Standard storage requirements:

Class Inventory keys and contents
inputs input:run-path for validated workflow inputs; trigger inputs use trigger:run-path.
arguments arguments:boundary for evaluated invocation arguments.
results result:boundary, details:boundary, and binding:step-path for nondeterministic results, error details and materialized step bindings. Superseded occurrences retain their generation-qualified identities.
events event:delivery-identity for accepted event payloads. Delivery identity and acceptance remain facts.
outputs output:run-path for the selected successful workflow result, including a stop result.

Additional retained payload occurrences use unique kind-qualified keys and the corresponding class. The inventory includes all representations available to the operation, including any comparison surrogate that could replace a removed value. An adapter cannot satisfy removal by hiding only its exported representation. An additional equivalent representation identifies its primary inventory key with occurrence; its class and availability/value match that primary entry. Aliases cannot point to other aliases. Selective removal follows these identities, so a remaining comparison surrogate cannot silently supply an erased occurrence. Independent payload occurrences, such as a result and its materialized binding, keep distinct primary keys even when their values happen to be equal.

Sensitive representations already withheld at acquisition use {class, withheld: true}. A withheld projection has no value; independently retained public results remain separate entries. This marker reports unavailable information, not a workflow value or permission to discard a live input. Reconstruction can inspect such a marker with the recorded public state; an operation needing the unavailable value instead reports missing-payload and its class. Suppression markers stay in facts; removed markers never become workflow values.

Action boundaries are <step>@<generation>#<attempt>. A timer boundary is <step>@<generation>#<purpose> for the single-occurrence timers used here. root.inputs, root.outputs, and <stop-path>.result identify pure boundaries; root identifies history/header or resolution validation before evaluation. Fixtures name the exact observed boundary for refusal, divergence and invalidation. These names are portable observation aliases, not native journal keys. Child and agent boundaries use <step>@<generation>#child and #agent. An ordinary fixture’s single event uses <step>@<generation>#1 as its delivery identity; a timeline event keeps the accepting control’s ID. Successive retry timers use #retry-1, #retry-2, and so on. A step guard is <step>.when; <step>.fault observes a reproduced evaluation fault as canonical {"code":...}. State subscription paths include both the entry position and subscription name, such as waiting[1].change. Entry position and recovery generation are distinct. When both generations must appear in one retained subscription map, keys append @generation. A state projection is <state-entry>.into or .accepts; a scope join is <construct>.result. Agent tool traces use agent-tools:<agent-boundary> payloads, with proposed/dispatched arguments and results in stable position order; canned tool responses keep the existing <agent-step>[N] fixture paths.

history.expect asserts selected source facts, payloads and optionally pins. Each asserted map entry must be present. Acquisition is also compared against the selected timeline snapshot, or the original terminal outcome and result. For example, retaining a stopped result means preserving its complete projection, suppression, and numeric kinds, not recomputing root outputs. Optional ranges: [{path, min, max}] asserts a recorded numeric fact by a list of map keys from the source root. The endpoints are inclusive. Retry-jitter cases use this to accept every permitted delay, including implementations that use no jitter, while exact operation-state comparisons preserve the chosen value.

Each operation has a distinct id, an operation, an expect map, and optional controls:

  • view: original is the default. view: reduced selects the reduced copy. A selective reduction ID selects its own retained copy, as described below.
  • definition names a Workfile in files and requests deployment-authorized successor continuation. It is permitted only for continue. Without it, continuation uses the same run and original pins. Dependencies of a successor are resolved and pinned normally; expected changed pins can be asserted in expect.pins. The supplied successor is also checked against the Workfile schema.
  • context: {dry_run: boolean} changes that deployment input for a successor continuation. It requires definition, even if the supplied definition is unchanged. The observation echoes the applied context; source metadata stays immutable, and resulting operation contexts report the actual dry-run mode. Other compatibility cases change the supplied successor itself, including timeout, attempts, numeric argument kind, operation name, and named connection.
  • environment supplies successor resolution inputs. artifacts is a list of {identity, path} selecting exact documents from files for connector or overlay identities. Their document digests become active pins after successful successor creation. Referenced sample and fixture files are supplied relative to their manifests. constraints lists {owner, identity, version} exact requirements; owner is root or a call path. Connector and filter-package identities use their kind-qualified namespace; tzdb uses IANA release identifiers. These controls model deployment/project resolution requirements, not a production lock format. Incompatible graph-wide requirements refuse with version-conflict and deployment.dependency; an overlay incompatible with the selected manifest refuses with invalid-artifact and dependency.invalid. The adapter applies normal resolution and validation, acknowledges the exact supplied inputs, and reports actual active pins. It cannot substitute an old source artifact to make the successor succeed. Refusal precedes evaluation, dispatch, or successor creation and leaves the source intact.
  • resolve lists continuation-time {boundary, actor, outcome, result?, code?} operator decisions. Successful results use canonical Workfile JSON; failure supplies a stable code. Each decision is presented only when the continued run is held for that boundary. The response acknowledges these controls and records matching resolve decisions after its durable live boundary, before resumed work. Reconciliation instead uses the successor’s declared steps and ordinary continuation action responses. Neither control treats an ambiguous original invocation as absent or as permission to repeat it.
  • dependencies temporarily changes availability at the operation’s resolver boundary, including caches. Entries are {path, source, content?} with the same retained, missing, and mismatch meanings as restoration controls. A missing executable connector does not prevent reconstruction when its recorded identity and all required state remain intelligible. Re-evaluation can use retained pinned content when the original source is unavailable.
  • override replaces named payload values in a disposable input history with supplied canonical JSON strings. The payload must already exist. The adapter preserves its declaration/type and structural history integrity, including native checksums. Argument/output overrides provide structurally valid alternate historical observations whose mismatch is discovered by pure re-evaluation. They are not changes to the pinned workflow or an instruction to force a branch.
  • damage: unsupported-standard changes only the input copy’s history version to unsupported 999.0.0. damage: inconsistent-history makes a definite invocation boundary contradictory by retaining both success and failure for that same identity. The original source and its normalized projection remain intact; the separate acknowledgement records the applied integrity control.
  • actions supplies continuation-only canned responses keyed by boundary: {boundary, outcome, result?, code?}. Results are canonical JSON strings. timers lists retained timer identities made eligible after live execution begins. These controls cannot supply missing recorded history or authorize reissue. Unlisted external work remains observable and fails the effects check.

An operation can use expect: {any_of: [...]} to list complete permitted observations. Each alternative has the ordinary expectation shape; the judge accepts one whole alternative, never a mixture of its evaluations, decisions, effects and results. An optional source: {facts?, payloads?} inside an expectation restricts it to the actual selected source history. Stop-race cases use this to preserve whichever result won during acquisition. Independent-branch cases admit both safe preservation and conservative invalidation, and both orders of independent live effects.

Optional compatibility assertions compare an exact list of {boundary, dimension, compatible} findings. Dimensions are generation, artifact, connection, arguments, dry-run, timeout, attempts, profile-inputs and result. These expose which part of a boundary matched, including when an artifact change also changes its result contract. A superseded candidate can report a generation mismatch without evaluating its old values. Missing findings fail comparison.

One observation per operation, in order, contains id, operation, status, input, source_after, run_id, pins, effects, evaluations, and decisions. input is the selected normalized source after payload overrides; dependency and damage controls are acknowledged separately with the exact supplied fields. source_after rereads the immutable acquisition source. Both projections are compared in full, so skipping a reduced view, ignoring an override, or mutating original history fails. The controlled copies cannot consult the intact source to fill unavailable payloads.

Successful reconstruction has status reconstructed; successful re-evaluation has status matched. Both return state: {facts, payloads} exactly matching the selected input. Argument payloads are inspectable history but are not needed to materialize recorded bindings; removed argument markers can therefore remain in a reconstructed observation. A required missing binding/input/result cannot be replaced by a marker or null to fabricate successful reconstruction.

evaluations reports pure boundary evaluations in order as {boundary, value, after}, using canonical JSON for evaluated values. It includes invocation argument projections, workflow results, stop results, guards and other pure decisions actually evaluated. after counts the decisions already recorded; evaluation of a reused boundary precedes its reuse decision. Successful cases assert an ordered subsequence, allowing additional pure evaluations. Divergence cases assert the complete boundary prefix through the first mismatch, and prohibit evaluation beyond it. Subexpression instrumentation is not part of this list.

effects reports every non-pure operation as {kind, boundary, arguments?, after}. Kinds include action, child, agent, clock, event, timer, retry, reconciliation and compensation; filter and expression also expose otherwise forbidden execution during reconstruction. Arguments use canonical JSON. after is the count of decisions already recorded when that operation occurred. Effects and decisions are exact ordered lists within each expectation. Cases with permitted scheduling or retention choices use any_of to admit complete alternative observations.

Reconstruction permits no evaluations, effects, or continuation decisions. Re-evaluation permits pure evaluations but no effects or continuation decisions. An adapter observes these prohibitions at execution boundaries, including clocks, filter invocations and event admission; an empty list certifies that none occurred.

Continuation returns continued after completing the fixture-controlled live suffix, or held when operator resolution is required. Its state is the complete resulting facts/payload inventory; expect.state compares selected members. decisions records ordered {kind: reuse, boundary}, {kind: invalidate, boundary, discarded: [...]}, and {kind: live, boundary, complete: true} observations. A live decision acknowledges that all required active state is durable before external work begins. The judge rejects an effect whose after places it before that decision, superseded reuse, reissue of a reused action/child/agent, and work while held. Expiring a reused timer is permitted after the live boundary; it does not create a new timer. The discarded suffix is exact. Original history remains intact even after invalidation. A successor has a distinct run identity, its actual pins, and lineage: {parent: source-run-id, boundary: source-boundary}.

refused reports unusable input; diverged reports a pure-evaluation mismatch. Neither yields a successful state or permits live execution. The exact diagnostic identifies the earliest boundary and reason. missing-payload includes its class; divergence includes canonical recorded and evaluated values. content-unavailable and digest-mismatch use registered code deployment.dependency; unsupported-standard uses deployment.version. This revision allocates no generic history-integrity or divergence error code: inconsistent-history, missing-payload, and divergence are protocol reasons, not new Standard codes or terminal run failures. When pure evaluation needs an unrecorded nondeterministic boundary, a diverged observation uses reason: unrecorded-boundary, its exact boundary, and the canonical evaluated invocation or timer request. There is no fabricated recorded value. The evaluation list ends at that boundary. This differs from known recorded information that has become unavailable through reduction or withholding, which produces a refused observation with missing-payload.

The existing top-level reduce control now requires history assertions. It is permitted only with at: terminal. After acquisition, the adapter removes every payload occurrence of the requested classes from a disposable retained copy and returns that complete copy as reduction. Every removed entry is {class, removed: true}; all other entries, pins, identity and facts remain unchanged. The judge derives and compares this projection from the actual source, rather than trusting a reported list of removed classes. At least one operation must request the reduced view. Each operation starts with its own copy of that view, so earlier inspection cannot restore removed data.

Optional history.reductions requests selective or successive reductions. Each entry is {id, from, remove}: from selects original, the class-reduced view, or an earlier successful reduction ID; remove lists distinct primary payload keys. IDs are unique and cannot be original or reduced. The response’s reductions map contains the complete resulting source for each successful ID. Every selected occurrence and its aliases become removed markers; other entries, identity, pins and facts stay unchanged. A second reduction cannot restore data removed by the first. Every successful view has an operation assertion. Unknown occurrences, alias selectors and forward references are fixture errors.

Nonterminal requests instead include refusal: {reason: nonterminal-required, boundary, class} and select information still required for resumption. The response reports reduction_refusals[id]: {diagnostic, source_after} with that diagnostic and the unchanged selected source; no reduced view is created. The case then continues from the original history to verify it remains usable. This tests the nonterminal retention minimum without instructing an adapter to destroy required history. This control currently covers refusal of an exact removal request, not partial approval or safe nonterminal garbage collection. Withheld entries remain withheld through either reduction form.

The cases establish operation-specific consequences. Removing arguments alone leaves recorded state reconstructible, while invocation compatibility cannot be checked after removing both the arguments and any equivalent comparison data. Removing inputs and materialized results can prevent reconstruction as well. Retaining successful outputs does not establish re-evaluability. The controls do not grant permission to reduce nonterminal history below resumption requirements or impose a terminal-retention policy on implementations.

Runner passes use evidence: adapter-reported-history-operations. This label reports the adapter’s observations, not independently verified storage behavior. check_history.mjs uses controlled positive and negative transcripts and actual runner subprocesses to test comparison and negotiation. They are judge tests, not implementation execution evidence. See PHASE5-COVERAGE.md for the individual Replay audit and remaining late-result, recovery, restart and reduction gaps. check_portable_history.mjs adds targeted comparison and transport checks for the broader coverage. These optional controls retain workfile.history/1; an adapter that cannot apply any requested control reports the existing harness_unsupported result. Acknowledgement followed by an ignored reduction, context change, required marker, or boundary observation fails comparison.

Cases in run-records/ negotiate workfile.run-records/1 with request kind run-records. They apply only to Core Executors claiming wf.run-records and any requires_profiles listed by the case. The response contains run_records: {protocol: "workfile.run-records/1", observations: [...]} or harness_unsupported: ["workfile.run-records/1"]. Missing acknowledgement is harness-unsupported; a wrong version, mixed acknowledgement, or incomplete observations after acknowledgement fails. A pass supplies evidence: adapter-reported-run-record-operations, not independent evidence of storage behavior. These controls do not define a production service binding.

Each case has case, requirements, capability: wf.run-records, workflows, runs, and operations, with optional requires_profiles and files. workflows maps fixture aliases to complete Workfiles. These aliases identify stable deployment workflows; two aliases remain distinct even when their names or definitions agree. runs maps run aliases to {workflow, definition?, inputs?, dry_run?, actions?, parent?}. Optional definition selects a supplied project file as a new revision of that same workflow identity; otherwise its workflows definition is used. Action scripts and project files use the trace-case conventions. Optional parent: {run, path} binds a child alias to the child actually started by that parent’s call occurrence. Child aliases are never started independently. The fixture catalog and default connections are the same as for trace cases. Adapters must preserve source numeric spellings when loading these Workfiles and payloads.

The adapter uses an isolated visibility scope with no unrelated runs and keeps records through the requested observations. Each operation has a unique id and operation. The following are suite controls over ordinary execution and retention, not additional capability operations:

Control Inputs Result
start Root run alias; optional at Start its configured workflow with the supplied inputs and scripts. Return a map from that alias and each child alias actually started to their actual current run summaries.
advance Started root run alias; optional at Allow the run and supplied action script to progress; return the same summary map updated for the reached boundary.
reduce Terminal run alias and nonempty classes Remove retained payloads of those classes and return the complete resulting run record. A preceding get-run supplies the comparison baseline.
evict Terminal run alias Evict the record, retain evidence of eviction, and return null.
forget Evicted run alias Delete the remaining evidence and return null.

at: started pauses immediately after run start, before the first root step begins; at: settled (the default) advances to a terminal outcome or a held or suspended boundary with no immediately runnable work. Between controls no run progress, timer delivery, or eviction occurs. Each child alias that has started must appear in its parent’s control result. Identity values come from the implementation, not the fixture aliases. reduce changes every retained representation of the selected optional payloads, preserving mandatory facts, withheld and not-recorded states, suppression paths, and unselected values. An implementation need not offer retention controls through its public API; an adapter unable to establish a requested fixture condition reports harness-unsupported. No case mandates recording optional payloads: assertions admit not-recorded where the implementation never retained them.

The three capability operations use list (optional workflow alias), get-run (run alias), and get-step (run alias and path). Before invoking them, the adapter resolves aliases to the actual identities observed by the controls. $unknown denotes a fresh run or workflow identity not assigned in this fixture scope. A list’s members is the complete set of expected run aliases, independent of order. refusal asserts exactly run-not-found, run-not-retained, step-not-found, or invalid-request; these are observation reasons, not additions to the Workfile Error Code Registry.

Every observation contains the matching id and operation, and an effects list of work dispatched while performing it (trace-style {kind, step?} observations). The three inspection operations require this list to be empty. Successful observations contain result, a string containing Workfile canonical JSON: the summary map for execution controls, the summary list for list, the document or step record for lookup, and null for eviction controls. Refusals instead contain refusal and no result. Encoding the result as text preserves exact numbers through the line protocol’s JSON parser.

expect asserts a recursive subset of document or step facts. entries asserts an ordered subsequence of entry patterns. A payloads assertion has path (a list of keys locating a slot in the returned document), states (the allowed availability states), and optional value (canonical JSON text encoding the expected typed value). If a value is present it is compared exactly, including numeric runtime kinds. Availability assertions always require the slot to exist. The judge checks schema shape, canonical spelling, metadata and identity stability, definition pins when present, complete listings, step/run agreement, ordered generation and attempt facts, and unchanged snapshots between controls. It checks reductions against every representation in the preceding record. Unknown structural x- extensions do not participate in comparisons; keys inside user payload maps retain their ordinary value semantics.

check_run_records.mjs tests these comparisons with positive and adversarial transcripts and runner subprocesses. Such transcripts test the suite; only observations supplied by an implementation exercise that implementation.