Skip to content

CF Durable Object drive pipeline (WA MR1) ​

This guide covers the release in which every run a Cloudflare Durable Object (DO) starts or recovers goes through one pipeline (runDrive): /start, /resume, /retry, the /subagent/:type/* protocol routes, the client-tool submit that continues a paused run, and the wake that recovers an evicted run. Each of them now resolves the agent, builds the executor, classifies a failure, fires the host hooks and settles the run the same way.

It is a major release of @helix-agents/runtime-cloudflare, @helix-agents/runtime-temporal (its AgentNotFoundError is removed), @helix-agents/runtime-js, @helix-agents/ai-sdk and @helix-agents/agent-server (its INVALID_RESULT answer and two typed /resume answers), and a minor (0.x) release of @helix-agents/core and the four stores. Nothing is deprecated: every removed API is deleted.

Design: 2026-10-08-do-drive-pipeline-design.md (read its "§0 As built (MR1)" first). Runtime page: Cloudflare Durable Objects.

Hooks ​

The seven server hooks collapse into three, and the hook-facing ExecutionState is read-only.

RemovedReplacement
beforeStart, beforeResume, beforeRetrybeforeEntry(ctx), one hook; ctx.drive is 'start' | 'resume' | 'retry'
afterStart, afterResume, afterRetryafterEntry(ctx), once per drive, wakes included
onCompleteonRunSettled(ctx), at least once per (runId, status)
ExecutionState.interrupt() / abort() / getHandle()concurrentEntry: 'interrupt_previous', or EndpointContext.stop()
BeforeStartContext … AfterRetryContext, CompleteContextBeforeEntryContext, AfterEntryContext, RunSettledContext, SettledStatus

Before:

typescript
hooks: {
  beforeStart: async ({ request, body, executionState }) => {
    if (!(await isAllowed(request))) {
      return { abort: true, response: new Response('Forbidden', { status: 403 }) };
    }
  },
  afterStart: async ({ sessionId, runId, agentType }) => log('started', sessionId, runId, agentType),
  onComplete: async ({ sessionId, runId, status, error }) => notify(sessionId, runId, status),
},

After:

typescript
hooks: {
  beforeEntry: async (ctx) => {
    // ctx.origin.request is the original request (headers intact).
    if (!(await isAllowed(ctx.origin.request))) {
      return { abort: true, response: new Response('Forbidden', { status: 403 }) };
    }
  },
  afterEntry: async ({ sessionId, runId, drive }) => log('entered', sessionId, runId, drive),
  onRunSettled: async ({ sessionId, runId, status, errorDetail, deliveryId }) => {
    // At least once per (runId, status): deduplicate on deliveryId.
    if (await alreadyProcessed(deliveryId)) return;
    await notify(sessionId, runId, status, errorDetail);
  },
},

What changed in the semantics:

  • beforeEntry is admission only (auth, rate limits, quotas). It runs for /start, /resume, /retry and the /subagent/:type/start|resume routes (ctx.origin.route is 'native' or 'subagent'), and only AFTER the busy check, the identity / type gate and agent materialization. A refused entry never reaches it, and it never sees a live run to interrupt. Its context is per drive:

    • start: body, the transformed, validated start body (transformRequest's result on the native /start). The raw body stays readable at ctx.origin.request.
    • resume: options (a DoResumeOptions; a with_message resume carries the message).
    • retry: checkpointId? and message?, the retry's new user message.
    • every drive: env, origin and identity (the session's userId, tags, metadata, parentSessionId, rootSessionId).

    Return { abort: true, response } to refuse: the response is answered verbatim and nothing is written. A throw answers framework_hook_failed, retryable (503) unless its cause is classified non-retryable.

  • Submits are outside beforeEntry. /submit-tool-result stores caller-supplied tool results and approval decisions, and the continuation it starts puts them in front of the model; no host hook admits it. Authorise submits in the worker in front of the DO (FU-DO-SUBMIT-ADMISSION-HOOK tracks a possible hook).

  • afterEntry fires once per drive, after the drive owns its run: after the run-start commit for an entry, after the in-memory claim for a wake's recover. Its context is { env, sessionId, runId, drive, handle }. It is not awaited by the drive; a throw is logged.

  • onRunSettled fires whenever a run reaches a status at which nothing executes for it (suspended_client_tool, suspended_awaiting_children, suspended_step_partial, completed, failed, interrupted, superseded), on every path: request drives, the submit continuation and wakes. Its context is { env, sessionId, runId, drive?, status, result?, errorDetail?, deliveryId }, with deliveryId = ${runId}:${status}.

    • At least once. A durable mark (DO KV __helix_settle_delivered:{runId}:{status}) is written after the hook returns, and a wake redelivers a settled run whose mark is missing (an eviction between the commit and the hook). Deduplicate on deliveryId.
    • A throwing hook is logged and still marked, so it is never retried forever.
    • A paused run settles once, suspended_*. The built-in routes cannot stop it (/interrupt answers 400 for a paused session and /abort writes nothing for it). A custom endpoint's EndpointContext.stop('interrupt') sets the durable interrupt flag on the paused session: nothing ends at once, and the next continuation (a submit or a resume) honours it, ending that run interrupted at its first boundary. Stopping a paused run at once is MR2.
    • It runs after the run's stream is finalized and after the execution registration is released, so a new entry is never refused because a slow hook is still running.
    • With no onRunSettled configured, no mark is written.
  • transformRequest / transformResponse are unchanged in purpose and apply to the native /start only. A transformRequest throw is classified: a HelixError keeps its code, a Zod failure answers 400 validation_error, anything else answers framework_hook_failed (503 unless its cause is non-retryable). Its result type is StartBody (formerly StartAgentRequestV2). transformResponse receives StartResponseBody ({ sessionId, streamId, runId, status: 'started' | 'resumed', startSequence? }).

  • ExecutionState keeps sessionId, runId, isExecuting and status, read-only. isExecuting is true from just before a drive's executor call until the settle path releases it, which happens before onRunSettled.

  • Custom endpoints stop a run with EndpointContext.stop(kind, reason?) ('interrupt' or 'abort'), the mechanism of the /interrupt and /abort routes. In this release it waits for the run to settle, as the routes do: a tool that ignores its abort signal keeps the call waiting (the stop MR changes this).

Before:

typescript
endpoints: {
  '/cancel': async (request, { executionState }) => {
    await executionState.abort('cancelled by user');
    return Response.json({ ok: true });
  },
},

After:

typescript
endpoints: {
  '/cancel': async (request, { stop }) => {
    await stop('abort', 'cancelled by user');
    return Response.json({ ok: true });
  },
},

Concurrent entries ​

"Last message wins" is a config option. It replaces interrupting the live run from beforeStart.

Before:

typescript
hooks: {
  beforeStart: async ({ executionState }) => {
    if (executionState.isExecuting) {
      await executionState.interrupt('New message received');
    }
  },
},

After:

typescript
createAgentServer<Env>({
  llmAdapter: ({ env }) => createAdapter(env),
  agents: registry,
  concurrentEntry: 'interrupt_previous', // default: 'reject'
  concurrentEntryTimeoutMs: 30_000, // default
});
  • 'reject' (the default): a /start, /resume or /retry that arrives while another drive executes a run on the DO answers 409 state_already_running.
  • 'interrupt_previous': the new entry first passes the identity / type gate, materialization and beforeEntry. Only then is the executing run interrupted (reason superseded_by_new_entry), and the entry waits for it to settle, bounded by concurrentEntryTimeoutMs. A refused or unadmitted entry never interrupts a live run. If the old run has not settled at the bound, the entry answers 409 state_already_running; the interrupt still applies.
  • A run whose result has already settled (its drive is in its tail) is waited for, not interrupted, under either policy.
  • Wake drives never interrupt a run.

Resolver and factories ​

The resolver may be async, and every factory receives one context object: the session's persisted identity, the drive kind, and the run id only when a run is about to be driven.

RemovedReplacement
LLMAdapterFactory = (env, ctx) => …LLMAdapterFactory = (ctx) => …, ctx: DriveFactoryContext & { agentType }
UsageStoreFactory = (env, ctx) => …UsageStoreFactory = (ctx) => …, ctx: DriveFactoryContext | UsageReadContext
LLMAdapterContext, UsageStoreContextDriveFactoryContext, UsageReadContext
AgentFactoryContext { env, sessionId, userId?, runId }DriveFactoryContext | LookupFactoryContext
a synchronous AgentResolverAgentResolver returning AgentConfig | Promise<AgentConfig>
typescript
interface AgentIdentity {
  userId?: string;
  tags?: string[];
  metadata?: Record<string, string>;
  parentSessionId?: string;
  rootSessionId?: string;
}
type DriveKind = 'start' | 'resume' | 'retry' | 'continue' | 'recover';
interface DriveFactoryContext<TEnv> {
  purpose: 'drive';
  env: TEnv;
  sessionId: string;
  runId: string;
  drive: DriveKind;
  identity: AgentIdentity;
}
interface LookupFactoryContext<TEnv> {
  purpose: 'lookup';
  env: TEnv;
  sessionId: string;
  identity: AgentIdentity;
  reason: 'submit_validation' | 'retention';
}
interface UsageReadContext<TEnv> {
  purpose: 'usage_read';
  env: TEnv;
  sessionId: string;
  identity: AgentIdentity;
}

Before:

typescript
const registry = new AgentRegistry();
registry.registerFactory<Env>('assistant', (ctx) =>
  createAssistant({ db: ctx.env.DB, userId: ctx.userId })
);

createAgentServer<Env>({
  llmAdapter: (env, ctx) => new VercelAIAdapter({ anthropic: makeAnthropic(env, ctx.agentType) }),
  agents: registry,
  usageStore: (env, ctx) => new BillingUsageStore(env.BILLING_DB, ctx.sessionId),
});

After:

typescript
const registry = new AgentRegistry();
registry.registerFactory<Env>('assistant', async (ctx) =>
  createAssistant({
    db: ctx.env.DB,
    userId: ctx.identity.userId,
    prompt: await loadPrompt(ctx.env),
  })
);

createAgentServer<Env>({
  llmAdapter: ({ env, agentType }) =>
    new VercelAIAdapter({ anthropic: makeAnthropic(env, agentType) }),
  agents: registry,
  usageStore: ({ env, sessionId }) => new BillingUsageStore(env.BILLING_DB, sessionId),
});
  • The resolver sees a run id only with purpose: 'drive'. A purpose: 'lookup' call (validating a submitted client-tool result against the owning tool's outputSchema) starts nothing. Narrow on ctx.purpose before you read ctx.runId.
  • The identity comes from the session row on every drive (wakes, resumes, retries and companion children included), not only on /start. Async setup that used to live in beforeStart belongs in the resolver.
  • A resolver failure is typed. Throw core AgentResolutionError.notFound(agentType, available) for an unknown type: the native routes answer 404 framework_agent_not_found, the /subagent routes 404 NOT_FOUND, and a wake ends the run typed. Any other resolver or adapter-factory throw becomes framework_agent_resolution_failed, with the original as cause and its retryability: an unclassified throw is transient (a request answers 503, a wake re-arms).
  • A sub-agent tool without outputSchema fails materialization with validation_error (400 on a request; a wake ends the run failed). It used to throw an untyped error, and on a wake it made every later request to the DO fail.
  • persistentAgents without subAgentNamespace fails with framework_not_supported (500) instead of a 400 with the ad hoc code persistent_agents_no_subagent_namespace.
  • The usage-store factory is also called to read usage, with no drive: the /subagent/:type/usage rollup, and the end-of-stream rollup of a run finished on a wake. That call gets UsageReadContext (purpose: 'usage_read').

Strict usage store ​

A consumer usage-store factory that throws stops the entry. It no longer falls back to the DO's internal store, and usageStoreFactoryNegativeCacheMs is deleted.

Before:

typescript
createAgentServer<Env>({
  // …
  usageStore: (env, ctx) => new BillingUsageStore(env.BILLING_DB, ctx.sessionId),
  // A throwing factory fell back to the internal DOUsageStore for 5 s.
  usageStoreFactoryNegativeCacheMs: 5_000,
});

After:

typescript
createAgentServer<Env>({
  // …
  // A throw is `framework_usage_store_unavailable` (retryable): a request
  // answers 503, a wake re-arms. Usage is never split across two stores.
  usageStore: ({ env, sessionId }) => new BillingUsageStore(env.BILLING_DB, sessionId),
});
  • With no usageStore, the internal DOUsageStore is used. It is bound to the DO's session: recordEntry for another session throws validation_error and writes nothing, and getEntries / getRollup read only that session's rows.
  • A consumer usage-DB outage now stops runs (503 / re-arm) instead of silently splitting usage.

Resume body ​

The native /resume body is { agentType, options: DoResumeOptions }, a strict union that mirrors core ResumeOptions. Every field it accepts is honoured; anything else is a 400 validation_error.

typescript
type DoResumeOptions =
  | { mode: 'continue' }
  | { mode: 'with_message'; message: string | UserInputMessage[] }
  | { mode: 'from_checkpoint'; checkpointId: string };

Before:

typescript
await stub.fetch('https://do/resume', {
  method: 'POST',
  body: JSON.stringify({
    agentType: 'assistant',
    options: { mode: 'continue', appendMessages: [{ role: 'user', content: 'Go on' }] },
  }),
});

After:

typescript
await stub.fetch('https://do/resume', {
  method: 'POST',
  body: JSON.stringify({
    agentType: 'assistant',
    options: { mode: 'with_message', message: 'Go on' },
  }),
});
RemovedReplacement
options.mode: 'retry'POST /retry { agentType, checkpointId?, message? }
options.mode: 'branch'none: the DO cannot branch (see below)
options.modifyStatePOST /start with input.state
options.appendMessages{ mode: 'with_message', message }
options.checkpointId beside mode: 'continue'{ mode: 'from_checkpoint', checkpointId } (it used to be ignored)
an omitted mode (e.g. options: {}), which used to mean continue{ mode: 'continue' } (an omitted mode is now a 400 validation_error)
ResumeAgentRequestV2, StartAgentRequestV2, RetryAgentRequestV2DoResumeOptions, StartBody, the /retry body above

DO branching is unsupported. A Durable Object cannot read another DO's session (DOStateStore.cloneSession refuses), so a POST /start whose body carries branch answers 500 framework_not_supported at parse, before any read or write. It used to reach the executor and fail there, often as a retryable 503. Branch on CF Workflows, or start a new session.

DOFrontendExecutor.resume and DOWorkflowExecutor.resume send this body. A with_confirmation resume is refused typed before any request: approvals go through /submit-tool-result (approval-response).

Identity is fixed when the session is created. A /start on an existing session whose body names a different userId, tags or metadata answers 400 validation_error (the error's field names which). Omitted fields keep the session's values. A host that sent per-turn metadata must stop.

The persisted agent type decides. /resume, /retry and a /start on an existing session whose agentType differs from the session's answer 409 state_agent_type_mismatch (the body's type used to be trusted).

Wire answers ​

The native routes (/start, /resume, /retry) answer every refused entry with core typedErrorResponse's body, { error, code, cause? }, and the status of core httpStatusForHelixError. A request drive that fails before its run-start commit writes nothing (no status, version, stream or hook write), including /resume, which used to answer { status: 'resumed' } at once and fail later.

SituationBeforeNow
the DO is executing a run (busy)409 { error: 'Agent already running', code: 'ALREADY_RUNNING', sessionId, streamId }409 { error, code: 'state_already_running' } (no streamId)
a persisted live run that nothing executes409 busy (the persisted-status precheck)accepted: the run-start commit supersedes the orphaned run
a rejected run start (RunStartRejectedError)409 state_run_start_conflict, no cause409 { error, code: 'state_run_start_conflict', cause }; cause is one of the six reject causes
a resume the session's status refuses409 busy, or a detached failure409 { error, code: 'state_not_resumable' }; a partial batch's message names the awaited call ids first
/retry on a session that is not failed409 { error, currentStatus }409 { error, code: 'state_not_resumable' } (no currentStatus)
an unknown agent type400 { error: "Unknown agent type: …" }404 { error, code: 'framework_agent_not_found' }
a resolver / factory throw, unclassified400 Unknown agent type503 framework_agent_resolution_failed (409 / 500 once the input committed)
/resume / /retry with no session on the DO400 { error: 'No agent state to resume' }409 { error, code: 'state_session_not_found' } (state category, not retryable)
an agent-type mismatchaccepted409 state_agent_type_mismatch
an unclassified failure before the commit500503 (retryable)
framework_not_supported (a server misconfiguration)400 / 500 ad hoc500

Once the entry's input committed, a retryable error is answered with its non-retryable row (409 for a state_* code, else 500), never 503: a re-sent request would append the input twice.

The /subagent/:type/start|resume routes keep the remote-protocol envelope (RemoteAgentErrorResponse), whose uppercase codes are the protocol's own. The typed code now rides as errorCode, and a rejected run start's reason as cause:

AnswerEnvelope
busy (the DO is executing a run)409 ALREADY_RUNNING, no streamId (it used to carry the persisted one)
a persisted live run, nothing executing200, a new run (it used to answer 409 ALREADY_RUNNING)
unknown agent type, unknown session404 NOT_FOUND
state_not_resumable, completed session409 ALREADY_COMPLETED, errorCode: 'state_not_resumable'
state_not_resumable, otherwise409 INVALID_REQUEST, errorCode: 'state_not_resumable'
a caller error (400)400 INVALID_REQUEST, errorCode
anything elsethe typed status (409 / 500 / 503), INTERNAL_ERROR, errorCode, cause?

A /subagent/:type/start body whose agentType differs from the path answers 400. The route parses the protocol body directly, so the original request's headers reach beforeEntry (they used to be dropped).

The /start answer is { sessionId, streamId, runId, status: 'started' | 'resumed', startSequence? } (/resume and /retry answer status: 'resumed'). startSequence is the stream head at the run-start commit, omitted only when the run record could not be read. /subagent/:type/start|resume answer the same body (a RemoteStartResponse plus status and startSequence).

Busy means the DO is executing a run. The persisted-status precheck is removed: the run-start commit is the gate. A /start (or /subagent/:type/start) on a session whose persisted run is live but which nothing executes (its execution was lost, e.g. to an eviction) is accepted, and its run-start commit supersedes the orphaned run. The busy 409 no longer carries streamId (it used to send the persisted one from that precheck); read /snapshot to attach.

/retry on a session that is not failed answers 409 { error, code: 'state_not_resumable' } (runtime-js retry() throws ResumeRefusedError). It used to answer 409 { error, currentStatus } from a precheck. The protocol routes have no retry; their resume refusals are the state_not_resumable rows above (ALREADY_COMPLETED for a completed session, else INVALID_REQUEST).

Every host that answers with core typedErrorResponse (agent-server's /start and /resume, handleChatStream) now answers framework_agent_not_found with 404 instead of 500, through the new row in httpStatusForHelixError.

agent-server's /resume answers two races typed. With a runtime-js executor, a /resume whose session is deleted after agent-server's own session check, or whose agents map key differs from that agent's name, now answers 409 { error, code: 'state_session_not_found' } / 409 { error, code: 'state_agent_type_mismatch' } (both non-retryable), through the typed runtime-js refusals (see Errors). Both used to answer 500 { error, code: 'INTERNAL_ERROR' }. The message text is unchanged.

The usage_store_factory_fallback counter is gone (there is no fallback): nothing emits it, and core's USAGE_STORE_FACTORY_FALLBACK_COUNTER export is removed. Remove dashboards or alerts built on it, and alert on framework_usage_store_unavailable instead.

A new entry right after a completed run is accepted. A /start that arrives between a run's terminal K5 commit and its terminal run status used to answer the busy 409; the run-start commit now finishes that run itself and starts the new one.

Before (a caller matching the old envelope):

typescript
const res = await stub.fetch('https://do/start', { method: 'POST', body: JSON.stringify(body) });
if (res.status === 409) {
  const { code, streamId } = await res.json();
  if (code === 'ALREADY_RUNNING') return attach(streamId);
}

After:

typescript
const res = await stub.fetch('https://do/start', { method: 'POST', body: JSON.stringify(body) });
if (!res.ok) {
  const { error, code, cause } = (await res.json()) as {
    error: string;
    code: string;
    cause?: string;
  };
  if (code === 'state_already_running') return retryLater(); // no streamId: read /snapshot to attach
  throw new Error(`${code}${cause ? ` (${cause})` : ''}: ${error}`);
}

Scrubbed error text ​

Every free-text field handleChatStream writes for a failure is scrubbed with core sanitizeErrorMessage / sanitizeErrorDetail (control characters stripped, < / > escaped, known secrets redacted, at most 256 characters). The code, category, retryable and cause structure are kept, so parseHelixChatError reads the same code.

AnswerWhat a client now sees
409 state_already_running (busy){ error, code, retryable: true }; error scrubbed (a plain AgentAlreadyRunningError's bytes are unchanged)
409 state_run_start_conflict{ error, code, cause }; cause is present only when it parses as one of the six reject causes (core RunStartRejectCauseSchema); free text is dropped
409 state_not_resumable and every other typed entry failure{ error, code } through core typedErrorResponse (scrubbed)
409 stream_failed (resuming a failed stream)message and errorDetail scrubbed at every level of the cause chain
the error chunk when the transformer itself throws{ type: 'error', errorText }, errorText scrubbed; the shape is unchanged

Before:

json
{
  "error": "Agent already running: postgres://app:hunter2@10.0.0.5/db\n    at /srv/app/store.js:12:5",
  "code": "state_already_running"
}

After:

json
{
  "error": "Agent already running: postgres://[REDACTED]    at [path][ip]",
  "code": "state_already_running",
  "retryable": true
}

The run's own error text the stream transformer forwards (a tool's error, a run's error chunk) is unchanged (FU-AISDK-RUN-ERROR-TEXT-UNSCRUBBED).

The CF Durable Object's own answers are scrubbed the same way: the /subagent/:type/usage 503 (the consumer usage store's error text) and the /snapshot 500 keep their code and shape with the message scrubbed.

INVALID_RESULT no longer reveals the tool's outputSchema, on every host. A /submit-tool-result INVALID_RESULT 400 carries the Zod issues and the detailed message only with verboseValidationErrors: true, the gate invalid_payload's details already used. By default it is { error: 'The submitted result failed the tool output validation', code: 'INVALID_RESULT', toolName, toolCallId } (the error is core's new INVALID_RESULT_MESSAGE). This holds on the CF Durable Object and on agent-server; the ai-sdk submitToolResultErrorResponse always answers this body (it has no verbose option). A client that read issues or parsed the message must turn the option on where it exists. The code, the status and toolName / toolCallId are unchanged.

handleChatStream's send, resume and client-tool submit paths now share one classification: any failure core can classify (a HelixError, a raw core error with a known code, a rejected run start, a busy refusal, or an error wrapping one) ends the request with the typed HTTP answer. A resume or submit that hit a busy refusal or a RunStartRejectedError used to answer 200 with a data-resume-rejected frame; it now answers the typed 409. Only an unclassifiable error keeps the per-intent data-resume-rejected frame.

Errors ​

AgentNotFoundError is deleted from @helix-agents/runtime-cloudflare and @helix-agents/runtime-temporal. Both registries throw core AgentResolutionError.notFound (framework_agent_not_found, non-retryable, with agentType and availableTypes). Match it with the core guards, which check name and code, so an error from a second copy of core still matches.

Before:

typescript
import { AgentNotFoundError } from '@helix-agents/runtime-cloudflare'; // or runtime-temporal

try {
  registry.get(agentType);
} catch (err) {
  if (err instanceof AgentNotFoundError) return notFound(err.availableTypes);
  throw err;
}

After:

typescript
import { isAgentNotFoundError } from '@helix-agents/core';

try {
  registry.get(agentType);
} catch (err) {
  if (isAgentNotFoundError(err)) return notFound(err.availableTypes);
  throw err;
}
  • isAgentResolutionError(e) matches both codes (framework_agent_not_found, framework_agent_resolution_failed).
  • AgentRegistry.get() returns static registrations only; for a factory-registered type it throws notFound with a reason that says to use resolve().
  • New ErrorCodes: framework_agent_not_found (404), framework_agent_resolution_failed, framework_auto_resume_exhausted, framework_usage_store_unavailable, framework_hook_failed, state_agent_type_mismatch (409). A switch over ErrorCode that must be exhaustive gains these cases.

ErrorDetail.cause is string | ErrorDetail, nested at most 8 deep (ERROR_DETAIL_MAX_CAUSE_DEPTH). It appears in stream fail frames, /status, the remote protocol and the persisted error_detail. resolveErrorDetail records a classified error's wrapped cause as its own detail, and toErrorDetail keeps a boolean retryable property of a raw Error.

Before:

typescript
const why: string | undefined = result.errorDetail?.cause;

After:

typescript
const cause = result.errorDetail?.cause;
const why = typeof cause === 'string' ? cause : cause?.message;
const causeCode = typeof cause === 'object' ? cause.code : undefined;

The DO clients throw typed errors. DOFrontendExecutor and DOWorkflowExecutor read a native answer with core PeerErrorBodySchema and no longer throw DOPeerError for an entry failure (DOPeerError is now only DOStateStoreClient's read error):

AnswerThrown
409 state_already_runningAgentAlreadyRunningError
404 framework_agent_not_foundAgentResolutionError for the requested type (no availableTypes)
409 state_run_start_conflict + a known causeRunStartRejectedError(sessionId, cause)
any other typed bodya HelixError with the body's code and cause (helixErrorFromPeerBody); retryable when the body says so, else exactly on a 503

Before:

typescript
import { DOPeerError } from '@helix-agents/runtime-cloudflare';

try {
  await frontend.execute(agent, input, { sessionId });
} catch (err) {
  // A typed body was a DOPeerError carrying the code, with no rejection cause.
  if (err instanceof DOPeerError && err.code === 'state_run_start_conflict')
    return reloadAndRetry();
  // An unknown agent was an untyped 400, matched by its text.
  if (err instanceof Error && err.message.includes('Unknown agent type')) return notFound();
  throw err;
}

After:

typescript
import { HelixError, isAgentNotFoundError } from '@helix-agents/core';

try {
  await frontend.execute(agent, input, { sessionId });
} catch (err) {
  // A RunStartRejectedError; its rejectCause names the gate that refused the entry.
  if (err instanceof Error && err.name === 'RunStartRejectedError') return reloadAndRetry();
  if (isAgentNotFoundError(err)) return notFound();
  if (HelixError.isInstance(err) && err.retryable) return retryLater();
  throw err;
}

DOPeerError loses its retryable constructor parameter. It is always retryable: it is a store read error.

Before:

typescript
new DOPeerError(message, code, reason, status, status === 503);

After:

typescript
new DOPeerError(message, code, reason, status); // always retryable: it is a store read error

The remote transports throw typed errors. HttpRemoteAgentTransport and DOStubTransport share core remoteProtocolError. Besides RemoteAgentAlreadyRunningError / RemoteAgentNotFoundError / RemoteAgentFeatureUnavailableError, an envelope with an errorCode (or any 503) is a HelixError with that code and cause (retryable exactly on a 503), and a non-envelope typed body (agent-server's { error, code }) is helixErrorFromPeerBody. HttpRemoteAgentTransport.resume() used to throw a generic Error for every answer.

Before:

typescript
try {
  await transport.start(request);
} catch (err) {
  if (err instanceof Error && err.message.includes('state_history_incomplete')) return reload();
  throw err;
}

After:

typescript
import { HelixError } from '@helix-agents/core';

try {
  await transport.start(request);
} catch (err) {
  if (HelixError.isInstance(err) && err.code === 'state_history_incomplete') return reload();
  throw err;
}

runtime-js refusals are typed. A resume() of a completed or failed session, or of a paused session still awaiting client-tool results, throws core ResumeRefusedError (state_not_resumable, with sessionStatus and awaitingToolCallIds; the partial-batch case is retryable). So does a retry() of a session that is not failed (non-retryable; its message still says retry() is only for failed sessions). An unknown session is a non-retryable HelixError state_session_not_found from resume(), retry() and a reconnect handle's resume (getHandle() still returns null for it). Another agent type's session is a non-retryable state_agent_type_mismatch from resume(), retry(), getHandle() and a reconnect handle's resume. resume({ mode: 'with_confirmation' }) is framework_not_supported. All of these were plain Errors with the same message text, which core classifyEntryFailure read as transient (a host answered a retryable 503). retry() honours options.usageStore.

An already-aborted signal rejects the entry. execute(), resume(), retry() and recover() honour options.abortSignal: an entry whose signal is aborted before its run-start commit (an already-aborted signal included) rejects with a HelixError whose code is framework_cancelled (non-retryable) and writes nothing: no run, no message, no stream chunk, no lifecycle hook. Only a session that never had a run (the one a fresh execute() just created) is recorded failed on its session (rule M1). Before, execute() with an already-aborted signal returned a handle whose result() was failed.

Before:

typescript
const handle = await executor.execute(agent, input, { sessionId, abortSignal });
const result = await handle.result();
if (result.status === 'failed') return showCancelled();

After:

typescript
try {
  const handle = await executor.execute(agent, input, { sessionId, abortSignal });
  const result = await handle.result();
} catch (err) {
  if (HelixError.isInstance(err) && err.code === 'framework_cancelled') return showCancelled();
  throw err;
}

Before:

typescript
try {
  await executor.resume(agent, sessionId);
} catch (err) {
  if (err instanceof Error && err.message.startsWith('Cannot resume')) return showRetry();
  throw err;
}

After:

typescript
try {
  await executor.resume(agent, sessionId);
} catch (err) {
  if (HelixError.isInstance(err) && err.code === 'state_not_resumable') return showRetry();
  throw err;
}

Companion delivery ​

On a DO with subAgentNamespace, the persistent-companion tools map the child DO's answer onto the CP-105 delivery contract (companion__sendMessage):

Child /resume answerBeforeNow
2xx RemoteStartResponse{ delivered: true }unchanged; the parent ref is written running after it
2xx that is not a RemoteStartResponse (a hook's own 200, a 204, non-JSON){ delivered: true }a RemoteProtocolValidationError tool error; ref untouched
409 ALREADY_RUNNING (the child is executing)a thrown ChildDOFetchError{ delivered: false, reason: 'child_busy' }; ref untouched
no answer within 60 s, a rejected fetch, or an unreadable bodythe tool hung, or threw{ delivered: false, reason: 'child_start_unconfirmed' }
any other answera thrown ChildDOFetchErrora HelixError with the envelope's errorCode (or code) and cause, retryable only on a 503

Before:

typescript
const out = await callTool('companion__sendMessage', { name: 'helper', message });
// a busy child threw "Child DO resume failed (409): …"

After:

typescript
const out = await callTool('companion__sendMessage', { name: 'helper', message });
if (!out.delivered && out.reason === 'child_busy') {
  // the child is running: try again after it settles
} else if (!out.delivered && out.reason === 'child_start_unconfirmed') {
  // unknown whether the child took the message: check getChildStatus before re-sending
}
  • The 60 s bound (CHILD_RESUME_TIMEOUT_MS) is internal, not an option. It exceeds the 30 s default concurrentEntryTimeoutMs, so a child under interrupt_previous that times out answers child_busy first; a child configured with concurrentEntryTimeoutMs of 60 s or more reports child_start_unconfirmed instead.
  • A spawn whose /start answers 409 attaches to the running child (a replay) only on the ALREADY_RUNNING envelope; any other 409 is a typed tool error.
  • The companion /start body carries the parent's userId, tags and metadata (it carried metadata only).
  • The spawn and continuation /start and the /status polls are not bounded yet (FU-DO-COMPANION-START-FETCH-UNBOUNDED).
  • ChildDOFetchError is deleted (it was module-internal).

Never-run sessions ​

A session that never had a run (an entry failed before its first run-start commit, e.g. a short history read, or the DO was evicted between createSession and the commit) is recorded failed on the SESSION only: the failed status and errorDetail, idle-fenced so a run that commits meanwhile wins. No stream is created or written and no lifecycle hook fires. Streams are written only by their runs.

  • runtime-js no longer fails the stream or fires onAgentFail for a never-run pre-commit failure (a state_history_incomplete first history read, a cancellation, any other class). A session that has a run is still left untouched.
  • The CF DO wake fails an active session with no run record instead of resuming it on an empty history.
  • The snapshot of a never-run failed session (readSnapshotView / buildSnapshot) reads status: 'failed' with the session's errorDetail (it read ended).
  • Core failSessionWithoutRun({ currentRunId: null }) is that write. With a run id it also fails the stream, fenced on that run.

Before:

typescript
hooks: {
  onAgentFail: async ({ sessionId, errorDetail }) => alert(sessionId, errorDetail), // fired for a never-run failure
},
// and a reader watched the stream:
const info = await streamManager.getStreamInfo(sessionId); // 'failed', with the error

After:

typescript
// The failure is on the session (and on the handle / the rejection the caller got):
const state = await stateStore.loadState(sessionId);
if (state?.status === 'failed' && (await stateStore.getCurrentRun(sessionId)) === null) {
  alert(sessionId, state.errorDetail); // never-run: no stream, no onAgentFail
}

Per runtime: runtime-js and the CF DO follow this rule. CF Workflows, Temporal and DBOS still write a never-run session's stream (or, on DBOS, record failed whenever no run is live); see CLAUDE.md and the follow-ups.

Wake recovery ​

The onStart / alarm wake re-drives a stranded run through the same pipeline, and every failure is classified:

ClassWake action
lostnothing (another drive owns the session)
refused, fatalthe run ends failed with the classified detail (K5 terminal commit), then settles
transientre-arm under the crash-loop budget; the failure is recorded as WakeRecoveryRecord.lastFailure
  • Exhaustion ends the run failed with framework_auto_resume_exhausted (retryable: /retry may succeed), whose cause is the last re-armed failure. It replaces the untyped auto_resume_exhausted, and the message no longer blames "a step that crashes the instance".
  • The wake's executor now has the DO's workspace providers, and the wake fires the host hooks (afterEntry, onRunSettled).
  • onStart and onAlarm never throw. A sub-agent without outputSchema used to escape onStart and fail every later request to the DO; it now ends the run validation_error.
  • A failed submit continuation (the run a completed client-tool batch continues) is classified the same way, so a continuation whose agent no longer resolves ends failed with framework_agent_not_found instead of exhausting.

Before:

typescript
if (state.status === 'failed' && state.error === 'auto_resume_exhausted') showRetry();

After:

typescript
if (state.status === 'failed' && state.errorDetail?.code === 'framework_auto_resume_exhausted') {
  const cause = state.errorDetail.cause; // the last failure (string or ErrorDetail)
  showRetry(typeof cause === 'object' ? cause.code : undefined);
}

A run whose owner is gone is settled through core's closed-owner reconciliation (planClosedOwnerReconciliation → reconcileClosedOwnerRun, the same pair CF Workflows' readers use) before the ladder above runs:

  • a run whose K5 landed is finished and settles once (onRunSettled), as before;
  • new: a run evicted between its suspension commit and its suspended_* run status (the session paused, the run still running) gets its suspended status and its stream paused, and settles suspended_*. It used to stay running until the next entry superseded it;
  • an ended run whose stream is still open has the stream ended or failed (no hook);
  • a run with nothing committed is re-driven, as before;
  • a store read that fails permanently (a corrupt row) ends a stranded active run failed typed instead of re-arming on every alarm. A transient one is retried on the next alarm.

autoResumeOnWake: false keeps the demote to paused for a stranded active run; a paused session whose whole client-tool batch was submitted stays paused until an explicit /resume.

The submit route (/submit-tool-result) resolves the owning agent and materializes the continuation BEFORE it writes: a failing factory answers typed (503 framework_agent_resolution_failed, framework_usage_store_unavailable, …) and the call stays pending. It used to answer 200 and lose the error to the wake.

A submit that completes its batch now answers after the continuation's entry has started (or another drive owns the run), including a submit that lands while the suspending run is still settling; it used to answer before any entry. That wait is bounded by concurrentEntryTimeoutMs (default 30 s): past it the DO answers 500 transport_error (the result is stored and the run continues via the unconsumed-submit alarm, so re-read the snapshot instead of re-submitting; the message says the continuation was not confirmed, since it may already have started). A duplicate submit (the result is already stored or consumed) answers 200 already_completed at once and never waits for the continuation; it means this call's result is recorded, not that the run has finished. A forwarded submit whose owning DO does not answer within that bound plus 1 s answers 503 transport_error; re-submitting it is safe. A forwarded submit relays the owning DO's typed { error, code } answer with its status, a 404 included (e.g. framework_agent_not_found; it used to be answered as a retryable 503); a 404 that is neither typed nor unknown_tool_call answers a non-retryable 500 transport_error.

A DO submit can now answer a typed failure or a timeout where it always answered 200. When the continuation the submit starts fails to enter, the answer follows the failure table's wake row, on the committed row (never 503): refused / fatal → the run's typed failure (it ended failed with that ErrorDetail); transient → its code on the non-retryable row (500, or 409 for a state error); the run is re-armed and continues, so re-anchor through the snapshot (useHelixChat().retry()) instead of re-submitting. lost (another drive owns the run) still answers 200 accepted.

The route's own store read and alarm writes never reach the untyped 500 { error: 'Internal server error' } either. A failed read of the owning session (nothing written yet) answers typed (an unclassified failure is a retryable 503). Once the result is stored, a failed alarm reschedule falls back to a backoff alarm and the answer stands, and a failed read of the batch (to decide whether to continue) answers its code on the non-retryable row (500, never 503; the unconsumed-submit alarm continues a complete batch, so re-read the snapshot instead of re-submitting).

Registries ​

Registryresolve()
@helix-agents/runtime-cloudflare AgentRegistryasync-compatible: a static registration synchronously, a factory's result as it returns it (a Promise for an async factory); AgentFactory may be async
@helix-agents/runtime-temporal AgentRegistrystays synchronous: Temporal activities resolve through your synchronous resolveAgent

Before:

typescript
const config = registry.resolve('assistant', { env, sessionId, runId, userId });

After:

typescript
const config = await registry.resolve('assistant', {
  purpose: 'drive',
  env,
  sessionId,
  runId,
  drive: 'start',
  identity: { userId },
});

Store and runtime implementers ​

An entry finishes a SUSPENDED run whose K5 landed (R66). applyRunStartRules verifies finishCurrent for any live current run (running or suspended_*), and every runtime's entry passes finishCurrent for one. A stop on a paused run whose run-status write was lost is now recorded with the K5's status (interrupted), not superseded. A custom store whose finish write pins status = 'running' must pin the live statuses instead (D1 did); the contract scenario run.k5-landed-finishes-suspended checks it.

A permanent store error is never wrapped. Core classifyEntryFailure classifies a corrupt-row error (InvalidStoredStateError, CorruptMessageRowError, UnknownCommitKindError) as fatal (core isPermanentStoreError), so a DO wake ends the run typed instead of re-arming, and a DO request (so a chat route backed by the CF DO) answers it typed and non-retryable (500) instead of 503. DOStateStore rethrows every permanent store error as itself: an InvalidStoredStateError used to reach callers wrapped in DOStateError. A custom store should not wrap one either.

assertRunStatusTransition is renamed decideRunStatusTransition and returns 'write' | 'noop': a terminal run written with its own status is a no-op (write nothing), a live run always writes, and any other transition of a terminal run still throws RunSupersededError. A custom store's updateRunStatus must return without writing on 'noop'. The run-start rule order also changed: applyRunStartRules verifies finishCurrent before the owner guard.

Before:

typescript
assertRunStatusTransition({ sessionId, run, to: status });
await writeRunStatus(runId, status, metadata);

After:

typescript
if (decideRunStatusTransition({ sessionId, run, to: status }) === 'noop') return;
await writeRunStatus(runId, status, metadata);

updateRunStatus to a live status never sets completedAt, even when the caller passes one; a terminal status defaults it to now. endSequence is stored only as the caller passes it (the D1 and DO stores used to default it to a message sequence). The contract scenario is run.live-status-no-completedAt in @helix-agents/core/testing.

The store-cloudflare stream Durable Object answers an /end, /fail or /pause whose expectRunId fence it cannot read with 400 validation_error and applies nothing (it used to apply an unparseable /pause unfenced); the binding throws StreamConnectionError.

runtime-js:

  • RunStartEntryState.build returns RunStartEntrySessionState, whose status ('active' | 'paused') is required; a custom runStartCommitter writes the status as given.

    Before:

    typescript
    build: (session) => ({ ...session, stepCount: 0 }),

    After:

    typescript
    build: (session) => ({ ...session, stepCount: 0, status: 'active' }),
  • JsClientToolResolver's getCompletedRetentionMsForRoot option is deleted. A call's completed-tombstone retention is recorded on its pending entry at registration (PendingClientToolCall.completedRetentionMs, from the agent's completedTombstoneRetentionMs); registerPending takes an optional completedRetentionMs.

    Before:

    typescript
    new JsClientToolResolver({
      stateStore,
      getCompletedRetentionMsForRoot: async (root) => lookup(root),
    });

    After:

    typescript
    new JsClientToolResolver({ stateStore }); // retention comes from each pending entry
  • JSAgentExecutor takes submitForwarder?: SubmitForwarder for submits another host owns, and its submitToolResult owns the whole submit. JsClientToolResolver.submit is internal; it maps an approval-response itself.

    Before:

    typescript
    // An approval mapped by hand (approved / reason were dropped if you got it wrong):
    await resolver.submit({
      kind: 'client-tool-result',
      sessionId,
      toolCallId,
      result: { approved: true },
    });

    After:

    typescript
    await executor.submitToolResult({
      kind: 'approval-response',
      sessionId,
      toolCallId,
      approvalId,
      approved: true,
    });
  • A submit whose submitting session's agent type is unknown to the configured registry fails framework_agent_not_found instead of skipping output validation.

  • defineAgent throws unless completedTombstoneRetentionMs is a finite number >= 0 (0 disables retention).

A new host or runtime embedding runtime-js the way the DO does classifies entry failures with core classifyEntryFailure (lost / refused / transient / fatal) and fails a never-run session with failSessionWithoutRun.

What to check ​

  • Replace every removed hook, ExecutionState method, factory signature and /resume option above. TypeScript reports each one.
  • Deduplicate onRunSettled on deliveryId.
  • Authorise /submit-tool-result in the worker in front of the DO.
  • A /submit-tool-result caller handles a typed failure or a 500 / 503 transport_error timeout (above) instead of treating every answer as accepted or waiting indefinitely.
  • Callers that matched the native ALREADY_RUNNING envelope, a 400 for an unknown agent, DOPeerError, a generic transport Error, or auto_resume_exhausted now get the typed answers above.
  • A consumer usage store that can fail now stops runs (503 / re-arm).
  • Stop sending per-turn metadata, userId or tags on a continuing /start.

Released under the MIT License.