CF Durable Object drive pipeline (WA MR1)
This guide covers the release in which every run a Cloudflare Durable Object (DO) starts or recovers goes through one pipeline (runDrive): /start, /resume, /retry, the /subagent/:type/* protocol routes, the client-tool submit that continues a paused run, and the wake that recovers an evicted run. Each of them now resolves the agent, builds the executor, classifies a failure, fires the host hooks and settles the run the same way.
It is a major release of @helix-agents/runtime-cloudflare, @helix-agents/runtime-temporal (its AgentNotFoundError is removed), @helix-agents/runtime-js, @helix-agents/ai-sdk and @helix-agents/agent-server (its INVALID_RESULT answer and two typed /resume answers), and a minor (0.x) release of @helix-agents/core and the four stores. Nothing is deprecated: every removed API is deleted.
Design: 2026-10-08-do-drive-pipeline-design.md (read its "§0 As built (MR1)" first). Runtime page: Cloudflare Durable Objects.
Hooks
The seven server hooks collapse into three, and the hook-facing ExecutionState is read-only.
| Removed | Replacement |
|---|---|
beforeStart, beforeResume, beforeRetry | beforeEntry(ctx), one hook; ctx.drive is 'start' | 'resume' | 'retry' |
afterStart, afterResume, afterRetry | afterEntry(ctx), once per drive, wakes included |
onComplete | onRunSettled(ctx), at least once per (runId, status) |
ExecutionState.interrupt() / abort() / getHandle() | concurrentEntry: 'interrupt_previous', or EndpointContext.stop() |
BeforeStartContext … AfterRetryContext, CompleteContext | BeforeEntryContext, AfterEntryContext, RunSettledContext, SettledStatus |
Before:
hooks: {
beforeStart: async ({ request, body, executionState }) => {
if (!(await isAllowed(request))) {
return { abort: true, response: new Response('Forbidden', { status: 403 }) };
}
},
afterStart: async ({ sessionId, runId, agentType }) => log('started', sessionId, runId, agentType),
onComplete: async ({ sessionId, runId, status, error }) => notify(sessionId, runId, status),
},After:
hooks: {
beforeEntry: async (ctx) => {
// ctx.origin.request is the original request (headers intact).
if (!(await isAllowed(ctx.origin.request))) {
return { abort: true, response: new Response('Forbidden', { status: 403 }) };
}
},
afterEntry: async ({ sessionId, runId, drive }) => log('entered', sessionId, runId, drive),
onRunSettled: async ({ sessionId, runId, status, errorDetail, deliveryId }) => {
// At least once per (runId, status): deduplicate on deliveryId.
if (await alreadyProcessed(deliveryId)) return;
await notify(sessionId, runId, status, errorDetail);
},
},What changed in the semantics:
beforeEntryis admission only (auth, rate limits, quotas). It runs for/start,/resume,/retryand the/subagent/:type/start|resumeroutes (ctx.origin.routeis'native'or'subagent'), and only AFTER the busy check, the identity / type gate and agent materialization. A refused entry never reaches it, and it never sees a live run to interrupt. Its context is per drive:start:body, the transformed, validated start body (transformRequest's result on the native/start). The raw body stays readable atctx.origin.request.resume:options(aDoResumeOptions; awith_messageresume carries the message).retry:checkpointId?andmessage?, the retry's new user message.- every drive:
env,originandidentity(the session'suserId,tags,metadata,parentSessionId,rootSessionId).
Return
{ abort: true, response }to refuse: the response is answered verbatim and nothing is written. A throw answersframework_hook_failed, retryable (503) unless its cause is classified non-retryable.Submits are outside
beforeEntry./submit-tool-resultstores caller-supplied tool results and approval decisions, and the continuation it starts puts them in front of the model; no host hook admits it. Authorise submits in the worker in front of the DO (FU-DO-SUBMIT-ADMISSION-HOOK tracks a possible hook).afterEntryfires once per drive, after the drive owns its run: after the run-start commit for an entry, after the in-memory claim for a wake'srecover. Its context is{ env, sessionId, runId, drive, handle }. It is not awaited by the drive; a throw is logged.onRunSettledfires whenever a run reaches a status at which nothing executes for it (suspended_client_tool,suspended_awaiting_children,suspended_step_partial,completed,failed,interrupted,superseded), on every path: request drives, the submit continuation and wakes. Its context is{ env, sessionId, runId, drive?, status, result?, errorDetail?, deliveryId }, withdeliveryId = ${runId}:${status}.- At least once. A durable mark (DO KV
__helix_settle_delivered:{runId}:{status}) is written after the hook returns, and a wake redelivers a settled run whose mark is missing (an eviction between the commit and the hook). Deduplicate ondeliveryId. - A throwing hook is logged and still marked, so it is never retried forever.
- A paused run settles once,
suspended_*. The built-in routes cannot stop it (/interruptanswers 400 for a paused session and/abortwrites nothing for it). A custom endpoint'sEndpointContext.stop('interrupt')sets the durable interrupt flag on the paused session: nothing ends at once, and the next continuation (a submit or a resume) honours it, ending that runinterruptedat its first boundary. Stopping a paused run at once is MR2. - It runs after the run's stream is finalized and after the execution registration is released, so a new entry is never refused because a slow hook is still running.
- With no
onRunSettledconfigured, no mark is written.
- At least once. A durable mark (DO KV
transformRequest/transformResponseare unchanged in purpose and apply to the native/startonly. AtransformRequestthrow is classified: aHelixErrorkeeps its code, a Zod failure answers 400validation_error, anything else answersframework_hook_failed(503 unless its cause is non-retryable). Its result type isStartBody(formerlyStartAgentRequestV2).transformResponsereceivesStartResponseBody({ sessionId, streamId, runId, status: 'started' | 'resumed', startSequence? }).ExecutionStatekeepssessionId,runId,isExecutingandstatus, read-only.isExecutingistruefrom just before a drive's executor call until the settle path releases it, which happens beforeonRunSettled.Custom endpoints stop a run with
EndpointContext.stop(kind, reason?)('interrupt'or'abort'), the mechanism of the/interruptand/abortroutes. In this release it waits for the run to settle, as the routes do: a tool that ignores its abort signal keeps the call waiting (the stop MR changes this).
Before:
endpoints: {
'/cancel': async (request, { executionState }) => {
await executionState.abort('cancelled by user');
return Response.json({ ok: true });
},
},After:
endpoints: {
'/cancel': async (request, { stop }) => {
await stop('abort', 'cancelled by user');
return Response.json({ ok: true });
},
},Concurrent entries
"Last message wins" is a config option. It replaces interrupting the live run from beforeStart.
Before:
hooks: {
beforeStart: async ({ executionState }) => {
if (executionState.isExecuting) {
await executionState.interrupt('New message received');
}
},
},After:
createAgentServer<Env>({
llmAdapter: ({ env }) => createAdapter(env),
agents: registry,
concurrentEntry: 'interrupt_previous', // default: 'reject'
concurrentEntryTimeoutMs: 30_000, // default
});'reject'(the default): a/start,/resumeor/retrythat arrives while another drive executes a run on the DO answers 409state_already_running.'interrupt_previous': the new entry first passes the identity / type gate, materialization andbeforeEntry. Only then is the executing run interrupted (reasonsuperseded_by_new_entry), and the entry waits for it to settle, bounded byconcurrentEntryTimeoutMs. A refused or unadmitted entry never interrupts a live run. If the old run has not settled at the bound, the entry answers 409state_already_running; the interrupt still applies.- A run whose result has already settled (its drive is in its tail) is waited for, not interrupted, under either policy.
- Wake drives never interrupt a run.
Resolver and factories
The resolver may be async, and every factory receives one context object: the session's persisted identity, the drive kind, and the run id only when a run is about to be driven.
| Removed | Replacement |
|---|---|
LLMAdapterFactory = (env, ctx) => … | LLMAdapterFactory = (ctx) => …, ctx: DriveFactoryContext & { agentType } |
UsageStoreFactory = (env, ctx) => … | UsageStoreFactory = (ctx) => …, ctx: DriveFactoryContext | UsageReadContext |
LLMAdapterContext, UsageStoreContext | DriveFactoryContext, UsageReadContext |
AgentFactoryContext { env, sessionId, userId?, runId } | DriveFactoryContext | LookupFactoryContext |
a synchronous AgentResolver | AgentResolver returning AgentConfig | Promise<AgentConfig> |
interface AgentIdentity {
userId?: string;
tags?: string[];
metadata?: Record<string, string>;
parentSessionId?: string;
rootSessionId?: string;
}
type DriveKind = 'start' | 'resume' | 'retry' | 'continue' | 'recover';
interface DriveFactoryContext<TEnv> {
purpose: 'drive';
env: TEnv;
sessionId: string;
runId: string;
drive: DriveKind;
identity: AgentIdentity;
}
interface LookupFactoryContext<TEnv> {
purpose: 'lookup';
env: TEnv;
sessionId: string;
identity: AgentIdentity;
reason: 'submit_validation' | 'retention';
}
interface UsageReadContext<TEnv> {
purpose: 'usage_read';
env: TEnv;
sessionId: string;
identity: AgentIdentity;
}Before:
const registry = new AgentRegistry();
registry.registerFactory<Env>('assistant', (ctx) =>
createAssistant({ db: ctx.env.DB, userId: ctx.userId })
);
createAgentServer<Env>({
llmAdapter: (env, ctx) => new VercelAIAdapter({ anthropic: makeAnthropic(env, ctx.agentType) }),
agents: registry,
usageStore: (env, ctx) => new BillingUsageStore(env.BILLING_DB, ctx.sessionId),
});After:
const registry = new AgentRegistry();
registry.registerFactory<Env>('assistant', async (ctx) =>
createAssistant({
db: ctx.env.DB,
userId: ctx.identity.userId,
prompt: await loadPrompt(ctx.env),
})
);
createAgentServer<Env>({
llmAdapter: ({ env, agentType }) =>
new VercelAIAdapter({ anthropic: makeAnthropic(env, agentType) }),
agents: registry,
usageStore: ({ env, sessionId }) => new BillingUsageStore(env.BILLING_DB, sessionId),
});- The resolver sees a run id only with
purpose: 'drive'. Apurpose: 'lookup'call (validating a submitted client-tool result against the owning tool'soutputSchema) starts nothing. Narrow onctx.purposebefore you readctx.runId. - The identity comes from the session row on every drive (wakes, resumes, retries and companion children included), not only on
/start. Async setup that used to live inbeforeStartbelongs in the resolver. - A resolver failure is typed. Throw core
AgentResolutionError.notFound(agentType, available)for an unknown type: the native routes answer 404framework_agent_not_found, the/subagentroutes 404NOT_FOUND, and a wake ends the run typed. Any other resolver or adapter-factory throw becomesframework_agent_resolution_failed, with the original ascauseand its retryability: an unclassified throw is transient (a request answers 503, a wake re-arms). - A sub-agent tool without
outputSchemafails materialization withvalidation_error(400 on a request; a wake ends the runfailed). It used to throw an untyped error, and on a wake it made every later request to the DO fail. persistentAgentswithoutsubAgentNamespacefails withframework_not_supported(500) instead of a 400 with the ad hoc codepersistent_agents_no_subagent_namespace.- The usage-store factory is also called to read usage, with no drive: the
/subagent/:type/usagerollup, and the end-of-stream rollup of a run finished on a wake. That call getsUsageReadContext(purpose: 'usage_read').
Strict usage store
A consumer usage-store factory that throws stops the entry. It no longer falls back to the DO's internal store, and usageStoreFactoryNegativeCacheMs is deleted.
Before:
createAgentServer<Env>({
// …
usageStore: (env, ctx) => new BillingUsageStore(env.BILLING_DB, ctx.sessionId),
// A throwing factory fell back to the internal DOUsageStore for 5 s.
usageStoreFactoryNegativeCacheMs: 5_000,
});After:
createAgentServer<Env>({
// …
// A throw is `framework_usage_store_unavailable` (retryable): a request
// answers 503, a wake re-arms. Usage is never split across two stores.
usageStore: ({ env, sessionId }) => new BillingUsageStore(env.BILLING_DB, sessionId),
});- With no
usageStore, the internalDOUsageStoreis used. It is bound to the DO's session:recordEntryfor another session throwsvalidation_errorand writes nothing, andgetEntries/getRollupread only that session's rows. - A consumer usage-DB outage now stops runs (503 / re-arm) instead of silently splitting usage.
Resume body
The native /resume body is { agentType, options: DoResumeOptions }, a strict union that mirrors core ResumeOptions. Every field it accepts is honoured; anything else is a 400 validation_error.
type DoResumeOptions =
| { mode: 'continue' }
| { mode: 'with_message'; message: string | UserInputMessage[] }
| { mode: 'from_checkpoint'; checkpointId: string };Before:
await stub.fetch('https://do/resume', {
method: 'POST',
body: JSON.stringify({
agentType: 'assistant',
options: { mode: 'continue', appendMessages: [{ role: 'user', content: 'Go on' }] },
}),
});After:
await stub.fetch('https://do/resume', {
method: 'POST',
body: JSON.stringify({
agentType: 'assistant',
options: { mode: 'with_message', message: 'Go on' },
}),
});| Removed | Replacement |
|---|---|
options.mode: 'retry' | POST /retry { agentType, checkpointId?, message? } |
options.mode: 'branch' | none: the DO cannot branch (see below) |
options.modifyState | POST /start with input.state |
options.appendMessages | { mode: 'with_message', message } |
options.checkpointId beside mode: 'continue' | { mode: 'from_checkpoint', checkpointId } (it used to be ignored) |
an omitted mode (e.g. options: {}), which used to mean continue | { mode: 'continue' } (an omitted mode is now a 400 validation_error) |
ResumeAgentRequestV2, StartAgentRequestV2, RetryAgentRequestV2 | DoResumeOptions, StartBody, the /retry body above |
DO branching is unsupported. A Durable Object cannot read another DO's session (DOStateStore.cloneSession refuses), so a POST /start whose body carries branch answers 500 framework_not_supported at parse, before any read or write. It used to reach the executor and fail there, often as a retryable 503. Branch on CF Workflows, or start a new session.
DOFrontendExecutor.resume and DOWorkflowExecutor.resume send this body. A with_confirmation resume is refused typed before any request: approvals go through /submit-tool-result (approval-response).
Identity is fixed when the session is created. A /start on an existing session whose body names a different userId, tags or metadata answers 400 validation_error (the error's field names which). Omitted fields keep the session's values. A host that sent per-turn metadata must stop.
The persisted agent type decides. /resume, /retry and a /start on an existing session whose agentType differs from the session's answer 409 state_agent_type_mismatch (the body's type used to be trusted).
Wire answers
The native routes (/start, /resume, /retry) answer every refused entry with core typedErrorResponse's body, { error, code, cause? }, and the status of core httpStatusForHelixError. A request drive that fails before its run-start commit writes nothing (no status, version, stream or hook write), including /resume, which used to answer { status: 'resumed' } at once and fail later.
| Situation | Before | Now |
|---|---|---|
| the DO is executing a run (busy) | 409 { error: 'Agent already running', code: 'ALREADY_RUNNING', sessionId, streamId } | 409 { error, code: 'state_already_running' } (no streamId) |
| a persisted live run that nothing executes | 409 busy (the persisted-status precheck) | accepted: the run-start commit supersedes the orphaned run |
a rejected run start (RunStartRejectedError) | 409 state_run_start_conflict, no cause | 409 { error, code: 'state_run_start_conflict', cause }; cause is one of the six reject causes |
| a resume the session's status refuses | 409 busy, or a detached failure | 409 { error, code: 'state_not_resumable' }; a partial batch's message names the awaited call ids first |
/retry on a session that is not failed | 409 { error, currentStatus } | 409 { error, code: 'state_not_resumable' } (no currentStatus) |
| an unknown agent type | 400 { error: "Unknown agent type: …" } | 404 { error, code: 'framework_agent_not_found' } |
| a resolver / factory throw, unclassified | 400 Unknown agent type | 503 framework_agent_resolution_failed (409 / 500 once the input committed) |
/resume / /retry with no session on the DO | 400 { error: 'No agent state to resume' } | 409 { error, code: 'state_session_not_found' } (state category, not retryable) |
| an agent-type mismatch | accepted | 409 state_agent_type_mismatch |
| an unclassified failure before the commit | 500 | 503 (retryable) |
framework_not_supported (a server misconfiguration) | 400 / 500 ad hoc | 500 |
Once the entry's input committed, a retryable error is answered with its non-retryable row (409 for a state_* code, else 500), never 503: a re-sent request would append the input twice.
The /subagent/:type/start|resume routes keep the remote-protocol envelope (RemoteAgentErrorResponse), whose uppercase codes are the protocol's own. The typed code now rides as errorCode, and a rejected run start's reason as cause:
| Answer | Envelope |
|---|---|
| busy (the DO is executing a run) | 409 ALREADY_RUNNING, no streamId (it used to carry the persisted one) |
| a persisted live run, nothing executing | 200, a new run (it used to answer 409 ALREADY_RUNNING) |
| unknown agent type, unknown session | 404 NOT_FOUND |
state_not_resumable, completed session | 409 ALREADY_COMPLETED, errorCode: 'state_not_resumable' |
state_not_resumable, otherwise | 409 INVALID_REQUEST, errorCode: 'state_not_resumable' |
| a caller error (400) | 400 INVALID_REQUEST, errorCode |
| anything else | the typed status (409 / 500 / 503), INTERNAL_ERROR, errorCode, cause? |
A /subagent/:type/start body whose agentType differs from the path answers 400. The route parses the protocol body directly, so the original request's headers reach beforeEntry (they used to be dropped).
The /start answer is { sessionId, streamId, runId, status: 'started' | 'resumed', startSequence? } (/resume and /retry answer status: 'resumed'). startSequence is the stream head at the run-start commit, omitted only when the run record could not be read. /subagent/:type/start|resume answer the same body (a RemoteStartResponse plus status and startSequence).
Busy means the DO is executing a run. The persisted-status precheck is removed: the run-start commit is the gate. A /start (or /subagent/:type/start) on a session whose persisted run is live but which nothing executes (its execution was lost, e.g. to an eviction) is accepted, and its run-start commit supersedes the orphaned run. The busy 409 no longer carries streamId (it used to send the persisted one from that precheck); read /snapshot to attach.
/retry on a session that is not failed answers 409 { error, code: 'state_not_resumable' } (runtime-js retry() throws ResumeRefusedError). It used to answer 409 { error, currentStatus } from a precheck. The protocol routes have no retry; their resume refusals are the state_not_resumable rows above (ALREADY_COMPLETED for a completed session, else INVALID_REQUEST).
Every host that answers with core typedErrorResponse (agent-server's /start and /resume, handleChatStream) now answers framework_agent_not_found with 404 instead of 500, through the new row in httpStatusForHelixError.
agent-server's /resume answers two races typed. With a runtime-js executor, a /resume whose session is deleted after agent-server's own session check, or whose agents map key differs from that agent's name, now answers 409 { error, code: 'state_session_not_found' } / 409 { error, code: 'state_agent_type_mismatch' } (both non-retryable), through the typed runtime-js refusals (see Errors). Both used to answer 500 { error, code: 'INTERNAL_ERROR' }. The message text is unchanged.
The usage_store_factory_fallback counter is gone (there is no fallback): nothing emits it, and core's USAGE_STORE_FACTORY_FALLBACK_COUNTER export is removed. Remove dashboards or alerts built on it, and alert on framework_usage_store_unavailable instead.
A new entry right after a completed run is accepted. A /start that arrives between a run's terminal K5 commit and its terminal run status used to answer the busy 409; the run-start commit now finishes that run itself and starts the new one.
Before (a caller matching the old envelope):
const res = await stub.fetch('https://do/start', { method: 'POST', body: JSON.stringify(body) });
if (res.status === 409) {
const { code, streamId } = await res.json();
if (code === 'ALREADY_RUNNING') return attach(streamId);
}After:
const res = await stub.fetch('https://do/start', { method: 'POST', body: JSON.stringify(body) });
if (!res.ok) {
const { error, code, cause } = (await res.json()) as {
error: string;
code: string;
cause?: string;
};
if (code === 'state_already_running') return retryLater(); // no streamId: read /snapshot to attach
throw new Error(`${code}${cause ? ` (${cause})` : ''}: ${error}`);
}Scrubbed error text
Every free-text field handleChatStream writes for a failure is scrubbed with core sanitizeErrorMessage / sanitizeErrorDetail (control characters stripped, < / > escaped, known secrets redacted, at most 256 characters). The code, category, retryable and cause structure are kept, so parseHelixChatError reads the same code.
| Answer | What a client now sees |
|---|---|
409 state_already_running (busy) | { error, code, retryable: true }; error scrubbed (a plain AgentAlreadyRunningError's bytes are unchanged) |
409 state_run_start_conflict | { error, code, cause }; cause is present only when it parses as one of the six reject causes (core RunStartRejectCauseSchema); free text is dropped |
409 state_not_resumable and every other typed entry failure | { error, code } through core typedErrorResponse (scrubbed) |
409 stream_failed (resuming a failed stream) | message and errorDetail scrubbed at every level of the cause chain |
the error chunk when the transformer itself throws | { type: 'error', errorText }, errorText scrubbed; the shape is unchanged |
Before:
{
"error": "Agent already running: postgres://app:hunter2@10.0.0.5/db\n at /srv/app/store.js:12:5",
"code": "state_already_running"
}After:
{
"error": "Agent already running: postgres://[REDACTED] at [path][ip]",
"code": "state_already_running",
"retryable": true
}The run's own error text the stream transformer forwards (a tool's error, a run's error chunk) is unchanged (FU-AISDK-RUN-ERROR-TEXT-UNSCRUBBED).
The CF Durable Object's own answers are scrubbed the same way: the /subagent/:type/usage 503 (the consumer usage store's error text) and the /snapshot 500 keep their code and shape with the message scrubbed.
INVALID_RESULT no longer reveals the tool's outputSchema, on every host. A /submit-tool-result INVALID_RESULT 400 carries the Zod issues and the detailed message only with verboseValidationErrors: true, the gate invalid_payload's details already used. By default it is { error: 'The submitted result failed the tool output validation', code: 'INVALID_RESULT', toolName, toolCallId } (the error is core's new INVALID_RESULT_MESSAGE). This holds on the CF Durable Object and on agent-server; the ai-sdk submitToolResultErrorResponse always answers this body (it has no verbose option). A client that read issues or parsed the message must turn the option on where it exists. The code, the status and toolName / toolCallId are unchanged.
handleChatStream's send, resume and client-tool submit paths now share one classification: any failure core can classify (a HelixError, a raw core error with a known code, a rejected run start, a busy refusal, or an error wrapping one) ends the request with the typed HTTP answer. A resume or submit that hit a busy refusal or a RunStartRejectedError used to answer 200 with a data-resume-rejected frame; it now answers the typed 409. Only an unclassifiable error keeps the per-intent data-resume-rejected frame.
Errors
AgentNotFoundError is deleted from @helix-agents/runtime-cloudflare and @helix-agents/runtime-temporal. Both registries throw core AgentResolutionError.notFound (framework_agent_not_found, non-retryable, with agentType and availableTypes). Match it with the core guards, which check name and code, so an error from a second copy of core still matches.
Before:
import { AgentNotFoundError } from '@helix-agents/runtime-cloudflare'; // or runtime-temporal
try {
registry.get(agentType);
} catch (err) {
if (err instanceof AgentNotFoundError) return notFound(err.availableTypes);
throw err;
}After:
import { isAgentNotFoundError } from '@helix-agents/core';
try {
registry.get(agentType);
} catch (err) {
if (isAgentNotFoundError(err)) return notFound(err.availableTypes);
throw err;
}isAgentResolutionError(e)matches both codes (framework_agent_not_found,framework_agent_resolution_failed).AgentRegistry.get()returns static registrations only; for a factory-registered type it throwsnotFoundwith a reason that says to useresolve().- New
ErrorCodes:framework_agent_not_found(404),framework_agent_resolution_failed,framework_auto_resume_exhausted,framework_usage_store_unavailable,framework_hook_failed,state_agent_type_mismatch(409). AswitchoverErrorCodethat must be exhaustive gains these cases.
ErrorDetail.cause is string | ErrorDetail, nested at most 8 deep (ERROR_DETAIL_MAX_CAUSE_DEPTH). It appears in stream fail frames, /status, the remote protocol and the persisted error_detail. resolveErrorDetail records a classified error's wrapped cause as its own detail, and toErrorDetail keeps a boolean retryable property of a raw Error.
Before:
const why: string | undefined = result.errorDetail?.cause;After:
const cause = result.errorDetail?.cause;
const why = typeof cause === 'string' ? cause : cause?.message;
const causeCode = typeof cause === 'object' ? cause.code : undefined;The DO clients throw typed errors. DOFrontendExecutor and DOWorkflowExecutor read a native answer with core PeerErrorBodySchema and no longer throw DOPeerError for an entry failure (DOPeerError is now only DOStateStoreClient's read error):
| Answer | Thrown |
|---|---|
409 state_already_running | AgentAlreadyRunningError |
404 framework_agent_not_found | AgentResolutionError for the requested type (no availableTypes) |
409 state_run_start_conflict + a known cause | RunStartRejectedError(sessionId, cause) |
| any other typed body | a HelixError with the body's code and cause (helixErrorFromPeerBody); retryable when the body says so, else exactly on a 503 |
Before:
import { DOPeerError } from '@helix-agents/runtime-cloudflare';
try {
await frontend.execute(agent, input, { sessionId });
} catch (err) {
// A typed body was a DOPeerError carrying the code, with no rejection cause.
if (err instanceof DOPeerError && err.code === 'state_run_start_conflict')
return reloadAndRetry();
// An unknown agent was an untyped 400, matched by its text.
if (err instanceof Error && err.message.includes('Unknown agent type')) return notFound();
throw err;
}After:
import { HelixError, isAgentNotFoundError } from '@helix-agents/core';
try {
await frontend.execute(agent, input, { sessionId });
} catch (err) {
// A RunStartRejectedError; its rejectCause names the gate that refused the entry.
if (err instanceof Error && err.name === 'RunStartRejectedError') return reloadAndRetry();
if (isAgentNotFoundError(err)) return notFound();
if (HelixError.isInstance(err) && err.retryable) return retryLater();
throw err;
}DOPeerError loses its retryable constructor parameter. It is always retryable: it is a store read error.
Before:
new DOPeerError(message, code, reason, status, status === 503);After:
new DOPeerError(message, code, reason, status); // always retryable: it is a store read errorThe remote transports throw typed errors. HttpRemoteAgentTransport and DOStubTransport share core remoteProtocolError. Besides RemoteAgentAlreadyRunningError / RemoteAgentNotFoundError / RemoteAgentFeatureUnavailableError, an envelope with an errorCode (or any 503) is a HelixError with that code and cause (retryable exactly on a 503), and a non-envelope typed body (agent-server's { error, code }) is helixErrorFromPeerBody. HttpRemoteAgentTransport.resume() used to throw a generic Error for every answer.
Before:
try {
await transport.start(request);
} catch (err) {
if (err instanceof Error && err.message.includes('state_history_incomplete')) return reload();
throw err;
}After:
import { HelixError } from '@helix-agents/core';
try {
await transport.start(request);
} catch (err) {
if (HelixError.isInstance(err) && err.code === 'state_history_incomplete') return reload();
throw err;
}runtime-js refusals are typed. A resume() of a completed or failed session, or of a paused session still awaiting client-tool results, throws core ResumeRefusedError (state_not_resumable, with sessionStatus and awaitingToolCallIds; the partial-batch case is retryable). So does a retry() of a session that is not failed (non-retryable; its message still says retry() is only for failed sessions). An unknown session is a non-retryable HelixError state_session_not_found from resume(), retry() and a reconnect handle's resume (getHandle() still returns null for it). Another agent type's session is a non-retryable state_agent_type_mismatch from resume(), retry(), getHandle() and a reconnect handle's resume. resume({ mode: 'with_confirmation' }) is framework_not_supported. All of these were plain Errors with the same message text, which core classifyEntryFailure read as transient (a host answered a retryable 503). retry() honours options.usageStore.
An already-aborted signal rejects the entry. execute(), resume(), retry() and recover() honour options.abortSignal: an entry whose signal is aborted before its run-start commit (an already-aborted signal included) rejects with a HelixError whose code is framework_cancelled (non-retryable) and writes nothing: no run, no message, no stream chunk, no lifecycle hook. Only a session that never had a run (the one a fresh execute() just created) is recorded failed on its session (rule M1). Before, execute() with an already-aborted signal returned a handle whose result() was failed.
Before:
const handle = await executor.execute(agent, input, { sessionId, abortSignal });
const result = await handle.result();
if (result.status === 'failed') return showCancelled();After:
try {
const handle = await executor.execute(agent, input, { sessionId, abortSignal });
const result = await handle.result();
} catch (err) {
if (HelixError.isInstance(err) && err.code === 'framework_cancelled') return showCancelled();
throw err;
}Before:
try {
await executor.resume(agent, sessionId);
} catch (err) {
if (err instanceof Error && err.message.startsWith('Cannot resume')) return showRetry();
throw err;
}After:
try {
await executor.resume(agent, sessionId);
} catch (err) {
if (HelixError.isInstance(err) && err.code === 'state_not_resumable') return showRetry();
throw err;
}Companion delivery
On a DO with subAgentNamespace, the persistent-companion tools map the child DO's answer onto the CP-105 delivery contract (companion__sendMessage):
Child /resume answer | Before | Now |
|---|---|---|
2xx RemoteStartResponse | { delivered: true } | unchanged; the parent ref is written running after it |
2xx that is not a RemoteStartResponse (a hook's own 200, a 204, non-JSON) | { delivered: true } | a RemoteProtocolValidationError tool error; ref untouched |
409 ALREADY_RUNNING (the child is executing) | a thrown ChildDOFetchError | { delivered: false, reason: 'child_busy' }; ref untouched |
| no answer within 60 s, a rejected fetch, or an unreadable body | the tool hung, or threw | { delivered: false, reason: 'child_start_unconfirmed' } |
| any other answer | a thrown ChildDOFetchError | a HelixError with the envelope's errorCode (or code) and cause, retryable only on a 503 |
Before:
const out = await callTool('companion__sendMessage', { name: 'helper', message });
// a busy child threw "Child DO resume failed (409): …"After:
const out = await callTool('companion__sendMessage', { name: 'helper', message });
if (!out.delivered && out.reason === 'child_busy') {
// the child is running: try again after it settles
} else if (!out.delivered && out.reason === 'child_start_unconfirmed') {
// unknown whether the child took the message: check getChildStatus before re-sending
}- The 60 s bound (
CHILD_RESUME_TIMEOUT_MS) is internal, not an option. It exceeds the 30 s defaultconcurrentEntryTimeoutMs, so a child underinterrupt_previousthat times out answerschild_busyfirst; a child configured withconcurrentEntryTimeoutMsof 60 s or more reportschild_start_unconfirmedinstead. - A spawn whose
/startanswers 409 attaches to the running child (a replay) only on theALREADY_RUNNINGenvelope; any other 409 is a typed tool error. - The companion
/startbody carries the parent'suserId,tagsandmetadata(it carriedmetadataonly). - The spawn and continuation
/startand the/statuspolls are not bounded yet (FU-DO-COMPANION-START-FETCH-UNBOUNDED). ChildDOFetchErroris deleted (it was module-internal).
Never-run sessions
A session that never had a run (an entry failed before its first run-start commit, e.g. a short history read, or the DO was evicted between createSession and the commit) is recorded failed on the SESSION only: the failed status and errorDetail, idle-fenced so a run that commits meanwhile wins. No stream is created or written and no lifecycle hook fires. Streams are written only by their runs.
- runtime-js no longer fails the stream or fires
onAgentFailfor a never-run pre-commit failure (astate_history_incompletefirst history read, a cancellation, any other class). A session that has a run is still left untouched. - The CF DO wake fails an
activesession with no run record instead of resuming it on an empty history. - The snapshot of a never-run failed session (
readSnapshotView/buildSnapshot) readsstatus: 'failed'with the session'serrorDetail(it readended). - Core
failSessionWithoutRun({ currentRunId: null })is that write. With a run id it also fails the stream, fenced on that run.
Before:
hooks: {
onAgentFail: async ({ sessionId, errorDetail }) => alert(sessionId, errorDetail), // fired for a never-run failure
},
// and a reader watched the stream:
const info = await streamManager.getStreamInfo(sessionId); // 'failed', with the errorAfter:
// The failure is on the session (and on the handle / the rejection the caller got):
const state = await stateStore.loadState(sessionId);
if (state?.status === 'failed' && (await stateStore.getCurrentRun(sessionId)) === null) {
alert(sessionId, state.errorDetail); // never-run: no stream, no onAgentFail
}Per runtime: runtime-js and the CF DO follow this rule. CF Workflows, Temporal and DBOS still write a never-run session's stream (or, on DBOS, record failed whenever no run is live); see CLAUDE.md and the follow-ups.
Wake recovery
The onStart / alarm wake re-drives a stranded run through the same pipeline, and every failure is classified:
| Class | Wake action |
|---|---|
lost | nothing (another drive owns the session) |
refused, fatal | the run ends failed with the classified detail (K5 terminal commit), then settles |
transient | re-arm under the crash-loop budget; the failure is recorded as WakeRecoveryRecord.lastFailure |
- Exhaustion ends the run
failedwithframework_auto_resume_exhausted(retryable:/retrymay succeed), whosecauseis the last re-armed failure. It replaces the untypedauto_resume_exhausted, and the message no longer blames "a step that crashes the instance". - The wake's executor now has the DO's workspace providers, and the wake fires the host hooks (
afterEntry,onRunSettled). onStartandonAlarmnever throw. A sub-agent withoutoutputSchemaused to escapeonStartand fail every later request to the DO; it now ends the runvalidation_error.- A failed submit continuation (the run a completed client-tool batch continues) is classified the same way, so a continuation whose agent no longer resolves ends
failedwithframework_agent_not_foundinstead of exhausting.
Before:
if (state.status === 'failed' && state.error === 'auto_resume_exhausted') showRetry();After:
if (state.status === 'failed' && state.errorDetail?.code === 'framework_auto_resume_exhausted') {
const cause = state.errorDetail.cause; // the last failure (string or ErrorDetail)
showRetry(typeof cause === 'object' ? cause.code : undefined);
}A run whose owner is gone is settled through core's closed-owner reconciliation (planClosedOwnerReconciliation → reconcileClosedOwnerRun, the same pair CF Workflows' readers use) before the ladder above runs:
- a run whose K5 landed is finished and settles once (
onRunSettled), as before; - new: a run evicted between its suspension commit and its
suspended_*run status (the sessionpaused, the run stillrunning) gets its suspended status and its stream paused, and settlessuspended_*. It used to stayrunninguntil the next entry superseded it; - an ended run whose stream is still open has the stream ended or failed (no hook);
- a run with nothing committed is re-driven, as before;
- a store read that fails permanently (a corrupt row) ends a stranded
activerunfailedtyped instead of re-arming on every alarm. A transient one is retried on the next alarm.
autoResumeOnWake: false keeps the demote to paused for a stranded active run; a paused session whose whole client-tool batch was submitted stays paused until an explicit /resume.
The submit route (/submit-tool-result) resolves the owning agent and materializes the continuation BEFORE it writes: a failing factory answers typed (503 framework_agent_resolution_failed, framework_usage_store_unavailable, …) and the call stays pending. It used to answer 200 and lose the error to the wake.
A submit that completes its batch now answers after the continuation's entry has started (or another drive owns the run), including a submit that lands while the suspending run is still settling; it used to answer before any entry. That wait is bounded by concurrentEntryTimeoutMs (default 30 s): past it the DO answers 500 transport_error (the result is stored and the run continues via the unconsumed-submit alarm, so re-read the snapshot instead of re-submitting; the message says the continuation was not confirmed, since it may already have started). A duplicate submit (the result is already stored or consumed) answers 200 already_completed at once and never waits for the continuation; it means this call's result is recorded, not that the run has finished. A forwarded submit whose owning DO does not answer within that bound plus 1 s answers 503 transport_error; re-submitting it is safe. A forwarded submit relays the owning DO's typed { error, code } answer with its status, a 404 included (e.g. framework_agent_not_found; it used to be answered as a retryable 503); a 404 that is neither typed nor unknown_tool_call answers a non-retryable 500 transport_error.
A DO submit can now answer a typed failure or a timeout where it always answered 200. When the continuation the submit starts fails to enter, the answer follows the failure table's wake row, on the committed row (never 503): refused / fatal → the run's typed failure (it ended failed with that ErrorDetail); transient → its code on the non-retryable row (500, or 409 for a state error); the run is re-armed and continues, so re-anchor through the snapshot (useHelixChat().retry()) instead of re-submitting. lost (another drive owns the run) still answers 200 accepted.
The route's own store read and alarm writes never reach the untyped 500 { error: 'Internal server error' } either. A failed read of the owning session (nothing written yet) answers typed (an unclassified failure is a retryable 503). Once the result is stored, a failed alarm reschedule falls back to a backoff alarm and the answer stands, and a failed read of the batch (to decide whether to continue) answers its code on the non-retryable row (500, never 503; the unconsumed-submit alarm continues a complete batch, so re-read the snapshot instead of re-submitting).
Registries
| Registry | resolve() |
|---|---|
@helix-agents/runtime-cloudflare AgentRegistry | async-compatible: a static registration synchronously, a factory's result as it returns it (a Promise for an async factory); AgentFactory may be async |
@helix-agents/runtime-temporal AgentRegistry | stays synchronous: Temporal activities resolve through your synchronous resolveAgent |
Before:
const config = registry.resolve('assistant', { env, sessionId, runId, userId });After:
const config = await registry.resolve('assistant', {
purpose: 'drive',
env,
sessionId,
runId,
drive: 'start',
identity: { userId },
});Store and runtime implementers
An entry finishes a SUSPENDED run whose K5 landed (R66). applyRunStartRules verifies finishCurrent for any live current run (running or suspended_*), and every runtime's entry passes finishCurrent for one. A stop on a paused run whose run-status write was lost is now recorded with the K5's status (interrupted), not superseded. A custom store whose finish write pins status = 'running' must pin the live statuses instead (D1 did); the contract scenario run.k5-landed-finishes-suspended checks it.
A permanent store error is never wrapped. Core classifyEntryFailure classifies a corrupt-row error (InvalidStoredStateError, CorruptMessageRowError, UnknownCommitKindError) as fatal (core isPermanentStoreError), so a DO wake ends the run typed instead of re-arming, and a DO request (so a chat route backed by the CF DO) answers it typed and non-retryable (500) instead of 503. DOStateStore rethrows every permanent store error as itself: an InvalidStoredStateError used to reach callers wrapped in DOStateError. A custom store should not wrap one either.
assertRunStatusTransition is renamed decideRunStatusTransition and returns 'write' | 'noop': a terminal run written with its own status is a no-op (write nothing), a live run always writes, and any other transition of a terminal run still throws RunSupersededError. A custom store's updateRunStatus must return without writing on 'noop'. The run-start rule order also changed: applyRunStartRules verifies finishCurrent before the owner guard.
Before:
assertRunStatusTransition({ sessionId, run, to: status });
await writeRunStatus(runId, status, metadata);After:
if (decideRunStatusTransition({ sessionId, run, to: status }) === 'noop') return;
await writeRunStatus(runId, status, metadata);updateRunStatus to a live status never sets completedAt, even when the caller passes one; a terminal status defaults it to now. endSequence is stored only as the caller passes it (the D1 and DO stores used to default it to a message sequence). The contract scenario is run.live-status-no-completedAt in @helix-agents/core/testing.
The store-cloudflare stream Durable Object answers an /end, /fail or /pause whose expectRunId fence it cannot read with 400 validation_error and applies nothing (it used to apply an unparseable /pause unfenced); the binding throws StreamConnectionError.
runtime-js:
RunStartEntryState.buildreturnsRunStartEntrySessionState, whosestatus('active' | 'paused') is required; a customrunStartCommitterwrites the status as given.Before:
typescriptbuild: (session) => ({ ...session, stepCount: 0 }),After:
typescriptbuild: (session) => ({ ...session, stepCount: 0, status: 'active' }),JsClientToolResolver'sgetCompletedRetentionMsForRootoption is deleted. A call's completed-tombstone retention is recorded on its pending entry at registration (PendingClientToolCall.completedRetentionMs, from the agent'scompletedTombstoneRetentionMs);registerPendingtakes an optionalcompletedRetentionMs.Before:
typescriptnew JsClientToolResolver({ stateStore, getCompletedRetentionMsForRoot: async (root) => lookup(root), });After:
typescriptnew JsClientToolResolver({ stateStore }); // retention comes from each pending entryJSAgentExecutortakessubmitForwarder?: SubmitForwarderfor submits another host owns, and itssubmitToolResultowns the whole submit.JsClientToolResolver.submitis internal; it maps anapproval-responseitself.Before:
typescript// An approval mapped by hand (approved / reason were dropped if you got it wrong): await resolver.submit({ kind: 'client-tool-result', sessionId, toolCallId, result: { approved: true }, });After:
typescriptawait executor.submitToolResult({ kind: 'approval-response', sessionId, toolCallId, approvalId, approved: true, });A submit whose submitting session's agent type is unknown to the configured registry fails
framework_agent_not_foundinstead of skipping output validation.defineAgentthrows unlesscompletedTombstoneRetentionMsis a finite number>= 0(0disables retention).
A new host or runtime embedding runtime-js the way the DO does classifies entry failures with core classifyEntryFailure (lost / refused / transient / fatal) and fails a never-run session with failSessionWithoutRun.
What to check
- Replace every removed hook,
ExecutionStatemethod, factory signature and/resumeoption above. TypeScript reports each one. - Deduplicate
onRunSettledondeliveryId. - Authorise
/submit-tool-resultin the worker in front of the DO. - A
/submit-tool-resultcaller handles a typed failure or a 500 / 503transport_errortimeout (above) instead of treating every answer asacceptedor waiting indefinitely. - Callers that matched the native
ALREADY_RUNNINGenvelope, a 400 for an unknown agent,DOPeerError, a generic transportError, orauto_resume_exhaustednow get the typed answers above. - A consumer usage store that can fail now stops runs (503 / re-arm).
- Stop sending per-turn
metadata,userIdortagson a continuing/start.