Typed chat resume and submit failures (DI-39)
This guide covers the release in which a chat resume (a handleChatStream request that carries client-tool or approval results) answers a typed failure as a typed HTTP error, as a chat send already does, and agent-server's POST /start answers a typed entry failure the same way. It is a major release of @helix-agents/ai-sdk (37.0.0) and of @helix-agents/agent-server (3.0.0): both change a wire answer.
The send half shipped earlier, in CF-D2 MR 1 (!316): see CF Workflows terminal settle → Over HTTP: a typed error, never started. A chat send whose execute() fails with a HelixError (a completion-support refusal included) already answers { error, code } with core typedErrorResponse's status. This release makes the resume path match it.
What changed
handleChatStream: a resume or submit that fails typed
Every handleChatStream host is affected: agent-server POST /chat, the Express adapter, createCloudflareChatHandler (the CF Durable Object chat route) and any custom host (a CF Workflows chat host is one).
| A chat resume whose … | Before | Now |
|---|---|---|
executor.submitToolResult() rejects with a HelixError (e.g. the CF DO's refusal) | 200 SSE, data-resume-rejected, no code | { error, code }, core status, no resume |
executor.resume() rejects with a HelixError (e.g. runtime-js's refusal) | 200 SSE, data-resume-rejected, no code | { error, code }, core status |
executor.resume() returns a handle with preCommitFailure (CF Workflows) | 200 stream of the handle | { error, code }, core status |
submit or resume rejects with a plain Error | 200 SSE, data-resume-rejected | unchanged |
| resume fails after the submit already woke the run (the CF DO's auto-resume race) | 200 stream of the woken run | unchanged |
The status is core httpStatusForHelixError's: 400 for a completion-support refusal (framework_completion_strategy_unavailable / framework_completion_schema_unsupported) or validation_error, 409 for a non-retryable state_* error (e.g. a pre-commit state_history_incomplete), 503 for a retryable error, else 500. The body is exactly { error, code } with the message scrubbed, plus the X-Session-Id header, before any stream opens.
A typed error replaces any per-intent rejections. When one intent's submit fails typed, the request ends there with that typed error alone: rejections gathered for other intents (a plain Error submit failure, a stale intent) are not added to it. The request failed as a whole; re-read the snapshot and re-send.
Never 503 once an answer is stored. When this request already stored a submitted result (its submit returned accepted), a retryable error is answered with its non-retryable row (409 / 500): re-sending the request would submit the answer again.
A chat resume refused after its answer was stored (runtime-js). A runtime-js submit never refuses: it stores the answer on the pending call (pendingClientToolCalls[id].submittedResult), then the resume is refused with the 400. The stored answer is kept. The next resume on an adapter that can finish the agent (the model fixed, or its capabilities declared) continues from it. The CF Durable Object refuses the submit itself, so there nothing is written and the call stays pending.
agent-server POST /start and POST /resume
POST /start and POST /resume answer an entry that failed typed with core typedErrorResponse, as the CF Durable Object's /start does:
AgentServer.startAgentandAgentServer.resumeAgentreject a handle whose entry failed before its run-start commit (preCommitFailure) withhelixErrorFromDetail(handle.preCommitFailure)instead of answering{ sessionId, streamId, runId }(started). On/resumethe old answer made the caller attach to the session's stream, which still held the previous turn. The thrown error keeps the handle's detail:resolveErrorDetail(error)returnspreCommitFailureverbatim.- A thrown
HelixErroranswers{ error, code }with the core status instead of500{ code: 'INTERNAL_ERROR', errorCode }. - A retryable error answers
503only when a re-sent request can succeed:- on
/start, only while the failed entry left no session behind. Once the session exists (the entry's run committed, or runtime-js / CF Workflows created a fresh session before a pre-commit failure), a re-sent/startis refused bystartAgent's pre-check, so the error gets its non-retryable row (409for astate_*error, else500); - on
/resume, only while the session's current run did not move during the call (the resume committed no new run and its message is not in the log).
- on
- The session or run read behind that decision answers "committed" when it fails, and is logged (
warn) through the server'slogger, now a public read-onlyAgentServer.loggerproperty. - A completion-support refusal already answered the typed
400; it is unchanged.
What to check
- A client that read
data-resume-rejectedfor these failures now receives a non-2xx response. A stockuseChat/useHelixChatends the request instatus: 'error'; read the code withparseHelixChatError(chat.error). A stock client needs no code change. - A custom transport that treated every chat POST as a
200stream must handle a4xx/5xxJSON body{ error, code }on a resume, as it already does for a send (409 state_already_running, the typed send failures). - A caller of agent-server
POST /startorPOST /resumethat matched500{ code: 'INTERNAL_ERROR', errorCode }for a typed executor failure now gets the core status and{ error, code }(the code incode, noterrorCode). A/resumethat used to answer200for a resume that failed before its run-start commit now answers that typed error. Follow a503with a retry; a409/500means re-read the session first. - Persisted state is unchanged by this release.