Skip to content

Migrating to Complete LLM History and Checkpoint Rewind ​

Overview ​

This release (@helix-agents/core 0.50) makes every runtime give the model the complete, persisted conversation history, and makes checkpoint rewind, branch and retry behave the same everywhere. It fixes several silent failures:

  • Temporal sent the model no history at all (RM-10). Every Temporal LLM call received only the system prompt. With a real provider every call failed with messages must not be empty, from the first turn on. Upgrade Temporal deployments first.
  • DBOS sent the oldest 10,000 messages of a longer session (RM-11), checkpointed that capped count (RM-12), and wrote a content-less user row on every retry() / resume() (RM-28).
  • resume({ mode: 'from_checkpoint' }) did not rewind to the target on JS / CF DO (RM-25) or CF Workflows (RM-31), was ignored on Temporal (CP-31) and DBOS (CP-32).
  • Branch ignored checkpointId / messageIndex on Temporal and CF Workflows (CP-33).
  • A checkpoint id from another session was accepted (CP-38); on CF Workflows it could overwrite the other session's state.
  • Stores reported stale message counts after a truncate (RM-16 Postgres, RM-22 Redis, RM-30 D1), the DO store's checkpoints claimed 0 messages (RM-36), and DBOS / CF Workflows wrote no checkpoint for turns that ran no tools (RM-33).
  • An interrupted persistent companion could never be resumed on JS / CF Workflows (CP-68).

This release also makes @helix-agents/store-redis session creation and appendMessages atomic, and appendMessages on a missing session now throws: see Redis: atomic session writes.

Most of this needs no code change. The items below do change behavior or need action. See the Checkpoints guide and Interrupt and Resume → from_checkpoint for the resulting semantics.


1. DBOS: drain in-flight workflows before deploying ​

What changed ​

The DBOS workflow body now calls different steps in a different order:

  • checkpoints read the message count with the getMessageCount step instead of getMessagesStep (whose output is now the complete history);
  • turns that end without running tools write an extra checkpoint step;
  • the initial appendMessages step is skipped when there is nothing to append (every retry() / resume()).

There is no DBOS patch marker for this change. A workflow that is pending when the new code deploys replays against the new step sequence on recovery and fails.

What you need to do ​

Before rolling out:

  1. Stop starting new DBOS runs on the old version (or accept they will be affected).
  2. Let running workflows finish, and resolve or cancel suspended ones (workflows waiting in DBOS.recv on a client tool or approval).
  3. Persistent-mode session workflows never finish on their own: cancel them with DBOS.cancelWorkflow. The next execute() on the session starts a fresh workflow.
  4. Deploy.

Sessions with content-less user rows fail loudly ​

Before this release, DBOS retry() and resume() appended a user row with no content (RM-28). The history loader now counts every row, so a session holding such a row fails its next run with state_history_incomplete and makes no model call. The error message names this cause. There is no repair tool and no compatibility path: start a new session for that conversation.


2. Temporal: RunLLMStepResult is delta-only ​

What changed ​

runLLMStep now loads the history inside the activity and returns only what the step produced:

typescript
interface RunLLMStepResult<TState, TOutput> {
  stepResult: StepResult<TOutput>;
  nextState: AgentState<TState, TOutput>; // nextState.messages is always []
  newMessages: Message[]; // NEW: messages this step appended, in order
  baseMessageCount: number; // NEW: length of the history the step was built from
  // ...other fields unchanged
}

History never travels through workflow code or activity payloads.

What you need to do ​

  • If you call runLLMStep yourself (a custom workflow), read newMessages for the step's messages and baseMessageCount + newMessages.length as the absolute index base. Do not read nextState.messages.
  • A history-read failure now fails the activity with an ApplicationFailure carrying the classified ErrorDetail. A persistent count mismatch is non-retryable (it is not retried three times). A store throw while loading is retryable, so the activity retry policy applies and a transient Redis / Postgres blip does not fail the run.

Deploying: drain, and never mix old and new workers ​

The workflow code and the runLLMStep result shape changed together, and there is no patched() guard. Mixing versions breaks in both directions:

  • New workflow code replaying an old runLLMStep result. The recorded result has no newMessages, so the workflow task throws TypeError: Cannot read properties of undefined (reading 'length') (stepResp.newMessages.length). Temporal retries the workflow task forever and the workflow stays Running without progressing. No history is lost, but the workflow must be terminated or reset.
  • Old workflow code receiving a new runLLMStep result (old and new workers on the same task queue, e.g. during a rolling deploy). The old code commits the step's messages from nextState.messages, which is now []: the assistant message is silently not persisted, and the tool results appended after it are orphaned.

What to do:

  1. Drain in-flight Temporal workflows before deploying. Let running workflows finish (sessions suspended at a HITL boundary have already exited their workflow and are not affected).
  2. Never run old and new workers on the same task queue. Cut over fully (stop every old worker before starting new ones), or isolate the versions with Worker Versioning or a new task queue.
  3. If a workflow is stuck with the workflow-task error Cannot read properties of undefined (reading 'length'), terminate or reset it.

3. New error codes ​

Three ErrorCode members were added, all category: 'state'. They are retryable: false, except state_history_incomplete raised for a store throw while loading, which is retryable: true (a transient connection blip; the runtime's retry policy applies — Temporal activity retry, DBOS / CF Workflows step retry):

CodeWhen
state_history_incompleteThe store returned fewer (or more) messages than its live row count, or threw while paging. The run fails; no model call.
state_checkpoint_not_foundAn explicit checkpointId for retry() / from_checkpoint / branch does not exist.
state_checkpoint_foreignThat checkpointId belongs to another session.

What you need to do ​

  • An exhaustive switch over ErrorCode needs the three new cases.
  • Code that matched the old messages (Checkpoint not found: …, Checkpoint … belongs to a different session) should match HelixError.isInstance(err) && err.code === '…' instead.
  • A state_history_incomplete message gives the counts read and expected, then the likely cause: a persisted row that failed validation or was skipped (naming the pre-upgrade DBOS content-less rows, RM-28), or the store error when a store call threw. Match on the code, not the text.
  • state_history_incomplete replaces silent truncation. A store that skipped a corrupt or unreadable row used to hand the model a shorter transcript. The run now fails with this code in errorDetail (persisted, on the failed stream chunk and in onAgentFail). If you see it, the store's rows and its getMessageCount disagree: fix or remove the bad row.

4. A checkpoint from another session is rejected ​

retry({ checkpointId }), resume({ mode: 'from_checkpoint', checkpointId }) and execute(..., { branch: { fromSessionId, checkpointId } }) now require the checkpoint to belong to the session they act on (for a branch, fromSessionId). A foreign id throws state_checkpoint_foreign before any state change: no CAS, no version bump, no clone, no workflow start. Previously JS, Temporal and CF Workflows accepted it; CF Workflows then saved the checkpoint's state under the OTHER session's id.

On every runtime, including Temporal and CF Workflows, the rejection comes from the execute() / resume() / retry() call itself, not from a failed workflow or instance.


5. from_checkpoint discards pending client-tool calls it rewinds past ​

resume({ mode: 'from_checkpoint' }) now truncates the durable message log to the checkpoint on every runtime. Client-tool or approval calls that were pending and were made after the checkpoint are discarded:

  • a tool_end chunk with errorCode: 'aborted' is emitted for each, so a UI waiting on the call can settle it;
  • a later submitToolResult for one returns { status: 'unknown_tool_call' }.

If your client keeps its own list of pending tool calls, clear an entry when you see an aborted tool_end for it. A call that was pending AT the checkpoint and answered after it is pending again after the rewind when the store recorded it on the checkpoint.

run_resumed.fromCheckpointId and onAgentResumed's resumedFromCheckpointId now report the target checkpoint (CF Workflows used to report the old latest one).

Because the rewind truncates the log, it is refused while the session's run is still executing. On every runtime, from_checkpoint on an active session with no pending client-tool or approval calls throws AgentAlreadyRunningError and changes nothing (previously JS, CF Workflows and Temporal claimed the session and truncated the log under the live run). Interrupt the run first, or wait for it to finish or suspend. Two concurrent from_checkpoint calls now have exactly one winner on every runtime; the loser throws AgentAlreadyRunningError (Temporal used to start a __resume-N workflow for both). So does a from_checkpoint racing a continue / with_message, by fencing rather than a clock: on Temporal / CF Workflows the run that takes held calls over (the rewind, or a continue) starts under …__rewind-<key> / …__rewind__<key> instead of …__resume-N / …__resume__N, so only one can run; on JS the rewind's takeover write and a continue's claim are pinned against each other. After a crashed rewind, a continue takes the calls over immediately. On CF Workflows a continue now claims (version-pinned) before its staging cleanup and truncate, and two concurrent continue calls have one winner (both used to create an instance).

A failed resume() (any mode) no longer strands the session: a failure after the claim and before the run starts rolls the session back to its previous status, on JS, Temporal and CF Workflows — unless a from_checkpoint rewind already changed the session (it saved and then failed, or it rewound a completed / failed session whose run could not start). Then the session ends failed with a typed errorDetail and a failed stream, on every runtime including DBOS, rather than showing its old status over a rewound log. A history-read failure during the rewind (it happens before the save) now restores the previous status on CF Workflows and JS (CF Workflows used to mark the session failed). On Temporal, from_checkpoint on a finished session whose stream has not ended yet (its old workflow is still finishing) throws AgentAlreadyRunningError; retry it once the stream has ended.

from_checkpoint on a completed / failed session now rewinds (behaviour change) ​

resume({ mode: 'from_checkpoint', checkpointId }) on a completed or failed session now rewinds to the target checkpoint and starts a run from there on every runtime (C7): the log is truncated to the checkpoint, its state restored, the old output / error cleared, and the ended / failed stream reactivated for the new run. continue and with_message on a finished session are unchanged (Temporal still returns the terminal result; the others throw; use retry() for a failed session).

RuntimeBeforeNow
JS (and CF Durable Object)Rewound to the checkpoint and ran.Unchanged.
CF WorkflowsThrew Cannot resume: Agent already completed / … has failed.Rewinds and runs.
DBOSThrew an untyped Error (the session is not in a resumable state).Rewinds and runs.
TemporalReturned a handle resolving to the OLD terminal result: nothing rewound, no workflow started, no error.Rewinds and runs.

What you need to do (Temporal / DBOS / CF Workflows): a from_checkpoint call on a finished session now changes it — it truncates the durable log and starts a run (an LLM call). If you relied on Temporal returning the old result, read the result with getHandle() / the session state instead; if you caught the DBOS / CF Workflows error to detect a finished session, check session.status before calling resume().


6. retry() after a forced-completion failure needs options.message ​

A turn that ends in a framework-forced terminal tail (maxSteps, or forced completion failing — for example a forced finishWith that fails) now writes a terminal checkpoint on every runtime (RM-33). That checkpoint already contains the user's message, so there is no "triggering" message after it:

  • CF Workflows: retry() without options.message then throws No user message found to retry with. Provide a message in options., as runtime-js and Temporal do (CP-58). Pass the message explicitly.
  • DBOS: retry() still ignores RetryOptions (CP-30). It continues from the persisted history, which already holds the user message, and no longer appends a content-less row.

An exception exit — a terminal LLM error, an error stop reason, a non-recoverable (content_filter / refusal) stop, an uncaught exception, a history-read failure, or an interrupt — writes no extra terminal checkpoint on any runtime (ruling R28). retry() truncates to the latest checkpoint; what it covers is runtime-specific and unchanged by this release: Temporal and DBOS keep the committed failing step, and on JS a first-step failure on a store whose saveState writes no checkpoint needs options.message.

This applies where the runtime classifies the finish as a failure: Temporal and DBOS currently treat a content_filter / refusal / max_tokens / error / unknown finish of an agent without outputSchema as completed, not failed (DI-26), so there it writes the completion checkpoint.


7. An interrupted companion cannot be re-spawned by name ​

On JS and CF Workflows, a persistent child that was interrupted (for example by interrupting its parent) now stays interrupted and resumable instead of failed:

  • companion__sendMessage resumes it on its preserved history;
  • companion__waitForResult returns { status: 'interrupted', message } instead of waiting until its timeout;
  • companion__spawnAgent with that child's name is refused: Agent '<name>' is interrupted. Use companion__sendMessage to resume it. Previously the child was failed, so a re-spawn started a fresh session.

If your prompts tell the model to re-spawn a companion after an interrupt, tell it to use companion__sendMessage instead.

On CF Workflows a resumed companion runs in a new instance agent__<type>__<childSession>__resume__<stepCount>__<toolCallId> and writes to its own stream (the child session id), not the parent's.

For custom runtimes built on executeCompanionToolDispatch: ChildTerminalResult.status may now be 'interrupted', which records the ref 'interrupted' rather than a terminal status.


8. Custom state stores ​

If you implement SessionStateStore yourself:

  • getMessageCount must return a live row count. loadLLMHistory compares every history read against it.
  • truncateMessages must update any denormalized message count in the same transaction.
  • Every checkpoint you write must record the live message count as messageCount.

Run stateStoreContractTests from @helix-agents/core/testing; it now includes truncate-then-checkpoint-count, truncate-then-list-count and a saveState-minted-checkpoint case (skip it with skip.saveStateCheckpoint if your saveState does not write checkpoints).


Validation checklist ​

  • [ ] DBOS: in-flight workflows drained or cancelled before the deploy.
  • [ ] Temporal: in-flight workflows drained; no old and new workers on one task queue (full cutover, Worker Versioning or a new task queue); any custom workflow reads newMessages / baseMessageCount.
  • [ ] Exhaustive ErrorCode switches handle the three new codes.
  • [ ] retry() calls pass message where a first-step failure is possible.
  • [ ] Client UIs settle a pending tool call on an aborted tool_end.
  • [ ] Prompts do not rely on re-spawning an interrupted companion by name.
  • [ ] Custom stores pass the updated contract suite.

Released under the MIT License.