Stop conditions, maxSteps failures and verbatim string tool results
This guide covers the release that gives every runtime one stop decision and stores string tool results verbatim (DI-29, DI-25, RM-32). It is a major release of @helix-agents/runtime-js, @helix-agents/runtime-cloudflare, @helix-agents/runtime-temporal, @helix-agents/runtime-dbos and @helix-agents/ai-sdk (a string tool result now renders as that string), and a minor release of @helix-agents/core (a 0.x package: minor marks an observable behaviour change), @helix-agents/store-cloudflare and @helix-agents/sdk.
What changed, by runtime
| Behaviour | JS / Cloudflare DO | Cloudflare Workflows | Temporal | DBOS |
|---|---|---|---|---|
stopWhen is evaluated | yes (unchanged) | yes (new) | not yet (CF-D1 MR 2) | not yet (CF-D1 MR 3) |
stopWhen is checked before maxSteps on the cap step | yes (new) | yes (new) | n/a | n/a |
A maxSteps cut-off (no outputSchema) | failed, typed | failed, typed | completed (unchanged) | completed (unchanged) |
retry() of a cut-off continues with a fresh budget | yes (new) | yes (new) | n/a (the run completed) | n/a (the run completed) |
RunMetadata.completionReason (stop_when / max_steps) | yes (new) | yes (new) | not written | its own reasons (DI-27) |
| A string tool result is stored verbatim, with a marker | yes (marker new) | yes (new) | yes (new) | yes (new) |
| Child instance ids keyed on the parent run | n/a | yes (new) | n/a | n/a |
1. A maxSteps cut-off now fails
On JS, the Cloudflare DO and Cloudflare Workflows, an agent without outputSchema that runs out of maxSteps before the model finished (it called a tool, or wrote a non-final step, on the last allowed step) ends failed:
result.status; // 'failed'
result.errorDetail;
// { message: 'Max steps (5) reached', code: 'framework_max_steps_exhausted',
// category: 'framework', retryable: false }The detail also reaches SessionState.errorDetail, the failed stream chunk and onAgentFail (which now fires instead of onAgentComplete), and the run record carries completionReason: 'max_steps'. A parent agent receives a cut-off sub-agent as a failed tool result. Before, the run ended completed with output: undefined (Cloudflare Workflows failed it with an untyped Max steps (N) reached).
Unchanged: a final answer on the last allowed step still completes; an agent with outputSchema still enters forced completion at the cap; a step that stops on an error stop reason (content_filter, refusal, …) still fails with that reason.
Before (code that treated a cut-off as a completion):
const result = await handle.result();
if (result.status === 'completed') {
// A finished answer, or a run that simply ran out of steps.
render(await loadTranscript(handle.sessionId));
} else {
reportError(result.error);
}After:
const result = await handle.result();
if (result.status === 'completed') {
render(await loadTranscript(handle.sessionId));
} else if (result.errorDetail?.code === 'framework_max_steps_exhausted') {
// Out of budget, not broken: continue it (see §3), or show the partial transcript.
const continued = await executor.retry(agent, handle.sessionId);
await continued.result();
} else {
reportError(result.errorDetail ?? result.error);
}Also check:
- Hooks and alerting.
onAgentFailnow fires for a cut-off. Classifyframework_max_steps_exhaustedas "budget", not as an outage. - Parents of sub-agents. A cut-off child is now a failed tool result (
{"error":"Max steps (N) reached"},helixToolFailed: true) that the parent's model sees. - Parents of persistent companions. A persistent companion without
outputSchemathat runs out ofmaxStepsnow endsfailed, notcompleted. A failed companion is not continuable:sendMessageto it throws "not active (status: failed)", and the nextspawnAgentre-spawns it fresh, so it loses its memory of earlier rounds. Before, it completed and the next re-consult continued its session. If a companion must stay continuable, give it astopWhenthat ends its tool loop successfully, or a largermaxSteps(FU-MAX-STEPS-COMPANION-CONTINUATION). See Re-consulting a persistent companion. - A bigger budget. Raise
maxSteps, or setmaxSteps: Infinityto opt out of the cap (the default is 50).
2. stopWhen is the way to stop successfully
To end a tool loop successfully before the budget runs out, use stopWhen (or a finishWith tool, or an outputSchema). The stop decision is the new core resolveStopDecision, checked in this order:
- terminal: structured output, or a text / error step that ends the run by itself;
stopWhen: the predicate returnstrue;maxSteps:stepCount >= maxSteps.
const agent = defineAgent({
name: 'indexer',
maxSteps: 20,
// Completes the run (completionReason: 'stop_when') once `done` was called.
stopWhen: (result) =>
result.type === 'tool_calls' && result.toolCalls.some((tc) => tc.name === 'done'),
// ...
});What changed for existing stopWhen users:
- Cloudflare Workflows now evaluates it. Before, the predicate was ignored there. The predicate runs inside the LLM step's
step.do, so a replay reuses its answer, and a throwing predicate fails the run before that step's tools run (on JS and the DO, after them). - It runs on the cap step.
stopWhenis now checked beforemaxSteps, so a predicate that fires on the last allowed step is a successfulstop_whenstop. Before, JS stopped atmaxStepswithout calling it on that step. Custom loops built onshouldStopExecution()see the same change: the predicate is now called on the cap step too. - A throwing predicate fails the run once with its error; it is never retried.
- With
outputSchema, a predicate that fires before any output enters forced completion, as atmaxSteps.
Custom loops that need the cause should switch from shouldStopExecution (still exported, now the boolean form) to resolveStopDecision:
import {
resolveStopDecision,
createMaxStepsExhaustedError,
resolveMaxSteps,
} from '@helix-agents/core';
const decision = resolveStopDecision(stepResult, state.stepCount, {
maxSteps: agent.maxSteps,
stopWhen: agent.stopWhen,
});
if (decision.stop && decision.cause === 'max_steps' && !agent.outputSchema) {
throw createMaxStepsExhaustedError(resolveMaxSteps(agent.maxSteps));
}3. retry() continues a cut-off run
After a framework_max_steps_exhausted failure, retry() with no options:
- restores the failed run's last committed step (its log,
customStateand forced-completion state); - restarts
stepCountat 0, so the retried run gets the fullmaxStepsbudget (also when the cut-off run was a resume that entered with a higher step count); - needs no message: the kept log ends in the last step's tool results, and the model carries on. A
messageyou pass is appended as new input.
This holds even when the cut-off run committed no step. For example, a client tool called on the last allowed step suspends the run, and the resume enters at the cap and is cut off before it calls the model. retry() then keeps the resume's entry, including the submitted client-tool result, and continues from it.
await executor.retry(agent, sessionId); // continues; no message needed
await executor.retry(agent, sessionId, { message: 'Focus on the src/ folder.' }); // continues with a hintBefore, on Cloudflare Workflows (the only runtime that failed a cut-off), retry() without a message threw validation_error ("No user message found to retry with") after a multi-step run, and a retry with a message restored the exhausted step count, so the retried run was cut off again at once. On the other runtimes the run had completed and could not be retried. An explicit checkpointId is still restored as given. planRetry recognises the cut-off from the failed run's completionReason, so custom stores must persist that field (every shipped store does; see §5).
4. String tool results are stored verbatim
A tool (or sub-agent) that succeeds with a string result is now stored exactly as that string, with metadata.helixContentEncoding: 'text' (COMMON_METADATA_KEYS.CONTENT_ENCODING). Everything else is unchanged: objects, numbers, booleans, null and every error result are stored as JSON, with no marker.
| Tool returns | content before (Temporal, DBOS, CF Workflows) | content now | Marker |
|---|---|---|---|
'hello' | "hello" (JSON-quoted) | hello | 'text' |
'{"a":1}' | "{\"a\":1}" | {"a":1} | 'text' |
{ a: 1 } | {"a":1} | {"a":1} | none |
| failure | {"error":"…"} | {"error":"…"} | none |
- The model input changes. On Temporal, DBOS and Cloudflare Workflows the model now sees
hello, not"hello". runtime-js and the Cloudflare DO already stored strings verbatim; they now add the marker. - Reading stored results. Never
JSON.parsea tool result'scontentdirectly: a tool that returned the string'42'or'{"a":1}'would decode as a number or an object. Use the new core helper:
// Before
const value = JSON.parse(toolMessage.content);
// After
import { decodeToolResultContent } from '@helix-agents/core';
const value = decodeToolResultContent(toolMessage); // marked → the string; else JSON; invalid JSON → raw text- Messages stored before the upgrade have no marker.
decodeToolResultContentparses them as JSON and falls back to the raw text, which matches the old readers. A pre-upgrade row whose string happens to look like JSON still decodes as the parsed value; new rows never do. @helix-agents/ai-sdk.convertToAISDKMessagesdecodes a tool part'soutputthis way, so a string result renders as that string. No change is needed in auseChatapp.- Custom stores must round-trip message
metadata(every shipped store does).
5. Migrations
| Store | Migration | Adds | Action |
|---|---|---|---|
D1 (@helix-agents/store-cloudflare) | V19 | nullable __agents_runs.completion_reason TEXT | runMigration() applies it; run it before the new code serves traffic |
DO-SQLite (DOStateStore) | 13 | nullable runs.completion_reason TEXT | none: applied lazily on first DO access |
| Postgres, Redis, memory | none | they already persisted completionReason | none |
Both columns are added only when missing, need no backfill, and read back as "no reason" for older runs.
6. Deploy notes
- Cloudflare Workflows: drain in-flight instances. The workflow body changed (the stop decision, new step names, and child instance ids keyed on the parent run:
…__spawn__<run suffix>, a companion's…__continue__<run suffix>-<step>-<callId>/…__resume__<run suffix>-<step>-<callId>). A replay of an instance started by the old code would build different child ids. Let in-flight instances finish, or terminate them, before you deploy. Code that rebuilt a child instance id by hand must useexecutor.getHandle()instead. - D1: migrate first. Run
runMigration()(V19) before, or as part of, the deploy that ships this release. - Temporal and DBOS. Only the string-result encoding changes there. As with any workflow-code release, let in-flight workflows finish (or use your usual versioning) before you deploy; a session that mixes old (quoted, unmarked) and new (verbatim, marked) string results decodes correctly with
decodeToolResultContent.
7. Temporal and DBOS caveat
Temporal and DBOS still ignore stopWhen (DI-29) and still end a maxSteps cut-off completed with no output (DI-25), so §1–§3 do not apply to them yet, and their retry() refuses such a session because it is not failed. CF-D1 MR 2 (Temporal, FU-CFD1C-DI29-DI25-TEMPORAL) and MR 3 (DBOS, FU-CFD1C-DI29-DI25-DBOS) wire resolveStopDecision and the run's completionReason into those workflows. Until then, end loops there with a finishWith tool or an outputSchema. See DI-25 and DI-29 in the conformance findings.