Skip to content

Migrating to the Default Step Cap and CF Workflows Multi-Turn ​

Overview ​

This release (@helix-agents/core 0.50) makes two changes. It applies the documented maxSteps default on every runtime, and it makes Cloudflare Workflows runs end truthfully and continue across turns.

  • An agent with no maxSteps now stops after 50 steps (DI-28). The default was documented but applied only on CF Workflows and DBOS. JS, the CF Durable Object and Temporal ran uncapped, so a model stuck in a tool loop kept calling the LLM. One live run made more than 580 calls.
  • CF Workflows:
    • A plain-text turn now ends completed (CP-91).
    • A second execute(sessionId) starts a new turn (CP-88).
    • execute() no longer races its own workflow (CP-89).
    • Any session id works (CP-90).
    • handle.resume() after an interrupt no longer loses the resumed turn (CP-67).

Most of this needs no code change. The items below change behaviour or need action.

1. maxSteps now defaults to 50 on every runtime ​

shouldStopExecution(), the shared step iterator (runtime-js and the CF DO), CF Workflows, DBOS and Temporal root workflows all resolve an unset maxSteps through resolveMaxSteps(), which returns DEFAULT_MAX_STEPS (50). An agent with outputSchema that reaches the cap mid-tool-loop enters forced completion there, as it already did with an explicit maxSteps.

Action: an agent that must run with no step cap now has to say so.

typescript
const agent = defineAgent({
  name: 'long-runner',
  maxSteps: Infinity, // no step cap
  // ...
});
  • Temporal and DBOS: their workflow input is JSON, so "no cap" is sent as Number.MAX_SAFE_INTEGER.
  • Custom loops: a loop built on shouldStopExecution() gets the cap too. Pass maxSteps: Infinity to keep it uncapped.
  • Exception: a Temporal sub-agent or companion child workflow never receives its agent's maxSteps. It stays uncapped, as before (FU-TEMPORAL-CHILD-MAX-STEPS).
  • Reporting: how a run that hits the cap is reported still differs per runtime (DI-25, unchanged in this release). A later release makes it a typed framework_max_steps_exhausted failure on JS, the CF DO and CF Workflows; see Stop conditions and string results.

Temporal deploy note ​

A Temporal root workflow that has already run past step 50, for an agent with no maxSteps, would take a different path when it replays under the new code. That is a non-determinism error.

Action: drain those in-flight root workflows before deploying, or give the agent an explicit maxSteps first. Workflows of agents with an explicit maxSteps are unaffected.

2. CF Workflows: multi-turn execute(sessionId), new instance ids ​

A second execute(agent, msg, { sessionId }) now continues the session into a new turn.

Which sessions continue:

  • completed, failed, interrupted and paused sessions continue, as on runtime-js and the CF DO. A paused session continues only when no client-tool calls are pending and no sub-agent is suspended.
  • A session with a live run is refused with AgentAlreadyRunningError.

Instance ids: every turn runs in its own write-once Workflows instance. If you look instances up by id yourself, note the new formats:

InstanceId
Every turnagent__<name>__<sessionId>__turn__<key> (key minted per call)
Resume / retry…__resume__<key> / …__retry__<key>
Respawned sub-agent…__respawn__<parent-run suffix> (was …__respawn-<suffix>)
Companion continuation…__continue__<step>-<toolCallId>

Encoded ids: an id that Cloudflare would reject, or whose parts contain __, is replaced by an encoded form, h_<hash>_<readable tail>. Cloudflare accepts at most 100 characters of [A-Za-z0-9_-]. waitForEvent / sendEvent types are encoded the same way.

Action: do not rebuild agent__<name>__<sessionId> to find a run. Call executor.getHandle(agent, sessionId): it finds whichever turn, resume, retry or rewind instance is live.

  • In-flight instances started before the upgrade keep their old ids.
  • getHandle() still finds a running instance at the base id.

3. CF Workflows: the run-start commit picks one winner ​

Before it creates the workflow instance, execute() writes only a write-once session row for a new session. The instance's run-start commit decides which call wins; a call that loses is refused with AgentAlreadyRunningError and writes nothing, and only the winner makes the stream live.

Action: none needed. execute() creates the session itself, and it refuses a session whose run is live as already running.

A session that a cancelled Worker left without an instance has no run record. The next execute() for that session runs it.

4. CF Workflows: handle.resume(options) is executor.resume() ​

Every handle's resume(options) now runs executor.resume(). As a result it:

  • reactivates the stream an interrupt ended (the winner, after its run-start commit);
  • truncates nothing before its run-start commit wins;
  • refuses a live run;
  • supports every resume mode (continue, with_message, from_checkpoint).

No code change is needed.

Reference ​

Released under the MIT License.