Skip to content

Migrating to Completion Strategies and Preserved Thinking ​

Overview ​

Anthropic's Claude Opus 5.5, Sonnet 5.5, Fable 5.1 and Mythos 5.1 reject forced tool use (HTTP 400 tool_choice: type "tool" and "any" are not supported for this model.) and run preserved thinking: a thinking block is bound to the exact system prompt, tool set and earlier messages it was produced with, and a request whose prefix changed is rejected. Before this release, every forced-completion call of an outputSchema agent on those models failed (it forces a named tool, narrows the tools and trims the system prompt), the 400 was swallowed as a "miss", and the run ended framework_forced_completion_exhausted with the provider's error lost. Multi-turn sessions also replayed thinking unfaithfully (blocks merged, empty and redacted blocks dropped).

This release (@helix-agents/core 0.50, @helix-agents/llm-vercel, the JS, Cloudflare DO and Cloudflare Workflows runtimes, the stores, @helix-agents/ai-sdk and @helix-agents/tracing-langfuse) adds:

  • Completion strategies. A forced call is enforced either by a named tool choice (forced_tool, unchanged) or by the provider's native structured output (native_output), chosen per call from capabilities the LLM adapter reports for the model. @helix-agents/llm-vercel ships a capability table for Anthropic models and a capabilities override. See Finishing Agents → Completion strategies.
  • Preserved thinking. Responses are recorded as ordered parts and replayed verbatim; on a model whose thinking is bound to history, a changed system prompt or tool set is refused with a typed error before the call. See Preserved thinking.

defineAgent and AgentConfig do not change, and there is no new agent option.

Who is affected:

You useEffect
An outputSchema agent on Opus 5.5, Sonnet 5.5, Fable 5.1 or Mythos 5.1 (JS, CF DO, CF Workflows)Through @ai-sdk/anthropic (direct, or a LiteLLM proxy behind it): forced completion now works (native_output), nothing to configure for a known model id. Bedrock, Vertex, AI Gateway and OpenRouter ids now resolve to the right profile, but native_output through those providers' own AI SDK packages is unverified (warning, FU-COMPLETION-BEDROCK-VERTEX-NATIVE-OUTPUT in follow-ups)
An outputSchema agent on a Claude model id the table does not know (a proxy alias such as a LiteLLM sonnet-latest, or a model newer than your package)Refused at execute() / resume() / retry() with framework_completion_strategy_unavailable (JS, CF DO, CF Workflows). Add a capabilities override
Any other model (OpenAI, Google, …, or Claude up to Opus 5 / Sonnet 5)No change to requests, except that a thinking session with multi-block, empty or redacted reasoning now replays it verbatim (§2); the behaviour changes in §3 still apply
A custom LLMAdapterCompiles and behaves as before (legacy); see §5 to opt in
Temporal or DBOSForced completion on the 5.5 generation still fails until MR 2 / MR 3; see §7

1. Configure proxy aliases and new models ​

What changed ​

The adapter maps the model to a capability profile (legacy, anthropic-structured, anthropic-5.5). An Anthropic model with no row in the table is unresolved: the framework cannot know whether a forced call would work, so an agent with outputSchema on it is refused before its run starts, typed and with nothing written. An agent without outputSchema is never refused and runs as before, and neither is one with maxCompletionRetries: 0 (it never makes a forced call). On runtime-js and the Cloudflare DO an agent with an llmConfigOverride is not refused at entry either (the override may pick the forced call's model); its forced call fails with the same code if the model that call uses cannot finish. Over HTTP the refusal is a typed 400 { error, code } (code framework_completion_strategy_unavailable or framework_completion_schema_unsupported) from both the Cloudflare DO (/start, /resume, /retry, /submit-tool-result) and @helix-agents/agent-server (/start, /resume, submit-tool-result), never a 500. A refused submit writes no result and leaves its call pending.

A model counts as Anthropic when its id starts with claude- or its provider id contains anthropic. So a LiteLLM alias served through createAnthropic({ baseURL }) (provider anthropic.messages, id sonnet-latest) is an unknown Anthropic model and is refused. Give a custom-named provider a name containing anthropic (createAnthropic({ name: 'anthropic-litellm', baseURL })); under { name: 'litellm' } an alias is not recognised as Anthropic and every forced call on a 5.5-generation backend fails at the provider.

What you need to do ​

Declare the alias, or a model newer than the table, on the adapter:

typescript
import { modelIdOf } from '@helix-agents/core';
import { VercelAIAdapter, resolveAnthropicCapabilities } from '@helix-agents/llm-vercel';

const LITELLM_ALIASES: Record<string, string> = { 'sonnet-latest': 'claude-sonnet-5-5' };

const adapter = new VercelAIAdapter({
  capabilities: (config) => {
    const alias = modelIdOf(config.model);
    const real = alias ? LITELLM_ALIASES[alias] : undefined;
    return real ? resolveAnthropicCapabilities(real) : undefined; // undefined → built-in table
  },
});

See Vercel adapter → Overriding capabilities and the completion strategies example. Install @ai-sdk/anthropic >=3.0.38 (now an optional peer dependency of @helix-agents/llm-vercel) for Anthropic models.


2. If you do nothing ​

  • Models other than the 5.5 generation send the same requests as before, including forced calls (the anthropic-structured profile still uses the named tool choice); only the hidden correction message's wording changed (§3). One exception: on any model with thinking enabled (not only the 5.5 generation), a session whose responses had several reasoning blocks, or empty or redacted ones, now replays each block verbatim and in order instead of merging them into one and dropping the empty and redacted ones, so the bodies of that session's later work requests change (the intended preserved-thinking fix).
  • 5.5-generation outputSchema agents with a known model id served through @ai-sdk/anthropic start finishing correctly on JS, CF DO and CF Workflows. Through Bedrock, Vertex, AI Gateway or OpenRouter providers the id resolves to anthropic-5.5, but whether native structured output works through those providers is unverified (see the Vercel adapter warning).
  • outputSchema agents on an unknown Claude id (aliases, unreleased-in-the-table models) start failing at entry with framework_completion_strategy_unavailable; before, they ran and failed (5.5 generation) or worked (older models behind an alias). Add the override.
  • The behaviour changes below apply.

3. Behaviour changes on every model ​

Provider errors inside a forced call end the run (JS, CF DO, CF Workflows) ​

A provider error on a forced call used to be a recoverable miss: another correction message, another forced call, and finally framework_forced_completion_exhausted with the provider's code lost. Now it fails the run with the provider's own errorDetail (for example provider_invalid_request, provider_auth_error). Transient errors are retried by the runtime's own mechanism instead of a forced attempt: the adapter's maxRetries on runtime-js and the CF DO, the step.do retry (llmStepRetry) on CF Workflows. If you assert on framework_forced_completion_exhausted after a provider failure, assert on the provider's code instead. Temporal and DBOS keep the old behaviour (§7).

The hidden correction message is reworded ​

The persisted hidden [completion] message now reads:

[completion] Provide your final answer now (attempt N) in the required structured format. Do not call any other tool and do not reply with plain prose.

(previously it named the tool to call). It is the same on every runtime and model, because it is persisted before the next call's strategy is known. Recorded-replay LLM fixtures keyed on the request body that cover a forced call need re-recording. Finish semantics do not change: on models that support it, enforcement is still the named tool choice.

Skills and memory fail closed ​

  • A skills provider whose listSkills throws, or whose getSkill throws for a preloadSkills name, used to be logged and skipped (the step ran without the catalog or the body). It now fails the step with the retryableframework_skills_unavailable on every runtime except DBOS. Durable runtimes retry the step (Temporal activity retry, CF Workflows step.do); runtime-js and the CF DO fail the run, and retry() re-runs it. DBOS keeps the old behaviour until MR 3: it resolves the catalog in the workflow body, outside any step, so it logs a warn and runs the turn without the catalog and without any preloaded body (one failing preloaded name used to skip only that body) (FU-COMPLETION-DBOS-SKILLS-RECOVERY-STRAND). An unknown preloadSkills name (getSkill returns null) is still skipped with a warning. resolveSkillsCatalog now rejects instead of returning ''.
  • On runtime-js and the CF DO, a throwing memoryManager.buildTools used to drop the memory tools silently; it now fails the step with a retryable framework_internal_error.

Silently running without part of the prompt would change the prompt prefix, which a 5.5-generation model rejects for the rest of the session, and it hid the failure on every model.

Prompt-prefix check on 5.5-generation models (JS, CF DO, CF Workflows) ​

Once a 5.5-generation model has produced signed thinking in a session, a call whose rendered system prompt or tool set differs from the one that thinking was produced with (for that model) is refused before the API call with framework_prompt_prefix_changed (non-retryable), naming which changed. The API would reject it with an opaque 400 anyway. Typical causes: a systemPrompt function whose output changes between steps, a deploy that changes prompt text, tool descriptions or schemas, skills or workspace config while sessions are live. Keep them stable for a session's lifetime, put changing context in an appended message, or start a new session. A beforeLLMCall hook that edits earlier messages is unsupported on these models and is not detected. Other models are not checked.

Other observable changes ​

  • New error codes: framework_completion_strategy_unavailable, framework_completion_schema_unsupported, framework_prompt_prefix_changed, framework_skills_unavailable. See Error handling.
  • native_output forced calls stream no text_delta (the JSON answer is not shown live; thinking still streams), and the persisted assistant message carries a completion-tool call with metadata.helixNativeCompletion: true. A native miss (text that does not parse, or a max_tokens cut-off) is persisted with metadata.hidden: true, so convertToUIMessages leaves it out after a reload as well; it still goes to the model on the next call (its signed thinking must replay).
  • forced_completion chunk: attempt_started gains optional strategy and capabilityProfile. No new enum values, so older readers just ignore them.
  • Hooks: beforeLLMCall / afterLLMCall payloads gain optional completionStrategy and capabilityProfile; toolChoice is absent under native_output. Key on executionPhase === 'forced_completion' to detect forced calls.
  • Langfuse: generation metadata gains executionPhase, completionStrategy and capabilityProfile.
  • AssistantMessage.parts (ordered reasoning / text / tool-call parts) is recorded for responses with reasoning, and metadata.helixPrefix (the prompt fingerprint) on messages from 5.5-generation models. @helix-agents/ai-sdk does not forward helixPrefix to UI message metadata; stripThinking also removes reasoning parts.

4. Removed and changed APIs (@helix-agents/core) ​

BeforeAfter
framework_completion_tool_choice_unsupported error code (documented, never raised)Removed. A model that rejects forced tool choice uses native_output; one with no strategy fails framework_completion_strategy_unavailable
isTerminalLLMStepError(result, { forcedCall })isTerminalLLMStepError(result): a forced call's error ends the run too
prepareForcedCompletionCall({ state, effectiveTools })prepareForcedCompletionCall({ state, effectiveTools, resolution, modelLabel }); may also fail with framework_completion_strategy_unavailable
ForcedCompletionDirective = { tools, toolChoice, reasoningMode, executionPhase }A union discriminated by strategy (forced_tool / native_output); apply it with forcedDirectiveGenerateInput
classifyForcedCompletionResult made an error step a recoverable missIt is nonRecoverable with the provider's errorDetail
ForcedCompletionResultClassification nonRecoverable.code: ErrorCode (required)code?: ErrorCode: absent for an unclassified provider error, which carries no code anywhere (never an invented one). A failed ForcedCompletionTransition's errorCode (already optional) is then absent too, and FailForcedCompletionInput.code is optional
resolveSkillsCatalog returned '' when the provider threwRejects with framework_skills_unavailable
PlanStepProcessingOptions.forcedCompletion: { expectedToolName }Also { expectedToolName, strategy: 'native_output', callIdSalt }

Added: ModelCapabilities, CapabilityResolution, LEGACY_CAPABILITIES, LEGACY_RESOLUTION, resolveCompletionCapabilities, modelIdOf, modelLabelOf, CompletionStrategy, COMPLETION_STRATEGIES, selectCompletionStrategy, checkCompletionSupport, isCompletionSupportRefusal, resolveForcedCallTransition, forcedDirectiveGenerateInput, assertForcedToolsUnchanged, FORCED_NATIVE_MIN_OUTPUT_TOKENS, nativeCompletionCallId, computePromptFingerprint, checkPrefixStability, stampPromptFingerprint, LLMOutputFormat, AssistantMessagePart (+ schema), COMMON_METADATA_KEYS.PROMPT_PREFIX / NATIVE_COMPLETION, and FaithfulProviderSimulator (@helix-agents/core/testing). See the core reference.


5. Custom adapters and bring-your-own-loop code ​

  • Custom LLMAdapters need no change: without resolveCapabilities they are legacy (named tool choice), as before. To support a model that rejects forced tool choice, implement resolveCapabilities and checkOutputFormat, honour outputFormat / minOutputTokens, and return the answer as a plain text result; see Custom adapters. Adapters for models with history-bound thinking should capture and replay ordered parts.
  • Adapter wrappers (an object that forwards generateStep to a real adapter) must also forward resolveCapabilities and checkOutputFormat, or every model reads as legacy.
  • Bring-your-own loops that run forced completion: resolve capabilities per forced call, call checkCompletionSupport before starting a run, pass the strategy and a replay-stable callIdSalt to planStepProcessing, use resolveForcedCallTransition, and run checkPrefixStability / stampPromptFingerprint around the LLM call. See Completion strategies (cross-runtime).
  • Tests: new MockLLMAdapter(responses, { capabilities, capabilityProfile }) simulates a model with those capabilities (rejects forced tool choice, signs and verifies scripted reasoning blocks).

6. Deploying ​

  • No store migration. parts is an optional field in MessageSchema (messages are JSON on every store), helixPrefix is message metadata, and ForcedCompletionState does not change.
  • Rolling deploys on preserved-thinking models. A process still on the old version strips parts when it reads history and replays the merged legacy form, which an account with preserved-thinking enforcement rejects with a 400 until every process runs the new version. Keep the mixed-version window short for 5.5-generation traffic.
  • A config-changing deploy (prompt text, tool descriptions or schemas, skills, workspace) now fails live 5.5-generation sessions with framework_prompt_prefix_changed on their next call instead of an opaque 400. Start new sessions after such a deploy.
  • Cloudflare Workflows: the completion-support decision is recorded in a step.do (<agentType>-completion-support), so a replay reuses it and a later change to how an outputSchema agent's model resolves (a capability-table change, a removed override) never fails an instance that was admitted. An instance started before this release has no recorded decision: it makes the check fresh on its next replay, so for those keep the resolution stable across the deploy, or drain affected instances first.
  • Temporal / DBOS: no new activities, steps or workflow branches in this release.

7. Temporal and DBOS (known gaps) ​

Temporal (MR 2) and DBOS (MR 3) are not migrated yet. On those runtimes:

  • every forced call uses the legacy request whatever the adapter reports, so forced completion on Opus 5.5, Sonnet 5.5, Fable 5.1 and Mythos 5.1 still fails (HTTP 400 on each attempt, then framework_forced_completion_exhausted);
  • a forced call's provider error stays a recoverable miss (the step result crosses a serialization boundary that loses its classification);
  • there is no early refusal and no prompt-prefix check; hook payloads and chunks carry no strategy fields;
  • DBOS additionally drops providerOptions (thinking, effort) and maxOutputTokens across its step boundary.

What does apply there: the reworded correction message, skills fail-closed on Temporal (the typed code survives the activity boundary as an ApplicationFailure; DBOS keeps a failing skills provider non-fatal until MR 3), the removed error code, and ordered parts capture and replay (it is adapter-side). Run outputSchema agents on 5.5-generation models on runtime-js or a Cloudflare runtime until MR 2 / MR 3 land. See Finishing Agents → Temporal and DBOS.

Released under the MIT License.