Skip to content

feat(ai, ai-harness): TanStack AI harness, phases 0 to 14, media, provider keys, durable sessions, host hooks, reasoning, skills, prompt caching, turn overrides, block order, mid-conversation changes, routing, and loop fixes (combined stack) - #1555

Open
AlemTuzlak wants to merge 256 commits into
mainfrom
feat/harness-p14-mcp-server

Conversation

@AlemTuzlak

@AlemTuzlak AlemTuzlak commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

This is the one PR to review and merge for the TanStack AI harness. It puts the full harness stack, phases 0 to 14, plus harness media, provider keys, durable sessions, shared session logs with turn hooks and turn leases, chat({ reasoning }) with the @tanstack/ai-models catalog, skills with live commands, automatic prompt caching, Claude block order kept end to end, mid-conversation tool and prompt changes that keep the prompt cache, routing of a turn to root agents, and a set of loop and adapter fixes, on one branch against main, so CI can test it together. A harness keeps one agent conversation open across many turns. It has typed agents, plugins, a CLI, a dashboard, MCP connectors, code mode, coding agents, a live session view, and an MCP server. Now every front door can also send images, audio, video, and documents to a turn, and every UI can show the media that agents make.

This PR replaces #1551 (P0 to P13) and the 15 stacked PRs (#1513 to #1554). Those PRs are closed. They stay as the review record of each phase, and every fix now lands here.

Note

main is merged into the stack. Every stack branch now has main (62bec34bb). The merges resolved these conflicts:

🎯 Changes

Each phase was reviewed in its own PR, now closed. The table links them:

Phase PR What it adds
P0 #1513 Subagents return a value and call any activity (@tanstack/ai)
P1 #1515 Sessions, typed agents, and plugins (@tanstack/ai-harness)
P2 #1518 Resume crashed turns and serve sessions to clients
P3 #1519 Commands, settings, auth, and first-party plugins
P4 #1520 Subagent tree limits, harness ctx.agents and children
P5 #1521 Build artifacts, worker mode, remote harnessText
P6 #1522 Self-hosted dashboard, --dashboard, runnable example
P7 #1523 MCP connectors with browser sign-in, run-time tools, media in the example
P8 #1524 Code mode with pluggable isolates
P9 #1538 Delegate to Claude Code, Codex, and other coding agents
P10 #1540 Goal plugin that works until a judge says the goal is met
P11 #1546 createSessionView, a live store of a session for any UI
P12 #1549 UI-agnostic CLI, runCli({ ui }), Ink screen in the example
P13 #1550 agentMiddleware for every agent run, usage() counts every agent
P14 #1554 Use a harness from any MCP client: createHarnessMcpServer at @tanstack/ai-mcp/harness, harness --mcp, /mcp on --serve
Media this PR Send files to a turn and show the media that agents make, from every front door (see below)
Durable this PR One session log per thread: retries that run once, attempt and time limits, durable tool steps, and ordered joins (see below)
Host this PR Many sessions on one log, turn hooks (retry a model call, keep a turn going, join rules), a recover hook, turn leases, and replay: 'never' for durable tools (see below)
Reasoning this PR chat({ reasoning }) for every adapter, and the @tanstack/ai-models runtime catalog (see below)
Loop fixes this PR, #1578, #1579 Parallel server tools, Anthropic replay fixes, overflow detection, Workers AI gateway models, and a fake test adapter (see below)
Skills this PR Skills in a harness as live /<skill> commands, live plugin commands, and lazy code-mode tools (see below)
Workspace this PR workspaceTools({ outside: 'ask' }): a path outside the workspace asks the user first (see below)
Block order this PR Claude's thinking, text, and tool calls keep their order through the stream, the stored message, the useChat wire, and the replay (see below)
Mid-conversation changes this PR A tool or system prompt added between model calls goes out as a change on GPT and Claude models that support it, so the cached prompt start stays the same (see below)

Packages. New: @tanstack/ai-harness, @tanstack/ai-harness-cli, and @tanstack/ai-dashboard. Changed: @tanstack/ai, @tanstack/ai-persistence, @tanstack/ai-acp, @tanstack/ai-mcp, @tanstack/ai-code-mode, and @tanstack/ai-sandbox. The media work also changes 9 provider packages: ai-openai, ai-anthropic, ai-gemini, ai-mistral, ai-groq, ai-byteplus, ai-grok, ai-openrouter, and ai-llmgateway. @tanstack/ai-isolate-daytona and @tanstack/ai-opencode get new tests only. The reasoning work adds @tanstack/ai-models and changes openai-base and 18 provider packages. The loop fixes change @tanstack/ai, @tanstack/ai-anthropic, and @tanstack/ai-cloudflare. The skills work changes @tanstack/ai-skills, @tanstack/ai-harness, and @tanstack/ai-code-mode. The host work changes @tanstack/ai-harness and @tanstack/ai-persistence. The block order and mid-conversation work changes @tanstack/ai, @tanstack/openai-base, @tanstack/ai-openai, @tanstack/ai-anthropic, and @tanstack/ai-persistence, with new tests in @tanstack/ai-harness and @tanstack/ai-skills.

MCP in both directions.

  • @tanstack/ai-mcp, client side (P7): MCP connectors with browser sign-in.
  • @tanstack/ai-mcp, server side (P14): createHarnessMcpServer at @tanstack/ai-mcp/harness. Any MCP client can chat with a harness, answer its approvals, and run its agents and commands.
  • @tanstack/ai-harness-cli (P14): --mcp serves the harness over stdio, and --yes approves every tool call in MCP mode. --serve also serves MCP at /mcp, behind the same bearer token.
  • @tanstack/ai-mcp is an optional peer of the CLI. Without it, --mcp fails with a clear message.

Media (new on this branch).

  • Agent media is kept. Every ctx.generateImage, ctx.generateSpeech, ctx.generateAudio, and ctx.generateVideo({ stream: true }) result is saved (it reuses withGenerationPersistence). Each file publishes a harness.media event. Its record is saved on the message, so it comes back after a restart. Agent code does not change.
  • Send files from every front door:
    • code: session.putMedia plus mediaPart(record)
    • web: client.upload and POST .../media
    • CLI: @path in a message
    • MCP: chat attachments
    • ACP: image, audio, and resource blocks
    • POST .../run and harnessText
      The transcript keeps a small harness-media:<id> URL. Only the model call gets the bytes.
  • Show media in any UI. The session view has MediaParts with a signed url (for <img>, <audio>, and <video>) and load(). The handler signs URLs with mediaSecret, supports Range, and sends nosniff and a sandbox CSP. The CLI saves files to ./<harness-name>-media, and MCP results carry small images and audio inline, with other files as harness-media:// links.
  • The model's inputs are checked. Text adapters get a runtime inputModalities from model-meta. The 9 providers above set it, and the model sync keeps it in step. defineHarness({ media: { maxBytes, kinds, accepts, transcribe } }) narrows the inputs. A file the model cannot read stops the turn with a clear error, or transcribe turns audio into text.
  • One @tanstack/ai-mcp/server addition. resourceDefinition({ uriTemplate, argsSchema }) now gives read the parsed template variables and the URI, and a read result can set its own mimeType.

Provider keys (new on this branch). A harness you ship to users does not need a .env file. Users connect a model provider inside the app:

  • /connect openai opens the page to make a key (the provider's keyUrl), then asks for the key and hides the typing. /connect openrouter signs in through the browser (openrouterSignIn(), PKCE with a 127.0.0.1 callback). /disconnect and /keys (masked) come with it.
  • Keys go in the harness credential store, one set per user. An env var is still the fallback.
  • keyedAdapter(provider, create) in @tanstack/ai builds an adapter from the key per turn, for the main model, /model choices, compact, and goal. Agents get ctx.keys (get, require, adapter).
  • A missing key stops the turn with harness.auth_required: "Sign in to openai. Run /connect openai."
  • Questions can be secret. A secret answer is never stored, printed, or put in an event.

Durable sessions (new on this branch). A host can keep the whole state of a session in one log per thread. Then a crash, a restart, or a second host rebuilds the same session. The log is a store contract, so any backend or framework can build on it:

  • One log. Give the host persistence.stores.log, a LogStore that appends at an expected seq, all or nothing. The session writes its events, merged text deltas, the transcript, and your own records (session.append) to it. A project: { record, version } function folds your records into the model context. memoryLogStore() and the log conformance cases are in @tanstack/ai-persistence, and logMessageStore gives any reader the transcript.
  • Inputs with ids. prompt, steer, followUp, and resolve take { inputId }. A retry with the same id gets the first receipt and runs nothing again. turn.receipt, session.settled(inputId), and the harness.input.settled event tell how the input ended, also after a restart.
  • Limits. defineHarness({ durability: { maxAttempts, timeoutMs } }). A turn fails with attempts_exhausted or timeout when it runs out. cancel() records the abort first, so recovery does not run the turn again.
  • Durable tools. durableTool(definition, execute) gives the tool step.do(name, fn), which saves a step result and replays it after a crash, and append(records), which adds records with the tool batch. After a crash, a finished call in a batch keeps its result. Only an unfinished call with replay: 'never' gets the crash note.
  • Ordered joins. A message that arrives during a turn joins it in one log write. It reaches the model at the next model call, in the order it arrived, and it settles with that turn.

Harness host features (new on this branch). Any host or framework can now build its own agent runtime on the harness. Every new option is off until the host sets it:

  • One log, many sessions. host.open(harness, { threadId, logId }). Sessions with the same logId share one log, and each record names its session in a thread field. createHarnessHost({ reduce }) folds the whole log into one state, and host.logState(logId) returns it. logMessageStore({ logId }) reads the threads of one log.
  • Turn hooks. defineHarness({ turn: { onModelError, beforeFinish, maxFinishCycles, canJoin, onJoin } }). onModelError can retry a failed model call in the same turn, and clients get harness.turn.retry. retryTransientErrors() is a ready policy for 429, 5xx, and overload errors. beforeFinish can add messages or records, and then the turn continues.
  • Join rule. A turn takes the waiting steers in order. It stops at the first steer that has a cancel request, or that canJoin refuses. onJoin adds messages or records in the same log write. This rule applies to every host. A cancel that arrives after the join is refused with not_running.
  • Recover hook and leases. durability.recover decides for each input that a crashed host left: run it again, or settle it. persistence.stores.leases, a new LeaseStore in @tanstack/ai-persistence, can replace stores.runs to tell which host owns a turn.
  • Durable tools. durableTool(definition, execute, { replay: 'never' }) gives the model a tool error after a crash, and does not run the call again. A durable tool that a middleware adds in onConfig now gets step and append.

Reasoning and the model catalog (new on this branch).

  • chat({ reasoning }) takes a level ('off', 'minimal', 'low', 'medium', 'high', 'xhigh', 'max') or { level, summary?, budgetTokens? }. The type allows only the levels of the adapter's model. Each provider adapter writes it to its own wire field, and it replaces the reasoning fields in modelOptions.
  • @tanstack/ai-models (new) is a runtime catalog of 1,932 models from 38 providers: context window, max tokens, cost, input kinds, and reasoning maps. It has getModel, supportedReasoningLevels, clampReasoningLevel, and modelCost, and one subpath per provider.
  • openaiCompatible gets compat for provider quirks: the thinking format, the developer role, the max-tokens field, and reasoning replay.

Loop and adapter fixes (new on this branch).

  • Parallel server tools (also fix(ai): run the server tools of one model turn at the same time #1578 against main, fixes Tool calls from one model step run one at a time, but the docs say they run in parallel #1547). The server tools of one model turn now start together, and the results keep the call order. chat({ toolExecution: 'sequential' }) and toolDefinition({ sequential: true }) keep the old order. A tool that has not started when the run aborts gets "Operation aborted".
  • Anthropic replay (also fix(ai, ai-anthropic): send redacted thinking and tool errors back to Claude #1579 against main). A redacted_thinking block is kept as a thinking part with redacted: true and sent back unchanged, and a failed tool result sends is_error: true.
  • isContextOverflow({ error, usage, finishReason, contextWindow, provider }) in @tanstack/ai tells you that a call overflowed the context window, from about 25 provider error messages, or from the usage and the finish reason.
  • cloudflareBindingFetch({ binding, vendor, gateway }) in @tanstack/ai-cloudflare sends the requests of createAnthropicChat and createOpenaiChat through env.AI.run to the AI Gateway anthropic/... and openai/... models.
  • fakeText() at the new @tanstack/ai/testing subpath answers chat() from a script, with ceil(characters / 4) usage estimates, cache estimates, pacing, and abort.

Skills, live commands, and the workspace (new on this branch).

  • Skills in a harness. skills({ dirs }) at the new @tanstack/ai-skills/harness subpath gives the model the skills of a list of folders. Each skill is a /<name> command that starts a turn with that skill. A skill with the name of another command is /skill:<name>, and /skills lists them.
  • Live commands. The plugin watches the folders. A skill that you add or remove while the session runs changes the commands at once. Plugins get ctx.commands (has, set, delete, ready) for this. Each change sends a harness.commands.changed event, and session views read the command list again.
  • Lazy code-mode tools. codeMode({ lazy: true }) lists only the names of the tools that move into code mode. The model asks discover_tools for the signatures before it calls them.
  • Paths outside the workspace. workspaceTools({ root, outside: 'ask' }) asks the user before a path outside the workspace. A yes allows that folder for the rest of the session. Without the option, such a path is refused, as before. list_files and grep take an optional path.

Prompt caching (new on this branch). chat() now asks the provider to cache the stable start of each request (system prompt, tools, and earlier messages). Repeated requests cost less and answer sooner. It is on by default, in plain chat() and in every harness session.

  • One option, promptCache: 'none' | 'short' | 'long', or { retention, key }. The key is the threadId that the caller gives. A made-up id is never the key.
  • Each adapter sends its own fields: Claude cache markers (Anthropic, Bedrock, OpenRouter, and openaiCompatible with cacheControlFormat: 'anthropic'), prompt_cache_key (OpenAI, Mistral), and sessionId (OpenRouter). A manual cache_control, cachePoint, prompt_cache_key, or sessionId wins.
  • Harness: defineHarness({ promptCache }) sets the default. host.open(harness, { threadId, promptCache }) overrides it for one session.
  • Usage: promptTokens is now the total input on Anthropic, Bedrock, and Claude Code. Mistral and the OpenRouter Responses adapter now report cache reads. The usage() plugin counts cache reads and writes.
  • Docs: the new docs/advanced/prompt-caching.md, and a prompt caching section in docs/harness/overview.md.

Harness routing (new on this branch). A harness can send a turn to its root agents, with the same picks as subagents.router. The main model does not answer a turn that a picked agent answers.

  • defineHarness({ routing: { router, order, strategy, limits, sandbox } }). The router gets the chat() router fields (messages, agents, abortSignal) plus session, input, operationId, inputId, and the turn's main adapter.
  • The router picks from the harness agents and the plugin agents (run plugins too), not from subagents. On 'main', the turn runs as before. On a pick, only the picked agents run, and the turn's text is their answer.
  • strategy: 'handoff' runs the main model after the agents, with its subagent tools. resolve continues a routed plan without the router. A steer during a routed turn runs as its own turn.
  • @tanstack/ai: a pick can give an agent its input as { name, input }. The routed path checks it with the agent's inputSchema, and a resumed plan keeps it. subagentRoute adds needsInput(result) and pick(result, { inputs }).
  • A subagents.router turn in the harness now gets the full history, holds a run lease, and returns the agents' text. Before, it saw only the new message and returned ''.
  • Docs: a new section "Route a turn to an agent" in docs/harness/subagents.md, and changes in docs/chat/subagents.md and docs/harness/plugins.md.

Block order and mid-conversation changes (new on this branch). Users change no code: every new field and option is optional, and the library writes and reads them.

  • Block order. A Claude answer like thinking, tool call, thinking, tool call keeps that order. chat() and the converters write ModelMessage.blockOrder only when the order differs from the default (thinking, text, tool calls), and the Anthropic replay sends the blocks in that order. A message in the default order is byte for byte the same as before.
  • The wire. useChat sends such a message as ordered AG-UI rows (reasoning, assistant, reasoning, assistant). A row that exists only for the order carries metadata.tanstack.continues, and our server joins the rows back. A tool result now ends a row, so a UIMessage that holds two model calls reaches the server as assistant, tool, assistant.
  • Stream fixes. Each Anthropic thinking block ends with its own REASONING_MESSAGE_END and REASONING_END. Text after a second thinking block starts a new text part; before, it repeated the first text.
  • Mid-conversation changes. Before each model call, chat() compares the tools and the system prompts with records it saves on assistant messages (midConversationChange). On a model with a channel, an added tool, or a prompt added at the end, goes out as a change, and the cached start of the request stays the same:
    • openaiText on 8 GPT models (gpt-5.4-mini, gpt-5.5, gpt-5.6-*, gpt-6-astra, gpt-6-luna, gpt-6-sol): an additional_tools input item and a developer message.
    • anthropicText on claude-opus-4-8, claude-opus-5, claude-opus-5-5, claude-fable-5, and claude-fable-5-1: a system message, the mid-conversation-tool-changes-2026-07-01 beta, a deferred placeholder tool, and tool_addition blocks. The automatic tool cache marker stays on the last start tool.
  • Defaults. The channels are on only with the provider's own API. A custom baseURL, a custom fetch, an injected Anthropic client, or the SDK's OPENAI_BASE_URL / ANTHROPIC_BASE_URL env var turns the default off. midConversationChannels: true or false on the adapter config overrides it. Every other model, and every request with no channel, is byte for byte the same as before.
  • Message stores. mergeStoredMessages keeps a stored change record when a useChat client sends the same message again without it, so the cache holds across turns with withPersistence.
  • Docs: a new docs/advanced/mid-conversation-changes.md, a "When the tools change" section in prompt-caching.md, sections on the OpenAI, Anthropic, and Cloudflare adapter pages, an adapter-author section in docs/advanced/extend-adapter.md, and changes in docs/chat/thinking-content.md and docs/harness/turn-control.md.

Per-turn overrides (new on this branch). One prompt can run with its own settings. The rest of the session keeps its defaults.

  • session.prompt(message, { overrides }) and followUp(message, { overrides }) take TurnOverrides: adapter, reasoning, promptCache, and extra tools. They apply to every model call of that turn: the tool loop, onModelError retries, beforeFinish cycles, and joins.
  • The overrides live in memory only. A queued turn keeps them, a joined steer uses the host turn's overrides, and a turn that recovery runs again uses the defaults.
  • A middleware that rebuilds the tools keeps the override tools. A durable tool among them gets step and append.
  • HarnessConfig.reasoning sets the default reasoning. The plugin adapter picker gets the turn. ChatMiddlewareConfig.promptCache lets onConfig change the cache of the next model call.
  • A live check of the mid-conversation shapes found a bug, now fixed: the OpenAI Responses adapter dropped the namespace of a call to a tool that came through additional_tools, so the next call failed with a 400.

Docs and example. 24 new pages in docs/harness/, with mcp-server.md from P14, media.md from the media work, and inputs.md, durable-tools.md, and session-log.md from the durable work. The durable work also rewrites durable-sessions.md, and adds the LogStore contract to docs/persistence/store-reference.md and build-your-own-adapter.md. custom-ui.md, cli.md, connect.md, subagents.md, docs/mcp/server-content.md, docs/chat/subagents.md, and docs/config.json change too. The reasoning work adds docs/chat/reasoning.md, docs/models/catalog.md, and docs/migration/reasoning-option.md. The loop fixes add docs/advanced/testing.md, and change docs/tools/tools.md, tool-architecture.md, docs/advanced/middleware.md, compaction.md, docs/chat/thinking-content.md, agentic-cycle.md (a "Stop the loop from a tool" recipe), and docs/adapters/cloudflare.md. The skills work adds docs/harness/skills.md, and changes plugins.md, code-mode.md, coding-agent.md, docs/skills/agent-skills.md, and docs/config.json. The host work adds docs/harness/turn-control.md and docs/harness/shared-logs.md, and changes durable-sessions.md, durable-tools.md, inputs.md, session-log.md, and docs/persistence/store-reference.md. The new example examples/harness-cli is an open-code style agent in the terminal, and examples/README.md lists it. It shows what the harness does out of the box:

  • Voice. Ctrl+R records the microphone with ffmpeg, and the transcript is sent. Spoken file names ("fox dot png", "the last image") are attached, /mic picks the microphone, and a silent recording is not sent.
  • Media. It makes images (with the images you send as references), speech, video (Grok Imagine or Sora), and songs and sound effects (fal). The screen numbers and saves each file, and /open and /play show them.
  • Models and agents. /model lists real model ids (gpt-6-astra, claude-opus-5-5, grok-4.7, openai/gpt-6-astra, and more) with their context size. /effort sets how hard the model thinks, as the reasoning option of its provider. Your local Claude Code and Codex turn on when their CLIs are on the PATH, and the screen shows their output.
  • Services. Notion and Linear work through MCP connectors.
  • Screen. The screen clears at startup, and the input stays at the bottom. Answers render as markdown, and Mermaid blocks as text charts. Type / for the commands. /connect, /model, /effort, and /mic open a picker, and the arrows go through the lines you sent. A footer shows the model, the effort, the context against its window (from the new usage() state field contextTokens), and the tokens.

The example also continues saved sessions (--resume <id> and /resume), and puts the voice transcript in the input line, so you can fix it before you send. It suggests files after @, shows a boot splash, makes each skill a / command, and asks before a path outside the folder where you start it. /model lists the models of the @tanstack/ai-models catalog.

The Cloudflare example (new). examples/harness-cli-cloudflare is the same terminal agent, with its state in one Cloudflare Worker:

  • A Durable Object for each session log, and an Account Durable Object for the settings, the encrypted keys, the media records, and the session list. R2 keeps the media bytes.
  • Code mode runs each program in a new isolate in the Worker, with @tanstack/ai-isolate-cloudflare.
  • The models go through AI Gateway with createCloudflareText. Images, speech, and voice use Workers AI.
  • The first start asks for the Worker URL and secret. /connect cloudflare saves the token in the Worker, so a session continues on another PC.
  • The Worker tests run in workerd with the createTestHarness of wrangler, and pass the @tanstack/ai-persistence conformance suite.

The Cloudflare example adds wrangler and @cloudflare/workers-types as its dev dependencies.

The example adds @tanstack/ai-fal (a workspace package) for songs, and marked, marked-terminal, and beautiful-mermaid for markdown and charts. ts-react-chat, ts-solid-chat, and ts-code-mode-web only move to @tanstack/store ^0.11.1.

Changesets. Each phase has its own changeset in .changeset/, harness-p0-agent-results.md to harness-p14-mcp-server.md. The media work adds five more:

  • text-adapter-input-modalities.md
  • provider-input-modalities-a.md
  • provider-input-modalities-b.md
  • mcp-resource-template-args.md
  • harness-media.md

The provider keys work adds core-provider-keys.md, harness-provider-keys.md, and openrouter-sign-in.md. The durable work adds harness-durable-log.md. The reasoning work adds ai-models-catalog.md and one reasoning-option-<package>.md per changed package. The loop fixes add:

  • parallel-server-tools.md
  • anthropic-thinking-replay.md
  • context-overflow.md
  • cloudflare-binding-fetch.md
  • fake-text-adapter.md

The skills work adds harness-skills.md and workspace-outside.md. The host work adds harness-host-generic.md. The background agent fix adds harness-agent-restart.md. The routing work adds harness-routing.md. The resolve after a restart adds harness-interrupt-restart.md. The block order and mid-conversation work adds block-order.md, mid-conversation-changes.md, and persistence-keep-change-record.md.

How to review.

  1. For the details of one phase, read its closed PR in the table (feat(ai): let a subagent return a value and call any activity #1513 first). Each one has its own Testing section.
  2. For media, start at docs/harness/media.md. Then read packages/ai-harness/src/media.ts, where the store, capture, and model-call parts live, and the session.ts wiring.
  3. For durable sessions and the host features, start at docs/harness/durable-sessions.md, turn-control.md, and shared-logs.md. Then read packages/ai-harness/src/log.ts (the log writer and the folds), turn.ts, durable-tool.ts, and the turn loop in session.ts.
  4. For reasoning, start at docs/chat/reasoning.md and docs/models/catalog.md. For the loop fixes, read fix(ai): run the server tools of one model turn at the same time #1578 and fix(ai, ai-anthropic): send redacted thinking and tool errors back to Claude #1579, which carry the two fixes to main alone, then packages/ai/src/utilities/context-overflow.ts, packages/ai/src/testing/fake-text.ts, and packages/ai-cloudflare/src/utils/fetch.ts.
  5. For skills, start at docs/harness/skills.md. Then read packages/ai-skills/src/harness.ts and ctx.commands in packages/ai-harness/src/plugins.ts. For the Cloudflare example, read its README, then worker/src/index.ts and src/cloudflare-stores.ts.
  6. For block order and mid-conversation changes, start at docs/advanced/mid-conversation-changes.md. Then read packages/ai/src/utilities/block-order.ts and mid-conversation.ts, the wire in packages/ai/src/utilities/ag-ui-wire.ts, and the rendering in packages/openai-base/src/adapters/responses-text.ts and packages/ai-anthropic/src/adapters/text.ts.
  7. Use this PR for the full picture. Its CI runs the full suite against main, with the coverage gate.
  8. Merge only this PR. The stacked PRs are closed. fix(ai): run the server tools of one model turn at the same time #1578 and fix(ai, ai-anthropic): send redacted thinking and tool errors back to Claude #1579 can merge to main first: this branch has the same commits.

Fixes made on this branch

These came from CI after main was merged into the stack, and from review. They are on this branch only.

  • .d.ts path. harness.d.ts imported the folder ./server. harness.ts now imports ./server/index, so the emit names a real file. scan-dangling-dts is clean.
  • Interrupt kinds in the MCP server. approve, reject, auto-approve, and inline elicitation answer tool approvals only. chat and status list every interrupt with its kind (approval, client-tool, generic) and response schema. resolve takes { interruptId, approved } or { interruptId, payload }. The result key is interrupts.
  • Linear sign-in. Linear sends iss on the sign-in callback (RFC 9207), and the MCP SDK refused the code without it ("Issuer mismatch"). startLoopbackReceiver().waitForCode() now resolves { code, iss }, and mcpConnector passes iss on. The session view also clears a pending sign-in when its connect:<id> command ends.
  • Unique tool names. Two commands whose names clash (connect:notion, connect_notion) get _2, _3, and so on, in name order, instead of an HTTP 500.
  • MCP connector on the v2 client. A new /connect saves only its own tokens. An auth failure asks the user to run /connect <id>. The issuer and the discovery state are saved, so there are no SEP-2352 warnings. The v1-only fallbacks are gone.
  • Background agents after a restart. A background agent held no lease, and restart recovery skipped it. After a crash, its run stayed running and the thread never woke. Now the run holds the leases of a chat turn. When the lease expired, the next host ends the run failed, adds a note to the transcript, and on a durable host settles the input and wakes the thread for wake: true. The agent does not run again, because it has no checkpoints.
  • Routed turns: handoff hooks and child resume. With routing.strategy: 'handoff', the main model ran without the turn hooks, so a 503 failed the turn instead of reaching onModelError, and beforeFinish never ran. And a subagents.router child that stopped for an approval could not resume: resolve failed with "Tool x is unavailable" ("unknown interrupt" on a durable host), because the stored thread has no card for the child. Now the hooks are off only while the root agents run, and the harness keeps the cards of every routed run.
  • A resolve after a restart. The interrupts that a turn waits on lived only in memory, so the next host refused every resolve with no_pending_interrupts, for a main-model approval and for a routed agent. The session now keeps the interrupted turn, with the agent cards of a routed turn, in stores.metadata (namespace harness:interrupted). open() reads it back, and a resolve removes it. Without a metadata store, the state stays in memory.
  • ai-skills coverage. The Coverage job failed: ai-skills function coverage fell from 84.86% to 83.42%. Five functions of the skills harness plugin had no test (load, listResources, listScripts, folderOf, and the watcher's error handler). Three new tests in harness.test.ts run them through the plugin.
  • Harness MCP server sessions. main serves spec 2025 clients without a session by default. Elicitation needs a session, so createHarnessMcpServer now asks for sessions: 'memory'. approvals: 'ask' asks the user over HTTP again.
  • Tests for coverage. harness.ts (38 tests), the connector (17 tests), and 5 small ai-mcp edges. The deterministic opencode and daytona timing tests are also on this branch. Credential in @tanstack/ai-persistence has a new optional issuer field.

✅ Checklist

  • I have followed the steps in the Contributing guide.
  • I have tested code changes locally with pnpm run test:pr, or these tests do not apply to this pull request.
  • I fully understand the code in this pull request, including any code generated with AI assistance.
  • Docs: I updated docs/ for this change, or this change is not user-facing.
  • Changeset: I added a changeset (pnpm changeset), or this PR does not change a published package.

Not ticked:

  • Contributing guide: the guide asks for E2E coverage on every feature. The media work adds an E2E spec, but P7 to P14 add none (see Testing).
  • pnpm test:pr: Nx fails in the local worktree (EISDIR: lstat 'F:' from the Nx Cloud path). I ran the same targets directly, one package at a time (see Testing), but not the test:pr command itself. With NX_NO_CLOUD=true, Nx works. For the durable work, I ran the test:pr targets with nx affected --base=9030d6990 (see Testing), so only the projects that the durable work affects. For the host work, I ran the test:pr targets with nx affected (42 projects, see Testing).
  • The understanding box is for the author to tick after review.

🚀 Release Impact

  • This change affects published code, and I have generated a changeset.
  • This change is docs/CI/dev-only (no release).

Testing

Commands run (media work, at fc24c4cfd). Each command ran on its own, and every one passed:

  1. For each changed package (ai, ai-persistence, ai-harness, ai-harness-cli, ai-mcp, ai-acp, and the 9 providers): build, vitest run, test:types, test:oxlint, and test:build (publint).
  2. The example and E2E type checks: examples/harness-cli tsc --noEmit, and testing/e2e test:types.
  3. Root checks: test:sherif, test:knip, test:docs, test:kiira (1693 snippets), test:maintainer, test:ai-review, test:dts, and test:react-native.
  4. E2E: pnpm --filter @tanstack/ai-e2e test:e2e -- --grep "harness media" gives 3 passed. --grep "harness" gives 14 passed and 2 skipped (the gated Claude Code smoke tests).

Coverage was not run locally, because it is CI-only.

The example, run by hand with real keys (at 10f008d57). Line mode (piped input), and the real Ink screen through a fake terminal:

  1. Media: an image, speech, a Grok video, a watercolor version of an attached image, a 15-second song (fal-ai/elevenlabs/music, MP3), and a 5-second sound effect (fal-ai/stable-audio-25, WAV) were saved.
  2. Notion and Linear: the agent listed the 3 latest Linear issue titles and searched Notion.
  3. Coding agents: "Hey, run a Codex agent" and a Claude Code edit both ran, with the local CLIs.
  4. Voice, with speech files as the voice: "describe fox dot png" attached playground/fox.png, "make a pencil sketch of the last image" attached the last saved image, and "run a Codex agent that counts the files" ran Codex. The screen showed each transcript and the agent output.
  5. Screen: /model claude switched the model. Ctrl+R recorded the real microphone. /mic listed 6 inputs, and a silent recording was not sent.

Provider keys, run by hand (at 6a90d924b) with a temporary home folder and no env keys:

  1. /keys listed 5 providers as missing, and a message stopped with "Sign in to openai. Run /connect openai."
  2. /connect anthropic and /connect fal saved real keys. The output showed only the last 4 characters, and the full keys were in no output.
  3. In a new process, /model claude answered with the saved key, and the sound effect agent made a real WAV with the fal key from ctx.keys.
  4. /disconnect anthropic brought the sign-in message back.
  5. On the Ink screen, the key prompt showed dots while typing, and the key was in no frame.

The OpenRouter browser sign-in is covered by unit tests only (a fake exchange). Nobody ran it against openrouter.ai yet.

The new screen, in a fake terminal (at 4b61d2bf5). A script drove the real Ink screen with key presses, with a temporary home folder and no env keys. 34 of 34 checks passed:

  1. / lists the commands. Typing filters them, Tab fills one, and Enter runs it.
  2. /model and /connect open pickers. OpenRouter shows "sign in with the browser, no key to paste". OpenAI shows "opens the page to make a key".
  3. The up and down arrows go through the sent lines, and history.json has them. A line from the history does not open the command list.
  4. The cursor block is before the placeholder, and the arrows move it in the line. The footer counted the model call and the context after a demo turn.
  5. A made-up xAI key showed as dots, was saved, and was removed from the /disconnect picker. The key was in no frame and not in the history.

Also run: ai-harness vitest (21 tests in the 2 changed files), tsc --noEmit for ai-harness and the example, test:oxlint, oxfmt, and test:kiira (1695 snippets). The startup screen clear runs only in a real terminal. Nobody checked it in a real terminal yet.

The connect fixes and the screen update (at fd49987a0). 17 of 17 fake-terminal checks passed:

  1. The frame is one line shorter than the screen, at startup and after a reply, so the input sits at the bottom.
  2. A demo reply renders bold and a list, and "draw me a chart" renders a Mermaid chart as text.
  3. /model lists the real models with provider and context. /effort opens a picker, and picking high shows a green ✓ Effort: high. and effort high in the footer.
  4. /connect notion shows the waiting spinner. A made-up xAI key stays hidden and saves in green.
  5. With Ink's own console (no debug mode), an SDK strict: false warning, its details, and a highlighter warning are hidden, and other console output still shows.

A script checked that /effort reaches the model call: reasoning.effort for OpenAI, output_config.effort for Anthropic, and xhigh for max on OpenRouter. The new connector test fails without the fix, with the same "Issuer mismatch" error. ai-harness (391 tests) and ai-mcp (359 tests) pass, with tsc, test:oxlint, test:sherif, and test:knip. Nobody signed in to the real Linear yet after the fix.

Durable sessions (at 740e3aea0). Each command ran one task at a time:

  1. nx affected --base=9030d6990 --head=HEAD with the test:pr targets: 194 of 195 tasks pass. The one failure is ai-sandbox-docker tests/sbx.test.ts, "measures whether kill() stops the in-VM process". It is a live test that runs only when the Docker sbx CLI is installed (here v0.38.0), and it fails the same way on a second run. This branch does not change ai-sandbox or ai-sandbox-docker, and CI skips the test.
  2. nx run-many --targets=test:types --projects=examples/**,testing/**: 28 projects pass.
  3. E2E with CI=1 and 1 worker: the 14 specs that use ai-harness or ai-persistence give 50 passed. The other specs did not run, because the machine was low on memory. They use only the core packages, which the durable work does not change.
  4. ai-harness has 466 unit tests, and ai-persistence has 308 (with the log conformance cases). Both pass.

Reasoning and the catalog (at f00762cd5). From that work's report:

  1. For each of the 21 changed packages (ai, openai-base, ai-models, and 18 providers): vitest run, test:types, test:oxlint, and test:build. All 84 pass.
  2. A full build, test:docs, kiira check on every docs page (1,711 snippets), sherif, knip, test:dts, test:maintainer, the model-sync tests, and tsc on the changed examples and test apps: pass.
  3. With the pi catalog snapshot, 1,425 of 1,425 models of the 38 covered providers resolve.
  4. Its E2E specs were not run.

Loop and adapter fixes (at 8a461decf). Each command ran one task at a time, on this branch after the merge of #1578 and #1579:

  1. vitest run, test:types, and test:oxlint for ai (2047 tests), ai-harness (467), ai-anthropic (170), and ai-cloudflare (32): pass. With the parallel tool default, ai-mcp (359), ai-harness-cli (46), ai-acp (74), and ai-code-mode (153) also pass.
  2. test:build (publint) for ai, ai-anthropic, and ai-cloudflare, with the new @tanstack/ai/testing subpath: pass. test:knip, test:sherif, test:docs, and kiira check on the 8 changed docs pages (73 snippets): pass.
  3. E2E with CI=1 and 1 worker: the parallel tool specs, both Anthropic wire specs, cloudflare-binding-wire.spec.ts, harness-protocol.spec.ts, and harness.spec.ts give 20 passed. testing/e2e test:types passes.
  4. Not run: the full nx affected target set, because the machine was low on memory. The Gate 1 repros of the two fixes are in fix(ai): run the server tools of one model turn at the same time #1578 and fix(ai, ai-anthropic): send redacted thinking and tool errors back to Claude #1579.

Skills, live commands, the workspace, and the Cloudflare example (at 562f88a06). Each command ran on its own:

  1. vitest run for ai-skills (70 tests) and ai-code-mode (154 tests): pass.
  2. vitest run for ai-harness: 470 of 478 pass in the full run. The other 8 are in build-stub.test.ts and the bash test of workspace-tools.test.ts. They start child processes, and they hit the 5 s timeout on a busy machine. With --testTimeout=60000, all 22 tests of those two files pass.
  3. The Cloudflare example: tsc for the CLI and the Worker, and oxlint, pass. The Worker tests in workerd give 30 passed and 16 skipped (the stores that the Worker does not have). test:sherif and test:knip pass.
  4. The Cloudflare example against a local wrangler dev Worker, in line mode with a temporary home folder. A demo message saved the session in the Worker (30 log records, and its title in /sessions). --resume loaded it from the Worker and kept the demo model, and the log grew to 47 records. A wrong secret gets 401.
  5. Not run: the Ink screen of the Cloudflare example, real Workers AI and AI Gateway calls (they need a token), and a deploy to Cloudflare.

Harness host features (at a89b5228a, after the merge of 562f88a06). Each command ran with NX_NO_CLOUD=true:

  1. vitest run for ai-harness: 575 tests pass.
  2. nx run-many with the test:pr targets for the 5 changed or touched packages: 28 tasks pass.
  3. nx affected with the test:pr targets (42 projects): pass, with test:kiira at 1726 snippets. The one failure is the same live sbx test of ai-sandbox-docker as in the durable work. This branch does not change that package.
  4. nx run-many --targets=test:types --projects=examples/**,testing/**: 29 projects pass. test:dts is clean.
  5. E2E: the harness protocol spec gives 8 passed, with the new tests for a retried 503 and a beforeFinish turn. One local coverage run of ai-react (26 files, 248 tests) passed. The ai-react coverage job failed in CI on the old head, and it did not fail here.

Background agents after a restart (at 40bce9675). Each command ran on its own. nx affected did not run.

  1. vitest run for ai-harness: 581 tests pass. The 6 new tests fail without the fix. tsc and test:oxlint pass.
  2. vitest run for ai-mcp (359), ai-dashboard (4), ai-harness-cli (46), and ai-skills (70): pass.
  3. E2E: harness.spec.ts gives 5 passed, with the new restart test. harness-protocol.spec.ts and harness-media.spec.ts give 11 passed. testing/e2e test:types passes.
  4. test:docs, kiira check (1726 snippets), test:sherif, and test:knip: pass.

Prompt caching (at b6a750075). Each command ran on its own, and every one passed:

  1. For each changed package (ai, ai-harness, ai-anthropic, ai-bedrock, ai-openai, ai-models, ai-openrouter, ai-mistral, ai-claude-code, ai-cloudflare): build, vitest run, test:types, test:oxlint, and test:build (publint). Nx ran no tasks in the worktree, so each target ran with pnpm --filter.
  2. Root checks: test:knip, test:sherif, and test:docs. kiira check on the 6 changed doc pages gave 71 snippets passed.
  3. E2E, before the merge of 40bce9675: the full suite gave 451 passed and 3 skipped. One spec first failed because ai-sandbox-cloudflare had no dist in the new worktree, and it passed after a build. After the merge, harness.spec.ts and prompt-cache-wire.spec.ts gave 7 passed.
  4. A live check with real keys sent the same 8,345-token prompt start more than once. On claude-sonnet-4-6, call 1 wrote 8,342 tokens to the cache, and call 2 read 8,342 tokens from it. With promptCache: 'none', nothing was read or written. On gpt-5-mini, call 2 read 7,296 of 7,386 input tokens from the cache.

Not run for this work: pnpm test:pr itself, test:types of the example apps, and coverage (CI-only).

The merge of main (at 79a8b37ba). Each command ran on its own, and every one passed:

  1. For the 19 packages that the merge or the prompt caching work changes: build (with dependencies), vitest run, test:types, test:oxlint, and test:build (publint).
  2. Root checks: test:knip, test:sherif, test:docs, and test:kiira.
  3. E2E: the full suite gave 453 passed and 3 skipped.

Harness routing (at a501c7b8d, on top of 79a8b37ba). Each command ran on its own, and every one passed. Nx could not load @tanstack/workspace-plugin in the worktree, so each target ran per package.

  1. vitest run, test:types, test:oxlint, and test:build for ai (2082 tests), ai-harness (630), ai-client (849), ai-mcp (374), ai-harness-cli (46), ai-dashboard (4), ai-skills (70), ai-acp (74), ai-code-mode (154), and ai-react (248).
  2. Root checks: test:sherif, test:knip, test:docs, test:kiira (1748 snippets), and test:dts.
  3. test:types for examples/harness-cli, examples/harness-cli-cloudflare, and testing/e2e.
  4. E2E: harness.spec.ts, harness-protocol.spec.ts, and harness-media.spec.ts give 17 passed, with the new routing test.

Not run for this work: the full E2E suite, pnpm test:pr itself, and coverage (CI-only).

Routed-turn fixes (at 5cc000e22). Each command ran on its own, and every one passed:

  1. vitest run for ai-harness: 636 tests pass. The 6 new tests fail without the fix: a 503 in the handoff main part, beforeFinish in the handoff, and a subagents.router child that resumes after an approval (no log and durable). tsc, test:oxlint, and test:build pass.
  2. vitest run for ai-mcp, ai-harness-cli, ai-dashboard, ai-skills, and ai-acp: pass.
  3. E2E: the three harness specs give 18 passed. The new test "a handoff retries a model error, and an approved routed agent resumes" fails on the old build, and passes on the new one. testing/e2e test:types passes.

Resolve after a restart, and ai-skills coverage (at 193ec947a). Each command ran on its own, and every one passed:

  1. vitest run for ai-harness: 644 tests pass. interrupt-restart.test.ts has 8 new tests: a main-model approval, a subagents.router agent, and a root-routed plan, each resolved on a new host, and a resolve that removes the stored copy (no log and durable). The first 6 fail without the fix, and the last 2 fail without the delete. tsc, test:oxlint, and test:build pass.
  2. vitest run for ai-skills (73 tests), ai-mcp (374), ai-harness-cli (46), ai-dashboard (4), ai-acp (74), and ai-code-mode (154): pass. ai-skills test:types and test:oxlint pass. Coverage itself runs only in CI. A count of the functions of src/harness.ts gives about 86% for ai-skills, above the 84.86% base.
  3. E2E: the three harness specs give 19 passed. The new test "a resolve on a new host continues the turn that stopped for approval" fails on the old build, and passes on the new one. testing/e2e test:types passes.
  4. test:docs and kiira check (1748 snippets): pass.

Block order and mid-conversation changes (at 71f259161, after the merge of 193ec947a). Each command ran on its own with NX_NO_CLOUD=true:

  1. For @tanstack/ai, openai-base, ai-openai, ai-anthropic, ai-persistence, ai-harness, and ai-skills: build, vitest run, test:types, test:oxlint, and test:build. All 35 commands pass (@tanstack/ai: 2,157 tests). After the last merge, ai-skills (78 tests, the coverage tests of both sides kept) and ai-harness (647 tests) pass again.
  2. nx affected --base=origin/main with the PR target set (109 projects): pass, except 2 tasks in packages this work does not touch: the known live Docker sbx test of ai-sandbox-docker, and ai-solid tests/chat-ui/text-part.test.tsx, which fails to load on Windows (file:///@solid-refresh, a vite-plugin-solid environment error). pnpm test:react-native passes.
  3. nx run-many --targets=test:types --projects=examples/**,testing/**: 29 projects pass. test:dts: clean. test:docs: no broken links. test:kiira: 1,755 of 1,755 snippets pass.
  4. Each change has a test that failed before it. The engine, the converters, the wire, the server join, and the Anthropic request were checked end to end with a script, and the old and new adapters were compared on 240 OpenAI and 254 Anthropic requests: with no channels, every body is byte for byte the same.
  5. E2E: not run on purpose. One full run on a heavily loaded machine (--workers=2) gave 427 passed and 28 failed. anthropic-thinking-order-wire.spec.ts and prompt-cache-wire.spec.ts passed. 27 failures were page-load timeouts, and 1 was summarize.spec.ts "openai: summarizes text" (an empty summary). They were not re-run. The 2 new E2E specs are written and type-check, but they have not run locally. CI runs them.
  6. Coverage: not run locally (CI only). The ai-skills function coverage gets new tests.

Per-turn overrides and live mid-conversation checks (at a8b8d2dc6). Each command ran on its own, and every one passed:

  1. For ai, ai-harness, openai-base, ai-openai, and ai-anthropic: build, vitest run, test:types, test:oxlint, and test:build. One timing test (build-stub.test.ts) timed out once under parallel load. It passed alone, and so did the full harness suite (662 tests).
  2. Root checks: test:knip, test:sherif, test:docs, and test:kiira.
  3. E2E: the full suite gave 462 passed and 3 skipped.
  4. Live checks with real keys, each a chat() run that adds a tool or a prompt after the first model call:
    • OpenAI additional_tools on gpt-5.5: passes after the namespace fix. All 3 calls returned 200, and each read 5,632 cached tokens.
    • OpenAI mid-conversation developer message: passes, and the model followed it.
    • Claude tool addition on claude-opus-4-8 (the beta header, the deferred placeholder, defer_loading, and tool_addition): passes. Calls 2 and 3 read 9,956 and 10,130 cached tokens.
    • Claude mid-conversation system message: passes, and the model followed it.
    • Cloudflare AI Gateway: not run, because there are no Cloudflare credentials.
  5. A live harness session on Anthropic. Turn 2 had overrides (claude-opus-4-8, reasoning, and a get_secret tool), and the model called the tool. Turns 1 and 3 used claude-sonnet-4-6 with no extra tool.

E2E. P0 to P6 add or extend E2E specs: harness.spec.ts, harness-protocol.spec.ts, dashboard.spec.ts, and subagents.spec.ts. The media work adds harness-media.spec.ts: upload, a signed URL with no auth headers, a changed signature (403), and a generated image served from its signed URL. The durable work adds a test to harness-protocol.spec.ts: a prompt sent twice with the same inputId to a durable host runs once. The loop fixes add three wire specs: server-client-sequence.spec.ts "parallel server tools run at the same time" (two 300 ms tools must overlap), anthropic-redacted-thinking-wire.spec.ts, and cloudflare-binding-wire.spec.ts (the real Anthropic and OpenAI adapters through a fake env.AI). P7 to P14 add no E2E spec. One existing E2E test is flaky: interrupts-test/batch.spec.ts "clear ignores a late interrupt submission failure". It passed on re-run.

Manual test: the terminal agent.

  1. From the repo root, run pnpm install, then pnpm build:all.
  2. Run pnpm --filter harness-cli-example start. The screen clears and shows only the harness.
  3. Type /connect and pick OpenRouter (browser sign-in), or pick OpenAI and paste a key. Then type /model and pick a model of that provider.
  4. Type create hello.txt with a short poem. The agent asks before write_file. Type y.
  5. Look at the footer: the context and the tokens grew. Press the up arrow to get your last line back.

Manual test: media and voice (needs OPENAI_API_KEY and ffmpeg).

  1. In the example, type make an image of a fox in the snow. Expect - [1] image ... saved: example-coder-media\...png.
  2. Press Ctrl+R, say "make a pencil sketch of the last image", and press Ctrl+R. Expect the transcript, then image [2].
  3. Press Ctrl+R, say "Hey, run a Codex agent that lists the files in the playground", and press Ctrl+R. Expect the codex agent line and its output.

Manual test: durable sessions.

  1. Do step 1 of the terminal test.
  2. Run pnpm --filter @tanstack/ai-harness exec vitest run tests/durable-session.test.ts tests/durable-inputs.test.ts tests/durable-joins.test.ts tests/durable-tools.test.ts. The tests stop a host in the middle of a turn or a tool batch, then rebuild the session from the log.
  3. Run pnpm --filter @tanstack/ai-e2e test:e2e -- --grep "same inputId". Expect 1 passed: two sends give one answer in the transcript, and another message with the same id is rejected.

Manual test: harness host features.

  1. Do step 1 of the terminal test.
  2. Run pnpm --filter @tanstack/ai-harness exec vitest run tests/shared-log.test.ts tests/turn-control.test.ts tests/leases.test.ts tests/durable-joins.test.ts. The tests run two sessions on one log, retry a failed model call, and stop a host in the middle of a turn.
  3. Run pnpm --filter @tanstack/ai-e2e test:e2e -- --grep "harness protocol". Expect 8 passed, with "retries a 503 from the model in the same turn".

Manual test: background agents after a restart.

  1. Do step 1 of the terminal test.
  2. Run pnpm --filter @tanstack/ai-e2e test:e2e -- --grep "stopped host". Expect 1 passed: a second host ends the run failed, and the wake turn answers.

Manual test: harness routing.

  1. Do step 1 of the terminal test.
  2. Run pnpm --filter @tanstack/ai-e2e test:e2e -- --grep "routing.router". Expect 1 passed: write goes to writer, hello goes to the main model, article runs two agents and then a third that reads their text, and price gives pricer the input { vendor: 'acme' }.
  3. Run pnpm --filter @tanstack/ai-harness exec vitest run tests/routing.test.ts tests/subagent-router.test.ts. Expect 2 files passed.

Manual test: MCP (P14). Nobody tried a real MCP client by hand yet. The P14 tests use the real MCP SDK client over server.fetch.

  1. Do step 1 of the terminal test.
  2. Run claude mcp add harness-example -- npx tsx <repo>/examples/harness-cli/src/cli.ts --mcp. Use the absolute path of your clone for <repo>.
  3. In Claude Code, ask "Ask harness-example to say hello". Expect a chat tool call and an answer that starts with (demo model).

Manual test: skills.

  1. Do step 1 of the terminal test, then run pnpm --filter harness-cli-example start.
  2. While it runs, make ./.agents/skills/release-notes/SKILL.md with a name and a description in its front matter.
  3. Type /. Expect /release-notes in the list, with no restart.

Manual test: the Cloudflare example.

  1. Do step 1 of the terminal test.
  2. In examples/harness-cli-cloudflare, make worker/.dev.vars with HARNESS_SECRET= and ENCRYPTION_KEY= lines, then run pnpm worker:dev.
  3. In a second terminal, run pnpm --filter harness-cli-cloudflare-example start. Give http://localhost:8787 and the HARNESS_SECRET value.
  4. Type /model demo, send hello, then type /exit.
  5. Run the start command again with --resume. Expect the same session, loaded from the Worker.

Manual test: prompt caching (needs ANTHROPIC_API_KEY).

  1. Write a chat() call with anthropicText('claude-sonnet-4-6'), a system prompt of more than 1,024 tokens, a threadId, and an onUsage middleware that logs usage.promptTokensDetails.
  2. Run it 2 times in 5 minutes. Call 1 shows cacheWriteTokens near the prompt size. Call 2 shows cachedTokens near the prompt size.
  3. Run it again with promptCache: 'none'. Both counts are 0.

Manual test: block order and mid-conversation changes.

  1. Do step 1 of the terminal test.
  2. Run pnpm --filter @tanstack/ai-anthropic exec vitest run tests/thinking-replay.test.ts tests/mid-conversation-changes.test.ts. Expect the "sends thinking, tool_use, thinking, tool_use back in that order" test and the tool-mode tests to pass.
  3. Run pnpm --filter @tanstack/openai-base exec vitest run tests/responses-mid-conversation.test.ts. Expect "sends added tools as additional_tools and keeps the start set in tools" to pass.
  4. Run pnpm --filter @tanstack/ai-e2e test:e2e -- tests/mid-conversation-changes-wire.spec.ts tests/anthropic-thinking-order-wire.spec.ts. Expect 6 passed. These specs have not run locally yet.

Manual test: per-turn overrides.

  1. Open a harness session whose adapter is anthropicText('claude-sonnet-4-6').
  2. Run session.prompt('Call get_secret.', { overrides: { adapter: anthropicText('claude-opus-4-8'), tools: [getSecret] } }). The request goes to the Opus model with the get_secret tool, and the model calls it.
  3. Run session.prompt('Hi'). It goes to the Sonnet model with no get_secret tool.

How this PR makes testing easy.

  • Unit tests in each changed package. Most are in packages/ai-harness/tests. The media tests are media.test.ts, session-media.test.ts, http-media.test.ts, view-media.test.ts, and client-media.test.ts, plus harness-media.test.ts in ai-mcp, agent-media.test.ts in ai-acp, and attach.test.ts and cli-media.test.ts in the CLI.
  • P14 adds packages/ai-mcp/tests/harness.test.ts, with a real MCP SDK client.
  • The durable work adds log.test.ts, durable-session.test.ts, durable-inputs.test.ts, durable-joins.test.ts, durable-tools.test.ts, durable-tool.test.ts, and edge-safety.test.ts in ai-harness, and the opt-in log conformance cases in @tanstack/ai-persistence/testkit, which any LogStore backend can run.
  • The loop fixes add parallel-tools.test.ts, redacted-thinking.test.ts, context-overflow.test.ts (one test per provider message), and fake-text.test.ts in packages/ai/tests, thinking-replay.test.ts in ai-anthropic, and binding-fetch tests in ai-cloudflare. fakeText() itself is a test tool: any chat() test can run with no network.
  • The host work adds shared-log.test.ts, turn-control.test.ts, transient-errors.test.ts, and leases.test.ts in ai-harness, and more cases in durable-joins.test.ts, durable-inputs.test.ts, durable-tools.test.ts, and log.test.ts.
  • The background agent fix adds cases to resume.test.ts, durable-inputs.test.ts, leases.test.ts, and session.test.ts in ai-harness, and a restart test to harness.spec.ts.
  • The prompt caching work adds prompt-cache.test.ts in ai, ai-harness, ai-anthropic, ai-openai, ai-openrouter, ai-mistral, and ai-bedrock (in tests/converse), prompt-cache-binding.test.ts in ai-cloudflare, more cases in compatible-quirks.test.ts, and the E2E spec prompt-cache-wire.spec.ts.
  • The routing work adds routing.test.ts and subagent-router.test.ts in ai-harness, router input cases in define-agent.test.ts and subagent-interrupts.test.ts, the type test subagent-route-types.test-d.ts in ai, and a routing test in harness.spec.ts.
  • The block order and mid-conversation work adds block-order.test.ts, chat-block-order.test.ts, mid-conversation.test.ts, and chat-mid-conversation.test.ts in packages/ai/tests, ordered assistant rows in ag-ui-wire.test.ts, responses-mid-conversation.test.ts in openai-base, mid-conversation-changes.test.ts in ai-openai, ai-anthropic, and ai-harness, the Anthropic replay block order cases in thinking-replay.test.ts, and the E2E specs mid-conversation-changes-wire.spec.ts and a client-tool case in anthropic-thinking-order-wire.spec.ts.
  • The per-turn overrides work adds turn-overrides.test.ts in ai-harness, middleware-prompt-cache.test.ts in ai, namespace cases in openai-base/tests/responses-mid-conversation.test.ts, and the E2E spec harness-turn-overrides.spec.ts.
  • The E2E specs above, in testing/e2e/tests.
  • The skills work adds packages/ai-skills/tests/harness.test.ts, live-commands.test.ts and more workspace-tools.test.ts cases in ai-harness, and lazy cases in ai-code-mode/tests/harness.test.ts.
  • Two runnable examples: examples/harness-cli, and examples/harness-cli-cloudflare with its Worker tests in workerd. Their READMEs have more steps.

Risk / rollback

  • This is a large change: 629 files and about 174,600 added lines. Most of it is in the new packages, and much of that is the generated @tanstack/ai-models catalog data. The loop fixes are 51 files and about 2,900 lines. The media work is 97 files and about 8,000 lines. The durable work is 43 files and about 5,800 lines. The host work is 37 files and about 4,700 lines. The block order and mid-conversation work is 52 files and about 6,250 lines.
  • Block order: one wire change reaches every useChat user. An assistant message out of the default block order goes out as more AG-UI rows, and a UIMessage that holds two model calls reaches the server as assistant, tool, assistant. A message in the default order sends the same rows as before.
  • Mid-conversation changes: on by default for 10 GPT and 5 Claude models on the provider's own API. On those Claude models, every request with tools gets the mid-conversation-tool-changes-2026-07-01 beta and one deferred placeholder tool. The request shapes are copied from pi and were not checked against the real APIs. To turn it off, set midConversationChannels: false on the adapter.
  • Host work: one change reaches every host. A join now stops at the first steer that has a cancel request, or that canJoin refuses. The other steers run as their own turns. Every other host option is off until a host sets it. The fold checkpoint is now v: 2, so an old log folds from the start once.
  • Breaking: reasoning options. chat({ reasoning }) replaces the reasoning fields in modelOptions of each provider (for example OpenAI reasoning and Anthropic thinking). docs/migration/reasoning-option.md shows the move.
  • Default change: parallel server tools. Server tools of one turn now run at the same time, also in harness sessions. Tools that depend on each other need sequential: true, or the run needs toolExecution: 'sequential'. onAfterToolCall fires in finish order.
  • Durable mode runs only when persistence.stores.log is set. Three durable changes also reach the default mode. A steer that arrives while the model writes its final answer now gets its answer in the same turn. Crash repair keeps the results of the finished calls in a batch. snapshot().queuedTurns also counts queued steers.
  • In existing packages, the changes are in @tanstack/ai (subagents, the chat middleware context, and TextAdapter.inputModalities), the 9 providers (a runtime input map), ai-persistence, ai-mcp, ai-acp, ai-code-mode, and ai-sandbox.
  • A signed media URL skips authorize by design, so it works in <img>. It is bound to the thread, the file id, and a 1-hour expiry. Set mediaSecret, or URLs stop working after a restart.
  • Background agents. A failed agent run now adds a [name failed] note to the transcript, unless attach: 'none' is set, and wake: true also wakes on a failure. Without a session log, a host cannot know wake, so a stopped agent run gets only the note.
  • The stdio loop of --mcp has no automated test. Only the flag parse has one.
  • workspaceTools still refuses a path outside the workspace, unless outside: 'ask' is set.
  • The Cloudflare example is private. Adding it moved undici (a dependency of cheerio) from 7.29.0 to 7.29.1 in the lockfile.
  • Default change: prompt caching. Every chat() call and harness session now sends cache fields. On Claude, a cache write costs about 1.25 times the input price. So a large prompt that is sent one time only costs up to 25% more input there. Pass promptCache: 'none' for that case. summarize() and direct adapter calls do not change.
  • Harness routing. While the routed agents run, the turn runs no turn hooks and no chat middleware of the harness or its plugins, the same as chat() with subagents.router. The main part of 'handoff' runs the hooks.
  • Interrupts after a restart need stores.metadata. The session keeps the interrupted turn there. Without a metadata store, a resolve after a restart is refused, as before. The stored copy can hold large agent cards for a routed turn.
  • Usage change: promptTokens. On Anthropic, Bedrock, and Claude Code, it is now the total input (uncached + cache read + cache write), the same as OpenAI and Gemini. Before, it was the uncached part only.
  • Per-turn overrides. A turn without overrides sends the same request as before. The overrides are not stored, so two prompts with the same inputId and other overrides count as one input. The OpenAI namespace fix only adds a field that the API asks for.
  • Rollback: revert the merge commit.

Public API change

Each phase PR shows its own API before and after. For the stack as a whole:

Before

import { chat } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'

// One request, one answer. The app keeps the conversation.
const stream = chat({
  adapter: openaiText('gpt-5.6'),
  messages: [{ role: 'user', content: 'Write a haiku about the sea.' }],
})

After

import { createHarnessHost, defineHarness, mediaPart } from '@tanstack/ai-harness'
import { memoryPersistence } from '@tanstack/ai-persistence'
import { openaiText } from '@tanstack/ai-openai'

const assistant = defineHarness({
  name: 'acme/assistant',
  adapter: openaiText('gpt-6-astra'),
  systemPrompts: ['You are a helpful assistant.'],
})

// One long-lived session per conversation.
const host = createHarnessHost({ persistence: memoryPersistence() })
const session = await host.open(assistant, { threadId: 'thread-1' })
const turn = await session.prompt('Write a haiku about the sea.')
console.log(turn.text)

// Send a file with a prompt.
const photo = await session.putMedia(bytes, { mimeType: 'image/png', name: 'sea.png' })
await session.prompt([{ type: 'text', content: 'Describe this.' }, mediaPart(photo)])

With a log store, the same session is durable, and a retry runs once:

import { memoryLogStore } from '@tanstack/ai-persistence'

const durableHost = createHarnessHost({
  persistence: {
    stores: { log: memoryLogStore(), runs: memoryPersistence().stores.runs },
  },
})
const durable = await durableHost.open(assistant, { threadId: 'thread-2' })
durable.prompt('Summarize the report.', { inputId: 'req-42' })
const settlement = await durable.settled('req-42')
console.log(settlement.outcome) // 'completed', 'failed', 'aborted', or 'interrupted'

The host work adds these options. Each one is off until the host sets it:

import {
  createHarnessHost,
  defineHarness,
  durableTool,
  retryTransientErrors,
} from '@tanstack/ai-harness'
import { memoryLogStore } from '@tanstack/ai-persistence'

const agent = defineHarness({
  name: 'acme/agent',
  adapter: openaiText('gpt-6-astra'),
  turn: {
    onModelError: retryTransientErrors({ maxRetries: 3 }), // retry a 429 or a 503 in the same turn
    beforeFinish: ({ cycle }) =>
      cycle === 0
        ? { messages: [{ role: 'user', content: 'Check your work first.' }] }
        : undefined,
    canJoin: ({ message }) => typeof message === 'string',
  },
  durability: { recover: ({ decision }) => decision }, // run again, or settle
})

// Many sessions on one log, with one state for the whole log.
const host = createHarnessHost({
  persistence: { stores: { log: memoryLogStore(), leases: myLeaseStore } }, // or runs
  reduce: { initial: 0, record: ({ state }) => state + 1 },
})
const parent = await host.open(agent, { threadId: 'task-1' })
const child = await host.open(agent, { threadId: 'task-1/child', logId: 'task-1' })
host.logState('task-1') // 0 plus one for each host record of the log

// After a crash, the model gets a tool error, and the email is not sent again.
const send = durableTool(sendEmailDefinition, execute, { replay: 'never' })

Reasoning is one option for every adapter:

// Before: modelOptions: { reasoning: { effort: 'high' } } on OpenAI,
// modelOptions: { thinking: { ... } } on Anthropic, and so on.
chat({ adapter: openaiText('gpt-6.1-sol'), messages, reasoning: 'high' })
chat({ adapter: anthropicText('claude-opus-5-5'), messages, reasoning: 'high' })

The loop fixes add these entry points:

import { chat, isContextOverflow } from '@tanstack/ai'
import { fakeText } from '@tanstack/ai/testing'
import { cloudflareBindingFetch } from '@tanstack/ai-cloudflare'

chat({ adapter, messages, tools, toolExecution: 'sequential' }) // default 'parallel'
isContextOverflow({ error: runErrorChunk }) // true for "prompt is too long: ..."
const fake = fakeText() // fake.setResponses([{ text: 'Hello' }])
createAnthropicChat('claude-opus-5-5', 'cloudflare-binding', {
  fetch: cloudflareBindingFetch({ binding: env.AI, vendor: 'anthropic' }),
})

The skills work adds these options:

import { defineHarness } from '@tanstack/ai-harness'
import { workspaceTools } from '@tanstack/ai-harness/plugins'
import { codeMode } from '@tanstack/ai-code-mode/harness'
import { skills } from '@tanstack/ai-skills/harness'

defineHarness({
  name: 'acme/coder',
  adapter,
  plugins: () => [
    skills({ dirs: ['./.agents/skills'] }), // each skill is a /<name> command
    codeMode({ driver, lazy: true }), // the model discovers the signatures
    workspaceTools({ root: process.cwd(), outside: 'ask' }),
  ],
})

Block order and mid-conversation changes need no code. The new fields and options are optional:

import { openaiText } from '@tanstack/ai-openai'
import { anthropicText } from '@tanstack/ai-anthropic'
import type { ChatMiddleware } from '@tanstack/ai'

// A middleware that adds a tool after the first tool batch.
const addForecast: ChatMiddleware = {
  name: 'add-forecast',
  onConfig: (ctx, config) =>
    ctx.phase === 'beforeModel' && ctx.iteration === 1
      ? { tools: [...config.tools, getForecast] }
      : undefined,
}

chat({ adapter: openaiText('gpt-6-astra'), messages, tools, middleware: [addForecast] })
// call 2: tools = [the start tools], input ends with { type: 'additional_tools', role: 'developer', tools: [get_forecast] }

chat({ adapter: anthropicText('claude-opus-5-5'), messages, tools, middleware: [addForecast] })
// call 2: tools = [start tools, placeholder, get_forecast (defer_loading)], then a system message with a tool_addition block

// Turn the channels off, or on through a gateway that passes them:
openaiText('gpt-6-astra', { midConversationChannels: false })
createAnthropicChat('claude-opus-5-5', key, { fetch, midConversationChannels: true })

// For adapter authors: ModelMessage.blockOrder, orderedAssistantBlocks(message),
// TextOptions.midConversationChanges, and splitMidConversationChanges(...) from '@tanstack/ai'.

With P14, any MCP client can use the same harness:

claude mcp add assistant -- npx tsx cli.ts --mcp

Prompt caching. Before

// Caching needed a manual marker on each request
chat({
  adapter: anthropicText('claude-sonnet-5-5'),
  messages: [
    {
      role: 'user',
      content: [
        { type: 'text', content: question, metadata: { cache_control: { type: 'ephemeral' } } },
      ],
    },
  ],
})

Prompt caching. After

chat({ adapter, messages, threadId }) // cached by default ('short'), key = threadId
chat({ adapter, messages, promptCache: 'none' }) // off
defineHarness({ name: 'acme/agent', adapter, promptCache: 'long' }) // every session
host.open(harness, { threadId, promptCache: { key: affinityKey } }) // one session

Harness routing. Before

// Every turn went to the main model. Root agents ran only from code.
await session.agents.pricer.run({ vendor: 'acme' })

Harness routing. After

defineHarness({
  name: 'acme/studio',
  adapter,
  agents: [writer, pricer],
  routing: {
    router: ({ messages }) =>
      String(messages.at(-1)?.content).includes('price')
        ? { name: 'pricer', input: { vendor: 'acme' } }
        : 'main',
  },
})

// With decide(): which picked agents need input, then attach it
const route = subagentRoute(agents)
route.needsInput(result) // [{ name: 'pricer', inputSchema }]
route.pick(result, { inputs: { pricer: { vendor: 'acme' } } })

Per-turn overrides. Before

// One adapter, reasoning, and tool set for every turn of the session
await session.prompt('Plan the migration')

Per-turn overrides. After

await session.prompt('Plan the migration', {
  overrides: {
    adapter: anthropicText('claude-opus-5-5'),
    reasoning: 'high',
    tools: [finishTool],
  },
})
defineHarness({ name: 'acme/agent', adapter, reasoning: 'medium' }) // default for every turn

AlemTuzlak and others added 30 commits September 26, 2026 12:41
Resume: a session holds a lease on each running turn, saves the transcript
around tool phases, and records tool calls that have no result. When a host
opens a thread whose turn lost its lease, it continues the turn in a new run
and emits harness.operation.resumed. toolDefinition({ replay }) decides
whether an unfinished tool runs again.

Protocol: createHarnessHandler serves capabilities, standard AG-UI runs, the
session event stream with cursors, control inputs with receipts, and
snapshots, behind a required authorize hook. handleHarnessSocket serves the
same session tier over a WebSocket. @tanstack/ai-harness/client adds a typed,
reconnecting client. harnessText runs a harness as the model of another
chat() call.

ACP: @tanstack/ai-acp/agent serves a harness as an ACP v2 agent (SDK 1.5),
with tool approvals as permission requests.

CLI: new @tanstack/ai-harness-cli. runCli(harness) gives an Ink terminal UI,
a print mode with exit codes, NDJSON output, --acp, and --serve.
Plugins can return commands (defineCommand), settings (configOption),
extension point items, and a main-model pick. The setup context adds
collect, typed events, persisted plugin state, settings, credentials, and a
session API with ask, prompt, transcript, and setConfig. The session adds
command, commands, setConfig, config, answer, and inspect, and the protocol
accepts command, answer, and config inputs.

Auth: oauthConnector adds connect/disconnect commands and gives tools a
fresh token. The OAuth runner does PKCE S256 with a single-use loopback on
127.0.0.1, and device-code sign-in. Missing credentials emit
harness.auth_required. ai-persistence adds the optional credentials store
and compare-and-set on the memory metadata store.

First-party plugins at @tanstack/ai-harness/plugins: permissions with modes,
workspaceTools, todos, modelPicker, projectInstructions, fileCommands,
compact, and usage. The CLI runs plugin commands, /config, /connect, answers
questions, and opens sign-in links.
@tanstack/ai: subagents.limits (maxDepth, maxConcurrent, maxCalls,
timeoutMs) for the tree of children the model starts through tools. One
SubagentBudget per root run is shared by every child and passed on through
ctx.chat({ subagents }), so a child cannot reset it. A refused start reaches
the model as a tool error.

@tanstack/ai-harness: plugins get ctx.agents.run, start, and group (with
cancel-siblings or collect). harnessAgent(harness) turns a harness into a
child agent for subagents.agents, and defineHarness takes a description.
A harness applies default limits (depth 2, 3 at once, 12 per tree), and
agents started from code count against them.
@tanstack/ai-harness/build: buildHarness bundles a harness with Bun into a
worker artifact (harness.js) plus harness.manifest.json (name, agents,
plugins, requirements, sha256 digest), and can compile a single executable.
artifactText(dir) checks the digest, starts one worker process per thread,
and uses it as a text adapter.

@tanstack/ai-harness/worker: runHarnessWorker serves session-tier frames as
NDJSON on stdin and stdout.

harnessText({ url, token }) uses a harness served on another machine (for
example with runCli --serve) as a text adapter.
…e example

New package @tanstack/ai-dashboard. `npx @tanstack/ai-dashboard` (or
startDashboard) runs a node:http server with no new dependencies. Hosts dial
out with connectDashboard, pair with a one-time code, and get a revocable
host token. The relay uses SSE plus POST and caches recent events per
session; inputs for an offline host wait and are delivered on reconnect.
The web app lists hosts and sessions, streams messages and tool calls,
shows approval cards and plugin questions, and sends prompts, steers, and
stops. It installs as a PWA on a phone.

@tanstack/ai-harness-cli adds --dashboard <url>.

examples/harness-cli: a small coding agent (permissions, workspace tools,
todos, model picker, typed agent) that runs with OpenAI, Anthropic, or a
demo model.
The Coverage job failed: ai-skills function coverage dropped from 84.86%
to 83.42%. Five functions of the skills harness plugin had no test:
`source.load`, `listResources`, `listScripts`, `folderOf`, and the
watcher's error handler. Three tests now run them through the plugin.
planMidConversationChanges folds the midConversationChange records of a transcript, compares them with the tools and system prompts of the current call, and gives the changes for the adapter and the record to save. promptHash is FNV-1a over a prompt.
With a valid blockOrder map, formatMessages sends the thinking, text, tool_use, and server tool blocks in the order Claude sent them. Unsigned thinking is still skipped at its place. A message without a valid map sends the same request as before.
When the adapter has channels and the engine passes midConversationChanges, tools keeps the start set, an added tool goes out as an additional_tools item, and an added prompt goes out as a developer message at its place. A provider tool keeps the full tools list. A subclass can override convertTools to use its own tool converter. An adapter without channels sends the same request as before.
For an adapter with mid-conversation channels, chat() compares the tools and system prompts of each model call with the records in the transcript, passes the change as TextOptions.midConversationChanges, and saves the record on the first assistant message of the call. An adapter without channels gets the same options as before, and no record is written.
An answer of thinking, tool call, thinking, tool call keeps its blockOrder in the next model call, the transcript, a fold of the log, and a second host.
… steer

With an adapter that has mid-conversation channels, the start tool set of a harness session holds across turns and a rebuild, and a tool added next to a steer goes out as a change before the next answer.
…lient sends the message again

A useChat client sends its history again on each turn without the engine's mid-conversation record. mergeStoredMessages now keeps the stored record when the incoming copy of the same message has none, so the prompt cache holds across turns with a message store.
…e server

uiMessagesToWire sends an assistant message whose blocks leave the default order as ordered AG-UI rows: a tool result ends a row, and a row that exists only for the order carries metadata.tanstack.continues. The server joins those rows back into one message with a blockOrder map, and the snapshot fold-back puts them into one UIMessage. A message in the default order sends the same rows as before.
… have them

A new model map gives the 10 GPT models from pi both channels. They are on by default only for OpenAI's own API: a baseURL, a fetch, or OPENAI_BASE_URL turns the default off. The config option midConversationChannels turns them on (true) or off (false). The adapter overrides convertTools, so the start set and additional_tools use the OpenAI tool converter, and the prompt cache fields stay as they are. Other models and custom endpoints send the same request as before.
…s that have them

A new model map gives the 5 Claude models from pi both channels. They are on by default only for Anthropic's own API: a baseURL, a fetch, ANTHROPIC_BASE_URL, or an injected client turns the default off. The config option midConversationChannels turns them on (true) or off (false). A change goes out as a system message before the next assistant message, or at the end. In tool mode the request gets the tool-change beta, a deferred placeholder tool, deferred added tools, and tool_addition blocks. The automatic tool cache marker goes on the last start tool, and a system message at the end takes the message marker. A provider tool keeps the full tools list. Other models, custom endpoints, and Vertex send the same request as before.
A new guide for tools and system prompts added during a conversation, sections under Prompt Caching on the OpenAI and Anthropic pages with the default rule, a When the tools change section in prompt-caching, a gateway note on the Cloudflare page, an adapter-author section in extend-adapter, the block order in thinking-content, and a link from turn-control.
Claude answers thinking, tool_use, thinking, tool_use with two client tools. The client shows that order, and after the client sends the results over the wire the next request replays the blocks in that order with their signatures.
A middleware adds a tool after the first tool batch. gpt-6-astra sends it as additional_tools, claude-opus-5-5 sends a tool_addition block with the beta and the placeholder, and models outside the map send the full tools. A custom fetch with no option also sends the full tools (the gateway case).
The record is saved only on a call that makes a start point or a change, and a provider tool keeps Claude out of tool mode, so that request has no beta and no placeholder.
…o feat/harness-host-generic

# Conflicts:
#	packages/ai-skills/tests/harness.test.ts
@AlemTuzlak AlemTuzlak changed the title feat(ai, ai-harness): TanStack AI harness, phases 0 to 14, media, provider keys, durable sessions, host hooks, reasoning, skills, prompt caching, routing, and loop fixes (combined stack) feat(ai, ai-harness): TanStack AI harness, phases 0 to 14, media, provider keys, durable sessions, host hooks, reasoning, skills, prompt caching, block order, mid-conversation changes, routing, and loop fixes (combined stack) Oct 2, 2026
ChatMiddlewareConfig has promptCache, next to reasoning. onConfig gets the current value, and a returned value goes to the adapter on the next model call. It stays until a middleware changes it.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…ses API

A tool that came through additional_tools comes back with a namespace. The adapter kept only the item id, so the next request failed with 400 Missing namespace for function_call. The tool call metadata now keeps the namespace, and the replayed function_call item sends it. Found by a live check on gpt-5.5.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
prompt() and followUp() take overrides: an adapter, reasoning, a prompt cache, and extra tools for one turn. They apply to every chat() call of the turn. They live in memory only: a queued turn keeps them, a joined steer uses the host turn's overrides, and a recovered turn uses the defaults. A middleware that rebuilds the tools keeps the override tools. HarnessConfig.reasoning sets the default reasoning, and the plugin adapter picker now gets the turn.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
A harness session runs 3 prompts on a capturing fetch. Only turn 2 has overrides: another model, reasoning, and an extra tool. The spec checks that turn 2 uses them and that turns 1 and 3 use the defaults.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
TurnOverrides.reasoning and HarnessConfig.reasoning now take ReasoningOption, as chat({ reasoning }) does: a level such as 'high', or { level, summary?, budgetTokens? }. ReasoningRequest is the normalized adapter type and needed summary.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
turn-control gets a section on overrides for one prompt. The harness overview shows HarnessConfig.reasoning. Plugins shows the adapter picker and its turn argument. Middleware shows reasoning and promptCache in onConfig.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
@AlemTuzlak AlemTuzlak changed the title feat(ai, ai-harness): TanStack AI harness, phases 0 to 14, media, provider keys, durable sessions, host hooks, reasoning, skills, prompt caching, block order, mid-conversation changes, routing, and loop fixes (combined stack) feat(ai, ai-harness): TanStack AI harness, phases 0 to 14, media, provider keys, durable sessions, host hooks, reasoning, skills, prompt caching, turn overrides, block order, mid-conversation changes, routing, and loop fixes (combined stack) Oct 2, 2026
@github-actions github-actions Bot added waiting-on: maintainer The ball is in the maintainers’ court and removed waiting-on: author Waiting for the author to respond or update merge-conflicts Conflicts with the base branch — needs a rebase labels Oct 2, 2026

@tannerlinsley tannerlinsley left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Alem, I focused this pass on permissions, explicit agent runs, and persistence failure handling at a8b8d2d. Three issues reproduced against the PR's package source that I'd like addressed before we merge.

Code mode drops explicit permission rules. At ai-code-mode/src/harness.ts:74, tool selection reads contributed PermissionRules, but not the rules passed to permissions({ rules }). A custom delete_record tool with an explicit deny rule executed through external_delete_record in code mode. Without code mode, the same tool stayed blocked: execution count 0 versus 1. Please apply the effective permission policy to moved tools/bindings, including explicit rules and defaults. This bypasses the configured tool policy; it doesn't establish an isolate escape.

Explicit agent runs lose approval continuation. At ai-harness/src/session.ts:2829, SUBAGENT_FINISHED is read for its result but its suspended outcome is ignored. A session.agents.child.run() whose ctx.chat() requested tool approval returned "" and recorded completed, even though its event contained outcome: suspended. Resolving the emitted approval ID returned no_pending_interrupts. The tool didn't execute. Please retain the continuation and expose a resolution path instead of recording success.

Store errors strand background operations and leases. The initial writes and terminal run update sit outside the execution catch. Injecting a rejection in either left the operation running, produced an unhandled rejection, and left both the result and host.close() pending at the probe deadline. The terminal failure also kept advancing the lease expiry. The normal-store control completed and closed. Please handle rejection across the whole lifecycle, settle the operation, and release leases in finally.

These were focused runtime probes with a scripted model, an injected code-mode driver invoking the real bindings, and wrapped memory stores. They weren't real-provider/real-isolate tests or a full review of all 711 files. Please add regression coverage for these cases.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

waiting-on: maintainer The ball is in the maintainers’ court

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Tool calls from one model step run one at a time, but the docs say they run in parallel

3 participants