Skip to main content

Assemble explicit capabilities

A production host should make its capabilities and security boundary visible in one builder chain. Installing tools makes them available; policy decides which calls are authorized.
  • Configure parent and delegated models independently when they need different providers or limits.
  • Give the parent and children separate step budgets.
  • Use an exact allowlist for high-risk presets; tool presence does not grant permission.
  • This example caps process count, retry attempts, and tool-result string size, but not wall-clock time. Set Limits.deadline for parent and child runs as needed, or enforce an outer host timeout for the whole operation.

Attach stateful services

Memory

harness.memory() opens the workspace-scoped SQLite store. MemoryConfig::new enables automatic recall and read-only search; MemoryConfig::read_write also exposes model-initiated mutation.

Skills

Skill installation or scaffolding changes disk. Call reload() before building or starting the next run so the catalog reflects those changes. Construct Skills directly when an embedded host should exclude user-home skill roots.

MCP

Connect servers before building the agent and disconnect them during cleanup. Use structured connect_stdio launches when arguments contain spaces or need explicit environment and working-directory fields.
MCP tools use names such as mcp__tickets__lookup. Skill and MCP catalogs are snapshotted at run boundaries; changing them does not rewrite an active run’s tool registry.

Choose a session lifecycle

Agent::run(request) opens a fresh ephemeral session for a one-shot request. For multiple turns or imported history, open a session explicitly. context(messages) seeds its conversation; persistent() saves that history so agent.resume_session(&id) can reopen it later. Resuming restores the transcript, not live built-in tool state such as processes.
Continuation adds no user turn and requires existing conversation beyond the system prompt. Attach RunRequest::cancellation(token) when the caller owns cancellation: cancelling the token stops the run, but cancelling the run does not cancel the caller’s token.

Preserve partial outcomes

Use run_outcome or RunHandle::outcome when accounting and partial transcript state matter. Record them before collapsing the outcome into a simple success/error result.
RunOutcome preserves execution status, accumulated usage, metered steps, partial messages, persistence status, and dropped-event count. into_result() intentionally gives up that richer failure view.

Retry and recover deliberately

RetryConfig::attempts(2) means two total attempts, not two retries after the initial call. Model retry targets transient transport and provider failures. Tool retry uses built-in classification, but hosts still need idempotency for externally mutating operations.
Truncation limits what enters model context while retaining the original tool result in a bounded session recovery store. When enabled, read_tool_result retrieves it by the original call ID. Persistent sessions save this store; forks snapshot it; clear and reset operations remove it.

Verify actions, not prose

A model saying “the check passed” is not proof. Match successful tool calls with their results, then inspect host-owned workflow and process state separately.
Also verify workflow terminal outcomes and process exit codes before reporting host-level success. The full pattern is implemented in crates/sdk/examples/live_host_assembly.rs.