deepseek-harness/docs/core-data-structures/subagent.md

100 lines
8.5 KiB
Markdown
Raw Normal View History

# Subagent
The subagent seam — an agent delegating work to a child agent. Like [bash](bash.md) it is **one optional capability**, not part of the agent-loop spine, so its vocabulary lives here rather than in [core.md](core.md). But it differs from every other seam on one axis: **multiple provider implementations coexist** in one context, registered by name (`ctx.subagents`), where bash allows only one executor. The registry shape mirrors the [LLM adapter registry](llm-streaming.md), not the single-service bash executor.
Add the ACP subagent backend: out-of-process delegation (PR3) The first OUT-OF-PROCESS subagent backend, proving the seam generalizes past the in-process backends. @deepseek-ai/dsh-subagent-acp runs each child agent in a spawned subprocess, driven over the Agent Client Protocol as the CLIENT — the direction-inverted twin of the dsh-acp server bridge. Point the configured command at the acp-agent example and the harness talks to its own process. - Fresh process per run: start spawns, runs one ACP session (initialize → newSession → prompt), dispose kills the subprocess and awaits its exit. - Minimal client stub: advertises no fs/terminal; accumulates agent_message_chunk text as the result output; auto-answers session/request_permission by a configured policy (reject default / allow). No start-time capabilities (an out-of-process child can't enforce the parent's depth/tool-filter); ignores request.parent; injects only `subagents`. - StopReason mapping (end_turn→completed, cancelled→aborted, …); result resolves error/aborted on a child failure, never rejects (seam contract). - Security: credential-shaped ambient env vars are scrubbed; the child's own key is forwarded only via explicit config.env. A spawn-level error (ENOENT) is captured and raced against the ACP drive so a bad command settles error rather than crashing the parent. Testing designed at every tier: keyless integration drives a scripted mock ACP server subprocess (cancellation incl. the pre-newSession race and a torn-pipe-after-cancel, permission auto-answer, non-message updates, spawn failure, HMR, export shape) at 100% coverage; a with-key e2e drives the REAL acp-agent example process (PONG + real file write, verified on disk) — the harness driving itself. Snapshot coverage of an ACP child is deferred as TODO(acp-subagent-replay) (each child is its own process with its own replay). Stayed on @agentclientprotocol/sdk 0.25.1: the proposed 0.28.x bump only deprecates the stable ClientSideConnection/AgentSideConnection API this layer uses (33 sites incl. the server bridge), turning no-deprecated red across code this PR shouldn't rewrite — that fluent-API migration is its own follow-up. The backend needs nothing 0.28.x adds. This completes the subagent seam stack (PR1 interface → PR2 in-process → PR2.5 snapshot infra → PR3 ACP); the seam RFC moves to implemented/, amended.
2026-06-22 10:47:02 +08:00
Interface: [dsh-subagent](../../packages/subagent/subagent) (`ctx.subagents` + the vocabulary below). Implementations are sibling packages (`dsh-subagent-spawn`, `-fork`, `-acp`); the model-facing consumer is [dsh-tool-subagent](../../packages/subagent/tool-subagent). The proposal and rationale: [the subagent RFC](../rfc/implemented/feature/2026-06-21-subagent-capability-seam.md).
Source: [`packages/subagent/subagent/src/types.ts`](../../packages/subagent/subagent/src/types.ts)
## Two kinds of capability, discovered two ways
A provider advertises its **start-time** features on a static descriptor the service checks BEFORE a run exists; a request that needs one the provider lacks is rejected loud (`SubagentError('UNSUPPORTED_CAPABILITY')`), never accepted-then-ignored. **Runtime** features (steering, resume) are instead optional methods on [`SubagentRun`](#a-live-run-subagentrun) — the method's presence IS the capability, and TS narrowing is the discovery mechanism.
```ts type-equiv
interface SubagentCapabilities {
readonly outputSchema: boolean
readonly depthLimit: boolean
readonly toolFilter: boolean
readonly persona: boolean
}
```
## The start request
What a caller asks for when starting a subagent. The tool layer builds this from the model's `{ description, prompt }` plus its own config; the service validates the start-time capabilities against the named provider, then passes it to `provider.start`. `parent` is REQUIRED — in-process backends read `parent.session.header` for the working directory, the `parentSession` lineage, and the delegation depth. The four optional fields (`outputSchema`, `maxDepth`, `toolFilter`, `persona`) each gate on the matching `SubagentCapabilities` flag — in-process backends realize `toolFilter` as a scoped `tools.restrict()` and `persona` as a scoped shadowing `deployment:persona` section, both composed in the child's creation window. `outputSchema` is an object-rooted JSON Schema within the subset `assertSupportedOutputSchema` (dsh-tools) enforces — a schema outside it is rejected loud at start; the in-process backends realize it with a forced `structured_output` capture tool (see the [driver README](../../packages/subagent/subagent-inprocess/README.md)).
```ts type-equiv
interface SubagentStartRequest {
readonly prompt: ContentBlock[]
readonly parent: Agent
readonly signal: AbortSignal
readonly agentOptions?: AgentOptions
readonly outputSchema?: StructuredOutputSchema
readonly maxDepth?: number
readonly toolFilter?: ToolRestriction
readonly persona?: string
}
```
`signal` is the single cancellation channel before and after readiness. The [subagent composition-controls RFC](../rfc/implemented/feature/2026-07-12-subagent-persona-tool-filter-and-depth.md) owns the persona, live global-tool filter, absolute-depth, and visibility-not-authority rationale.
## The terminal result: `SubagentResult`
2026-07-11 22:55:40 +08:00
The outcome of a run, resolved by `SubagentRun.result`. `structured` is present only after a requested `outputSchema` was successfully satisfied; requesting a schema does not guarantee it, and a provider may return `stopReason: 'error'` when the child fails or finishes without a valid capture. A non-`completed` `stopReason` means `output` may be partial — the consumer maps it to an `isError` tool result rather than reporting partial output as success.
```ts type-equiv
interface SubagentResult {
readonly output: ContentBlock[]
readonly structured?: unknown
readonly stopReason: SubagentStopReason
}
```
`SubagentStopReason` is a [merge-extensible derived union](core.md#the-map--derived-union-pattern) — a backend may add variants, so consumers branch on the known cases and treat an unknown terminal reason as a failure:
```ts type-equiv
interface SubagentStopReasonMap {
completed: 'completed'
aborted: 'aborted'
error: 'error'
'max-tokens': 'max-tokens'
refusal: 'refusal'
}
```
## A live run: `SubagentRun`
The handle the consumer holds after a provider has established a ready child. The consumer awaits `result` and MUST `dispose` on every path to cancel remaining work and reach child quiescence. `result` does NOT reject on a child-level failure — a model/transport failure resolves with `stopReason: 'error'` — so the consumer maps a non-`completed` reason to an `isError` result; it rejects only on an infrastructure fault the seam cannot represent. `sendMessage` and `resume` are OPTIONAL: a provider that supports the runtime capability defines the method; one that doesn't omits it.
```ts type-equiv
interface SubagentRun {
readonly id: AgentId
readonly result: Promise<SubagentResult>
dispose(): Promise<void>
sendMessage?(content: ContentBlock[]): void
resume?(content: ContentBlock[]): Promise<SubagentRun>
}
```
## The provider seam: `SubagentProvider`
One transport for running a child agent. Implementations register under a unique name via `SubagentService.registerProvider`; multiple coexist in one context. The service validates every requested start-time capability before calling `start`, so an implementation may assume e.g. `request.maxDepth` is honorable when present. `inheritsParentContext` is a DESCRIPTIVE fact beside the capabilities (nothing validates against it): whether a child sees the parent conversation (`fork`: true, `spawn`/`acp`: false) — the model-facing consumer derives truthful tool wording from it. It describes conversation history only, not tool registrations, injected services, or authority inheritance.
```ts type-equiv
interface SubagentProvider {
readonly name: string
readonly capabilities: SubagentCapabilities
readonly inheritsParentContext: boolean
start(request: SubagentStartRequest): Promise<SubagentRun>
}
```
`SubagentProvider.start()` and `ctx.subagents.start()` are the publication boundary: their promises fulfill only with a ready run. The service attaches result observation, emits `subagent/start`, and returns the same holder-owned run; a rejected start has already cleaned provider-owned partial resources and emits neither lifecycle event. For an in-process provider, a start listener can resolve the live child with `ctx.agents.get(info.id)`; a remote provider need not publish into the local registry. `subagent/end` carries `lastAssistantMessage` (the child's final `output`) on the settle path and reports `error` on infrastructure rejection. Both lifecycle events are observe-only emits with per-listener exception containment.
Add in-process subagent backends: spawn (fresh) and fork (seeded) The second PR of the subagent seam: the two in-process backends that run a child agent on the same cordis context, reusing the agent factory's quiescent AgentHandle teardown. Both register on ctx.subagents (PR1's named-provider registry) and share one run driver. - dsh-subagent-spawn: a FRESH child via ctx.agents.create — own session, the parent's model by default (overridable), zero inherited conversation. Also exports the shared in-process run driver (startInProcessRun): mint ids, stamp cwd/parentSession-lineage/depth, drive the one-shot (send → whenIdle), read the last assistant/message + turn/end reason, dispose to quiescence. - dsh-subagent-fork: a child SEEDED with the parent's balanced completed-turn prefix (the log up to and including its last turn/end), so the child inherits context. The in-flight unbalanced turn is excluded — a raw seed would fail the invariants replay. Proven: a regression test goes red if the boundary seeds the open turn. - Seam extension: CreateAgentOptions.seed, threaded through AgentLoop.createAgent → ctx.sessions.prepare({ seed }) (the primitive resume already used). This is the fork-lineage path the TODO(sub-agents) markers anticipated. - Depth: a merge-extensible AgentOptions.subagentDepth (0 top-level, parent+1 for a child); the depthLimit capability refuses a spawn past request.maxDepth. Tests: real-loop unit tests for both backends (mock MODEL only, real loop + invariants), a multi-subagent test (one parent drives a fork AND a spawn child then keeps working), and a with-key e2e (a real parent delegates via the `subagent` tool to a real child that writes a file on disk — world-verified). 100% per-file coverage. The coding-agent demo wires the spawn backend + tool. Snapshot coverage of nested agents is deferred to a stacked follow-up (TODO(subagent-snapshots)): dsh-llm-replay is a single global positional cursor that cannot route calls to a parent vs. a child on one context. Recorded in the RFC's deferrals and a new AGENTS.md rule: designing a subsystem must design its test infrastructure END TO END up front, verifying the snapshot/e2e harness can express the new shape — a gap this plan hit.
2026-06-22 05:58:40 +08:00
## In-process backends: depth and seed
The two in-process backends ([dsh-subagent-spawn](../../packages/subagent/subagent-spawn) fresh, [dsh-subagent-fork](../../packages/subagent/subagent-fork) seeded) run the child as an ordinary `Agent` in the same application. The provider creates it directly through `parent.ctx`, passes the required signal into the core creation transaction, and delegates quiescent disposal to the returned `AgentHandle`. Provider removal prevents new starts but does not revoke an accepted run. The child receives a flat new scope rather than inheriting the parent's registrations. Two pieces of vocabulary ride on the existing agent/session types rather than new core types:
Add in-process subagent backends: spawn (fresh) and fork (seeded) The second PR of the subagent seam: the two in-process backends that run a child agent on the same cordis context, reusing the agent factory's quiescent AgentHandle teardown. Both register on ctx.subagents (PR1's named-provider registry) and share one run driver. - dsh-subagent-spawn: a FRESH child via ctx.agents.create — own session, the parent's model by default (overridable), zero inherited conversation. Also exports the shared in-process run driver (startInProcessRun): mint ids, stamp cwd/parentSession-lineage/depth, drive the one-shot (send → whenIdle), read the last assistant/message + turn/end reason, dispose to quiescence. - dsh-subagent-fork: a child SEEDED with the parent's balanced completed-turn prefix (the log up to and including its last turn/end), so the child inherits context. The in-flight unbalanced turn is excluded — a raw seed would fail the invariants replay. Proven: a regression test goes red if the boundary seeds the open turn. - Seam extension: CreateAgentOptions.seed, threaded through AgentLoop.createAgent → ctx.sessions.prepare({ seed }) (the primitive resume already used). This is the fork-lineage path the TODO(sub-agents) markers anticipated. - Depth: a merge-extensible AgentOptions.subagentDepth (0 top-level, parent+1 for a child); the depthLimit capability refuses a spawn past request.maxDepth. Tests: real-loop unit tests for both backends (mock MODEL only, real loop + invariants), a multi-subagent test (one parent drives a fork AND a spawn child then keeps working), and a with-key e2e (a real parent delegates via the `subagent` tool to a real child that writes a file on disk — world-verified). 100% per-file coverage. The coding-agent demo wires the spawn backend + tool. Snapshot coverage of nested agents is deferred to a stacked follow-up (TODO(subagent-snapshots)): dsh-llm-replay is a single global positional cursor that cannot route calls to a parent vs. a child on one context. Recorded in the RFC's deferrals and a new AGENTS.md rule: designing a subsystem must design its test infrastructure END TO END up front, verifying the snapshot/e2e harness can express the new shape — a gap this plan hit.
2026-06-22 05:58:40 +08:00
- **Delegation depth** is a merge-extensible `AgentOptions.subagentDepth` field (`0` for a top-level agent, parent + 1 for a child). Only `undefined` means top level; every stored present value must be a non-negative safe integer. The seam owns it — the loop neither sets nor reads it — so a nested spawn validates its parent's stored depth, rejects a derived child depth outside the safe-integer domain, and applies a defined absolute `request.maxDepth` cap to that child.
Add in-process subagent backends: spawn (fresh) and fork (seeded) The second PR of the subagent seam: the two in-process backends that run a child agent on the same cordis context, reusing the agent factory's quiescent AgentHandle teardown. Both register on ctx.subagents (PR1's named-provider registry) and share one run driver. - dsh-subagent-spawn: a FRESH child via ctx.agents.create — own session, the parent's model by default (overridable), zero inherited conversation. Also exports the shared in-process run driver (startInProcessRun): mint ids, stamp cwd/parentSession-lineage/depth, drive the one-shot (send → whenIdle), read the last assistant/message + turn/end reason, dispose to quiescence. - dsh-subagent-fork: a child SEEDED with the parent's balanced completed-turn prefix (the log up to and including its last turn/end), so the child inherits context. The in-flight unbalanced turn is excluded — a raw seed would fail the invariants replay. Proven: a regression test goes red if the boundary seeds the open turn. - Seam extension: CreateAgentOptions.seed, threaded through AgentLoop.createAgent → ctx.sessions.prepare({ seed }) (the primitive resume already used). This is the fork-lineage path the TODO(sub-agents) markers anticipated. - Depth: a merge-extensible AgentOptions.subagentDepth (0 top-level, parent+1 for a child); the depthLimit capability refuses a spawn past request.maxDepth. Tests: real-loop unit tests for both backends (mock MODEL only, real loop + invariants), a multi-subagent test (one parent drives a fork AND a spawn child then keeps working), and a with-key e2e (a real parent delegates via the `subagent` tool to a real child that writes a file on disk — world-verified). 100% per-file coverage. The coding-agent demo wires the spawn backend + tool. Snapshot coverage of nested agents is deferred to a stacked follow-up (TODO(subagent-snapshots)): dsh-llm-replay is a single global positional cursor that cannot route calls to a parent vs. a child on one context. Recorded in the RFC's deferrals and a new AGENTS.md rule: designing a subsystem must design its test infrastructure END TO END up front, verifying the snapshot/e2e harness can express the new shape — a gap this plan hit.
2026-06-22 05:58:40 +08:00
- **Fork seeding** uses `CreateAgentOptions.seed` (a `SessionEvent[]` prefix threaded through `AgentLoop.createAgent` → `ctx.sessions.prepare({ seed })`, the same primitive `resume` uses). The fork backend passes a *balanced completed-turn prefix* of the parent's log — the parent's events up to and including its last `turn/end` — so the seed is contiguous-from-0 and the [invariants](../../packages/support/invariants) replay accepts it (the in-flight, unbalanced turn is excluded).