fix(subagent): preserve Claude Code failure facts
This commit is contained in:
parent
1c405d41a5
commit
cd4f8b7f46
24 changed files with 648 additions and 129 deletions
|
|
@ -2,5 +2,5 @@
|
|||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-08-04-claude-code-and-codex-subagent-backends.md
|
||||
2026-08-04-claude-code-and-codex-subagent-backends.md: f65c0626ad22db8f3e7d2a543c7aa87e58df54d4
|
||||
2026-08-04-claude-code-and-codex-subagent-backends.zh.md: 97ac527b8e89cc07d65aa28102ba43d648b1b64c
|
||||
2026-08-04-claude-code-and-codex-subagent-backends.md: 829dca8dbd79b408fcfcfd1d88490d793ad4b5ee
|
||||
2026-08-04-claude-code-and-codex-subagent-backends.zh.md: 063bd8c9a59a1b1eccf4893a00efe12a80fbe03f
|
||||
|
|
|
|||
|
|
@ -12,7 +12,7 @@ The product integrations must not become second owners for task text, cwd, cance
|
|||
|
||||
## Decision
|
||||
|
||||
The harness publishes two sibling one-shot provider packages: `codex` and `claude-code`. This note owns their product protocols, result mapping, and process lifecycle; the [production-install exclusion decision](../simplification/2026-08-12-production-dsh-excludes-product-subagent-providers.md) owns their explicit Profile installation and host-plane placement, the [product one-shot background decision](2026-08-12-product-subagent-one-shot-background-tasks.md) owns the model-visible scheduling choice, and the [non-interactive permissions decision](2026-08-15-product-subagent-noninteractive-permissions.md) owns each product Provider's Profile-selected mode and diagnostic production. Loading either provider starts no product process, and each tool accepts only a standalone text task; product selection remains deployment configuration.
|
||||
The harness publishes two sibling one-shot provider packages: `codex` and `claude-code`. This note owns their product protocols, result mapping, and process lifecycle; the [production-install exclusion decision](../simplification/2026-08-12-production-dsh-excludes-product-subagent-providers.md) owns their explicit Profile installation and host-plane placement, the [product one-shot background decision](2026-08-12-product-subagent-one-shot-background-tasks.md) owns the model-visible scheduling choice, the [non-interactive permissions decision](2026-08-15-product-subagent-noninteractive-permissions.md) owns each product Provider's Profile-selected mode and safe permission decisions, and the [structured failure-facts decision](2026-08-18-product-subagent-failure-facts.md) owns version-pinned product categories, lifecycle stages, and process outcomes exposed through the same diagnostic. Loading either provider starts no product process, and each tool accepts only a standalone text task; product selection remains deployment configuration.
|
||||
|
||||
Both providers report `inheritsParentContext: false`, advertise no optional start capabilities, and pass the parent Session cwd without copying the parent conversation. Their documented tools use `backgroundMode: 'one-shot'` and `maxDepth: 'provider-managed'`: the consumer keeps foreground collection as the default and may place the same run in the generic Job runtime, while recursion policy stays with the out-of-process product. Every call creates a fresh product process and a non-resumable product conversation. `ctx.subagents` owns named-request resolution and paired lifecycle events; `dsh-tool-subagent` owns model-visible scheduling and foreground-versus-Job adaptation; `ctx.jobs` and `dsh-tool-jobs` own Job ids, state, output, controls, notices, and parent-owner cancellation; each product provider owns native result mapping, while `dsh-subprocess` owns credential scrubbing, process-tree termination, and whole-tree exit observation.
|
||||
|
||||
|
|
@ -52,9 +52,9 @@ Codex 0.147.0 speaks the Responses protocol, while DeepSeek's public OpenAI-comp
|
|||
|
||||
The public configuration contains an explicit `env` overlay, a positive finite `disposeGraceMs` no greater than the repository's shared `MAX_TIMER_DELAY_MS`, and a five-value native `permissionMode` that defaults to `dontAsk`. Each run creates its own `AbortController`, sets `persistSession: false`, disables `AskUserQuestion`, and passes the resolved mode to the SDK; only `bypassPermissions` receives the SDK's explicit dangerous confirmation. The provider deliberately omits `settingSources`, so the SDK reads the host's normal user, project, and local Claude settings relative to the parent Session cwd. It neither copies nor filters those settings and does not create or modify login state. Remaining permission prompts are denied, MCP elicitation is declined, and blocking dialogs fail closed instead of waiting for a user interface the provider does not own.
|
||||
|
||||
The provider publishes only after both the SDK `Query` and a live managed CLI handle exist. It consumes the complete SDK stream and completes only when a `result` message has `subtype: "success"`, `is_error: false`, and a nonblank `result`, and the iterator then ends normally. Every SDK error subtype, an error-marked success, a missing result, iterator failure, protocol failure, or process failure becomes `error`. When a permission denial or unattended callback contributes to that failure, the result may additionally carry the bounded, non-assistant diagnostic owned by the non-interactive permissions decision. SDK turn, budget, and structured-output limits are not token-window facts, and the SDK exposes no native refusal terminal, so this provider produces neither `max-tokens` nor `refusal`. Local cancellation wins and becomes `aborted` without permission detail.
|
||||
The provider publishes only after both the SDK `Query` and a live managed CLI handle exist. It consumes the complete SDK stream and completes only when a `result` message has `subtype: "success"`, `is_error: false`, and a nonblank `result`, and the iterator then ends normally. Every other result remains `error`, but its bounded diagnostic preserves the four exact SDK error subtypes, fixed categories for invalid success and missing result, a safe `unknown` fallback, the current `query-start`, `query-run`, `process`, or `teardown` stage, and any observed exit code and signal. A contributing permission decision follows that structured failure line. SDK turn, budget, and structured-output limits are not token-window facts, and the SDK exposes no native refusal terminal, so this provider produces neither `max-tokens` nor `refusal`. Local cancellation wins and becomes `aborted` without either diagnostic fact.
|
||||
|
||||
Startup rollback and published disposal close the SDK query, abort the per-run controller, invoke shared process-tree termination, and wait for whole-tree exit. `Query.close()` expresses graceful protocol intent but does not replace the subprocess owner's exit proof. Query-close failure, process failure, and teardown failure remain independently observable.
|
||||
Startup rollback and published disposal close the SDK query, abort the per-run controller, invoke shared process-tree termination, and wait for whole-tree exit. `Query.close()` expresses graceful protocol intent but does not replace the subprocess owner's exit proof. An unpublished failure exposes only fixed `query-start` facts; a published process failure can expose its independent exit code and signal; an independent cleanup rejection exposes `teardown`. Original SDK, Host, and cleanup errors remain on internal cause chains and logs rather than entering the diagnostic.
|
||||
|
||||
The credentialed Claude Code e2e uses the official DeepSeek Claude Code contract directly: the runtime-only DeepSeek key becomes `ANTHROPIC_AUTH_TOKEN`, the fixed official base gains `/anthropic`, and the main and subagent model variables select the documented DeepSeek models. It starts the production provider and real SDK/CLI, requires one random nonce as the complete answer, persists no credential in settings, and waits for every managed handle to exit.
|
||||
|
||||
|
|
@ -66,7 +66,7 @@ The Codex evidence pins `@openai/codex@0.147.0` and `codex-cli 0.147.0`. Its rea
|
|||
|
||||
The Codex credentialed e2e registers the production provider, starts the same real app-server, and requests one random nonce through the test-private bridge described above. It fixes the external endpoint and model, stores no credential or request payload, requires exactly one completed upstream response, compares the trimmed product answer byte-for-byte with the nonce, and waits for every managed handle to exit.
|
||||
|
||||
The Claude Code evidence pins Agent SDK 0.3.220 and uses its platform-distributed Claude Code 2.1.220 CLI as the deterministic compatibility fixture, routed through the same native executable-resolution path production uses. Its real-product spec observes the exact `x-api-key`, original task, byte-exact final answer, an inherited interactive host setting overridden by the safe Provider mode, denied and bypassed writes in suite-owned temporary directories, safe permission diagnostics, process failure, local cancellation, whole-tree exit, and a real Windows batch shim under a path containing percent, ampersand, and exclamation metacharacters. This evidence proves the official SDK/CLI integration path, not compatibility with every independently installed product version. The Loader and shipped-profile evidence resolve both product packages by name while starting neither product, and the provider suite proves that the SDK receives the executable resolved from the host `PATH`.
|
||||
The Claude Code evidence pins Agent SDK 0.3.220 and uses its platform-distributed Claude Code 2.1.220 CLI as the deterministic compatibility fixture, routed through the same native executable-resolution path production uses. Its real-product spec observes the exact `x-api-key`, original task, byte-exact final answer, an inherited interactive host setting overridden by the safe Provider mode, denied and bypassed writes in suite-owned temporary directories, a real `error_max_turns` result, a process exit with its outcome, safe permission diagnostics, local cancellation, whole-tree exit, and a real Windows batch shim under a path containing percent, ampersand, and exclamation metacharacters. Package tests pin the complete SDK error union, all four stages, unknown fallback, independent code and signal fields, sanitization, success and cancellation omission, and concurrent-run isolation. This evidence proves the official SDK/CLI integration path, not compatibility with every independently installed product version. The Loader and shipped-profile evidence resolve both product packages by name while starting neither product, and the provider suite proves that the SDK receives the executable resolved from the host `PATH`.
|
||||
|
||||
The Claude Code credentialed e2e maps the key and fixed official endpoint only in the provider's in-memory environment, uses the documented `deepseek-v4-pro[1m]` and `deepseek-v4-flash` model variables, and traverses the production provider, official SDK, and real CLI. It compares the trimmed result with a random nonce and proves whole-tree exit without calling the Messages API directly from the test.
|
||||
|
||||
|
|
@ -90,6 +90,6 @@ The project owner's distribution authorization is scoped to the official `@anthr
|
|||
|
||||
Users delegate through two stable one-shot tools backed by the official product integrations. Explicit Profile installation and host-plane provider placement are owned by the [production-install exclusion decision](../simplification/2026-08-12-production-dsh-excludes-product-subagent-providers.md); per-Preset tool exposure and foreground-default optional Job scheduling are owned by the [product one-shot background decision](2026-08-12-product-subagent-one-shot-background-tasks.md). This note's provider lifecycle keeps native settings and behavior while shared services retain the sole ownership of job settlement and process-tree quiescence.
|
||||
|
||||
Every delegation pays for a fresh product process and independent model context. Successful product payload remains final assistant text; a failed product run may separately expose the shared safe diagnostic. Background scheduling additionally exposes generic Job ids, status, completion notices, and collection or cancellation results. Product-native configuration makes behavior depend on the deployment's installed product, account state, workspace settings, and selected Provider mode. Credentialed e2e runs also spend external API quota and depend on the official DeepSeek endpoint; deterministic protocol, failure, cancellation, and approval coverage remains in the keyless tier. The providers do not resume sessions, stream progress, accept new human interaction, roll back tool or file side effects, or impose a wall-clock timeout.
|
||||
Every delegation pays for a fresh product process and independent model context. Successful product payload remains final assistant text; a failed product run may separately expose the shared safe diagnostic containing provider-owned permission facts or version-pinned structured failure facts. Background scheduling additionally exposes generic Job ids, status, completion notices, and collection or cancellation results. Product-native configuration makes behavior depend on the deployment's installed product, account state, workspace settings, and selected Provider mode. Credentialed e2e runs also spend external API quota and depend on the official DeepSeek endpoint; deterministic protocol, failure, cancellation, and approval coverage remains in the keyless tier. The providers do not resume sessions, stream progress, accept new human interaction, roll back tool or file side effects, or impose a wall-clock timeout.
|
||||
|
||||
Compatibility is pinned by package-level unit coverage, keyless real-product loopback tests, credentialed DeepSeek nonce tests, public Loader composition, built-package and NodeNext consumer checks, generated documentation and notices, and the repository CI matrix. A supported product or DeepSeek endpoint/model baseline change must refresh those facts; production performs no separate runtime version probe.
|
||||
|
|
|
|||
|
|
@ -12,7 +12,7 @@ Status: implemented
|
|||
|
||||
## 决策
|
||||
|
||||
harness 交付两个同级的一次性提供方包:`codex` 与 `claude-code`。本说明负责它们的产品协议、结果映射和进程生命周期;[生产安装排除决策](../simplification/2026-08-12-production-dsh-excludes-product-subagent-providers.md)负责显式 Profile 安装与 host plane(宿主平面)放置,[产品一次性后台任务决策](2026-08-12-product-subagent-one-shot-background-tasks.md)负责模型可见的调度选择,[非交互权限决策](2026-08-15-product-subagent-noninteractive-permissions.md)则负责各产品提供方的 Profile 模式选择与诊断生产。加载任一提供方都不会启动产品进程,而且每个工具只接受独立文本任务;产品选择仍属于部署配置。
|
||||
harness 交付两个同级的一次性提供方包:`codex` 与 `claude-code`。本说明负责它们的产品协议、结果映射和进程生命周期;[生产安装排除决策](../simplification/2026-08-12-production-dsh-excludes-product-subagent-providers.md)负责显式 Profile 安装与 host plane(宿主平面)放置,[产品一次性后台任务决策](2026-08-12-product-subagent-one-shot-background-tasks.md)负责模型可见的调度选择,[非交互权限决策](2026-08-15-product-subagent-noninteractive-permissions.md)负责各产品提供方的 Profile 模式选择与安全权限决定,[结构化失败事实决策](2026-08-18-product-subagent-failure-facts.md)则负责通过同一诊断公开锁定产品版本的类别、生命周期阶段与进程结果。加载任一提供方都不会启动产品进程,而且每个工具只接受独立文本任务;产品选择仍属于部署配置。
|
||||
|
||||
这两个提供方都报告 `inheritsParentContext: false`,不声明任何可选的启动能力,并传递父会话 cwd,但不会复制父级对话。文档所示的工具使用 `backgroundMode: 'one-shot'` 与 `maxDepth: 'provider-managed'`:消费方默认在前台收集结果,也可把同一次运行放入通用 Job 运行时,而递归策略仍由进程外产品负责。每次调用都会创建一个全新的产品进程和一次不可续接的产品对话。`ctx.subagents` 负责具名请求解析与成对生命周期事件;`dsh-tool-subagent` 负责模型可见的调度以及前台与 Job 适配;`ctx.jobs` 和 `dsh-tool-jobs` 负责 Job id、状态、输出、控制、通知与父级 owner 取消;各产品提供方负责原生结果映射,`dsh-subprocess` 则负责凭证清洗、进程树终止以及整棵进程树的退出观测。
|
||||
|
||||
|
|
@ -52,9 +52,9 @@ Codex 0.147.0 使用 Responses 协议,而 DeepSeek 的公开 OpenAI 兼容端
|
|||
|
||||
公开配置包含显式的 `env` 覆盖项、须为正有限值且不得大于仓库共享 `MAX_TIMER_DELAY_MS` 的 `disposeGraceMs`,以及默认使用 `dontAsk` 的五值原生 `permissionMode`。每次运行都会创建自己的 `AbortController`,设置 `persistSession: false`、禁用 `AskUserQuestion`,并把已解析模式传给 SDK;只有 `bypassPermissions` 会取得 SDK 的显式危险确认。提供方故意省略 `settingSources`,因此 SDK 会相对于父会话 cwd 读取宿主机常规的用户、项目和本地 Claude 设置。它既不复制也不过滤这些设置,也不会创建或修改登录状态。其余权限提示会被拒绝,MCP elicitation 会被拒绝,阻塞对话会快速失败,而不会等待本提供方不负责的用户界面。
|
||||
|
||||
只有在 SDK `Query` 与受管的活动 CLI 句柄都已存在后,提供方才会发布运行。它会消费完整的 SDK 流;只有 `result` 消息具有 `subtype: "success"`、`is_error: false` 和非空白 `result`,且迭代器随后正常结束时,运行才会完成。所有 SDK 错误子类型、标记为错误的成功消息、结果缺失、迭代器失败、协议失败或进程失败都会成为 `error`。当权限拒绝或无人值守回调参与了该失败时,结果还可以携带由非交互权限决策负责的有界、非 assistant 诊断。SDK 的轮次、预算和结构化输出限制不表示 token 窗口耗尽,而且 SDK 没有原生的拒绝终止状态,因此本提供方不会产生 `max-tokens` 或 `refusal`。本地取消会胜出并成为 `aborted`,且不附带权限说明。
|
||||
只有在 SDK `Query` 与受管的活动 CLI 句柄都已存在后,提供方才会发布运行。它会消费完整的 SDK 流;只有 `result` 消息具有 `subtype: "success"`、`is_error: false` 和非空白 `result`,且迭代器随后正常结束时,运行才会完成。其他所有结果仍成为 `error`,但其有界诊断会保留四种准确 SDK 错误子类型、标记为错误的成功消息与结果缺失所对应的固定类别、安全的 `unknown` 回退、当前 `query-start`、`query-run`、`process` 或 `teardown` 阶段,以及已观测到的退出码和信号。若权限决定也参与失败,它会跟在结构化失败行之后。SDK 的轮次、预算和结构化输出限制不表示 token 窗口耗尽,而且 SDK 没有原生的拒绝终止状态,因此本提供方不会产生 `max-tokens` 或 `refusal`。本地取消会胜出并成为 `aborted`,且不附带这两类诊断事实。
|
||||
|
||||
启动回滚和已发布运行的资源释放都会关闭 SDK query、中止该次运行的控制器、调用共享的进程树终止机制,并等待整棵进程树退出。`Query.close()` 表达优雅的协议关闭意图,但不能取代子进程责任方的退出证明。Query 关闭失败、进程失败和清理失败仍可彼此独立地观察。
|
||||
启动回滚和已发布运行的资源释放都会关闭 SDK query、中止该次运行的控制器、调用共享的进程树终止机制,并等待整棵进程树退出。`Query.close()` 表达优雅的协议关闭意图,但不能取代子进程责任方的退出证明。未发布失败只公开固定的 `query-start` 事实;已发布进程失败可以分别公开退出码与信号;独立清理拒绝则公开 `teardown`。原始 SDK、Host 与清理错误只保留在内部 cause 链和日志中,不进入诊断。
|
||||
|
||||
带密钥 Claude Code e2e 直接使用官方 DeepSeek Claude Code 约定:仅在运行时提供的 DeepSeek 密钥会映射为 `ANTHROPIC_AUTH_TOKEN`,固定的官方基础 URL 会追加 `/anthropic`,主模型与 subagent 模型变量会选择文档所示的 DeepSeek 模型。该测试会启动生产提供方与真实 SDK 和 CLI,要求一个随机数作为完整答案,不会把任何凭据持久化到设置中,并等待所有受管句柄退出。
|
||||
|
||||
|
|
@ -66,7 +66,7 @@ Codex 证据锁定 `@openai/codex@0.147.0` 与 `codex-cli 0.147.0`。其真实
|
|||
|
||||
带密钥 Codex e2e 会注册生产提供方,启动同样的真实 app-server,并通过上述测试专用桥接层请求一个随机数。该测试固定外部端点与模型,不存储任何凭据或请求载荷,要求上游恰好完成一次响应,将去除首尾空白后的产品答案与该随机数逐字节比较,并等待所有受管句柄退出。
|
||||
|
||||
Claude Code 证据锁定 Agent SDK 0.3.220,并使用 SDK 按平台分发的 Claude Code 2.1.220 CLI 作为确定性兼容性 fixture(测试前置数据),且该 fixture 经生产环境所用的同一原生可执行文件解析路径运行。其真实产品测试会观测确切的 `x-api-key`、原始任务、逐字节完全一致的最终回答、安全提供方模式对继承的交互式宿主设置的覆盖、测试所拥有临时目录中的拒绝写入与 bypass 写入、安全权限诊断、进程失败、本地取消、整棵进程树退出,以及位于同时含百分号、与号和感叹号路径中的真实 Windows batch shim。这项证据证明官方 SDK/CLI 集成路径,而不证明它与每个独立安装的产品版本兼容。Loader 与随附 profile 证据会按名称解析两个产品包且不启动产品,provider 测试则证明 SDK 收到由宿主 `PATH` 解析出的可执行文件。
|
||||
Claude Code 证据锁定 Agent SDK 0.3.220,并使用 SDK 按平台分发的 Claude Code 2.1.220 CLI 作为确定性兼容性 fixture(测试前置数据),且该 fixture 经生产环境所用的同一原生可执行文件解析路径运行。其真实产品测试会观测确切的 `x-api-key`、原始任务、逐字节完全一致的最终回答、安全提供方模式对继承的交互式宿主设置的覆盖、测试所拥有临时目录中的拒绝写入与 bypass 写入、真实的 `error_max_turns` 结果、携带进程结果的提前退出、安全权限诊断、本地取消、整棵进程树退出,以及位于同时含百分号、与号和感叹号路径中的真实 Windows batch shim。包测试固定完整 SDK 错误联合、四个阶段、unknown 回退、相互独立的退出码与信号字段、脱敏、成功与取消时省略诊断,以及并发运行隔离。这项证据证明官方 SDK/CLI 集成路径,而不证明它与每个独立安装的产品版本兼容。Loader 与随附 profile 证据会按名称解析两个产品包且不启动产品,provider 测试则证明 SDK 收到由宿主 `PATH` 解析出的可执行文件。
|
||||
|
||||
带密钥 Claude Code e2e 仅在提供方的内存环境中映射密钥与固定的官方端点,把模型变量设为文档所示的 `deepseek-v4-pro[1m]` 与 `deepseek-v4-flash`,并实际经过生产提供方、官方 SDK 与真实 CLI。它将去除首尾空白后的结果与一个随机数比较,并证明整棵进程树退出,且测试不会直接调用 Messages API。
|
||||
|
||||
|
|
@ -90,6 +90,6 @@ Claude Code 证据锁定 Agent SDK 0.3.220,并使用 SDK 按平台分发的 Cl
|
|||
|
||||
用户通过官方产品集成支持的两个稳定一次性工具进行委派。显式 Profile 安装与 host plane 提供方放置由[生产安装排除决策](../simplification/2026-08-12-production-dsh-excludes-product-subagent-providers.md)负责;按 Preset 暴露工具以及默认前台且可选通用 Job 的调度方式由[产品一次性后台任务决策](2026-08-12-product-subagent-one-shot-background-tasks.md)负责。本说明规定的提供方生命周期会保留原生设置与行为,而共享服务继续独占作业结算与进程树完全停稳的责任。
|
||||
|
||||
每次委派都要承担新建产品进程和独立模型上下文的开销。成功的产品载荷仍只有最终 assistant 文本;失败的产品运行可以另行公开共享安全诊断。后台调度还会额外公开通用 Job id、状态、完成通知以及收集或取消结果。产品原生配置使行为取决于部署环境中安装的产品、账户状态、工作区设置和所选提供方模式。带密钥 e2e 运行还会消耗外部 API 配额,并依赖 DeepSeek 官方端点;对协议、失败、取消与审批的确定性覆盖仍由无密钥层级承担。提供方不会恢复会话、以流式方式传送进度、接受新的人工交互、回滚工具或文件副作用,也不会施加按实际经过时间触发的超时。
|
||||
每次委派都要承担新建产品进程和独立模型上下文的开销。成功的产品载荷仍只有最终 assistant 文本;失败的产品运行可以另行公开共享安全诊断,其中包含由提供方拥有的权限事实,或锁定版本产品提供的结构化失败事实。后台调度还会额外公开通用 Job id、状态、完成通知以及收集或取消结果。产品原生配置使行为取决于部署环境中安装的产品、账户状态、工作区设置和所选提供方模式。带密钥 e2e 运行还会消耗外部 API 配额,并依赖 DeepSeek 官方端点;对协议、失败、取消与审批的确定性覆盖仍由无密钥层级承担。提供方不会恢复会话、以流式方式传送进度、接受新的人工交互、回滚工具或文件副作用,也不会施加按实际经过时间触发的超时。
|
||||
|
||||
兼容性由包级单元测试覆盖率、无密钥真实产品回环测试、带密钥 DeepSeek 随机数测试、公开 Loader 组合、已构建包与 NodeNext 消费方检查、生成的文档与声明以及仓库 CI 矩阵共同锁定。更改受支持的产品基线或 DeepSeek 端点/模型基线时必须刷新这些事实;生产环境不会另行执行运行时版本探测。
|
||||
|
|
|
|||
|
|
@ -2,5 +2,5 @@
|
|||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-08-15-product-subagent-noninteractive-permissions.md
|
||||
2026-08-15-product-subagent-noninteractive-permissions.md: df1f0d9939e951f16070729615a3779f1f7c2ddc
|
||||
2026-08-15-product-subagent-noninteractive-permissions.zh.md: 982b4409e08a506dec828db15c8c4aa5fcc36883
|
||||
2026-08-15-product-subagent-noninteractive-permissions.md: 9327412cfdd306f7867f989c8cfc091941cb26e6
|
||||
2026-08-15-product-subagent-noninteractive-permissions.zh.md: a7992b58a14aff94f93397bfa0aa21b9fe727fb2
|
||||
|
|
|
|||
|
|
@ -44,9 +44,9 @@ The Provider overrides only those thread fields. `CODEX_HOME`, project configura
|
|||
|
||||
### Failure diagnostic
|
||||
|
||||
`SubagentResult` carries an optional `diagnostic` for provider-authored, non-assistant failure detail. A Provider removes tool inputs, file contents, environment values, credentials, and raw protocol payloads before producing it. The shared out-of-process result boundary limits the complete text to 4096 UTF-8 bytes and marks truncation without splitting a character.
|
||||
`SubagentResult` carries an optional `diagnostic` for provider-authored, non-assistant failure detail. A Provider removes tool inputs, file contents, environment values, credentials, and raw protocol payloads before producing it. The shared out-of-process result boundary limits the complete text to 4096 UTF-8 bytes and marks truncation without splitting a character. The [structured failure-facts decision](2026-08-18-product-subagent-failure-facts.md) owns non-permission product categories, lifecycle stages, and process outcomes carried by the same field.
|
||||
|
||||
Each product records only the effective mode, request category, unattended decision, and a fixed safe reason. Claude Code derives those facts from SDK callbacks and `permission_denied` messages. Codex derives them from app-server requests, declined items, `sandboxError`, and two fixed permission signatures in a bounded stderr tail; raw stderr is still forwarded to the Host but never copied into the diagnostic. A successful result returns only the strict final answer; local cancellation remains `aborted` without permission detail; an unpublished startup failure still rejects `start()`. When a permission fact contributes to a published run that settles as `error`, the Provider attaches the diagnostic without adding it to assistant output, structured output, or `subagent/end.lastAssistantMessage`.
|
||||
Each product's permission fact contains only the effective mode, request category, unattended decision, and a fixed safe reason. Claude Code derives those facts from SDK callbacks and `permission_denied` messages. Codex derives them from app-server requests, declined items, `sandboxError`, and two fixed permission signatures in a bounded stderr tail; raw stderr is still forwarded to the Host but never copied into the diagnostic. Claude Code places its structured failure line before the latest contributing permission fact; Codex retains its permission-only diagnostic in this product version. A successful result returns only the strict final answer; local cancellation remains `aborted` without permission detail; an unpublished startup failure still rejects `start()`. The Provider never adds either diagnostic fact to assistant output, structured output, or `subagent/end.lastAssistantMessage`.
|
||||
|
||||
The foreground consumer presents the stop-reason headline, then the optional diagnostic, then any partial assistant output. The one-shot background adapter stores the same diagnostic beside the stop reason in the failed Job detail. Providers that omit the field retain their previous behavior.
|
||||
|
||||
|
|
@ -63,7 +63,7 @@ The foreground consumer presents the stop-reason headline, then the optional dia
|
|||
|
||||
## Verification
|
||||
|
||||
Package tests pin every allowed and rejected Config value, the exact SDK and app-server field mappings, dangerous confirmations, unattended terminal responses, diagnostic sanitization and UTF-8 bound, successful-result omission, concurrent-run isolation, foreground ordering, Job detail, stderr observer disposal, and process cleanup. The real Claude Agent SDK/CLI fixture proves its safe default, restricted denial, explicit bypass, and whole-tree quiescence. The real Codex app-server fixture proves that thread-level `never` overrides ambient `on-request`, automatic review starts, dangerous bypass writes only inside suite-owned temporary storage, fixed stderr signatures produce safe diagnostics, and the wrapper/native tree exits. Loader composition proves non-default modes can be published without starting either product, and keyless ACP snapshots record the shared diagnostic presentation while the model-facing product tool schemas contain no permission parameter.
|
||||
Package tests pin every allowed and rejected Config value, the exact SDK and app-server field mappings, dangerous confirmations, unattended terminal responses, diagnostic sanitization and UTF-8 bound, successful-result omission, concurrent-run isolation, foreground ordering, Job detail, stderr observer disposal, and process cleanup. The real Claude Agent SDK/CLI fixture proves its safe default, restricted denial, explicit bypass, and whole-tree quiescence. The real Codex app-server fixture proves that thread-level `never` overrides ambient `on-request`, automatic review starts, dangerous bypass writes only inside suite-owned temporary storage, fixed stderr signatures produce safe diagnostics, and the wrapper/native tree exits. Loader composition proves non-default modes can be published without starting either product, and keyless ACP snapshots record the shared foreground and Job diagnostic presentation while the model-facing product tool schemas contain no permission parameter.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
|
|
@ -83,6 +83,6 @@ Package tests pin every allowed and rejected Config value, the exact SDK and app
|
|||
|
||||
Profiles can select each product's native restricted, automatic, planning/edit-accepting where supported, or bypass behavior before the Provider starts, while both safe defaults never ask a person. Broader modes remain explicit deployment choices and retain their native sandbox consequences.
|
||||
|
||||
Permission failures become visible to both foreground parents and one-shot background Jobs without turning infrastructure text into an assistant answer. That diagnostic can enter model context, Job notices, API projections, and Job UI through the ordinary consumer paths, so the Provider must sanitize and bound it before result settlement.
|
||||
Permission failures become visible to both foreground parents and one-shot background Jobs without turning infrastructure text into an assistant answer. The same field can also carry the separately owned structured failure facts. It can enter model context, Job notices, API projections, and Job UI through the ordinary consumer paths, so the Provider must sanitize and bound the complete text before result settlement.
|
||||
|
||||
The change adds no product session persistence, human approval channel, dynamic permission operation, progress stream, retry policy, or rollback. Other Providers remain valid without producing a diagnostic or exposing a permission-mode Config.
|
||||
|
|
|
|||
|
|
@ -44,9 +44,9 @@ Codex 默认使用 `never`,并接受 Codex 0.147.0 公开的三种原生非交
|
|||
|
||||
### 失败诊断
|
||||
|
||||
`SubagentResult` 携带可选的 `diagnostic`,用于提供方产生且不属于 assistant 内容的失败说明。提供方在生成它之前会排除工具输入、文件内容、环境值、凭证与原始协议载荷。共享的进程外结果边界会把完整文本限制在 4096 个 UTF-8 字节以内,并在不切断字符的前提下标记截断。
|
||||
`SubagentResult` 携带可选的 `diagnostic`,用于提供方产生且不属于 assistant 内容的失败说明。提供方在生成它之前会排除工具输入、文件内容、环境值、凭证与原始协议载荷。共享的进程外结果边界会把完整文本限制在 4096 个 UTF-8 字节以内,并在不切断字符的前提下标记截断。[结构化失败事实决策](2026-08-18-product-subagent-failure-facts.md)负责由同一字段承载的非权限产品类别、生命周期阶段与进程结果。
|
||||
|
||||
每个产品都只记录有效模式、请求类别、无人值守决定与固定的安全原因。Claude Code 从 SDK 回调和 `permission_denied` 消息取得这些事实。Codex 从 app-server 请求、被拒绝的 item、`sandboxError` 与每次运行有界 stderr 尾部中的两个固定权限签名取得事实;原始 stderr 仍会转发给 Host,但绝不会复制进诊断。成功结果只返回严格的最终答案;本地取消仍以 `aborted` 结算且不附带权限说明;未发布的启动失败仍会拒绝 `start()`。当一项权限事实参与了已经发布、最终以 `error` 结算的运行时,提供方会附加诊断,但不会把它写入 assistant 输出、结构化输出或 `subagent/end.lastAssistantMessage`。
|
||||
每个产品的权限事实都只包含有效模式、请求类别、无人值守决定与固定的安全原因。Claude Code 从 SDK 回调和 `permission_denied` 消息取得这些事实。Codex 从 app-server 请求、被拒绝的 item、`sandboxError` 与每次运行有界 stderr 尾部中的两个固定权限签名取得事实;原始 stderr 仍会转发给 Host,但绝不会复制进诊断。Claude Code 会把结构化失败行放在最新参与失败的权限事实之前;当前产品版本中的 Codex 仍只生成权限诊断。成功结果只返回严格的最终答案;本地取消仍以 `aborted` 结算且不附带权限说明;未发布的启动失败仍会拒绝 `start()`。提供方绝不会把任一诊断事实写入 assistant 输出、结构化输出或 `subagent/end.lastAssistantMessage`。
|
||||
|
||||
前台消费方依次呈现终止原因标题、可选诊断和任何部分 assistant 输出。一次性后台适配器会在失败 Job 的 detail 中,把同一诊断与终止原因一起保存。没有填写该字段的提供方保持原有行为。
|
||||
|
||||
|
|
@ -63,7 +63,7 @@ Codex 默认使用 `never`,并接受 Codex 0.147.0 公开的三种原生非交
|
|||
|
||||
## Verification
|
||||
|
||||
包测试固定所有允许与拒绝的 Config 值、准确的 SDK 与 app-server 字段映射、危险确认、无人值守终态、诊断脱敏与 UTF-8 上限、成功结果不携带诊断、并发运行隔离、前台顺序、Job detail、stderr observer 释放和进程清理。真实 Claude Agent SDK/CLI fixture 证明其安全默认、受限拒绝、显式 bypass 与整棵进程树完全停稳。真实 Codex app-server fixture 证明线程级 `never` 覆盖环境中的 `on-request`、自动评审可以启动、危险绕过只在测试拥有的临时存储中写入、固定 stderr 签名产生安全诊断,而且 wrapper/native 进程树会退出。Loader 组装证明非默认模式可以在不启动任一产品的情况下发布;无密钥 ACP snapshot 则记录共享诊断呈现,同时面向模型的产品工具 schema 不包含权限参数。
|
||||
包测试固定所有允许与拒绝的 Config 值、准确的 SDK 与 app-server 字段映射、危险确认、无人值守终态、诊断脱敏与 UTF-8 上限、成功结果不携带诊断、并发运行隔离、前台顺序、Job detail、stderr observer 释放和进程清理。真实 Claude Agent SDK/CLI fixture 证明其安全默认、受限拒绝、显式 bypass 与整棵进程树完全停稳。真实 Codex app-server fixture 证明线程级 `never` 覆盖环境中的 `on-request`、自动评审可以启动、危险绕过只在测试拥有的临时存储中写入、固定 stderr 签名产生安全诊断,而且 wrapper/native 进程树会退出。Loader 组装证明非默认模式可以在不启动任一产品的情况下发布;无密钥 ACP snapshot 则记录前台与 Job 共享的诊断呈现,同时面向模型的产品工具 schema 不包含权限参数。
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
|
|
@ -83,6 +83,6 @@ Codex 默认使用 `never`,并接受 Codex 0.147.0 公开的三种原生非交
|
|||
|
||||
Profile 可以在提供方启动前选择各产品原生的受限、自动、在产品支持时仅规划/编辑放行,或 bypass 行为,而两个安全默认值都绝不会询问人员。更宽松的模式仍是显式部署选择,并保留其原生沙箱后果。
|
||||
|
||||
权限失败会同时到达前台父 agent 和一次性后台 Job,且不会把基础设施文本伪装成 assistant 回答。该诊断可以沿普通消费路径进入模型上下文、Job 通知、API 投影与 Job UI,因此提供方必须在结果结算前完成脱敏和限长。
|
||||
权限失败会同时到达前台父 agent 和一次性后台 Job,且不会把基础设施文本伪装成 assistant 回答。同一字段还可以承载由另一项决策负责的结构化失败事实。它可以沿普通消费路径进入模型上下文、Job 通知、API 投影与 Job UI,因此提供方必须在结果结算前对完整文本完成脱敏和限长。
|
||||
|
||||
本改动不增加产品会话持久化、人工审批通道、动态权限操作、进度流、重试策略或回滚。其他提供方无需产生诊断或公开权限模式 Config,仍然保持合法。
|
||||
|
|
|
|||
|
|
@ -0,0 +1,6 @@
|
|||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-08-18-product-subagent-failure-facts.md
|
||||
2026-08-18-product-subagent-failure-facts.md: 4380e36d172f395692d2d84c97d45cba95701f9f
|
||||
2026-08-18-product-subagent-failure-facts.zh.md: d601becdf14bd74ae871a66d4798ffe9c49b6490
|
||||
|
|
@ -0,0 +1,74 @@
|
|||
# Agent Note: Product subagents expose bounded structured failure facts
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-08-18-product-subagent-failure-facts.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The [Claude Code and Codex product providers](2026-08-04-claude-code-and-codex-subagent-backends.md) receive structured product failures, but a published run historically flattened most of them to the shared `error` stop reason. Product logs retained detail that the foreground parent and a [one-shot background Job](2026-08-12-product-subagent-one-shot-background-tasks.md) could not use to distinguish a product limit, an execution failure, or an early process exit.
|
||||
|
||||
Copying SDK error text, app-server payloads, or stderr into the result would expose task text, paths, environment values, credentials, or product internals. Adding shared error fields would also make the provider-neutral [subagent seam](2026-06-21-subagent-capability-seam.md) own product version vocabularies that change independently.
|
||||
|
||||
## Decision
|
||||
|
||||
Each product Provider owns the mapping from its pinned official error union, current operation, and managed process outcome to one fixed safe diagnostic line. `SubagentResult` remains unchanged: consumers receive the existing bounded `diagnostic` string and do not parse its product-private fields.
|
||||
|
||||
### Safe diagnostic
|
||||
|
||||
The structured line has this fixed order:
|
||||
|
||||
```text
|
||||
Product subagent failure (product: <product>; stage: <stage>; category: <category>; exit code: <code>; signal: <signal>)
|
||||
```
|
||||
|
||||
The Provider omits unavailable exit fields. Exit code and signal are independent facts and are each retained when observed. A contributing permission decision from the [non-interactive permissions decision](2026-08-15-product-subagent-noninteractive-permissions.md) follows the structured line; the latest safe permission fact remains operation-local. The shared result boundary limits the complete text to 4096 UTF-8 bytes.
|
||||
|
||||
Successful results and local cancellation expose no failure fact. Raw product errors, stderr, tool input, paths, environment values, credentials, and protocol payloads never enter the diagnostic. Startup and cleanup rejections use the same safe line in their Error message while retaining the original failure only on the internal cause chain and in Host logging.
|
||||
|
||||
### Claude Code facts
|
||||
|
||||
Agent SDK 0.3.220 defines four error subtypes: `error_during_execution`, `error_max_turns`, `error_max_budget_usd`, and `error_max_structured_output_retries`. The Claude Code Provider preserves each exact subtype as the category while keeping the shared stop reason `error`. An error-marked or blank success uses `invalid-success`, a missing result uses `missing-result`, a process exit before an SDK terminal result uses `process-exit`, and an unrecognized value or exception uses `unknown` without copying the value.
|
||||
|
||||
| Stage | Owned operation | Observable failure |
|
||||
| --- | --- | --- |
|
||||
| `query-start` | Native executable resolution, SDK query construction, and unpublished rollback | `start()` rejects with fixed safe facts and any process outcome observed before rollback |
|
||||
| `query-run` | Published SDK message iteration and strict terminal-result validation | The run resolves as `error` with the exact known subtype or a fixed result category |
|
||||
| `process` | Managed CLI exits before the SDK supplies a terminal result | The run resolves as `error` with `process-exit` and the available exit code and signal |
|
||||
| `teardown` | Query close and managed process-tree release | `dispose()` rejects independently with fixed safe facts after cleanup still reaches its final exit wait |
|
||||
|
||||
The Codex Provider retains its existing result mapping: `contextWindowExceeded` is `max-tokens`, other turn failures remain `error`, and permission-related paths may carry their existing safe diagnostic. Other Codex error-info members are not represented as shared categories by this decision's current implementation.
|
||||
|
||||
### Ownership and lifecycle
|
||||
|
||||
| Fact or resource | Owner | Consumer behavior |
|
||||
| --- | --- | --- |
|
||||
| Product error category | Pinned official SDK or app-server version | The Provider maps only the declared structured union and uses `unknown` outside it |
|
||||
| Current failure stage | Product Provider operation | Derived at the failure site; never persisted or used as a recovery state |
|
||||
| Exit code and signal | `dsh-subprocess` process handle | The Provider displays observed values without inferring missing ones |
|
||||
| Diagnostic bytes and delivery | `dsh-subagent`, foreground tool, and Job runtime | The same bounded text is presented separately from assistant output in both scheduling modes |
|
||||
| Raw product failure | Product runtime and Host log | It remains internal and never becomes model-visible result text |
|
||||
|
||||
## Verification
|
||||
|
||||
Claude Code package tests pin all four SDK subtypes, invalid success, missing result, unknown values and exceptions, all four stages, independent exit code and signal fields, permission-fact ordering, sanitization, successful-result and cancellation omission, concurrent-run isolation, and cleanup completion. The real SDK/CLI fixture produces an actual `error_max_turns` result and an actual early process exit while proving whole-tree quiescence. The keyless ACP snapshot records the same failure diagnostic in foreground error output, the background completion notice, and `job_output`.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Return raw SDK errors, app-server payloads, or stderr.** These values can contain commands, paths, workspace content, environment values, credentials, or upstream prose. A fixed allowlisted mapping preserves actionable facts without expanding the model-visible trust boundary.
|
||||
|
||||
**Add a shared product-error enum or structured result fields.** Claude Code and Codex version their error unions independently. A shared enum would duplicate those authorities and force unrelated Providers and consumers to track product releases.
|
||||
|
||||
**Parse generic stderr and exception messages.** Free-form text is neither stable nor safe. Only pinned structured product fields and the managed process outcome qualify as diagnostic input.
|
||||
|
||||
**Persist stages or add a recovery controller.** The stage is derived from the current call site only when a failure is reported. Persistence, retries, resume, and remediation need separate ownership and user contracts.
|
||||
|
||||
**Map product limits to new shared stop reasons.** Claude Code turn and budget limits are not token-window exhaustion, and an error category does not establish refusal semantics. Existing stop reasons remain unchanged.
|
||||
|
||||
## Consequences
|
||||
|
||||
The parent can distinguish important Claude Code product limits, invalid terminal results, unknown query failures, and early process exits without receiving raw product text. Foreground and background scheduling preserve the same fact because both consume one `SubagentResult`.
|
||||
|
||||
The diagnostic is display text rather than a new public protocol. Callers may present it but must not branch on its punctuation or product-private category names. A pinned product-version upgrade must update the Provider mapping and evidence when its official error union changes.
|
||||
|
||||
This decision adds no product session persistence, retry policy, recovery state, stderr classifier, authentication or configuration taxonomy, progress stream, or human interaction path.
|
||||
|
|
@ -0,0 +1,74 @@
|
|||
# Agent Note: 产品 subagent 公开有界结构化失败事实
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-08-18-product-subagent-failure-facts.md) | 中文
|
||||
|
||||
## Problem
|
||||
|
||||
[Claude Code 与 Codex 产品提供方](2026-08-04-claude-code-and-codex-subagent-backends.md)会收到结构化产品失败,但已发布运行以往会把其中大多数压成共享的 `error` 终止原因。产品日志保留了细节,前台父 agent 与[一次性后台 Job](2026-08-12-product-subagent-one-shot-background-tasks.md)却无法据此区分产品限制、执行失败或进程提前退出。
|
||||
|
||||
若把 SDK 错误文本、app-server payload 或 stderr 复制进结果,就会暴露任务文本、路径、环境值、凭证或产品内部信息。若增加共享错误字段,又会让提供方无关的 [subagent seam](2026-06-21-subagent-capability-seam.md)拥有彼此独立变化的产品版本词汇。
|
||||
|
||||
## Decision
|
||||
|
||||
每个产品提供方分别拥有从锁定版本官方错误联合、当前操作和受管进程结果到一行固定安全诊断的映射。`SubagentResult` 保持不变:消费方仍接收现有的有界 `diagnostic` 字符串,而且不解析其中由产品私有的字段。
|
||||
|
||||
### 安全诊断
|
||||
|
||||
结构化行采用以下固定顺序:
|
||||
|
||||
```text
|
||||
Product subagent failure (product: <product>; stage: <stage>; category: <category>; exit code: <code>; signal: <signal>)
|
||||
```
|
||||
|
||||
提供方会省略不可用的退出字段。退出码与信号是相互独立的事实,只要已观测到就分别保留。来自[非交互权限决策](2026-08-15-product-subagent-noninteractive-permissions.md)且参与失败的权限决定会跟在结构化行之后;最新的安全权限事实仍只属于当前操作。共享结果边界会把完整文本限制在 4096 个 UTF-8 字节以内。
|
||||
|
||||
成功结果与本地取消都不公开失败事实。原始产品错误、stderr、工具输入、路径、环境值、凭证和协议 payload 绝不会进入诊断。启动与清理拒绝会在 Error 消息中使用同一安全行,而原始失败只保留在内部 cause 链与 Host 日志中。
|
||||
|
||||
### Claude Code 事实
|
||||
|
||||
Agent SDK 0.3.220 定义四种错误子类型:`error_during_execution`、`error_max_turns`、`error_max_budget_usd` 和 `error_max_structured_output_retries`。Claude Code 提供方会把每种准确子类型保留为类别,同时维持共享终止原因 `error`。标记为错误或内容空白的成功消息使用 `invalid-success`,缺失结果使用 `missing-result`,SDK 给出终态结果前发生的进程退出使用 `process-exit`,无法识别的值或异常使用 `unknown`,且不会复制原值。
|
||||
|
||||
| 阶段 | 归属操作 | 可观察失败 |
|
||||
| --- | --- | --- |
|
||||
| `query-start` | 原生可执行文件解析、SDK query 构造与未发布回滚 | `start()` 以固定安全事实和回滚前已观测到的进程结果拒绝 |
|
||||
| `query-run` | 已发布 SDK 消息迭代与严格终态结果校验 | 运行以 `error` 兑现,并携带准确已知子类型或固定结果类别 |
|
||||
| `process` | SDK 提供终态结果之前受管 CLI 已退出 | 运行以 `error` 兑现,并携带 `process-exit` 以及可用的退出码和信号 |
|
||||
| `teardown` | Query 关闭与受管进程树释放 | `dispose()` 独立拒绝并携带固定安全事实,同时清理仍会完成最终退出等待 |
|
||||
|
||||
Codex 提供方保留既有结果映射:`contextWindowExceeded` 是 `max-tokens`,其他轮次失败仍是 `error`,权限相关路径可以携带既有安全诊断。本决策的当前实现不会把其他 Codex error-info 成员表示为共享类别。
|
||||
|
||||
### 所有权与生命周期
|
||||
|
||||
| 事实或资源 | Owner | 消费方行为 |
|
||||
| --- | --- | --- |
|
||||
| 产品错误类别 | 锁定版本的官方 SDK 或 app-server | 提供方只映射已声明的结构化联合,并对联合外值使用 `unknown` |
|
||||
| 当前失败阶段 | 产品提供方操作 | 只在失败点派生;绝不持久化,也不作为恢复状态 |
|
||||
| 退出码与信号 | `dsh-subprocess` 进程句柄 | 提供方展示已观测值,不推测缺失值 |
|
||||
| 诊断字节与送达 | `dsh-subagent`、前台工具与 Job 运行时 | 两种调度模式都把同一份有界文本与 assistant 输出分开呈现 |
|
||||
| 原始产品失败 | 产品运行时与 Host 日志 | 只保留在内部,绝不成为模型可见的结果文本 |
|
||||
|
||||
## Verification
|
||||
|
||||
Claude Code 包测试固定四种 SDK 子类型、无效成功、缺失结果、未知值与异常、四个阶段、相互独立的退出码与信号字段、权限事实顺序、脱敏、成功结果与取消时省略诊断、并发运行隔离和清理完成。真实 SDK/CLI fixture 会产生真实的 `error_max_turns` 结果与真实的进程提前退出,并证明整棵进程树完全停稳。无密钥 ACP snapshot 会在前台错误输出、后台完成通知和 `job_output` 中记录同一份失败诊断。
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**返回原始 SDK 错误、app-server payload 或 stderr。** 这些值可能包含命令、路径、工作区内容、环境值、凭证或上游文本。固定白名单映射可以保留可操作事实,同时不扩大模型可见的信任边界。
|
||||
|
||||
**增加共享产品错误 enum 或结构化结果字段。** Claude Code 与 Codex 各自独立版本化错误联合。共享 enum 会复制这些权威,并迫使无关提供方和消费方跟随产品版本。
|
||||
|
||||
**解析通用 stderr 与异常消息。** 自由文本既不稳定也不安全。只有锁定版本产品提供的结构化字段和受管进程结果可以成为诊断输入。
|
||||
|
||||
**持久化阶段或增加恢复控制器。** 阶段只在报告失败时从当前调用点派生。持久化、重试、resume 与修复需要独立的所有权和用户约定。
|
||||
|
||||
**把产品限制映射为新的共享终止原因。** Claude Code 的轮次和预算限制并不表示 token 窗口耗尽,错误类别也不能证明拒绝语义。既有终止原因保持不变。
|
||||
|
||||
## Consequences
|
||||
|
||||
父 agent 可以区分重要的 Claude Code 产品限制、无效终态结果、未知 query 失败和进程提前退出,而不会收到原始产品文本。前台与后台调度会保留同一事实,因为二者都消费同一个 `SubagentResult`。
|
||||
|
||||
诊断只是展示文本,不是新的公开协议。调用方可以呈现它,但不得根据其标点或产品私有类别名称进行分支。锁定产品版本升级并改变官方错误联合时,必须同步更新提供方映射与证据。
|
||||
|
||||
本决策不增加产品会话持久化、重试策略、恢复状态、stderr 分类器、身份验证或配置分类体系、进度流或人工交互路径。
|
||||
|
|
@ -2,5 +2,5 @@
|
|||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write docs/config-catalog.md
|
||||
config-catalog.md: ebef0167d5ecbe0c71d201cd2cc07962ba89c48d
|
||||
config-catalog.zh.md: 4b2ffba0e931c4c515097950e3e69b5744cb5f37
|
||||
config-catalog.md: c4c70a9bc8a0ae964189b7b1a22443ec18b4b8a2
|
||||
config-catalog.zh.md: b034155795efad1c808d3147c8223b29160c6283
|
||||
|
|
|
|||
|
|
@ -2103,7 +2103,7 @@ export interface Config {
|
|||
export type ClaudeCodePermissionMode = typeof CLAUDE_CODE_PERMISSION_MODES[number]
|
||||
```
|
||||
|
||||
Source: [`packages/subagent/subagent-claude-code/src/index.ts:35`](../packages/subagent/subagent-claude-code/src/index.ts)
|
||||
Source: [`packages/subagent/subagent-claude-code/src/index.ts:36`](../packages/subagent/subagent-claude-code/src/index.ts)
|
||||
|
||||
<a id="deepseek-aidsh-subagent-codex"></a>
|
||||
|
||||
|
|
|
|||
|
|
@ -2105,7 +2105,7 @@ export interface Config {
|
|||
export type ClaudeCodePermissionMode = typeof CLAUDE_CODE_PERMISSION_MODES[number]
|
||||
```
|
||||
|
||||
来源:[`packages/subagent/subagent-claude-code/src/index.ts:35`](../packages/subagent/subagent-claude-code/src/index.ts)
|
||||
来源:[`packages/subagent/subagent-claude-code/src/index.ts:36`](../packages/subagent/subagent-claude-code/src/index.ts)
|
||||
|
||||
<a id="deepseek-aidsh-subagent-codex"></a>
|
||||
|
||||
|
|
|
|||
|
|
@ -11,7 +11,7 @@ import { SessionId } from '@deepseek-ai/dsh-session'
|
|||
export const name = 'subagent-result-diagnostic'
|
||||
export const inject = ['subagents']
|
||||
|
||||
const DIAGNOSTIC = 'Claude Code unattended decision (mode: dontAsk; request: tool permission; decision: denied): Claude Code denied the request before an interactive prompt'
|
||||
const DIAGNOSTIC = 'Product subagent failure (product: Claude Code; stage: query-run; category: error_max_budget_usd)'
|
||||
|
||||
class DiagnosticProvider implements SubagentProvider {
|
||||
readonly name = 'snapshot-diagnostic'
|
||||
|
|
|
|||
|
|
@ -15,7 +15,7 @@
|
|||
{"type":"assistant/chunk","seq":13,"time":1783600630852,"data":{"turn":1,"step":1,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}}
|
||||
{"type":"assistant/message","seq":14,"time":1786781990608,"data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"tool-call","id":"call_diagnostic_foreground","name":"subagent_codex","arguments":"{\"description\":\"Observe foreground diagnostic\",\"prompt\":\"Return the diagnostic failure.\",\"run_in_background\":false}"}],"source":{"kind":"model","provider":"deepseek-official","model":"deepseek-v4-pro"},"id":"92e33995-2f02-4ad5-aec1-9df82cf4d583"},"usage":{"inputTokens":10,"outputTokens":5}},"sourceEventSeqs":[9,10,11,12,13],"surfaceOp":"append"}
|
||||
{"type":"tool/call","seq":15,"time":1786781990608,"data":{"turn":1,"step":1,"callId":"call_diagnostic_foreground","name":"subagent_codex","arguments":"{\"description\":\"Observe foreground diagnostic\",\"prompt\":\"Return the diagnostic failure.\",\"run_in_background\":false}"}}
|
||||
{"type":"tool/result","seq":16,"time":1786781990613,"data":{"turn":1,"step":1,"message":{"source":{"kind":"tool","callId":"call_diagnostic_foreground"},"content":[{"type":"tool-result","toolCallId":"call_diagnostic_foreground","content":[{"type":"text","text":"Error: subagent run failed\nDiagnostic: Claude Code unattended decision (mode: dontAsk; request: tool permission; decision: denied): Claude Code denied the request before an interactive prompt\nPartial output before the run ended:\npartial assistant text"}],"isError":true}],"role":"user","id":"4e84e7b3-40c1-488e-b119-45e8bd7ce448"}},"sourceEventSeqs":[15],"surfaceOp":"append"}
|
||||
{"type":"tool/result","seq":16,"time":1786781990613,"data":{"turn":1,"step":1,"message":{"source":{"kind":"tool","callId":"call_diagnostic_foreground"},"content":[{"type":"tool-result","toolCallId":"call_diagnostic_foreground","content":[{"type":"text","text":"Error: subagent run failed\nDiagnostic: Product subagent failure (product: Claude Code; stage: query-run; category: error_max_budget_usd)\nPartial output before the run ended:\npartial assistant text"}],"isError":true}],"role":"user","id":"f63d21ab-ccdc-44f2-9a96-4c60b46e5318"}},"sourceEventSeqs":[15],"surfaceOp":"append"}
|
||||
{"type":"step/end","seq":17,"time":1786781990613,"data":{"turn":1,"step":1}}
|
||||
{"type":"step/start","seq":18,"time":1786781990618,"data":{"turn":1,"step":2}}
|
||||
{"type":"assistant/chunk","seq":19,"time":1783600630926,"data":{"turn":1,"step":2,"chunk":{"type":"block-start","index":0,"blockType":"tool-call"}}}
|
||||
|
|
@ -25,12 +25,12 @@
|
|||
{"type":"assistant/chunk","seq":23,"time":1783600630944,"data":{"turn":1,"step":2,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}}
|
||||
{"type":"assistant/message","seq":24,"time":1786781990622,"data":{"turn":1,"step":2,"message":{"role":"assistant","content":[{"type":"tool-call","id":"call_diagnostic_background","name":"subagent_codex","arguments":"{\"description\":\"Observe background diagnostic\",\"prompt\":\"Return the diagnostic failure.\",\"run_in_background\":true}"}],"source":{"kind":"model","provider":"deepseek-official","model":"deepseek-v4-pro"},"id":"2fb444e2-7a52-4963-988e-b1ecbc3744d5"},"usage":{"inputTokens":10,"outputTokens":5}},"sourceEventSeqs":[19,20,21,22,23],"surfaceOp":"append"}
|
||||
{"type":"tool/call","seq":25,"time":1786781990623,"data":{"turn":1,"step":2,"callId":"call_diagnostic_background","name":"subagent_codex","arguments":"{\"description\":\"Observe background diagnostic\",\"prompt\":\"Return the diagnostic failure.\",\"run_in_background\":true}"}}
|
||||
{"type":"agent/inbox/spliced","seq":26,"time":1786781990627,"data":{"target":"next-step","start":0,"inserted":[{"content":[{"type":"text","text":"background job subagent-1 (subagent: Observe background diagnostic) finished [status: failed, error; diagnostic: Claude Code unattended decision (mode: dontAsk; request: tool permission; decision: denied): Claude Code denied the request before an interactive prompt]. Read its output with job_output."}],"source":{"kind":"plugin","plugin":"tool-jobs","form":"notice","summary":"subagent Observe background diagnostic [status: failed, error; diagnostic: Claude Code unattended decision (mode: dontA…"},"role":"user","id":"de606545-e637-4d9a-ba17-4c722a7331fd"}]}}
|
||||
{"type":"agent/inbox/spliced","seq":26,"time":1786781990627,"data":{"target":"next-step","start":0,"inserted":[{"content":[{"type":"text","text":"background job subagent-1 (subagent: Observe background diagnostic) finished [status: failed, error; diagnostic: Product subagent failure (product: Claude Code; stage: query-run; category: error_max_budget_usd)]. Read its output with job_output."}],"source":{"kind":"plugin","plugin":"tool-jobs","form":"notice","summary":"subagent Observe background diagnostic [status: failed, error; diagnostic: Product subagent failure (product: Claude Co…"},"role":"user","id":"05f93dde-37e9-40e7-92d5-f8de526a8bee"}]}}
|
||||
{"type":"tool/result","seq":27,"time":1786781990627,"data":{"turn":1,"step":2,"message":{"source":{"kind":"tool","callId":"call_diagnostic_background"},"content":[{"type":"tool-result","toolCallId":"call_diagnostic_background","content":[{"type":"text","text":"started background subagent job subagent-1"}],"isError":false}],"role":"user","id":"3377f724-b4a7-4ce1-bed7-774f174917d6"}},"sourceEventSeqs":[25],"surfaceOp":"append"}
|
||||
{"type":"step/end","seq":28,"time":1786781990627,"data":{"turn":1,"step":2}}
|
||||
{"type":"agent/inbox/spliced","seq":29,"time":1786781990627,"data":{"target":"next-step","start":0,"removedCount":1,"inserted":[]}}
|
||||
{"type":"step/start","seq":30,"time":1786781990632,"data":{"turn":1,"step":3}}
|
||||
{"type":"user/message","seq":31,"time":1786781990632,"data":{"content":[{"type":"text","text":"background job subagent-1 (subagent: Observe background diagnostic) finished [status: failed, error; diagnostic: Claude Code unattended decision (mode: dontAsk; request: tool permission; decision: denied): Claude Code denied the request before an interactive prompt]. Read its output with job_output."}],"source":{"kind":"plugin","plugin":"tool-jobs","form":"notice","summary":"subagent Observe background diagnostic [status: failed, error; diagnostic: Claude Code unattended decision (mode: dontA…"},"role":"user","id":"de606545-e637-4d9a-ba17-4c722a7331fd"},"surfaceOp":"append"}
|
||||
{"type":"user/message","seq":31,"time":1786781990632,"data":{"content":[{"type":"text","text":"background job subagent-1 (subagent: Observe background diagnostic) finished [status: failed, error; diagnostic: Product subagent failure (product: Claude Code; stage: query-run; category: error_max_budget_usd)]. Read its output with job_output."}],"source":{"kind":"plugin","plugin":"tool-jobs","form":"notice","summary":"subagent Observe background diagnostic [status: failed, error; diagnostic: Product subagent failure (product: Claude Co…"},"role":"user","id":"05f93dde-37e9-40e7-92d5-f8de526a8bee"},"surfaceOp":"append"}
|
||||
{"type":"assistant/chunk","seq":32,"time":1783600631009,"data":{"turn":1,"step":3,"chunk":{"type":"block-start","index":0,"blockType":"tool-call"}}}
|
||||
{"type":"assistant/chunk","seq":33,"time":1783600631009,"data":{"turn":1,"step":3,"chunk":{"type":"tool-call-delta","index":0,"id":"call_diagnostic_output","name":"job_output","argumentsDelta":"{\"job_id\":\"subagent-1\",\"wait\":true}"}}}
|
||||
{"type":"assistant/chunk","seq":34,"time":1783600631009,"data":{"turn":1,"step":3,"chunk":{"type":"block-end","index":0,"block":{"type":"tool-call","id":"call_diagnostic_output","name":"job_output","arguments":"{\"job_id\":\"subagent-1\",\"wait\":true}"}}}}
|
||||
|
|
@ -38,7 +38,7 @@
|
|||
{"type":"assistant/chunk","seq":36,"time":1785730415297,"data":{"turn":1,"step":3,"chunk":{"type":"finish","reason":{"kind":"tool-calls"}}}}
|
||||
{"type":"assistant/message","seq":37,"time":1785730415298,"data":{"turn":1,"step":3,"message":{"role":"assistant","content":[{"type":"tool-call","id":"call_diagnostic_output","name":"job_output","arguments":"{\"job_id\":\"subagent-1\",\"wait\":true}"}],"source":{"kind":"model","provider":"deepseek-official","model":"deepseek-v4-pro"},"id":"f43f988b-bc08-4811-8671-8edc0613f0d0"},"usage":{"inputTokens":10,"outputTokens":5}},"sourceEventSeqs":[32,33,34,35,36],"surfaceOp":"append"}
|
||||
{"type":"tool/call","seq":38,"time":1786781990636,"data":{"turn":1,"step":3,"callId":"call_diagnostic_output","name":"job_output","arguments":"{\"job_id\":\"subagent-1\",\"wait\":true}"}}
|
||||
{"type":"tool/result","seq":39,"time":1786781990640,"data":{"turn":1,"step":3,"message":{"source":{"kind":"tool","callId":"call_diagnostic_output"},"content":[{"type":"tool-result","toolCallId":"call_diagnostic_output","content":[{"type":"text","text":"(no new output)\n[status: failed, error; diagnostic: Claude Code unattended decision (mode: dontAsk; request: tool permission; decision: denied): Claude Code denied the request before an interactive prompt]"}],"isError":false}],"role":"user","id":"6785120f-ae46-48d0-9f3f-d6cd1e6fc5d7"}},"sourceEventSeqs":[38],"surfaceOp":"append"}
|
||||
{"type":"tool/result","seq":39,"time":1786781990640,"data":{"turn":1,"step":3,"message":{"source":{"kind":"tool","callId":"call_diagnostic_output"},"content":[{"type":"tool-result","toolCallId":"call_diagnostic_output","content":[{"type":"text","text":"(no new output)\n[status: failed, error; diagnostic: Product subagent failure (product: Claude Code; stage: query-run; category: error_max_budget_usd)]"}],"isError":false}],"role":"user","id":"10ac5635-c519-4f69-ab5a-cb0f930e9df0"}},"sourceEventSeqs":[38],"surfaceOp":"append"}
|
||||
{"type":"step/end","seq":40,"time":1786781990640,"data":{"turn":1,"step":3}}
|
||||
{"type":"step/start","seq":41,"time":1786781990645,"data":{"turn":1,"step":4}}
|
||||
{"type":"assistant/chunk","seq":42,"time":1786781990649,"data":{"turn":1,"step":4,"chunk":{"type":"block-start","index":0,"blockType":"text"}}}
|
||||
|
|
|
|||
|
|
@ -2,5 +2,5 @@
|
|||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write packages/subagent/subagent-claude-code/README.md
|
||||
README.md: be3b2262addc487e545fed1f792600a9a5ca24c0
|
||||
README.zh.md: 7ea1b8ca7243790afd387b04d776088cea012718
|
||||
README.md: 21beb0a0534e601f9dd26d36f53d1fb5f09b4e07
|
||||
README.zh.md: c83999ee7b2d38e9c4ee57ba74df25217998bffd
|
||||
|
|
|
|||
|
|
@ -8,15 +8,15 @@ This package registers the fixed `claude-code` subagent provider. Each accepted
|
|||
|
||||
`start(request)` accepts only a non-empty sequence of text blocks and derives the child cwd from the parent Session. It creates one private `AbortController`, calls the official SDK `query()`, and publishes the run only after the SDK's `spawnClaudeCodeProcess` hook has supplied a live CLI handle owned by [`dsh-subprocess`](../../subprocess/subprocess/README.md). A failure or cancellation before publication closes the query, terminates any acquired process tree, waits for it to exit, and rejects `start()`.
|
||||
|
||||
The SDK receives the exact concatenated text task. The provider iterates the complete SDK message stream and accepts only a `result` message with `subtype: "success"`, `is_error: false`, and a nonblank `result`, followed by normal iterator completion. Every SDK error subtype, an error-marked success, a missing answer, iterator failure, protocol failure, or process failure maps to `error`; the provider produces neither `max-tokens` nor `refusal`.
|
||||
The SDK receives the exact concatenated text task. The provider iterates the complete SDK message stream and accepts only a `result` message with `subtype: "success"`, `is_error: false`, and a nonblank `result`, followed by normal iterator completion. Every failure still maps to `error`: the four error subtypes in Agent SDK 0.3.220 retain their exact category, an error-marked or blank success becomes `invalid-success`, a missing result becomes `missing-result`, an unclassified query failure becomes `unknown`, and an early CLI exit becomes `process-exit`. The diagnostic also names the current `query-start`, `query-run`, `process`, or `teardown` stage and independently includes an observed exit code and signal. The provider produces neither `max-tokens` nor `refusal`.
|
||||
|
||||
Local cancellation wins the result race and maps to `aborted`. `dispose()` is idempotent: it aborts the run, asks the SDK query to close, invokes the shared process-tree termination escalation, and waits for whole-tree exit. SDK graceful close expresses protocol intent; the subprocess handle remains the authority for process quiescence. Result failure and independent teardown failure remain separate.
|
||||
Local cancellation wins the result race and maps to `aborted` without a failure diagnostic. `dispose()` is idempotent: it aborts the run, asks the SDK query to close, invokes the shared process-tree termination escalation, and waits for whole-tree exit. SDK graceful close expresses protocol intent; the subprocess handle remains the authority for process quiescence. Startup and teardown rejections expose the same fixed safe stage and process facts through their Error message, while the original product or Host error remains only on the internal cause chain. Result failure and independent teardown failure remain separate.
|
||||
|
||||
## Native settings and interaction
|
||||
|
||||
The provider deliberately omits the SDK `settingSources` option. The official SDK therefore reads the host's normal user, project, and local Claude settings relative to the parent Session cwd, including native account state and product configuration. The provider neither copies nor filters those files and does not create or modify login state. The Profile-selected `permissionMode` is the one query-level override: Claude Code still owns its settings and sandbox, while the selected native mode decides how this unattended query handles permission checks.
|
||||
|
||||
Each query sets `persistSession: false` and disables `AskUserQuestion`. Except in bypass mode, `canUseTool` immediately denies requests that still require human approval. Plan mode also places `ExitPlanMode` in the SDK's `disallowedTools`, so native settings cannot pre-approve a transition back to execution and the model must return the completed plan as its final answer. MCP elicitation is declined, the known refusal fallback dialog is cancelled, and undeclared dialog kinds use the SDK's no-dialog failure behavior. These decisions never wait for a user interface. A permission denial or unattended callback that contributes to a failed run produces an optional `SubagentResult.diagnostic` containing only the product, effective mode, request category, decision, and fixed safe reason; the shared result boundary limits the complete text to 4096 UTF-8 bytes. Successful and locally cancelled runs do not expose the captured failure detail.
|
||||
Each query sets `persistSession: false` and disables `AskUserQuestion`. Except in bypass mode, `canUseTool` immediately denies requests that still require human approval. Plan mode also places `ExitPlanMode` in the SDK's `disallowedTools`, so native settings cannot pre-approve a transition back to execution and the model must return the completed plan as its final answer. MCP elicitation is declined, the known refusal fallback dialog is cancelled, and undeclared dialog kinds use the SDK's no-dialog failure behavior. These decisions never wait for a user interface. When both facts contribute to a failed run, `SubagentResult.diagnostic` contains the structured failure line first and the latest safe permission decision second; the shared result boundary limits the complete text to 4096 UTF-8 bytes. Successful and locally cancelled runs expose neither captured fact.
|
||||
|
||||
## Capabilities and context
|
||||
|
||||
|
|
@ -93,7 +93,7 @@ Independent of the parent request cache. Reuse depends only on Claude Code's own
|
|||
|
||||
#### What the model sees
|
||||
|
||||
Through `dsh-tool-subagent`, a foreground call gives the parent the strict final Claude Code answer or an error containing the stop reason and optional safe diagnostic for a non-completed result. A background call first returns a Job id; the generic job controls later deliver a completion notice, expose the final answer or failed status detail through `job_output`, and let `job_kill` request cancellation. Claude Code reasoning, tool activity, intermediate messages, stderr, workspace diffs, usage, product ids, tool inputs, and raw protocol payloads are not copied into the parent Session.
|
||||
Through `dsh-tool-subagent`, a foreground call gives the parent the strict final Claude Code answer or an error containing the stop reason and optional safe diagnostic for a non-completed result. That diagnostic can distinguish the fixed SDK error category, lifecycle stage, and observed process outcome without copying raw product text. A background call first returns a Job id; the generic job controls later deliver a completion notice, expose the same final answer or failed status detail through `job_output`, and let `job_kill` request cancellation. Claude Code reasoning, tool activity, intermediate messages, stderr, workspace diffs, usage, product ids, tool inputs, and raw protocol payloads are not copied into the parent Session.
|
||||
|
||||
#### Token effect
|
||||
|
||||
|
|
@ -107,7 +107,7 @@ Append-only: foreground adds one result after the reusable parent prefix, while
|
|||
|
||||
- **One fresh query and process per run** — there is no continuation, resume, pooling, progress stream, or product-session persistence.
|
||||
- **Host settings are intentionally authoritative** — project and user settings can change model, tools, and behavior; the provider does not provide a filtered or hermetic production mode.
|
||||
- **Product installation and account state remain native** — a missing or incompatible `claude`, configuration error, or authentication failure is surfaced as a startup or run error; the plugin provides no installer or login flow.
|
||||
- **Product installation and account state remain native** — a missing or incompatible `claude`, configuration error, or authentication failure is surfaced with its lifecycle stage and the safe `unknown` fallback rather than a separate public classification; the plugin provides no installer or login flow.
|
||||
- **The SDK platform CLI remains in the install closure** — production ignores it in favor of the host `claude`, but the current SDK optional dependency is still installed and supplies the keyless compatibility fixture. Removing that payload belongs to the separate product installation-closure follow-up.
|
||||
- **No human interaction path** — `AskUserQuestion` is disabled, permission prompts are denied, MCP elicitation is declined, and blocking dialogs fail closed instead of suspending.
|
||||
- **Assistant payload is final text only** — a failed run may additionally expose the separate safe diagnostic; reasoning, intermediate messages, tool traffic, usage, stderr, and workspace diffs remain product-local, while generic Job ids, notices, and status come from the shared job runtime.
|
||||
|
|
|
|||
|
|
@ -8,15 +8,15 @@
|
|||
|
||||
`start(request)` 只接受非空的文本块序列,并根据父会话确定子级 cwd。它会创建一个私有 `AbortController`,调用官方 SDK 的 `query()`,并仅在 SDK 的 `spawnClaudeCodeProcess` 钩子已经提供由 [`dsh-subprocess`](../../subprocess/subprocess/README.md) 管理的活动 CLI 句柄后发布此次运行。若在发布前发生失败或取消,它会关闭 query、终止所有已取得的进程树并等待其退出,然后拒绝 `start()` 调用。
|
||||
|
||||
SDK 接收由文本块原样拼接成的任务。提供方会完整迭代 SDK 消息流,而且只接受满足以下条件的 `result` 消息:其 `subtype: "success"`、`is_error: false` 且 `result` 非空白,之后迭代器还须正常结束。所有 SDK 错误子类型、标记为错误的成功消息、缺失答案、迭代器失败、协议失败或进程失败都映射为 `error`;该提供方不会产生 `max-tokens` 或 `refusal`。
|
||||
SDK 接收由文本块原样拼接成的任务。提供方会完整迭代 SDK 消息流,而且只接受满足以下条件的 `result` 消息:其 `subtype: "success"`、`is_error: false` 且 `result` 非空白,之后迭代器还须正常结束。所有失败仍映射为 `error`:Agent SDK 0.3.220 的四种错误子类型保留准确类别;标记为错误或内容空白的成功消息成为 `invalid-success`;缺失结果成为 `missing-result`;未分类的 query 失败成为 `unknown`;CLI 提前退出成为 `process-exit`。诊断还会注明当前 `query-start`、`query-run`、`process` 或 `teardown` 阶段,并分别保留已观测到的退出码与信号。该提供方不会产生 `max-tokens` 或 `refusal`。
|
||||
|
||||
本地取消会在结果竞态中胜出并映射为 `aborted`。`dispose()`(资源释放)具有幂等性:它会中止此次运行、请求 SDK query 关闭、调用共享的进程树逐级终止机制,并等待整棵进程树退出。SDK 的优雅关闭只表达协议意图;进程是否完全停稳仍以子进程句柄为准。结果失败与独立的清理失败仍彼此分离。
|
||||
本地取消会在结果竞态中胜出并映射为 `aborted`,且不附带失败诊断。`dispose()`(资源释放)具有幂等性:它会中止此次运行、请求 SDK query 关闭、调用共享的进程树逐级终止机制,并等待整棵进程树退出。SDK 的优雅关闭只表达协议意图;进程是否完全停稳仍以子进程句柄为准。启动与清理拒绝会在 Error 消息中公开同样固定的安全阶段和进程事实,而原始产品或 Host 错误只保留在内部 cause 链上。结果失败与独立的清理失败仍彼此分离。
|
||||
|
||||
## 原生设置与交互
|
||||
|
||||
提供方故意省略 SDK 的 `settingSources` 选项。因此,官方 SDK 会相对于父会话 cwd 读取宿主机常规的用户、项目和本地 Claude 设置,包括原生账户状态与产品配置。提供方既不复制也不过滤这些文件,也不会创建或修改登录状态。Profile 选择的 `permissionMode` 是唯一的 query 级覆盖:Claude Code 仍拥有其设置与沙箱,而所选原生模式决定这个无人值守 query 如何处理权限检查。
|
||||
|
||||
每次 query 都设置 `persistSession: false` 并禁用 `AskUserQuestion`。除 bypass 模式外,`canUseTool` 会立即拒绝仍需人工审批的请求。Plan 模式还会把 `ExitPlanMode` 放入 SDK 的 `disallowedTools`,因此原生 settings 无法预先放行回到执行模式的转换,模型必须把完整计划作为最终答案返回。MCP elicitation 会被拒绝,已知的拒绝回退对话会被取消,未声明的对话类型则使用 SDK 的无对话失败行为。这些决定都不会等待用户界面。若权限拒绝或无人值守回调参与了一次失败运行,提供方会生成可选的 `SubagentResult.diagnostic`,其中只包含产品、有效模式、请求类别、决定与固定的安全原因;共享结果边界会把完整文本限制在 4096 个 UTF-8 字节以内。成功运行与本地取消不会公开已捕获的失败说明。
|
||||
每次 query 都设置 `persistSession: false` 并禁用 `AskUserQuestion`。除 bypass 模式外,`canUseTool` 会立即拒绝仍需人工审批的请求。Plan 模式还会把 `ExitPlanMode` 放入 SDK 的 `disallowedTools`,因此原生 settings 无法预先放行回到执行模式的转换,模型必须把完整计划作为最终答案返回。MCP elicitation 会被拒绝,已知的拒绝回退对话会被取消,未声明的对话类型则使用 SDK 的无对话失败行为。这些决定都不会等待用户界面。当两类事实共同参与一次失败运行时,`SubagentResult.diagnostic` 会先写入结构化失败行,再写入最新的安全权限决定;共享结果边界会把完整文本限制在 4096 个 UTF-8 字节以内。成功运行与本地取消都不会公开已捕获的事实。
|
||||
|
||||
## 能力与上下文
|
||||
|
||||
|
|
@ -93,7 +93,7 @@ Claude Code 子级会在一个全新的 SDK query 中接收独立文本任务。
|
|||
|
||||
#### 模型看到的内容
|
||||
|
||||
通过 `dsh-tool-subagent`,前台调用会让父级模型看到符合严格成功条件的 Claude Code 最终答案;若结果未完成,错误中会包含终止原因和可选的安全诊断。后台调用会先返回 Job id;随后通用作业控制面会送达完成通知,通过 `job_output` 公开最终答案或失败状态 detail,并允许 `job_kill` 请求取消。Claude Code 的推理、工具活动、中间消息、stderr、工作区差异、用量信息、产品标识符、工具输入和原始协议载荷均不会复制到父会话。
|
||||
通过 `dsh-tool-subagent`,前台调用会让父级模型看到符合严格成功条件的 Claude Code 最终答案;若结果未完成,错误中会包含终止原因和可选的安全诊断。该诊断可以区分固定 SDK 错误类别、生命周期阶段和已观测的进程结果,而不复制原始产品文本。后台调用会先返回 Job id;随后通用作业控制面会送达完成通知,通过 `job_output` 公开同一最终答案或失败状态 detail,并允许 `job_kill` 请求取消。Claude Code 的推理、工具活动、中间消息、stderr、工作区差异、用量信息、产品标识符、工具输入和原始协议载荷均不会复制到父会话。
|
||||
|
||||
#### 对 token 的影响
|
||||
|
||||
|
|
@ -107,7 +107,7 @@ Claude Code 子级会在一个全新的 SDK query 中接收独立文本任务。
|
|||
|
||||
- **每次运行均新建一个 query 和一个进程**:不支持续接、恢复、池化、进度流或产品会话持久化。
|
||||
- **宿主设置有意保持权威**:项目和用户设置可以改变模型、工具与行为;本提供方不提供经过筛选或与宿主环境隔离的生产模式。
|
||||
- **产品安装与账户状态仍由原生机制管理**:`claude` 缺失或不兼容、配置错误或身份验证失败都会呈现为启动错误或运行错误;本插件不提供安装程序或登录流程。
|
||||
- **产品安装与账户状态仍由原生机制管理**:`claude` 缺失或不兼容、配置错误或身份验证失败会公开其生命周期阶段与安全的 `unknown` 回退,而不会增加单独的公开分类;本插件不提供安装程序或登录流程。
|
||||
- **SDK 平台 CLI 仍在安装闭包内**:生产环境会忽略它,改用宿主提供的 `claude`,但当前 SDK 的可选依赖仍会安装,并提供无密钥兼容性 fixture。移除该载荷属于独立的产品安装闭包后续项。
|
||||
- **没有人工交互路径**:`AskUserQuestion` 被禁用,权限提示会被拒绝,MCP elicitation 会被拒绝,阻塞对话会快速失败而不会挂起。
|
||||
- **assistant 载荷仅包含最终文本**:失败运行可以额外公开独立的安全诊断;推理、中间消息、工具通信、用量信息、stderr 和工作区差异仍只保留在产品内部,通用 Job id、通知与状态来自共享作业运行时。
|
||||
|
|
|
|||
|
|
@ -21,6 +21,7 @@ import {
|
|||
CLAUDE_CODE_PERMISSION_MODES,
|
||||
DEFAULT_CLAUDE_CODE_PERMISSION_MODE,
|
||||
DEFAULT_DISPOSE_GRACE_MS,
|
||||
claudeCodeStartupFailure,
|
||||
startClaudeCodeRun,
|
||||
type ClaudeCodePermissionMode,
|
||||
type ClaudeCodeRunSpec,
|
||||
|
|
@ -78,17 +79,29 @@ class ClaudeCodeProvider implements SubagentProvider {
|
|||
'subagent-claude-code: no working directory for the child — delegate from a parent session that has one',
|
||||
)
|
||||
}
|
||||
const executable = await this.ctx.subprocess.resolveExecutable(
|
||||
'claude',
|
||||
this.config.env,
|
||||
request.signal,
|
||||
)
|
||||
const spec: ClaudeCodeRunSpec = {
|
||||
cwd: resolveChildCwd(
|
||||
let cwd: string
|
||||
let executable: string
|
||||
try {
|
||||
cwd = resolveChildCwd(
|
||||
'subagent-claude-code',
|
||||
undefined,
|
||||
parentCwd,
|
||||
),
|
||||
)
|
||||
executable = await this.ctx.subprocess.resolveExecutable(
|
||||
'claude',
|
||||
this.config.env,
|
||||
request.signal,
|
||||
)
|
||||
} catch (error: unknown) {
|
||||
if (request.signal.aborted) {
|
||||
throw new Error(
|
||||
'subagent-claude-code: request was aborted before SDK startup',
|
||||
)
|
||||
}
|
||||
throw claudeCodeStartupFailure(error)
|
||||
}
|
||||
const spec: ClaudeCodeRunSpec = {
|
||||
cwd,
|
||||
executable,
|
||||
permissionMode: this.config.permissionMode,
|
||||
env: this.config.env,
|
||||
|
|
@ -96,7 +109,8 @@ class ClaudeCodeProvider implements SubagentProvider {
|
|||
spawn: spawnSpec => this.ctx.subprocess.spawn(spawnSpec),
|
||||
onError: (error, stopReason) => {
|
||||
this.ctx.logger.warn(
|
||||
`subagent-claude-code: child run failed (${stopReason}): ${error.message}`,
|
||||
`subagent-claude-code: child run failed (${stopReason}): %o`,
|
||||
error,
|
||||
)
|
||||
},
|
||||
}
|
||||
|
|
|
|||
|
|
@ -28,6 +28,7 @@ import {
|
|||
import {
|
||||
scrubbedParentEnv,
|
||||
type SubprocessHandle,
|
||||
type SubprocessOutcome,
|
||||
type SubprocessSpawnSpec,
|
||||
} from '@deepseek-ai/dsh-subprocess'
|
||||
import {
|
||||
|
|
@ -57,6 +58,83 @@ const SUPPORTED_UNATTENDED_DIALOG_KINDS = [
|
|||
'refusal_fallback_prompt',
|
||||
] satisfies NonNullable<Options['supportedDialogKinds']>
|
||||
|
||||
type ClaudeCodeErrorSubtype = Exclude<SDKResultMessage['subtype'], 'success'>
|
||||
|
||||
type ClaudeCodeFailureStage =
|
||||
| 'query-start'
|
||||
| 'query-run'
|
||||
| 'process'
|
||||
| 'teardown'
|
||||
|
||||
type ClaudeCodeFailureCategory =
|
||||
| ClaudeCodeErrorSubtype
|
||||
| 'invalid-success'
|
||||
| 'missing-result'
|
||||
| 'process-exit'
|
||||
| 'unknown'
|
||||
|
||||
interface ClaudeCodeFailureFacts {
|
||||
readonly stage: ClaudeCodeFailureStage
|
||||
readonly category: ClaudeCodeFailureCategory
|
||||
readonly outcome?: SubprocessOutcome | undefined
|
||||
}
|
||||
|
||||
function failureDiagnostic(facts: ClaudeCodeFailureFacts): string {
|
||||
const fields = [
|
||||
'product: Claude Code',
|
||||
`stage: ${facts.stage}`,
|
||||
`category: ${facts.category}`,
|
||||
]
|
||||
const exitCode = facts.outcome?.exitCode
|
||||
if (exitCode !== null && exitCode !== undefined) {
|
||||
fields.push(`exit code: ${exitCode}`)
|
||||
}
|
||||
const signal = facts.outcome?.signal
|
||||
if (signal !== null && signal !== undefined) {
|
||||
fields.push(`signal: ${signal}`)
|
||||
}
|
||||
return `Product subagent failure (${fields.join('; ')})`
|
||||
}
|
||||
|
||||
class ClaudeCodeFailure extends Error {
|
||||
constructor(
|
||||
readonly facts: ClaudeCodeFailureFacts,
|
||||
cause?: unknown,
|
||||
) {
|
||||
super(
|
||||
`subagent-claude-code: ${failureDiagnostic(facts)}`,
|
||||
cause === undefined ? undefined : { cause },
|
||||
)
|
||||
this.name = 'ClaudeCodeFailure'
|
||||
}
|
||||
}
|
||||
|
||||
function sdkFailureCategory(
|
||||
subtype: string,
|
||||
): ClaudeCodeErrorSubtype | 'unknown' {
|
||||
switch (subtype) {
|
||||
case 'error_during_execution':
|
||||
case 'error_max_turns':
|
||||
case 'error_max_budget_usd':
|
||||
case 'error_max_structured_output_retries':
|
||||
return subtype
|
||||
default:
|
||||
return 'unknown'
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Hide an unpublished product startup failure behind fixed safe facts.
|
||||
* @param cause - original host-side failure retained only on the Error cause chain.
|
||||
* @returns a rejection safe to expose through the subagent start boundary.
|
||||
*/
|
||||
export function claudeCodeStartupFailure(cause: unknown): Error {
|
||||
return new ClaudeCodeFailure({
|
||||
stage: 'query-start',
|
||||
category: 'unknown',
|
||||
}, cause)
|
||||
}
|
||||
|
||||
function unattendedDiagnostic(
|
||||
mode: ClaudeCodePermissionMode,
|
||||
request: 'tool permission' | 'MCP elicitation' | 'user dialog',
|
||||
|
|
@ -90,6 +168,10 @@ function thrown(value: unknown): Error {
|
|||
/* v8 ignore next -- typed SDK and subprocess failures reject with Error. */
|
||||
return value instanceof Error ? value : new Error(String(value))
|
||||
}
|
||||
|
||||
function isAborted(signal: AbortSignal): boolean {
|
||||
return signal.aborted
|
||||
}
|
||||
/* jscpd:ignore-end */
|
||||
|
||||
/**
|
||||
|
|
@ -120,15 +202,23 @@ export function textTask(prompt: readonly ContentBlock[]): string {
|
|||
* @returns exact final text for a successful, non-error result.
|
||||
*/
|
||||
export function successfulResult(message: SDKResultMessage): string {
|
||||
if (
|
||||
message.subtype !== 'success'
|
||||
|| message.is_error
|
||||
|| message.result.trim().length === 0
|
||||
) {
|
||||
const detail = message.subtype === 'success'
|
||||
? 'success result was marked as an error or contained no answer'
|
||||
: message.errors.join('; ') || message.subtype
|
||||
throw new Error(`subagent-claude-code: Claude Code failed: ${detail}`)
|
||||
if (message.subtype !== 'success') {
|
||||
const category = sdkFailureCategory(message.subtype)
|
||||
const detail = category === 'unknown'
|
||||
? undefined
|
||||
: message.errors.join('; ')
|
||||
throw new ClaudeCodeFailure(
|
||||
{ stage: 'query-run', category },
|
||||
detail === undefined || detail.length === 0
|
||||
? undefined
|
||||
: new Error(detail),
|
||||
)
|
||||
}
|
||||
if (message.is_error || message.result.trim().length === 0) {
|
||||
throw new ClaudeCodeFailure({
|
||||
stage: 'query-run',
|
||||
category: 'invalid-success',
|
||||
})
|
||||
}
|
||||
return message.result
|
||||
}
|
||||
|
|
@ -154,7 +244,10 @@ export async function consumeClaudeQuery(
|
|||
answer = successfulResult(message)
|
||||
}
|
||||
if (answer === undefined) {
|
||||
throw new Error('subagent-claude-code: Claude Code ended without a result')
|
||||
throw new ClaudeCodeFailure({
|
||||
stage: 'query-run',
|
||||
category: 'missing-result',
|
||||
})
|
||||
}
|
||||
return {
|
||||
output: [{ type: 'text', text: answer }],
|
||||
|
|
@ -173,6 +266,7 @@ export async function disposeClaudeCodeChild(
|
|||
child: SubprocessHandle,
|
||||
): Promise<void> {
|
||||
const failures: Error[] = []
|
||||
let outcome: SubprocessOutcome | undefined
|
||||
try {
|
||||
query?.close()
|
||||
} catch (error: unknown) {
|
||||
|
|
@ -188,17 +282,24 @@ export async function disposeClaudeCodeChild(
|
|||
}
|
||||
}
|
||||
try {
|
||||
await child.done
|
||||
outcome = await child.done
|
||||
} catch (error: unknown) {
|
||||
failures.push(thrown(error))
|
||||
}
|
||||
|
||||
const firstFailure = failures[0]
|
||||
if (failures.length === 1 && firstFailure !== undefined) throw firstFailure
|
||||
if (failures.length > 1) {
|
||||
if (firstFailure !== undefined) {
|
||||
const facts = {
|
||||
stage: 'teardown',
|
||||
category: 'unknown',
|
||||
outcome,
|
||||
} as const
|
||||
if (failures.length === 1) {
|
||||
throw new ClaudeCodeFailure(facts, firstFailure)
|
||||
}
|
||||
throw new AggregateError(
|
||||
failures,
|
||||
'subagent-claude-code: query and process cleanup failed',
|
||||
failures.map(failure => new ClaudeCodeFailure(facts, failure)),
|
||||
`subagent-claude-code: ${failureDiagnostic(facts)}`,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
|
@ -296,9 +397,21 @@ export async function startClaudeCodeRun(
|
|||
|
||||
let child: SubprocessHandle | undefined
|
||||
let query: Query | undefined
|
||||
let diagnostic: string | undefined
|
||||
const captureDiagnostic = (value: string): void => {
|
||||
diagnostic = value
|
||||
let processOutcome: SubprocessOutcome | undefined
|
||||
let failureDetail: string | undefined
|
||||
let permissionDetail: string | undefined
|
||||
const capturePermissionDiagnostic = (value: string): void => {
|
||||
permissionDetail = value
|
||||
}
|
||||
const collectDiagnostic = (): string => [failureDetail, permissionDetail]
|
||||
.filter((value): value is string => value !== undefined)
|
||||
.join('\n')
|
||||
const captureChild = (captured: SubprocessHandle): void => {
|
||||
child = captured
|
||||
void captured.done.then(
|
||||
(outcome: SubprocessOutcome) => { processOutcome = outcome },
|
||||
() => undefined,
|
||||
)
|
||||
}
|
||||
try {
|
||||
query = officialQuery({
|
||||
|
|
@ -306,10 +419,8 @@ export async function startClaudeCodeRun(
|
|||
options: claudeQueryOptions(
|
||||
spec,
|
||||
controller,
|
||||
(captured) => {
|
||||
child = captured
|
||||
},
|
||||
captureDiagnostic,
|
||||
captureChild,
|
||||
capturePermissionDiagnostic,
|
||||
),
|
||||
})
|
||||
if (child === undefined || child.pid <= 0) {
|
||||
|
|
@ -323,46 +434,83 @@ export async function startClaudeCodeRun(
|
|||
} catch (error: unknown) {
|
||||
request.signal.removeEventListener('abort', onAbort)
|
||||
const cancelledBeforeCleanup = controller.signal.aborted
|
||||
await Promise.resolve()
|
||||
const startupOutcome = processOutcome
|
||||
const startupFacts = {
|
||||
stage: 'query-start',
|
||||
category: 'unknown',
|
||||
outcome: startupOutcome,
|
||||
} as const
|
||||
const startupFailure = (): ClaudeCodeFailure => new ClaudeCodeFailure(
|
||||
startupFacts,
|
||||
thrown(error),
|
||||
)
|
||||
requestCancel()
|
||||
if (child !== undefined) {
|
||||
try {
|
||||
await disposeClaudeCodeChild(query, child)
|
||||
} catch (disposeError: unknown) {
|
||||
const failure = startupFailure()
|
||||
throw new AggregateError(
|
||||
[thrown(error), thrown(disposeError)],
|
||||
'subagent-claude-code: startup failed and CLI cleanup also failed',
|
||||
[failure, thrown(disposeError)],
|
||||
`${failure.message}; startup cleanup also failed`,
|
||||
)
|
||||
}
|
||||
} else if (query !== undefined) {
|
||||
try {
|
||||
query.close()
|
||||
} catch (disposeError: unknown) {
|
||||
const failure = startupFailure()
|
||||
throw new AggregateError(
|
||||
[thrown(error), thrown(disposeError)],
|
||||
'subagent-claude-code: startup failed and query cleanup also failed',
|
||||
[
|
||||
failure,
|
||||
new ClaudeCodeFailure({
|
||||
stage: 'teardown',
|
||||
category: 'unknown',
|
||||
}, thrown(disposeError)),
|
||||
],
|
||||
`${failure.message}; startup cleanup also failed`,
|
||||
)
|
||||
}
|
||||
}
|
||||
// oxlint-disable-next-line typescript/no-unnecessary-condition -- the request can abort while process cleanup is awaited.
|
||||
if (cancelledBeforeCleanup || request.signal.aborted) {
|
||||
if (cancelledBeforeCleanup || isAborted(request.signal)) {
|
||||
throw new Error('subagent-claude-code: request was aborted before SDK startup')
|
||||
}
|
||||
throw thrown(error)
|
||||
throw startupFailure()
|
||||
}
|
||||
|
||||
const publishedQuery = query
|
||||
const publishedChild = child
|
||||
const result = settleRunResult({
|
||||
attempt: () => consumeClaudeQuery(publishedQuery, () => {
|
||||
captureDiagnostic(unattendedDiagnostic(
|
||||
spec.permissionMode,
|
||||
'tool permission',
|
||||
'denied',
|
||||
'Claude Code denied the request before an interactive prompt',
|
||||
))
|
||||
}),
|
||||
attempt: async () => {
|
||||
try {
|
||||
return await consumeClaudeQuery(publishedQuery, () => {
|
||||
capturePermissionDiagnostic(unattendedDiagnostic(
|
||||
spec.permissionMode,
|
||||
'tool permission',
|
||||
'denied',
|
||||
'Claude Code denied the request before an interactive prompt',
|
||||
))
|
||||
})
|
||||
} catch (error: unknown) {
|
||||
await Promise.resolve()
|
||||
const facts = error instanceof ClaudeCodeFailure
|
||||
? { ...error.facts, outcome: processOutcome }
|
||||
: processOutcome === undefined
|
||||
? { stage: 'query-run', category: 'unknown' } as const
|
||||
: {
|
||||
stage: 'process',
|
||||
category: 'process-exit',
|
||||
outcome: processOutcome,
|
||||
} as const
|
||||
failureDetail = failureDiagnostic(facts)
|
||||
throw error instanceof ClaudeCodeFailure
|
||||
? error
|
||||
: new ClaudeCodeFailure(facts, thrown(error))
|
||||
}
|
||||
},
|
||||
collectOutput: () => [],
|
||||
collectDiagnostic: () => diagnostic,
|
||||
collectDiagnostic,
|
||||
cancelled: () => controller.signal.aborted,
|
||||
onError: spec.onError,
|
||||
signal: request.signal,
|
||||
|
|
|
|||
|
|
@ -21,7 +21,11 @@ import { Context } from '@deepseek-ai/cordis'
|
|||
import { afterAll, afterEach, beforeAll, describe, expect, it, vi } from 'vitest'
|
||||
import type { Agent } from '@deepseek-ai/dsh-agent'
|
||||
import SubagentRuntime from '@deepseek-ai/dsh-subagent'
|
||||
import type { SubprocessHandle, SubprocessSpawnSpec } from '@deepseek-ai/dsh-subprocess'
|
||||
import type {
|
||||
SubprocessHandle,
|
||||
SubprocessOutcome,
|
||||
SubprocessSpawnSpec,
|
||||
} from '@deepseek-ai/dsh-subprocess'
|
||||
import LocalSubprocessRuntime from '@deepseek-ai/dsh-subprocess-local'
|
||||
import * as claudeCode from '../src/index.ts'
|
||||
import type { ClaudeCodePermissionMode } from '../src/run.ts'
|
||||
|
|
@ -32,6 +36,7 @@ import {
|
|||
} from './messages-fixture.ts'
|
||||
|
||||
const observedSdkMessages = vi.hoisted((): SDKMessage[] => [])
|
||||
const sdkTestOverrides = vi.hoisted((): { maxTurns?: number } => ({}))
|
||||
|
||||
vi.mock('@anthropic-ai/claude-agent-sdk', async (importOriginal) => {
|
||||
const actual = await importOriginal<
|
||||
|
|
@ -39,8 +44,13 @@ vi.mock('@anthropic-ai/claude-agent-sdk', async (importOriginal) => {
|
|||
>()
|
||||
return {
|
||||
...actual,
|
||||
query(options: Parameters<typeof actual.query>[0]): Query {
|
||||
const query = actual.query(options)
|
||||
query(params: Parameters<typeof actual.query>[0]): Query {
|
||||
const query = actual.query(sdkTestOverrides.maxTurns === undefined
|
||||
? params
|
||||
: {
|
||||
...params,
|
||||
options: { ...params.options, maxTurns: sdkTestOverrides.maxTurns },
|
||||
})
|
||||
// Observe the real SDK stream without replacing its protocol or CLI.
|
||||
return new Proxy(query, {
|
||||
get(target, property) {
|
||||
|
|
@ -112,6 +122,7 @@ afterEach(async () => {
|
|||
await rm(root, { recursive: true, force: true, maxRetries: 10, retryDelay: 100 })
|
||||
}
|
||||
observedSdkMessages.length = 0
|
||||
delete sdkTestOverrides.maxTurns
|
||||
})
|
||||
|
||||
interface RealHarness {
|
||||
|
|
@ -216,6 +227,17 @@ async function expectQuiescent(
|
|||
}
|
||||
}
|
||||
|
||||
function expectedProcessFailure(outcome: SubprocessOutcome): string {
|
||||
const fields = [
|
||||
'product: Claude Code',
|
||||
'stage: process',
|
||||
'category: process-exit',
|
||||
]
|
||||
if (outcome.exitCode !== null) fields.push(`exit code: ${outcome.exitCode}`)
|
||||
if (outcome.signal !== null) fields.push(`signal: ${outcome.signal}`)
|
||||
return `Product subagent failure (${fields.join('; ')})`
|
||||
}
|
||||
|
||||
function startRequest(
|
||||
harness: RealHarness,
|
||||
prompt: string,
|
||||
|
|
@ -294,14 +316,49 @@ describe('real Claude Agent SDK 0.3.220 and its distributed Claude Code 2.1.220
|
|||
await expectQuiescent(harness.handles)
|
||||
})
|
||||
|
||||
it('maps a real CLI process failure to error', async () => {
|
||||
it('maps a real SDK max-turns result to safe query-run facts', async () => {
|
||||
const root = mkdtempSync(join(tmpdir(), 'dsh-claude-code-max-turns-'))
|
||||
roots.push(root)
|
||||
const target = join(root, 'max-turns.txt')
|
||||
sdkTestOverrides.maxTurns = 1
|
||||
const { harness, fixture } = await realHarness({
|
||||
kind: 'tool-use',
|
||||
toolName: 'Write',
|
||||
input: {
|
||||
file_path: target,
|
||||
content: 'real-sdk-max-turns',
|
||||
},
|
||||
}, 'bypassPermissions')
|
||||
const run = await startRequest(harness, 'Exercise the SDK max-turns result.')
|
||||
const result = await run.result
|
||||
expect(observedSdkMessages
|
||||
.filter(message => message.type === 'result')
|
||||
.map(message => message.subtype)).toEqual(['error_max_turns'])
|
||||
expect(result).toMatchObject({
|
||||
output: [],
|
||||
stopReason: 'error',
|
||||
})
|
||||
expect(result.diagnostic).toContain(
|
||||
'product: Claude Code; stage: query-run; category: error_max_turns',
|
||||
)
|
||||
expect(readFileSync(target, 'utf8')).toBe('real-sdk-max-turns')
|
||||
expect(result.diagnostic).not.toContain(target)
|
||||
expect(result.diagnostic).not.toContain('real-sdk-max-turns')
|
||||
await run.dispose()
|
||||
expect(fixture.requests).toHaveLength(1)
|
||||
await expectQuiescent(harness.handles)
|
||||
})
|
||||
|
||||
it('maps a real CLI process failure to its exit outcome', async () => {
|
||||
const { harness, fixture } = await realHarness({ kind: 'hold' })
|
||||
const run = await startRequest(harness, 'Exercise the failure path.')
|
||||
await fixture.requestStarted
|
||||
expect(harness.handles).toHaveLength(1)
|
||||
harness.handles[0]!.terminate()
|
||||
const outcome = await harness.handles[0]!.done
|
||||
await expect(run.result).resolves.toEqual({
|
||||
output: [],
|
||||
diagnostic: expectedProcessFailure(outcome),
|
||||
stopReason: 'error',
|
||||
})
|
||||
await run.dispose()
|
||||
|
|
@ -330,10 +387,11 @@ describe('real Claude Agent SDK 0.3.220 and its distributed Claude Code 2.1.220
|
|||
}, { timeout: 30_000 })
|
||||
expect(existsSync(target)).toBe(false)
|
||||
harness.handles[0]!.terminate()
|
||||
const outcome = await harness.handles[0]!.done
|
||||
const result = await run.result
|
||||
expect(result).toEqual({
|
||||
output: [],
|
||||
diagnostic: 'Claude Code unattended decision (mode: dontAsk; request: tool permission; decision: denied): Claude Code denied the request before an interactive prompt',
|
||||
diagnostic: `${expectedProcessFailure(outcome)}\nClaude Code unattended decision (mode: dontAsk; request: tool permission; decision: denied): Claude Code denied the request before an interactive prompt`,
|
||||
stopReason: 'error',
|
||||
})
|
||||
expect(result.diagnostic).not.toContain(target)
|
||||
|
|
|
|||
|
|
@ -192,6 +192,25 @@ function failure(
|
|||
} as SDKResultMessage
|
||||
}
|
||||
|
||||
function expectedFailureDiagnostic(
|
||||
stage: 'query-start' | 'query-run' | 'process' | 'teardown',
|
||||
category: string,
|
||||
outcome?: Partial<SubprocessOutcome>,
|
||||
): string {
|
||||
const fields = [
|
||||
'product: Claude Code',
|
||||
`stage: ${stage}`,
|
||||
`category: ${category}`,
|
||||
]
|
||||
if (outcome?.exitCode !== null && outcome?.exitCode !== undefined) {
|
||||
fields.push(`exit code: ${outcome.exitCode}`)
|
||||
}
|
||||
if (outcome?.signal !== null && outcome?.signal !== undefined) {
|
||||
fields.push(`signal: ${outcome.signal}`)
|
||||
}
|
||||
return `Product subagent failure (${fields.join('; ')})`
|
||||
}
|
||||
|
||||
function permissionDenied(): SDKPermissionDeniedMessage {
|
||||
return {
|
||||
type: 'system',
|
||||
|
|
@ -397,7 +416,18 @@ describe('task admission and package contracts', () => {
|
|||
|
||||
resolveExecutable.mockRejectedValueOnce(new Error('claude missing from PATH'))
|
||||
await expect(ctx.subagents.start('claude-code', request()))
|
||||
.rejects.toThrow('claude missing from PATH')
|
||||
.rejects.toThrow(expectedFailureDiagnostic('query-start', 'unknown'))
|
||||
expect(queryMock).not.toHaveBeenCalled()
|
||||
|
||||
const resolutionAbort = new AbortController()
|
||||
resolveExecutable.mockImplementationOnce(async () => {
|
||||
resolutionAbort.abort(new Error('parent cancelled executable resolution'))
|
||||
throw new Error('SECRET_TOKEN from executable resolution')
|
||||
})
|
||||
await expect(ctx.subagents.start(
|
||||
'claude-code',
|
||||
request(undefined, resolutionAbort.signal),
|
||||
)).rejects.toThrow('aborted before SDK startup')
|
||||
expect(queryMock).not.toHaveBeenCalled()
|
||||
|
||||
const run = await ctx.subagents.start('claude-code', request())
|
||||
|
|
@ -405,11 +435,13 @@ describe('task admission and package contracts', () => {
|
|||
child.stdout.end()
|
||||
await expect(run.result).resolves.toEqual({
|
||||
output: [],
|
||||
diagnostic: expectedFailureDiagnostic('query-run', 'missing-result'),
|
||||
stopReason: 'error',
|
||||
})
|
||||
expect(warn).toHaveBeenCalledWith(expect.stringContaining(
|
||||
'subagent-claude-code: child run failed (error):',
|
||||
))
|
||||
expect(warn).toHaveBeenCalledWith(
|
||||
expect.stringContaining('subagent-claude-code: child run failed (error):'),
|
||||
expect.any(Error),
|
||||
)
|
||||
expect(resolveExecutable).toHaveBeenCalledWith(
|
||||
'claude',
|
||||
expect.objectContaining({ ANTHROPIC_API_KEY: 'provider-fake-key' }),
|
||||
|
|
@ -711,17 +743,34 @@ describe('query options and result mapping', () => {
|
|||
it('accepts only a non-error success with a non-blank final result', () => {
|
||||
expect(successfulResult(success('exact final'))).toBe('exact final')
|
||||
expect(() => successfulResult(success('answer', true)))
|
||||
.toThrow('marked as an error')
|
||||
.toThrow(expectedFailureDiagnostic('query-run', 'invalid-success'))
|
||||
expect(() => successfulResult(success(' \n ')))
|
||||
.toThrow('contained no answer')
|
||||
expect(() => successfulResult(failure(
|
||||
.toThrow(expectedFailureDiagnostic('query-run', 'invalid-success'))
|
||||
const sdkFailure = () => successfulResult(failure(
|
||||
'error_during_execution',
|
||||
['first', 'second'],
|
||||
))).toThrow('first; second')
|
||||
['SECRET_TOKEN', '/private/secret.txt'],
|
||||
))
|
||||
expect(sdkFailure).toThrow(expectedFailureDiagnostic(
|
||||
'query-run',
|
||||
'error_during_execution',
|
||||
))
|
||||
expect(sdkFailure).not.toThrow('SECRET_TOKEN')
|
||||
expect(sdkFailure).not.toThrow('/private/secret.txt')
|
||||
expect(() => successfulResult(failure(
|
||||
'error_max_turns',
|
||||
[],
|
||||
))).toThrow('error_max_turns')
|
||||
))).toThrow(expectedFailureDiagnostic('query-run', 'error_max_turns'))
|
||||
|
||||
const unknown = {
|
||||
type: 'result',
|
||||
subtype: 'future_failure',
|
||||
is_error: true,
|
||||
errors: ['SECRET_TOKEN'],
|
||||
} as unknown as SDKResultMessage
|
||||
expect(() => successfulResult(unknown))
|
||||
.toThrow(expectedFailureDiagnostic('query-run', 'unknown'))
|
||||
expect(() => successfulResult(unknown)).not.toThrow('future_failure')
|
||||
expect(() => successfulResult(unknown)).not.toThrow('SECRET_TOKEN')
|
||||
})
|
||||
|
||||
it('consumes the complete stream and keeps the latest strict success', async () => {
|
||||
|
|
@ -736,7 +785,7 @@ describe('query options and result mapping', () => {
|
|||
})
|
||||
await expect(consumeClaudeQuery(
|
||||
queryFrom([{ type: 'system', subtype: 'init' } as SDKMessage]),
|
||||
)).rejects.toThrow('ended without a result')
|
||||
)).rejects.toThrow(expectedFailureDiagnostic('query-run', 'missing-result'))
|
||||
|
||||
const onPermissionDenied = vi.fn()
|
||||
await expect(consumeClaudeQuery(queryFrom([
|
||||
|
|
@ -790,6 +839,7 @@ describe('run publication, cancellation, and settlement', () => {
|
|||
)
|
||||
await expect(run.result).resolves.toEqual({
|
||||
output: [],
|
||||
diagnostic: expectedFailureDiagnostic('query-run', subtype),
|
||||
stopReason: 'error',
|
||||
})
|
||||
expect(onError).toHaveBeenCalledWith(
|
||||
|
|
@ -809,7 +859,7 @@ describe('run publication, cancellation, and settlement', () => {
|
|||
const result = await run.result
|
||||
expect(result).toEqual({
|
||||
output: [],
|
||||
diagnostic: 'Claude Code unattended decision (mode: dontAsk; request: tool permission; decision: denied): Claude Code denied the request before an interactive prompt',
|
||||
diagnostic: `${expectedFailureDiagnostic('query-run', 'error_during_execution')}\nClaude Code unattended decision (mode: dontAsk; request: tool permission; decision: denied): Claude Code denied the request before an interactive prompt`,
|
||||
stopReason: 'error',
|
||||
})
|
||||
expect(result.diagnostic).not.toContain('SECRET_TOKEN')
|
||||
|
|
@ -851,6 +901,10 @@ describe('run publication, cancellation, and settlement', () => {
|
|||
})
|
||||
await expect(failed.result).resolves.toEqual({
|
||||
output: [],
|
||||
diagnostic: expectedFailureDiagnostic(
|
||||
'query-run',
|
||||
'error_during_execution',
|
||||
),
|
||||
stopReason: 'error',
|
||||
})
|
||||
await Promise.all([completed.dispose(), failed.dispose()])
|
||||
|
|
@ -864,26 +918,70 @@ describe('run publication, cancellation, and settlement', () => {
|
|||
const run = await startClaudeCodeRun(request(), fixture.spec)
|
||||
await expect(run.result).resolves.toEqual({
|
||||
output: [],
|
||||
diagnostic: expectedFailureDiagnostic('query-run', 'unknown'),
|
||||
stopReason: 'error',
|
||||
})
|
||||
await run.dispose()
|
||||
})
|
||||
|
||||
it('maps invalid success and missing result to error', async () => {
|
||||
for (const messages of [
|
||||
[success('answer', true)],
|
||||
[success('')],
|
||||
[{ type: 'system', subtype: 'init' } as SDKMessage],
|
||||
]) {
|
||||
it('maps invalid success and missing result to fixed query-run facts', async () => {
|
||||
for (const [messages, category] of [
|
||||
[[success('answer', true)], 'invalid-success'],
|
||||
[[success('')], 'invalid-success'],
|
||||
[[{ type: 'system', subtype: 'init' } as SDKMessage], 'missing-result'],
|
||||
] as const) {
|
||||
const fixture = fakeRun(messages)
|
||||
const run = await startClaudeCodeRun(request(), fixture.spec)
|
||||
await expect(run.result).resolves.toMatchObject({
|
||||
await expect(run.result).resolves.toEqual({
|
||||
output: [],
|
||||
diagnostic: expectedFailureDiagnostic('query-run', category),
|
||||
stopReason: 'error',
|
||||
})
|
||||
await run.dispose()
|
||||
}
|
||||
})
|
||||
|
||||
it('reports an early process exit with independent code and signal facts', async () => {
|
||||
const outcomes: SubprocessOutcome[] = [
|
||||
{ exitCode: 23, signal: null },
|
||||
{ exitCode: null, signal: 'SIGABRT' },
|
||||
{ exitCode: 23, signal: 'SIGABRT' },
|
||||
{ exitCode: null, signal: null },
|
||||
]
|
||||
for (const outcome of outcomes) {
|
||||
const child = fakeChild()
|
||||
async function* stream(): AsyncGenerator<SDKMessage, void> {
|
||||
child.settle(outcome)
|
||||
await Promise.resolve()
|
||||
throw new Error('SECRET_TOKEN from process transport')
|
||||
}
|
||||
queryMock.mockImplementation(({ options }) => {
|
||||
options.spawnClaudeCodeProcess!(sdkSpawnOptions())
|
||||
return Object.assign(stream(), { close: vi.fn() }) as unknown as Query
|
||||
})
|
||||
const run = await startClaudeCodeRun(request(), {
|
||||
cwd: '/workspace',
|
||||
executable: '/native/claude',
|
||||
permissionMode: DEFAULT_CLAUDE_CODE_PERMISSION_MODE,
|
||||
env: {},
|
||||
disposeGraceMs: 5,
|
||||
spawn: () => child.handle,
|
||||
})
|
||||
const result = await run.result
|
||||
expect(result).toEqual({
|
||||
output: [],
|
||||
diagnostic: expectedFailureDiagnostic(
|
||||
'process',
|
||||
'process-exit',
|
||||
outcome,
|
||||
),
|
||||
stopReason: 'error',
|
||||
})
|
||||
expect(result.diagnostic).not.toContain('SECRET_TOKEN')
|
||||
await run.dispose()
|
||||
}
|
||||
})
|
||||
|
||||
it('gives local cancellation precedence and isolates overlapping controllers', async () => {
|
||||
const firstChild = fakeChild()
|
||||
const secondChild = fakeChild()
|
||||
|
|
@ -974,7 +1072,7 @@ describe('run publication, cancellation, and settlement', () => {
|
|||
)
|
||||
await expect(startClaudeCodeRun(request(), {
|
||||
...unused.spec,
|
||||
})).rejects.toThrow('did not publish a controllable')
|
||||
})).rejects.toThrow(expectedFailureDiagnostic('query-start', 'unknown'))
|
||||
expect(noChildClose).toHaveBeenCalledOnce()
|
||||
|
||||
const closeFailure = vi.fn(() => { throw new Error('close boom') })
|
||||
|
|
@ -984,6 +1082,8 @@ describe('run publication, cancellation, and settlement', () => {
|
|||
const noChild = startClaudeCodeRun(request(), {
|
||||
...unused.spec,
|
||||
})
|
||||
await expect(noChild)
|
||||
.rejects.toThrow(expectedFailureDiagnostic('query-start', 'unknown'))
|
||||
await expect(noChild).rejects.toBeInstanceOf(AggregateError)
|
||||
|
||||
const startupAbort = new AbortController()
|
||||
|
|
@ -1006,12 +1106,40 @@ describe('run publication, cancellation, and settlement', () => {
|
|||
expect(abortedClose).toHaveBeenCalledOnce()
|
||||
expect(abortedChild.terminate).toHaveBeenCalledOnce()
|
||||
|
||||
const cleanupAbort = new AbortController()
|
||||
const cleanupFailedChild = fakeChild({
|
||||
waitForExitError: new Error('SECRET_TOKEN cleanup wait failure'),
|
||||
})
|
||||
queryMock.mockImplementationOnce(({ options }) => {
|
||||
options.spawnClaudeCodeProcess!(sdkSpawnOptions())
|
||||
cleanupAbort.abort(new Error('startup cancelled'))
|
||||
return queryFrom([])
|
||||
})
|
||||
const cancelledCleanupFailure = startClaudeCodeRun(
|
||||
request(undefined, cleanupAbort.signal),
|
||||
{
|
||||
...unused.spec,
|
||||
spawn: () => cleanupFailedChild.handle,
|
||||
},
|
||||
)
|
||||
await expect(cancelledCleanupFailure)
|
||||
.rejects.toBeInstanceOf(AggregateError)
|
||||
await expect(cancelledCleanupFailure)
|
||||
.rejects.toThrow(expectedFailureDiagnostic('query-start', 'unknown'))
|
||||
await expect(cancelledCleanupFailure)
|
||||
.rejects.not.toThrow('SECRET_TOKEN')
|
||||
|
||||
queryMock.mockImplementationOnce(() => {
|
||||
throw new Error('query failed before resource creation')
|
||||
})
|
||||
await expect(startClaudeCodeRun(request(), {
|
||||
const queryFailure = startClaudeCodeRun(request(), {
|
||||
...unused.spec,
|
||||
})).rejects.toThrow('query failed before resource creation')
|
||||
})
|
||||
await expect(queryFailure)
|
||||
.rejects.toThrow(expectedFailureDiagnostic('query-start', 'unknown'))
|
||||
await expect(queryFailure).rejects.not.toThrow(
|
||||
'query failed before resource creation',
|
||||
)
|
||||
|
||||
const spawned = fakeChild()
|
||||
const spawnSpecs: SubprocessSpawnSpec[] = []
|
||||
|
|
@ -1019,6 +1147,7 @@ describe('run publication, cancellation, and settlement', () => {
|
|||
queryMock.mockImplementationOnce(({ options }) => {
|
||||
factoryController = options.abortController
|
||||
options.spawnClaudeCodeProcess!(sdkSpawnOptions())
|
||||
spawned.settle({ exitCode: 17, signal: 'SIGABRT' })
|
||||
throw new Error('query construction failed')
|
||||
})
|
||||
const factoryFailure = startClaudeCodeRun(request(), {
|
||||
|
|
@ -1028,7 +1157,12 @@ describe('run publication, cancellation, and settlement', () => {
|
|||
return spawned.handle
|
||||
},
|
||||
})
|
||||
await expect(factoryFailure).rejects.toThrow('query construction failed')
|
||||
await expect(factoryFailure).rejects.toThrow(expectedFailureDiagnostic(
|
||||
'query-start',
|
||||
'unknown',
|
||||
{ exitCode: 17, signal: 'SIGABRT' },
|
||||
))
|
||||
await expect(factoryFailure).rejects.not.toThrow('query construction failed')
|
||||
expect(spawnSpecs).toHaveLength(1)
|
||||
expect(factoryController?.signal.aborted).toBe(true)
|
||||
expect(spawned.terminate).toHaveBeenCalledOnce()
|
||||
|
|
@ -1038,8 +1172,10 @@ describe('run publication, cancellation, and settlement', () => {
|
|||
doneError: new Error('spawn failed'),
|
||||
})
|
||||
const failed = fakeRun([], undefined, failedSpawn)
|
||||
await expect(startClaudeCodeRun(request(), failed.spec))
|
||||
.rejects.toBeInstanceOf(AggregateError)
|
||||
const failedStartup = startClaudeCodeRun(request(), failed.spec)
|
||||
await expect(failedStartup)
|
||||
.rejects.toThrow(expectedFailureDiagnostic('query-start', 'unknown'))
|
||||
await expect(failedStartup).rejects.toBeInstanceOf(AggregateError)
|
||||
expect(failed.close).toHaveBeenCalledOnce()
|
||||
})
|
||||
})
|
||||
|
|
@ -1080,10 +1216,16 @@ describe('query and process disposal', () => {
|
|||
waitForExitError: new Error('wait boom'),
|
||||
})
|
||||
const closeFailure = vi.fn(() => { throw new Error('close boom') })
|
||||
await expect(disposeClaudeCodeChild(
|
||||
const waitAndClose = disposeClaudeCodeChild(
|
||||
{ close: closeFailure },
|
||||
waitFailure.handle,
|
||||
)).rejects.toBeInstanceOf(AggregateError)
|
||||
)
|
||||
await expect(waitAndClose).rejects.toThrow(expectedFailureDiagnostic(
|
||||
'teardown',
|
||||
'unknown',
|
||||
{ exitCode: 0, signal: null },
|
||||
))
|
||||
await expect(waitAndClose).rejects.toBeInstanceOf(AggregateError)
|
||||
expect(waitFailure.terminate).toHaveBeenCalledOnce()
|
||||
|
||||
const doneFailure = fakeChild({
|
||||
|
|
@ -1093,15 +1235,18 @@ describe('query and process disposal', () => {
|
|||
await expect(disposeClaudeCodeChild(
|
||||
{ close: vi.fn() },
|
||||
doneFailure.handle,
|
||||
)).rejects.toThrow('spawn boom')
|
||||
)).rejects.toThrow(expectedFailureDiagnostic('teardown', 'unknown'))
|
||||
|
||||
const both = fakeChild({
|
||||
pid: -1,
|
||||
doneError: new Error('spawn boom'),
|
||||
})
|
||||
await expect(disposeClaudeCodeChild(
|
||||
const bothFailures = disposeClaudeCodeChild(
|
||||
{ close: () => { throw new Error('close boom') } },
|
||||
both.handle,
|
||||
)).rejects.toBeInstanceOf(AggregateError)
|
||||
)
|
||||
await expect(bothFailures)
|
||||
.rejects.toThrow(expectedFailureDiagnostic('teardown', 'unknown'))
|
||||
await expect(bothFailures).rejects.toBeInstanceOf(AggregateError)
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -2,5 +2,5 @@
|
|||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write packages/subagent/subagent/README.md
|
||||
README.md: 161159264ffadf32cc769d65f19caf6d74dc862d
|
||||
README.zh.md: 561206ae56684329ca54b1a524b224a73e4f30b3
|
||||
README.md: f976466610f37a4744c5a2f1e2fff590ea7587b9
|
||||
README.zh.md: bcd4e4cc7f1dc56db7f278f4ff32c148d1d6e872
|
||||
|
|
|
|||
|
|
@ -64,7 +64,7 @@ Both in-process delegation paths fix the child's permission scope at the delegat
|
|||
|
||||
`provider.start(request): Promise<SubagentRun>` is the ownership-transfer boundary; the delegation tool also uses it inside its one-shot Task-backed background path. Before fulfillment, the provider owns setup and must cancel, roll back, and quiesce unpublished resources on every failure. After fulfillment, the caller owns the run and must call `dispose()` on every path; remaining prompt and turn work belongs to `SubagentRun.result`.
|
||||
|
||||
`SubagentRun.result` resolves to `{ output, structured?, diagnostic?, stopReason }`. Child-level failures resolve with a non-`completed` reason; only an infrastructure fault that the seam cannot represent may reject. A provider may add a safe `diagnostic` to a non-completed result after removing tool inputs, file contents, environment values, credentials, and raw protocol payloads and limiting the complete text to 4096 UTF-8 bytes. The field is not assistant output: consumers present it separately, and it does not enter `subagent/end.lastAssistantMessage`. `dispose()` is idempotent, cancels remaining work, and waits for both result settlement and child-resource quiescence. A result rejection remains on `result`; `dispose()` rejects only for an independent resource-release failure. `output` and the `subagent/end` event's `lastAssistantMessage` use the exported `AssistantOutputFold`/`finalAssistantOutput` helpers to select the child's last non-empty assistant message, or its accumulated assistant text when no such message exists. `output` is `[]` and the event field is absent when the child produced neither ([`SubagentResult`](../../../docs/subsystems/subagent.md#the-terminal-result-subagentresult) owns the terminal result contract).
|
||||
`SubagentRun.result` resolves to `{ output, structured?, diagnostic?, stopReason }`. Child-level failures resolve with a non-`completed` reason; only an infrastructure fault that the seam cannot represent may reject. A provider may add a safe `diagnostic` to a non-completed result after removing tool inputs, file contents, environment values, credentials, and raw protocol payloads and limiting the complete text to 4096 UTF-8 bytes. The common result type does not define provider categories or lifecycle stages: an out-of-process provider may derive fixed display text from its version-pinned structured product facts and an observed process outcome, while consumers render that text without parsing it. The field is not assistant output: consumers present it separately, and it does not enter `subagent/end.lastAssistantMessage`. `dispose()` is idempotent, cancels remaining work, and waits for both result settlement and child-resource quiescence. A result rejection remains on `result`; `dispose()` rejects only for an independent resource-release failure. `output` and the `subagent/end` event's `lastAssistantMessage` use the exported `AssistantOutputFold`/`finalAssistantOutput` helpers to select the child's last non-empty assistant message, or its accumulated assistant text when no such message exists. `output` is `[]` and the event field is absent when the child produced neither ([`SubagentResult`](../../../docs/subsystems/subagent.md#the-terminal-result-subagentresult) owns the terminal result contract).
|
||||
|
||||
A local run publishes an ordinary child agent/session before `start()` fulfills, returns that shared session id as `SubagentRun.id`, exposes the exact child as `SubagentRun.localAgent`, records `request.parent.session.id` in the child's `parentSession` header, and appends the resolved descriptor inside its initial turn. Remote providers instead mint a parent-scoped lifecycle id and return `localAgent: undefined`; without a local child session, their one-shot runs are not part of trace-backed enumeration.
|
||||
|
||||
|
|
|
|||
|
|
@ -64,7 +64,7 @@ subagent seam 允许一个 agent(智能体)通过具名提供方把工作委
|
|||
|
||||
`provider.start(request): Promise<SubagentRun>` 是所有权转移边界;委派工具也会在其由 Task 支撑的一次性后台路径中使用它。兑现前,提供方拥有设置过程,并且在任何失败路径上都必须取消、回滚并使尚未发布的资源完全停稳。兑现后,run 的所有权转移给调用方;调用方必须在每条路径上调用 `dispose()`。剩余提示词和轮次工作属于 `SubagentRun.result`。
|
||||
|
||||
`SubagentRun.result` 兑现为 `{ output, structured?, diagnostic?, stopReason }`。子 agent 级失败会以非 `completed` 原因兑现;只有 seam 无法表示的基础设施故障才可以拒绝。提供方可以为非完成结果附加安全的 `diagnostic`:它会先排除工具输入、文件内容、环境值、凭证与原始协议载荷,并把完整文本限制在 4096 个 UTF-8 字节以内。该字段不是 assistant 输出;消费方会将它分开呈现,它也不会进入 `subagent/end.lastAssistantMessage`。`dispose()` 是幂等的,会取消剩余工作,并等待结果结算以及子 agent 资源完全停稳。result 的拒绝只通过 `result` 本身报告;只有独立的资源释放失败,才会使 `dispose()` 被拒绝。`output` 与 `subagent/end` 事件的 `lastAssistantMessage` 使用导出的 `AssistantOutputFold`/`finalAssistantOutput` 辅助函数选取子 agent 最后一条非空 assistant 消息;若没有这类消息,则选取其累积的 assistant 文本。子 agent 两种输出均未产生时,`output` 为 `[]`,该事件字段缺省(终态结果约定归 [`SubagentResult`](../../../docs/subsystems/subagent.md#the-terminal-result-subagentresult) 所有)。
|
||||
`SubagentRun.result` 兑现为 `{ output, structured?, diagnostic?, stopReason }`。子 agent 级失败会以非 `completed` 原因兑现;只有 seam 无法表示的基础设施故障才可以拒绝。提供方可以为非完成结果附加安全的 `diagnostic`:它会先排除工具输入、文件内容、环境值、凭证与原始协议载荷,并把完整文本限制在 4096 个 UTF-8 字节以内。共享结果类型不定义提供方类别或生命周期阶段:进程外提供方可以从锁定版本产品提供的结构化事实与已观测的进程结果派生固定展示文本,而消费方只负责原样呈现,不解析该文本。该字段不是 assistant 输出;消费方会将它分开呈现,它也不会进入 `subagent/end.lastAssistantMessage`。`dispose()` 是幂等的,会取消剩余工作,并等待结果结算以及子 agent 资源完全停稳。result 的拒绝只通过 `result` 本身报告;只有独立的资源释放失败,才会使 `dispose()` 被拒绝。`output` 与 `subagent/end` 事件的 `lastAssistantMessage` 使用导出的 `AssistantOutputFold`/`finalAssistantOutput` 辅助函数选取子 agent 最后一条非空 assistant 消息;若没有这类消息,则选取其累积的 assistant 文本。子 agent 两种输出均未产生时,`output` 为 `[]`,该事件字段缺省(终态结果约定归 [`SubagentResult`](../../../docs/subsystems/subagent.md#the-terminal-result-subagentresult) 所有)。
|
||||
|
||||
本地运行会在 `start()` 兑现前发布普通的子 agent/会话,把该共享会话 id 作为 `SubagentRun.id` 返回,以 `SubagentRun.localAgent` 公开准确的子 agent,把 `request.parent.session.id` 记录到子 agent 的 `parentSession` header,并在其初始轮次内追加已解析的描述符。远程提供方则生成 parent 作用域的生命周期 id,并返回 `localAgent: undefined`;由于没有本地 child 会话,其一次性运行不会进入基于追踪的枚举结果。
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue