deepseek-harness/docs/module-graph.md

179 lines
5.5 KiB
Markdown
Raw Normal View History

<!-- Generated by scripts/gen-module-graph.ts — do not edit by hand.
Run `pnpm run gen-module-graph` to regenerate. -->
# Module dependency graph
Inter-package dependencies among the `@deepseek-ai/dsh-*` harness packages, derived from each package's `peerDependencies` (the canonical runtime-dependency signal). An edge `a --> b` means package `a` depends on package `b`. Names have the `@deepseek-ai/dsh-` prefix stripped.
```mermaid
graph TD
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes the two gaps in the "brand ids that cross package boundaries" policy and fixes the dependency direction so a capability package never pulls in an unrelated one. - Extract the `Branded<B>` primitive into a new standalone type-only package `@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps. dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session, dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a generic execution backend must not couple to the LLM or session vocabulary). - Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id, the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary that casts SessionId -> OwnerToken. - Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters at the config boundary and the inner create()/resume casts disappear (only the genuinely-new per-run session-id string is cast). - Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store Map keys and public params/exports (SessionStore, AgentRegistry + factory options, the ACP session-id surface + ToolPresenter CallId map, the persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps). - Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point the Branded type-equiv at dsh-brand, fix stale param types in the session/ agent/bash READMEs, regenerate the cordis catalog + module graph. Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
bash --> brand
llm --> brand
bash-local --> bash
fs --> brand
fs --> llm
llm-deepseek --> llm
llm-pi-ai --> llm
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes the two gaps in the "brand ids that cross package boundaries" policy and fixes the dependency direction so a capability package never pulls in an unrelated one. - Extract the `Branded<B>` primitive into a new standalone type-only package `@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps. dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session, dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a generic execution backend must not couple to the LLM or session vocabulary). - Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id, the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary that casts SessionId -> OwnerToken. - Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters at the config boundary and the inner create()/resume casts disappear (only the genuinely-new per-run session-id string is cast). - Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store Map keys and public params/exports (SessionStore, AgentRegistry + factory options, the ACP session-id surface + ToolPresenter CallId map, the persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps). - Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point the Branded type-equiv at dsh-brand, fix stale param types in the session/ agent/bash READMEs, regenerate the cordis catalog + module graph. Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
session --> brand
session --> llm
system-prompt --> llm
web --> llm
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes the two gaps in the "brand ids that cross package boundaries" policy and fixes the dependency direction so a capability package never pulls in an unrelated one. - Extract the `Branded<B>` primitive into a new standalone type-only package `@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps. dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session, dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a generic execution backend must not couple to the LLM or session vocabulary). - Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id, the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary that casts SessionId -> OwnerToken. - Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters at the config boundary and the inner create()/resume casts disappear (only the genuinely-new per-run session-id string is cast). - Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store Map keys and public params/exports (SessionStore, AgentRegistry + factory options, the ACP session-id surface + ToolPresenter CallId map, the persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps). - Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point the Branded type-equiv at dsh-brand, fix stale param types in the session/ agent/bash READMEs, regenerate the cordis catalog + module graph. Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
agent --> brand
agent --> llm
agent --> session
compact --> llm
compact --> session
fs-local --> fs
fix(fs): address review — rename to dsh-fs-policy, fs/*-intent events, RFC currency, ENOTDIR Rename per review naming decisions: - package dsh-file-context → dsh-fs-policy (dir, package name, plugin name, tsconfig refs, importers, type-equiv manifest, generated catalog + module-graph) - events fs/write-expectation → fs/write-intent, fs/edit-expectation → fs/edit-intent (fs/observed unchanged); type FsWriteExpectation → FsWriteIntent, "expectation" wording → "intent" throughout - exported FileContextExec → FsPolicyExec Make the implemented RFCs describe what shipped, not the superseded designs: the 2026-06-17 capability-seam + tool-schemas RFCs no longer place policy on ctx.fs or use full/partial-view authorization, and the fsspec RFC's ctx.fileContext service prose is rewritten to the fs/* event-gate reality (freshness-based auth). Sharpen docs/rfc/implemented/AGENTS.md: a rename is a fact to fix IN PLACE — the "new RFC" escape hatch is for macro decision reversals only, not renames. Code fixes from review: - fsio.ts resolveLocalTarget/probe translate ENOTDIR (a parent path segment is a file) into the structured FsError taxonomy instead of leaking a raw Node error; resolve reports FS_NOT_FOUND, probe reports absent. Regression tests proven to fail on the unfixed code. - tool-fs HMR test now asserts prompt sections (not just tool schemas) are withdrawn on disposal. - fs/observed is a plain (unguarded) ctx.emit: correct the fs-policy comment, filesystem.md, and tool-fs module doc that wrongly claimed the tool "contains" a throwing listener; a throw surfaces as the tool's isError result. - drop the false "loaded by the default product config" claim (no config wires the fs tools yet), the duplicate ctx.bash service-map row, the stale FileReadRequest catalog link-map entry, and the fs/fs README EOF blank line; correct the dsh-fs package.json description.
2026-07-02 03:12:38 +08:00
fs-policy --> fs
feat(hooks): dsh-hook-protocol — shared Claude Code / Codex hook wire-protocol core The two hook bridges (dsh-hooks-claude, dsh-hooks-codex) would otherwise duplicate the bulk of the protocol — Codex deliberately reimplements a SUBSET of the Claude Code protocol (same hooks.json shape, exit-code/stdout contract, command-hook model). This library holds the genuinely-identical primitives; each bridge owns only what differs (per-event stdin payload, env/substitution, decision mapping). New packages/hooks/ group; hook-protocol is a LIBRARY (no plugin, registers/injects nothing): - matcher: matchesMatcher(pattern, query, mode) — the one dialect axis collapsed to a mode param (claude = literal-or-regex with pipe alternation; codex = always unanchored regex). Match-all on absent/''/'*'; invalid regex matches nothing. - codec: parseHookOutput(exit, stdout, stderr) → dialect-neutral HookOutput. Exit 0 → lenient JSON; exit 2 → blocking error (stderr = reason, surfaced as decision:'block'); other → non-blocking. Parses the CC superset (continue/stopReason/decision/hookSpecificOutput.{permissionDecision, additionalContext,updatedInput}/systemMessage); permissionDecision overrides the legacy top-level decision. - runner: runHook(bash, hook, opts, now) — runs a command hook via ctx.bash (stdin payload + trusted-plugin env), honors timeoutSec, never throws (executor reject → non-blocking-error HookOutput). Injected clock for testable durations. - merge: mergeHookOutputs — most-restrictive fold (deny>ask>allow, sticky stop, block reasons joined, context/system-messages accumulated). - hook/* session events (declaration-merged into SessionEventMap, log-only like compact/*) + appendHookInvoked/appendHookResult helpers. updatedInput is parsed but NOT honored (deferred pre-tool-input-rewrite RFC); a bridge logs+warns. 47 unit tests at per-file 100% (matcher per-mode, codec per exit-code/field, runner plumbing w/ stub executor, merge precedence, hook/* helpers). RFC: implemented/feature/2026-06-30-hook-protocol-lib.md.
2026-07-01 00:38:06 +08:00
hook-protocol --> bash
hook-protocol --> session
refactor(examples): extract reusable logic into tested packages Logic that lived under examples/ was outside the per-file 100% coverage gate (examples/ are not workspaces) and, in the stdio-UI case, duplicated across two examples. Move it into packages/ so it is gated and de-duped. - packages/ui-stdio (new): unify the two diverged stdio-chat.ts copies into one @deepseek-ai/dsh-ui-stdio plugin (welcome/agent Config). A test-only I/O seam (createStdioChat(ctx, config, runtime)) keeps process streams out of the serializable config and makes every render/EOF/disposal branch unit-testable. Per-file 100%. echo/coding cordis.yml now load the package; both src/stdio-chat.ts deleted. - packages/llm-replay (new): move examples/acp-agent/src/llm-replay.ts (+ its spec) here so its derive/parse/replay branches fall under the coverage gate. cordis.snapshot.yml + README rewired to the package name; added apply/env /assertNever/abort tests to reach per-file 100%. - examples/{echo,coding}-agent: keyless Loader-path e2e smokes that boot the real cordis.yml (no key) — the guard a hand-mounted unit test cannot be for the unwrapExports/export-shape class (postmortem 0001). examples/AGENTS.md codifies the keyless+with-key smoke convention (keyless-by-nature exception for echo-agent). - AGENTS.md: a scoped, removal-triggered pre-release stance (foundation over blast radius). packages/README.md: new rows + a FIXME to later regroup ALL packages into a hierarchy. Wiring: tsconfig paths/refs, publint, knip, module-graph. Verified: typecheck, lint, test:coverage (887 tests, 100%), build, hygiene, doc-sync, test:snapshot (10), test:e2e (6 keyless pass, with-key self-skip).
2026-06-19 12:42:28 +08:00
llm-replay --> llm
llm-replay --> session
session-persistence --> session
web-fetch-local --> web
web-search-deepseek --> web
web-search-exa --> web
web-search-perplexity --> web
compact-basic --> agent
compact-basic --> compact
compact-basic --> llm
compact-basic --> session
invariants --> agent
invariants --> llm
invariants --> session
session-persistence-jsonl --> session
session-persistence-jsonl --> session-persistence
session-persistence-sqlite --> session
session-persistence-sqlite --> session-persistence
tools --> agent
tools --> llm
tools --> system-prompt
feat(acp): tool-owned tool-call UI presentation (title/command/output) In Zed the tool-call card showed only "bash" — the bare tool name — instead of what the command does. Fix it by letting each TOOL own how its calls render, rather than the bridge special-casing names. dsh-tools: add an optional two-state presentation seam to ToolDefinition / defineTool — `presentCall(args)` (pending: title, kind, rawInput) and `presentResult(args, result)` (completed: title?, content?). Provider-neutral `ToolCallKind`/`ToolCallPresentation`/`ToolResultPresentation` vocabulary so tools never depend on ACP. defineTool soft-validates args (display runs on log replay, so a malformed/old shape returns undefined instead of throwing). dsh-tool-bash: bash declares presentCall (model `description` → title, exact `command` → rawInput, kind execute) and presentResult (wrap output in a fenced ```console block — a UI-only affordance kept out of the model-facing result); bash_output/bash_kill present task-scoped titles. dsh-acp: inject `tools`; a per-session `ToolPresenter` looks the tool up by name and maps its neutral presentation to the ACP tool_call/tool_call_update wire shape, with a generic fallback (title = name) for tools that declare nothing. Because the `tool/result` event carries only {callId, content, isError}, the presenter keeps a small bridge-local map of ONLY in-flight calls' (name, args), keyed by callId and removed as each result is presented — no event-schema or core change. Replay uses a throwaway presenter so loaded sessions render identically to live ones. Tests: dsh-tools defineTool presenters (typed args, soft-validate), tool-bash bash/bash_output/bash_kill presenters, acp ToolPresenter (tool-owned mapping, unknown-callId fallback, in-flight-only map), and an end-to-end turn through the bridge. The key-gated e2e now asserts a real bash call's title is the model description (not "bash") and rawInput is the command — verified against the real DeepSeek model. The test harness derives its inject from the bridge's exported `inject` so it can't drift again.
2026-06-18 09:01:36 +08:00
acp --> agent
acp --> llm
acp --> session
acp --> session-persistence
acp --> tools
agent-loop --> agent
agent-loop --> llm
agent-loop --> session
agent-loop --> session-persistence
agent-loop --> system-prompt
agent-loop --> tools
feat(hooks): dsh-hooks-claude + dsh-hooks-codex bridges (hooks stack PR-F) The two bridge plugins that run a user's existing Claude Code / Codex hook config on the harness's typed interception seams, built on the shared dsh-hook-protocol library. A bridge is a faithfulness adapter, not a power tool: anything it does a native cordis plugin does more powerfully — the bridge exists only to run UNMODIFIED external hooks. - dsh-hooks-claude: CC dialect. Seven hook points (SessionStart, UserPromptSubmit, PreToolUse, PostToolUse, Stop, SubagentStart, SubagentStop), CC per-event stdin payloads, env + ${CLAUDE_PLUGIN_ROOT}/ ${CLAUDE_PROJECT_DIR} substitution, literal-or-regex matcher. - dsh-hooks-codex: Codex dialect — a deliberate subset. Five hook points, always-regex matcher, snake_case payloads (turn_id/model, no trailing newline), no env/substitution, block-only decisions. Both map the neutral merged outcome onto the seam's typed Decision and stamp an explicit {kind:'plugin'} source on injected context (so it is never mislabeled as a user prompt). Config parse-failure is contained; only command hooks run. updatedInput is logged+warned (input rewrite deferred); the Stop loop-guard is deferred (TODO). Tests: per-file 100% — config-parse unit branches + per-seam mappings end-to-end through the REAL loop + REAL bash + REAL shell scripts (scripted mock model only) + a real-Loader export-shape guard. A keyless ACP snapshot scenario (hook-prompt-block) proves a UserPromptSubmit hook blocks a prompt end-to-end (rejected turn -> ACP cancelled, hook/* events in the log); a with-key e2e (hooks.e2e.ts) proves a PreToolUse hook blocks real bash (verified on disk). The snapshot normalizer now scrubs hook/result.durationMs. RFC: docs/rfc/implemented/feature/2026-06-30-hook-bridges.md
2026-07-01 04:22:00 +08:00
hooks-codex --> agent
hooks-codex --> hook-protocol
hooks-codex --> llm
hooks-codex --> session
hooks-codex --> tools
2026-06-21 22:31:56 +08:00
subagent --> agent
subagent --> llm
subagent --> tools
tool-bash --> agent
tool-bash --> bash
tool-bash --> llm
tool-bash --> tools
tool-fs --> fs
tool-fs --> llm
tool-fs --> session
tool-fs --> system-prompt
tool-fs --> tools
tool-todo --> agent
tool-todo --> session
tool-todo --> tools
tool-web --> llm
tool-web --> system-prompt
tool-web --> tools
tool-web --> web
refactor(examples): extract the app spine into dsh-agent-core + app packages Implements docs/rfc/.../2026-06-20-extract-example-app-packages.md. Each example was thick — a hand-rolled start.ts, an infra preamble, nested base.yml/base-core.yml/acp-tail.yml includes, and a coupled front-door cluster enforced only by prose. This moves the composition into packages so each example is a thin leaf cordis.yml: pick the swappable backends, load one app package. New packages: - @deepseek-ai/dsh-agent-core (packages/core/agent-core): one bundle plugin that loads the providerless/executor-less/UI-less spine (timer + llm + sessions + system-prompt + tools + agents + invariants + tool-bash + agent-loop) via ctx.plugin(...) inside apply(), and forwards agent-loop's `agents` list as its own Config (export const Config = AgentLoop.Config, default []). - @deepseek-ai/dsh-stdio-agent (packages/ui/stdio-agent): terminal chat APP — agent-core + console logger + readline UI + a pre-created `main` agent, with a bin. The demo:echo/coding front door. - @deepseek-ai/dsh-acp-agent (packages/ui/acp-agent): ACP server APP — agent-core + JSONL persistence + the acp bridge, NO stdout logger, with a bin. The stdout-purity footgun is structurally unreachable from the leaf. Amendment to the RFC: hmr stays a LEAF cordis.yml entry, not baked into dsh-stdio-agent. hmr is a Loader-only dev plugin (throws without --expose-internals; the in-process test tier can't even import its decorator form), so a package statically importing it could never carry the per-file coverage gate. Unlike the console logger, a stray hmr is not a stdout-purity footgun, so leaving it at the leaf costs no safety. With hmr out, all three new packages carry in-process unit specs at 100%. Boot glue (Loader tail, .env load, snapshot-mode selection, stdin-dispose lifecycle) moves into each app's bin; start.ts and base.yml/base-core.yml/ acp-tail.yml are deleted. Each app package gets a keyless real-load-path test that boots through its bin + the cordis Loader (guarding the unwrapExports export-shape bug class, postmortem 0001). ACP snapshot replay stays green against the existing committed goldens (pure boot restructuring). RFC moved proposed->implemented with the amendment recorded; package/example/architecture docs and the module graph updated.
2026-06-21 12:03:44 +08:00
agent-core --> agent
agent-core --> agent-loop
agent-core --> invariants
agent-core --> llm
agent-core --> session
agent-core --> system-prompt
agent-core --> tool-bash
agent-core --> tools
feat(hooks): dsh-hooks-claude + dsh-hooks-codex bridges (hooks stack PR-F) The two bridge plugins that run a user's existing Claude Code / Codex hook config on the harness's typed interception seams, built on the shared dsh-hook-protocol library. A bridge is a faithfulness adapter, not a power tool: anything it does a native cordis plugin does more powerfully — the bridge exists only to run UNMODIFIED external hooks. - dsh-hooks-claude: CC dialect. Seven hook points (SessionStart, UserPromptSubmit, PreToolUse, PostToolUse, Stop, SubagentStart, SubagentStop), CC per-event stdin payloads, env + ${CLAUDE_PLUGIN_ROOT}/ ${CLAUDE_PROJECT_DIR} substitution, literal-or-regex matcher. - dsh-hooks-codex: Codex dialect — a deliberate subset. Five hook points, always-regex matcher, snake_case payloads (turn_id/model, no trailing newline), no env/substitution, block-only decisions. Both map the neutral merged outcome onto the seam's typed Decision and stamp an explicit {kind:'plugin'} source on injected context (so it is never mislabeled as a user prompt). Config parse-failure is contained; only command hooks run. updatedInput is logged+warned (input rewrite deferred); the Stop loop-guard is deferred (TODO). Tests: per-file 100% — config-parse unit branches + per-seam mappings end-to-end through the REAL loop + REAL bash + REAL shell scripts (scripted mock model only) + a real-Loader export-shape guard. A keyless ACP snapshot scenario (hook-prompt-block) proves a UserPromptSubmit hook blocks a prompt end-to-end (rejected turn -> ACP cancelled, hook/* events in the log); a with-key e2e (hooks.e2e.ts) proves a PreToolUse hook blocks real bash (verified on disk). The snapshot normalizer now scrubs hook/result.durationMs. RFC: docs/rfc/implemented/feature/2026-06-30-hook-bridges.md
2026-07-01 04:22:00 +08:00
hooks-claude --> agent
hooks-claude --> hook-protocol
hooks-claude --> llm
hooks-claude --> session
hooks-claude --> subagent
hooks-claude --> tools
Add the ACP subagent backend: out-of-process delegation (PR3) The first OUT-OF-PROCESS subagent backend, proving the seam generalizes past the in-process backends. @deepseek-ai/dsh-subagent-acp runs each child agent in a spawned subprocess, driven over the Agent Client Protocol as the CLIENT — the direction-inverted twin of the dsh-acp server bridge. Point the configured command at the acp-agent example and the harness talks to its own process. - Fresh process per run: start spawns, runs one ACP session (initialize → newSession → prompt), dispose kills the subprocess and awaits its exit. - Minimal client stub: advertises no fs/terminal; accumulates agent_message_chunk text as the result output; auto-answers session/request_permission by a configured policy (reject default / allow). No start-time capabilities (an out-of-process child can't enforce the parent's depth/tool-filter); ignores request.parent; injects only `subagents`. - StopReason mapping (end_turn→completed, cancelled→aborted, …); result resolves error/aborted on a child failure, never rejects (seam contract). - Security: credential-shaped ambient env vars are scrubbed; the child's own key is forwarded only via explicit config.env. A spawn-level error (ENOENT) is captured and raced against the ACP drive so a bad command settles error rather than crashing the parent. Testing designed at every tier: keyless integration drives a scripted mock ACP server subprocess (cancellation incl. the pre-newSession race and a torn-pipe-after-cancel, permission auto-answer, non-message updates, spawn failure, HMR, export shape) at 100% coverage; a with-key e2e drives the REAL acp-agent example process (PONG + real file write, verified on disk) — the harness driving itself. Snapshot coverage of an ACP child is deferred as TODO(acp-subagent-replay) (each child is its own process with its own replay). Stayed on @agentclientprotocol/sdk 0.25.1: the proposed 0.28.x bump only deprecates the stable ClientSideConnection/AgentSideConnection API this layer uses (33 sites incl. the server bridge), turning no-deprecated red across code this PR shouldn't rewrite — that fluent-API migration is its own follow-up. The backend needs nothing 0.28.x adds. This completes the subagent seam stack (PR1 interface → PR2 in-process → PR2.5 snapshot infra → PR3 ACP); the seam RFC moves to implemented/, amended.
2026-06-22 10:47:02 +08:00
subagent-acp --> agent
subagent-acp --> llm
subagent-acp --> subagent
subagent-inprocess --> agent
subagent-inprocess --> llm
subagent-inprocess --> session
subagent-inprocess --> subagent
2026-06-21 22:31:56 +08:00
subagent-mock --> agent
subagent-mock --> llm
subagent-mock --> subagent
tool-subagent --> agent
tool-subagent --> llm
tool-subagent --> subagent
tool-subagent --> tools
refactor(examples): extract the app spine into dsh-agent-core + app packages Implements docs/rfc/.../2026-06-20-extract-example-app-packages.md. Each example was thick — a hand-rolled start.ts, an infra preamble, nested base.yml/base-core.yml/acp-tail.yml includes, and a coupled front-door cluster enforced only by prose. This moves the composition into packages so each example is a thin leaf cordis.yml: pick the swappable backends, load one app package. New packages: - @deepseek-ai/dsh-agent-core (packages/core/agent-core): one bundle plugin that loads the providerless/executor-less/UI-less spine (timer + llm + sessions + system-prompt + tools + agents + invariants + tool-bash + agent-loop) via ctx.plugin(...) inside apply(), and forwards agent-loop's `agents` list as its own Config (export const Config = AgentLoop.Config, default []). - @deepseek-ai/dsh-stdio-agent (packages/ui/stdio-agent): terminal chat APP — agent-core + console logger + readline UI + a pre-created `main` agent, with a bin. The demo:echo/coding front door. - @deepseek-ai/dsh-acp-agent (packages/ui/acp-agent): ACP server APP — agent-core + JSONL persistence + the acp bridge, NO stdout logger, with a bin. The stdout-purity footgun is structurally unreachable from the leaf. Amendment to the RFC: hmr stays a LEAF cordis.yml entry, not baked into dsh-stdio-agent. hmr is a Loader-only dev plugin (throws without --expose-internals; the in-process test tier can't even import its decorator form), so a package statically importing it could never carry the per-file coverage gate. Unlike the console logger, a stray hmr is not a stdout-purity footgun, so leaving it at the leaf costs no safety. With hmr out, all three new packages carry in-process unit specs at 100%. Boot glue (Loader tail, .env load, snapshot-mode selection, stdin-dispose lifecycle) moves into each app's bin; start.ts and base.yml/base-core.yml/ acp-tail.yml are deleted. Each app package gets a keyless real-load-path test that boots through its bin + the cordis Loader (guarding the unwrapExports export-shape bug class, postmortem 0001). ACP snapshot replay stays green against the existing committed goldens (pure boot restructuring). RFC moved proposed->implemented with the amendment recorded; package/example/architecture docs and the module graph updated.
2026-06-21 12:03:44 +08:00
acp-agent --> acp
acp-agent --> agent-core
acp-agent --> app-boot
refactor(examples): extract the app spine into dsh-agent-core + app packages Implements docs/rfc/.../2026-06-20-extract-example-app-packages.md. Each example was thick — a hand-rolled start.ts, an infra preamble, nested base.yml/base-core.yml/acp-tail.yml includes, and a coupled front-door cluster enforced only by prose. This moves the composition into packages so each example is a thin leaf cordis.yml: pick the swappable backends, load one app package. New packages: - @deepseek-ai/dsh-agent-core (packages/core/agent-core): one bundle plugin that loads the providerless/executor-less/UI-less spine (timer + llm + sessions + system-prompt + tools + agents + invariants + tool-bash + agent-loop) via ctx.plugin(...) inside apply(), and forwards agent-loop's `agents` list as its own Config (export const Config = AgentLoop.Config, default []). - @deepseek-ai/dsh-stdio-agent (packages/ui/stdio-agent): terminal chat APP — agent-core + console logger + readline UI + a pre-created `main` agent, with a bin. The demo:echo/coding front door. - @deepseek-ai/dsh-acp-agent (packages/ui/acp-agent): ACP server APP — agent-core + JSONL persistence + the acp bridge, NO stdout logger, with a bin. The stdout-purity footgun is structurally unreachable from the leaf. Amendment to the RFC: hmr stays a LEAF cordis.yml entry, not baked into dsh-stdio-agent. hmr is a Loader-only dev plugin (throws without --expose-internals; the in-process test tier can't even import its decorator form), so a package statically importing it could never carry the per-file coverage gate. Unlike the console logger, a stray hmr is not a stdout-purity footgun, so leaving it at the leaf costs no safety. With hmr out, all three new packages carry in-process unit specs at 100%. Boot glue (Loader tail, .env load, snapshot-mode selection, stdin-dispose lifecycle) moves into each app's bin; start.ts and base.yml/base-core.yml/ acp-tail.yml are deleted. Each app package gets a keyless real-load-path test that boots through its bin + the cordis Loader (guarding the unwrapExports export-shape bug class, postmortem 0001). ACP snapshot replay stays green against the existing committed goldens (pure boot restructuring). RFC moved proposed->implemented with the amendment recorded; package/example/architecture docs and the module graph updated.
2026-06-21 12:03:44 +08:00
acp-agent --> session-persistence-jsonl
stdio-agent --> agent
stdio-agent --> agent-core
stdio-agent --> app-boot
stdio-agent --> llm
refactor(examples): extract the app spine into dsh-agent-core + app packages Implements docs/rfc/.../2026-06-20-extract-example-app-packages.md. Each example was thick — a hand-rolled start.ts, an infra preamble, nested base.yml/base-core.yml/acp-tail.yml includes, and a coupled front-door cluster enforced only by prose. This moves the composition into packages so each example is a thin leaf cordis.yml: pick the swappable backends, load one app package. New packages: - @deepseek-ai/dsh-agent-core (packages/core/agent-core): one bundle plugin that loads the providerless/executor-less/UI-less spine (timer + llm + sessions + system-prompt + tools + agents + invariants + tool-bash + agent-loop) via ctx.plugin(...) inside apply(), and forwards agent-loop's `agents` list as its own Config (export const Config = AgentLoop.Config, default []). - @deepseek-ai/dsh-stdio-agent (packages/ui/stdio-agent): terminal chat APP — agent-core + console logger + readline UI + a pre-created `main` agent, with a bin. The demo:echo/coding front door. - @deepseek-ai/dsh-acp-agent (packages/ui/acp-agent): ACP server APP — agent-core + JSONL persistence + the acp bridge, NO stdout logger, with a bin. The stdout-purity footgun is structurally unreachable from the leaf. Amendment to the RFC: hmr stays a LEAF cordis.yml entry, not baked into dsh-stdio-agent. hmr is a Loader-only dev plugin (throws without --expose-internals; the in-process test tier can't even import its decorator form), so a package statically importing it could never carry the per-file coverage gate. Unlike the console logger, a stray hmr is not a stdout-purity footgun, so leaving it at the leaf costs no safety. With hmr out, all three new packages carry in-process unit specs at 100%. Boot glue (Loader tail, .env load, snapshot-mode selection, stdin-dispose lifecycle) moves into each app's bin; start.ts and base.yml/base-core.yml/ acp-tail.yml are deleted. Each app package gets a keyless real-load-path test that boots through its bin + the cordis Loader (guarding the unwrapExports export-shape bug class, postmortem 0001). ACP snapshot replay stays green against the existing committed goldens (pure boot restructuring). RFC moved proposed->implemented with the amendment recorded; package/example/architecture docs and the module graph updated.
2026-06-21 12:03:44 +08:00
stdio-agent --> session
stdio-agent --> session-persistence-jsonl
Add in-process subagent backends: spawn (fresh) and fork (seeded) The second PR of the subagent seam: the two in-process backends that run a child agent on the same cordis context, reusing the agent factory's quiescent AgentHandle teardown. Both register on ctx.subagents (PR1's named-provider registry) and share one run driver. - dsh-subagent-spawn: a FRESH child via ctx.agents.create — own session, the parent's model by default (overridable), zero inherited conversation. Also exports the shared in-process run driver (startInProcessRun): mint ids, stamp cwd/parentSession-lineage/depth, drive the one-shot (send → whenIdle), read the last assistant/message + turn/end reason, dispose to quiescence. - dsh-subagent-fork: a child SEEDED with the parent's balanced completed-turn prefix (the log up to and including its last turn/end), so the child inherits context. The in-flight unbalanced turn is excluded — a raw seed would fail the invariants replay. Proven: a regression test goes red if the boundary seeds the open turn. - Seam extension: CreateAgentOptions.seed, threaded through AgentLoop.createAgent → ctx.sessions.prepare({ seed }) (the primitive resume already used). This is the fork-lineage path the TODO(sub-agents) markers anticipated. - Depth: a merge-extensible AgentOptions.subagentDepth (0 top-level, parent+1 for a child); the depthLimit capability refuses a spawn past request.maxDepth. Tests: real-loop unit tests for both backends (mock MODEL only, real loop + invariants), a multi-subagent test (one parent drives a fork AND a spawn child then keeps working), and a with-key e2e (a real parent delegates via the `subagent` tool to a real child that writes a file on disk — world-verified). 100% per-file coverage. The coding-agent demo wires the spawn backend + tool. Snapshot coverage of nested agents is deferred to a stacked follow-up (TODO(subagent-snapshots)): dsh-llm-replay is a single global positional cursor that cannot route calls to a parent vs. a child on one context. Recorded in the RFC's deferrals and a new AGENTS.md rule: designing a subsystem must design its test infrastructure END TO END up front, verifying the snapshot/e2e harness can express the new shape — a gap this plan hit.
2026-06-22 05:58:40 +08:00
subagent-fork --> agent
subagent-fork --> session
subagent-fork --> subagent
subagent-fork --> subagent-inprocess
subagent-spawn --> subagent
subagent-spawn --> subagent-inprocess
```
| Package | Depends on |
| --- | --- |
| `app-boot` | — |
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes the two gaps in the "brand ids that cross package boundaries" policy and fixes the dependency direction so a capability package never pulls in an unrelated one. - Extract the `Branded<B>` primitive into a new standalone type-only package `@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps. dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session, dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a generic execution backend must not couple to the LLM or session vocabulary). - Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id, the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary that casts SessionId -> OwnerToken. - Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters at the config boundary and the inner create()/resume casts disappear (only the genuinely-new per-run session-id string is cast). - Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store Map keys and public params/exports (SessionStore, AgentRegistry + factory options, the ACP session-id surface + ToolPresenter CallId map, the persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps). - Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point the Branded type-equiv at dsh-brand, fix stale param types in the session/ agent/bash READMEs, regenerate the cordis catalog + module graph. Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
| `brand` | — |
| `bash` | `brand` |
| `llm` | `brand` |
| `bash-local` | `bash` |
| `fs` | `brand`, `llm` |
| `llm-deepseek` | `llm` |
| `llm-pi-ai` | `llm` |
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes the two gaps in the "brand ids that cross package boundaries" policy and fixes the dependency direction so a capability package never pulls in an unrelated one. - Extract the `Branded<B>` primitive into a new standalone type-only package `@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps. dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session, dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a generic execution backend must not couple to the LLM or session vocabulary). - Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id, the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary that casts SessionId -> OwnerToken. - Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters at the config boundary and the inner create()/resume casts disappear (only the genuinely-new per-run session-id string is cast). - Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store Map keys and public params/exports (SessionStore, AgentRegistry + factory options, the ACP session-id surface + ToolPresenter CallId map, the persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps). - Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point the Branded type-equiv at dsh-brand, fix stale param types in the session/ agent/bash READMEs, regenerate the cordis catalog + module graph. Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
| `session` | `brand`, `llm` |
| `system-prompt` | `llm` |
| `web` | `llm` |
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes the two gaps in the "brand ids that cross package boundaries" policy and fixes the dependency direction so a capability package never pulls in an unrelated one. - Extract the `Branded<B>` primitive into a new standalone type-only package `@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps. dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session, dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a generic execution backend must not couple to the LLM or session vocabulary). - Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id, the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary that casts SessionId -> OwnerToken. - Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters at the config boundary and the inner create()/resume casts disappear (only the genuinely-new per-run session-id string is cast). - Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store Map keys and public params/exports (SessionStore, AgentRegistry + factory options, the ACP session-id surface + ToolPresenter CallId map, the persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps). - Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point the Branded type-equiv at dsh-brand, fix stale param types in the session/ agent/bash READMEs, regenerate the cordis catalog + module graph. Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
| `agent` | `brand`, `llm`, `session` |
| `compact` | `llm`, `session` |
| `fs-local` | `fs` |
fix(fs): address review — rename to dsh-fs-policy, fs/*-intent events, RFC currency, ENOTDIR Rename per review naming decisions: - package dsh-file-context → dsh-fs-policy (dir, package name, plugin name, tsconfig refs, importers, type-equiv manifest, generated catalog + module-graph) - events fs/write-expectation → fs/write-intent, fs/edit-expectation → fs/edit-intent (fs/observed unchanged); type FsWriteExpectation → FsWriteIntent, "expectation" wording → "intent" throughout - exported FileContextExec → FsPolicyExec Make the implemented RFCs describe what shipped, not the superseded designs: the 2026-06-17 capability-seam + tool-schemas RFCs no longer place policy on ctx.fs or use full/partial-view authorization, and the fsspec RFC's ctx.fileContext service prose is rewritten to the fs/* event-gate reality (freshness-based auth). Sharpen docs/rfc/implemented/AGENTS.md: a rename is a fact to fix IN PLACE — the "new RFC" escape hatch is for macro decision reversals only, not renames. Code fixes from review: - fsio.ts resolveLocalTarget/probe translate ENOTDIR (a parent path segment is a file) into the structured FsError taxonomy instead of leaking a raw Node error; resolve reports FS_NOT_FOUND, probe reports absent. Regression tests proven to fail on the unfixed code. - tool-fs HMR test now asserts prompt sections (not just tool schemas) are withdrawn on disposal. - fs/observed is a plain (unguarded) ctx.emit: correct the fs-policy comment, filesystem.md, and tool-fs module doc that wrongly claimed the tool "contains" a throwing listener; a throw surfaces as the tool's isError result. - drop the false "loaded by the default product config" claim (no config wires the fs tools yet), the duplicate ctx.bash service-map row, the stale FileReadRequest catalog link-map entry, and the fs/fs README EOF blank line; correct the dsh-fs package.json description.
2026-07-02 03:12:38 +08:00
| `fs-policy` | `fs` |
feat(hooks): dsh-hook-protocol — shared Claude Code / Codex hook wire-protocol core The two hook bridges (dsh-hooks-claude, dsh-hooks-codex) would otherwise duplicate the bulk of the protocol — Codex deliberately reimplements a SUBSET of the Claude Code protocol (same hooks.json shape, exit-code/stdout contract, command-hook model). This library holds the genuinely-identical primitives; each bridge owns only what differs (per-event stdin payload, env/substitution, decision mapping). New packages/hooks/ group; hook-protocol is a LIBRARY (no plugin, registers/injects nothing): - matcher: matchesMatcher(pattern, query, mode) — the one dialect axis collapsed to a mode param (claude = literal-or-regex with pipe alternation; codex = always unanchored regex). Match-all on absent/''/'*'; invalid regex matches nothing. - codec: parseHookOutput(exit, stdout, stderr) → dialect-neutral HookOutput. Exit 0 → lenient JSON; exit 2 → blocking error (stderr = reason, surfaced as decision:'block'); other → non-blocking. Parses the CC superset (continue/stopReason/decision/hookSpecificOutput.{permissionDecision, additionalContext,updatedInput}/systemMessage); permissionDecision overrides the legacy top-level decision. - runner: runHook(bash, hook, opts, now) — runs a command hook via ctx.bash (stdin payload + trusted-plugin env), honors timeoutSec, never throws (executor reject → non-blocking-error HookOutput). Injected clock for testable durations. - merge: mergeHookOutputs — most-restrictive fold (deny>ask>allow, sticky stop, block reasons joined, context/system-messages accumulated). - hook/* session events (declaration-merged into SessionEventMap, log-only like compact/*) + appendHookInvoked/appendHookResult helpers. updatedInput is parsed but NOT honored (deferred pre-tool-input-rewrite RFC); a bridge logs+warns. 47 unit tests at per-file 100% (matcher per-mode, codec per exit-code/field, runner plumbing w/ stub executor, merge precedence, hook/* helpers). RFC: implemented/feature/2026-06-30-hook-protocol-lib.md.
2026-07-01 00:38:06 +08:00
| `hook-protocol` | `bash`, `session` |
refactor(examples): extract reusable logic into tested packages Logic that lived under examples/ was outside the per-file 100% coverage gate (examples/ are not workspaces) and, in the stdio-UI case, duplicated across two examples. Move it into packages/ so it is gated and de-duped. - packages/ui-stdio (new): unify the two diverged stdio-chat.ts copies into one @deepseek-ai/dsh-ui-stdio plugin (welcome/agent Config). A test-only I/O seam (createStdioChat(ctx, config, runtime)) keeps process streams out of the serializable config and makes every render/EOF/disposal branch unit-testable. Per-file 100%. echo/coding cordis.yml now load the package; both src/stdio-chat.ts deleted. - packages/llm-replay (new): move examples/acp-agent/src/llm-replay.ts (+ its spec) here so its derive/parse/replay branches fall under the coverage gate. cordis.snapshot.yml + README rewired to the package name; added apply/env /assertNever/abort tests to reach per-file 100%. - examples/{echo,coding}-agent: keyless Loader-path e2e smokes that boot the real cordis.yml (no key) — the guard a hand-mounted unit test cannot be for the unwrapExports/export-shape class (postmortem 0001). examples/AGENTS.md codifies the keyless+with-key smoke convention (keyless-by-nature exception for echo-agent). - AGENTS.md: a scoped, removal-triggered pre-release stance (foundation over blast radius). packages/README.md: new rows + a FIXME to later regroup ALL packages into a hierarchy. Wiring: tsconfig paths/refs, publint, knip, module-graph. Verified: typecheck, lint, test:coverage (887 tests, 100%), build, hygiene, doc-sync, test:snapshot (10), test:e2e (6 keyless pass, with-key self-skip).
2026-06-19 12:42:28 +08:00
| `llm-replay` | `llm`, `session` |
| `session-persistence` | `session` |
| `web-fetch-local` | `web` |
| `web-search-deepseek` | `web` |
| `web-search-exa` | `web` |
| `web-search-perplexity` | `web` |
| `compact-basic` | `agent`, `compact`, `llm`, `session` |
| `invariants` | `agent`, `llm`, `session` |
| `session-persistence-jsonl` | `session`, `session-persistence` |
| `session-persistence-sqlite` | `session`, `session-persistence` |
| `tools` | `agent`, `llm`, `system-prompt` |
feat(acp): tool-owned tool-call UI presentation (title/command/output) In Zed the tool-call card showed only "bash" — the bare tool name — instead of what the command does. Fix it by letting each TOOL own how its calls render, rather than the bridge special-casing names. dsh-tools: add an optional two-state presentation seam to ToolDefinition / defineTool — `presentCall(args)` (pending: title, kind, rawInput) and `presentResult(args, result)` (completed: title?, content?). Provider-neutral `ToolCallKind`/`ToolCallPresentation`/`ToolResultPresentation` vocabulary so tools never depend on ACP. defineTool soft-validates args (display runs on log replay, so a malformed/old shape returns undefined instead of throwing). dsh-tool-bash: bash declares presentCall (model `description` → title, exact `command` → rawInput, kind execute) and presentResult (wrap output in a fenced ```console block — a UI-only affordance kept out of the model-facing result); bash_output/bash_kill present task-scoped titles. dsh-acp: inject `tools`; a per-session `ToolPresenter` looks the tool up by name and maps its neutral presentation to the ACP tool_call/tool_call_update wire shape, with a generic fallback (title = name) for tools that declare nothing. Because the `tool/result` event carries only {callId, content, isError}, the presenter keeps a small bridge-local map of ONLY in-flight calls' (name, args), keyed by callId and removed as each result is presented — no event-schema or core change. Replay uses a throwaway presenter so loaded sessions render identically to live ones. Tests: dsh-tools defineTool presenters (typed args, soft-validate), tool-bash bash/bash_output/bash_kill presenters, acp ToolPresenter (tool-owned mapping, unknown-callId fallback, in-flight-only map), and an end-to-end turn through the bridge. The key-gated e2e now asserts a real bash call's title is the model description (not "bash") and rawInput is the command — verified against the real DeepSeek model. The test harness derives its inject from the bridge's exported `inject` so it can't drift again.
2026-06-18 09:01:36 +08:00
| `acp` | `agent`, `llm`, `session`, `session-persistence`, `tools` |
| `agent-loop` | `agent`, `llm`, `session`, `session-persistence`, `system-prompt`, `tools` |
feat(hooks): dsh-hooks-claude + dsh-hooks-codex bridges (hooks stack PR-F) The two bridge plugins that run a user's existing Claude Code / Codex hook config on the harness's typed interception seams, built on the shared dsh-hook-protocol library. A bridge is a faithfulness adapter, not a power tool: anything it does a native cordis plugin does more powerfully — the bridge exists only to run UNMODIFIED external hooks. - dsh-hooks-claude: CC dialect. Seven hook points (SessionStart, UserPromptSubmit, PreToolUse, PostToolUse, Stop, SubagentStart, SubagentStop), CC per-event stdin payloads, env + ${CLAUDE_PLUGIN_ROOT}/ ${CLAUDE_PROJECT_DIR} substitution, literal-or-regex matcher. - dsh-hooks-codex: Codex dialect — a deliberate subset. Five hook points, always-regex matcher, snake_case payloads (turn_id/model, no trailing newline), no env/substitution, block-only decisions. Both map the neutral merged outcome onto the seam's typed Decision and stamp an explicit {kind:'plugin'} source on injected context (so it is never mislabeled as a user prompt). Config parse-failure is contained; only command hooks run. updatedInput is logged+warned (input rewrite deferred); the Stop loop-guard is deferred (TODO). Tests: per-file 100% — config-parse unit branches + per-seam mappings end-to-end through the REAL loop + REAL bash + REAL shell scripts (scripted mock model only) + a real-Loader export-shape guard. A keyless ACP snapshot scenario (hook-prompt-block) proves a UserPromptSubmit hook blocks a prompt end-to-end (rejected turn -> ACP cancelled, hook/* events in the log); a with-key e2e (hooks.e2e.ts) proves a PreToolUse hook blocks real bash (verified on disk). The snapshot normalizer now scrubs hook/result.durationMs. RFC: docs/rfc/implemented/feature/2026-06-30-hook-bridges.md
2026-07-01 04:22:00 +08:00
| `hooks-codex` | `agent`, `hook-protocol`, `llm`, `session`, `tools` |
2026-06-21 22:31:56 +08:00
| `subagent` | `agent`, `llm`, `tools` |
| `tool-bash` | `agent`, `bash`, `llm`, `tools` |
| `tool-fs` | `fs`, `llm`, `session`, `system-prompt`, `tools` |
| `tool-todo` | `agent`, `session`, `tools` |
| `tool-web` | `llm`, `system-prompt`, `tools`, `web` |
refactor(examples): extract the app spine into dsh-agent-core + app packages Implements docs/rfc/.../2026-06-20-extract-example-app-packages.md. Each example was thick — a hand-rolled start.ts, an infra preamble, nested base.yml/base-core.yml/acp-tail.yml includes, and a coupled front-door cluster enforced only by prose. This moves the composition into packages so each example is a thin leaf cordis.yml: pick the swappable backends, load one app package. New packages: - @deepseek-ai/dsh-agent-core (packages/core/agent-core): one bundle plugin that loads the providerless/executor-less/UI-less spine (timer + llm + sessions + system-prompt + tools + agents + invariants + tool-bash + agent-loop) via ctx.plugin(...) inside apply(), and forwards agent-loop's `agents` list as its own Config (export const Config = AgentLoop.Config, default []). - @deepseek-ai/dsh-stdio-agent (packages/ui/stdio-agent): terminal chat APP — agent-core + console logger + readline UI + a pre-created `main` agent, with a bin. The demo:echo/coding front door. - @deepseek-ai/dsh-acp-agent (packages/ui/acp-agent): ACP server APP — agent-core + JSONL persistence + the acp bridge, NO stdout logger, with a bin. The stdout-purity footgun is structurally unreachable from the leaf. Amendment to the RFC: hmr stays a LEAF cordis.yml entry, not baked into dsh-stdio-agent. hmr is a Loader-only dev plugin (throws without --expose-internals; the in-process test tier can't even import its decorator form), so a package statically importing it could never carry the per-file coverage gate. Unlike the console logger, a stray hmr is not a stdout-purity footgun, so leaving it at the leaf costs no safety. With hmr out, all three new packages carry in-process unit specs at 100%. Boot glue (Loader tail, .env load, snapshot-mode selection, stdin-dispose lifecycle) moves into each app's bin; start.ts and base.yml/base-core.yml/ acp-tail.yml are deleted. Each app package gets a keyless real-load-path test that boots through its bin + the cordis Loader (guarding the unwrapExports export-shape bug class, postmortem 0001). ACP snapshot replay stays green against the existing committed goldens (pure boot restructuring). RFC moved proposed->implemented with the amendment recorded; package/example/architecture docs and the module graph updated.
2026-06-21 12:03:44 +08:00
| `agent-core` | `agent`, `agent-loop`, `invariants`, `llm`, `session`, `system-prompt`, `tool-bash`, `tools` |
feat(hooks): dsh-hooks-claude + dsh-hooks-codex bridges (hooks stack PR-F) The two bridge plugins that run a user's existing Claude Code / Codex hook config on the harness's typed interception seams, built on the shared dsh-hook-protocol library. A bridge is a faithfulness adapter, not a power tool: anything it does a native cordis plugin does more powerfully — the bridge exists only to run UNMODIFIED external hooks. - dsh-hooks-claude: CC dialect. Seven hook points (SessionStart, UserPromptSubmit, PreToolUse, PostToolUse, Stop, SubagentStart, SubagentStop), CC per-event stdin payloads, env + ${CLAUDE_PLUGIN_ROOT}/ ${CLAUDE_PROJECT_DIR} substitution, literal-or-regex matcher. - dsh-hooks-codex: Codex dialect — a deliberate subset. Five hook points, always-regex matcher, snake_case payloads (turn_id/model, no trailing newline), no env/substitution, block-only decisions. Both map the neutral merged outcome onto the seam's typed Decision and stamp an explicit {kind:'plugin'} source on injected context (so it is never mislabeled as a user prompt). Config parse-failure is contained; only command hooks run. updatedInput is logged+warned (input rewrite deferred); the Stop loop-guard is deferred (TODO). Tests: per-file 100% — config-parse unit branches + per-seam mappings end-to-end through the REAL loop + REAL bash + REAL shell scripts (scripted mock model only) + a real-Loader export-shape guard. A keyless ACP snapshot scenario (hook-prompt-block) proves a UserPromptSubmit hook blocks a prompt end-to-end (rejected turn -> ACP cancelled, hook/* events in the log); a with-key e2e (hooks.e2e.ts) proves a PreToolUse hook blocks real bash (verified on disk). The snapshot normalizer now scrubs hook/result.durationMs. RFC: docs/rfc/implemented/feature/2026-06-30-hook-bridges.md
2026-07-01 04:22:00 +08:00
| `hooks-claude` | `agent`, `hook-protocol`, `llm`, `session`, `subagent`, `tools` |
Add the ACP subagent backend: out-of-process delegation (PR3) The first OUT-OF-PROCESS subagent backend, proving the seam generalizes past the in-process backends. @deepseek-ai/dsh-subagent-acp runs each child agent in a spawned subprocess, driven over the Agent Client Protocol as the CLIENT — the direction-inverted twin of the dsh-acp server bridge. Point the configured command at the acp-agent example and the harness talks to its own process. - Fresh process per run: start spawns, runs one ACP session (initialize → newSession → prompt), dispose kills the subprocess and awaits its exit. - Minimal client stub: advertises no fs/terminal; accumulates agent_message_chunk text as the result output; auto-answers session/request_permission by a configured policy (reject default / allow). No start-time capabilities (an out-of-process child can't enforce the parent's depth/tool-filter); ignores request.parent; injects only `subagents`. - StopReason mapping (end_turn→completed, cancelled→aborted, …); result resolves error/aborted on a child failure, never rejects (seam contract). - Security: credential-shaped ambient env vars are scrubbed; the child's own key is forwarded only via explicit config.env. A spawn-level error (ENOENT) is captured and raced against the ACP drive so a bad command settles error rather than crashing the parent. Testing designed at every tier: keyless integration drives a scripted mock ACP server subprocess (cancellation incl. the pre-newSession race and a torn-pipe-after-cancel, permission auto-answer, non-message updates, spawn failure, HMR, export shape) at 100% coverage; a with-key e2e drives the REAL acp-agent example process (PONG + real file write, verified on disk) — the harness driving itself. Snapshot coverage of an ACP child is deferred as TODO(acp-subagent-replay) (each child is its own process with its own replay). Stayed on @agentclientprotocol/sdk 0.25.1: the proposed 0.28.x bump only deprecates the stable ClientSideConnection/AgentSideConnection API this layer uses (33 sites incl. the server bridge), turning no-deprecated red across code this PR shouldn't rewrite — that fluent-API migration is its own follow-up. The backend needs nothing 0.28.x adds. This completes the subagent seam stack (PR1 interface → PR2 in-process → PR2.5 snapshot infra → PR3 ACP); the seam RFC moves to implemented/, amended.
2026-06-22 10:47:02 +08:00
| `subagent-acp` | `agent`, `llm`, `subagent` |
| `subagent-inprocess` | `agent`, `llm`, `session`, `subagent` |
2026-06-21 22:31:56 +08:00
| `subagent-mock` | `agent`, `llm`, `subagent` |
| `tool-subagent` | `agent`, `llm`, `subagent`, `tools` |
| `acp-agent` | `acp`, `agent-core`, `app-boot`, `session-persistence-jsonl` |
| `stdio-agent` | `agent`, `agent-core`, `app-boot`, `llm`, `session`, `session-persistence-jsonl` |
| `subagent-fork` | `agent`, `session`, `subagent`, `subagent-inprocess` |
| `subagent-spawn` | `subagent`, `subagent-inprocess` |