deepseek-harness/docs/architecture.md

163 lines
17 KiB
Markdown
Raw Normal View History

# DeepSeek Harness Architecture
This document describes the architecture of the DeepSeek Harness — the foundation of **DeepSeek Code**. The governing principle, from the [microkernel design discussion][microkernel-doc]: **everything is a plugin**. The core is deliberately tiny — a handful of abstract services plus one concrete loop plugin (`dsh-agent-loop`) — and every product feature is a plugin against the extension surface described here, without modifying the loop.
This document covers **behavior**; type shapes live in [core-data-structures/](core-data-structures/core.md), the per-event/service reference in the generated [events](cordis-catalog/events.md) / [services](cordis-catalog/services.md) catalogs, per-package contracts in the package READMEs ([map](../packages/README.md)). Requirement context: [Coding Harness MVP 需求分析][mvp-doc].
[microkernel-doc]: https://trtgsjkv6r.feishu.cn/wiki/VS9Lw1kQki6mDJk2UHocyuphnsc
[mvp-doc]: https://trtgsjkv6r.feishu.cn/wiki/ZwK6wfBE9i91V6kzMGYcgRGanxg
## Layering
```
┌────────────────────────────────────────────────────────────────┐
│ extension + implementation plugins │
│ dsh-agent-loop — THE concrete loop plugin │
│ LLM adapters · executors/backends · model-facing tools │
│ subagent providers · hook bridges · UI bridges │
├────────────────────────────────────────────────────────────────┤
│ interface/service packages (each owns a ctx key + vocabulary) │
│ dsh-agent · dsh-tools · dsh-system-prompt · dsh-session │
│ dsh-llm · dsh-bash · dsh-fs · dsh-web · dsh-compact │
│ dsh-subagent · dsh-session-persistence │
├────────────────────────────────────────────────────────────────┤
│ vendor/: pinned Cordis framework source (cordis, loader, …) │
└────────────────────────────────────────────────────────────────┘
```
Dependency rule: extension plugins depend on interfaces, never on `dsh-agent-loop` (the loop is swappable); the sanctioned exception is the composition bundle `dsh-agent-core`, whose job is assembling the concrete spine ([full rule + generated graph](../packages/README.md#dependencies)).
## Service map
| ctx key | Package | Role |
|---|---|---|
| `ctx.llm` | dsh-llm | adapter registry; `stream()` |
| `ctx.sessions` | dsh-session | creates/holds event-sourced `Session`s |
| `ctx.sessionPersistence` | dsh-session-persistence | durable persistence: create/append/load/list |
| `ctx.systemPrompt` | dsh-system-prompt | ordered sections + tool schemas → `assemble()` |
| `ctx.tools` | dsh-tools | tool definitions; `execute()` through waterfall |
| `ctx.agents` | dsh-agent | live `Agent` handles + create/resume factory (returns `AgentHandle { agent, dispose() }`) |
| `ctx.agentLoop` | dsh-agent-loop | creates and drives `ReactLoopAgent`s |
| `ctx.bash` | dsh-bash | bash execution: foreground runs + background tasks |
| `ctx.fs` | dsh-fs | filesystem provider: read/stream, atomic writes/edits; owns the `fs/*` policy events |
| `ctx.compact` | dsh-compact | compaction: detect pressure, summarize an older range |
| `ctx.web` | dsh-web | search/fetch provider registries + `WebError` taxonomy |
| `ctx.subagents` | dsh-subagent | named provider registry for delegating to child agents |
workflow: dynamic workflows — script-driven multi-agent orchestration A new capability family at packages/workflow/ in the bash seam shape, modeled on Claude Code's dynamic workflows: the model writes a JavaScript orchestration script (export const meta = {...} + plain-JS body), a runtime executes it, and the script — not the conversation — holds the loop, the branching, and the intermediate results. - dsh-workflow (ctx.workflows): abstract WorkflowService + run vocabulary (WorkflowRun whose result NEVER rejects) + observe-only workflow/* events carrying data snapshots (id + meta, never the live run), per-listener contained like subagent/*. - dsh-workflow-vm: in-process node:vm engine. Meta extraction via a string/comment-aware scanner (template interpolation rejected; literal evaluated alone in an empty timed context; statement blanked line- preservingly so stacks keep script line numbers). Hooks: agent(prompt, {label, phase, schema, model}) over ctx.subagents, parallel(), pipeline() (no cross-stage barrier), phase(), log(), args. Fatal-vs-null discipline: hook misuse (unknown/deferred options, bad arguments, unsupported schemas, tripped caps, seam start failures, cancellation) throws fatal WorkflowErrors the combinators RE-THROW — never dissolved into the per-item null reserved for child failures. Realm boundary: inbound values materialized by descriptor walks that never invoke accessors (defineProperty copies, __proto__-safe); outbound values rebuilt in-realm via the context's own JSON.parse. Determinism bans (Date.now/Math.random/argless new Date) kept so future resume support cannot break scripts. Caps and timeouts are validated Config. Every hook promise carries a no-op rejection consumer (app-boot exits on unhandled rejections). - dsh-tool-workflow: the model-facing workflow tool, synchronous like dsh-tool-subagent (start → await → try/finally dispose; abort bridged; non-completed → isError). Generic render card titled by a textual meta.name sniff. The tool description carries the authoring contract. Wired into examples/{coding-agent,acp-agent} with explicit-ask-only guidance. Coverage at every tier: unit (meta scanner, materializer incl. counting-getter and __proto__ regressions, combinator semantics, concurrency ceiling, caps, cancellation, no-unhandled-rejection abandon), integration over the real spawn stack, with-key e2e (real two-phase run + the tool through the registry pipeline), and a recorded ACP snapshot scenario (workflow-run, 1 child session). RFC: docs/rfc/implemented/feature/2026-07-05-dynamic-workflows.md (deferred work explicitly listed). AGENTS.md budget 1575 → 1590 for the new group's layout line.
2026-07-05 13:29:35 +08:00
| `ctx.workflows` | dsh-workflow | script-driven multi-agent orchestration: `start()` runs a workflow script |
All registrations go through `ctx.effect()` and return disposers, so hot-reload and fiber disposal clean up automatically (full service interfaces: the generated [services catalog](cordis-catalog/services.md)).
## Capability seams: interface / implementation / consumer
Swappable capabilities split into three packages — **interface** (abstract service + vocabulary, owns the ctx key), **implementation** (a concrete subclass loaded as a plugin), **consumer** (what the model and plugins program against) — so each evolves independently; the bash trio is the template ([capability seams RFC](rfc/implemented/architecture/2026-06-13-capability-seams.md)). Keep interface + consumer together when they are one concern (the LLM seam: `dsh-llm` carries both, adapters implement); don't split preemptively.
Two seams bend the template deliberately:
- **Filesystem** adds a policy layer as an **event gate**, not a method service: `dsh-tool-fs` (the `read`/`write`/`edit` tools AND executor) dispatches `fs/*` intent events that `dsh-fs-policy` decides, so dropping the policy plugin degrades to the bare provider instead of breaking an injection ([event-gate RFC](rfc/implemented/architecture/2026-06-26-file-context-as-event-gate.md)). Paths resolve against the caller's session cwd, matching bash ([per-session cwd RFC](rfc/implemented/architecture/2026-07-02-fs-per-session-cwd.md)).
- **Web** folds search and fetch onto one seam: `ctx.web` is a provider REGISTRY (`registerSearchProvider`/`registerFetchProvider`, registration-order-independent selection); providers register like LLM adapters, and `dsh-tool-web` is the single consumer owning the tool schemas ([web seam RFC](rfc/implemented/architecture/2026-06-24-web-capability-seam.md)).
> The seam pattern is plain Cordis services + `inject` (a consumer's fiber stays pending until the service exists). Despite the name, `@cordisjs/plugin-capability` is unrelated — a permission-security service (a candidate for the deferred permissions work), not a mechanism for swapping implementations.
## The vocabulary (dsh-llm)
Messages are arrays of typed **content blocks** (`text`, `reasoning`, `tool-call`, `tool-result`); the union derives from the merge-extensible `ContentBlockMap`; the same pattern types `MessageSource`, `FinishReason`, `TurnTrigger`, `TurnEndReason`. The core set is limited to blocks every shipping path honors — multimodal content (images, audio, …) has no core block type; a feature that needs one adds it via the map in the same coordinated change that maps it in the adapters, surfaces it in the UI bridges, and prices it in compaction ([the drop-image RFC](rfc/implemented/simplification/2026-07-04-drop-image-content-block.md)). Streaming is a raw chunk protocol (`block-start` … `finish`) with `BlockAssembler` as the single shared chunk→block assembler; the loop logs raw chunks (replay fidelity) while assembling them. `LlmAdapter` is the provider seam: subclass, implement `stream()`, register via `ctx.llm.registerAdapter(models, adapter)`; `dsh-llm-deepseek` and `dsh-llm-pi-ai` implement the one contract as deliberate design twins ([twin RFC](rfc/implemented/architecture/2026-06-13-twin-llm-adapters.md)). The StreamChunk conventions (usage/finish ordering, raw-string tool arguments, the two sanctioned error paths) are pinned in `dsh-llm/src/types.ts` and [llm-streaming.md](core-data-structures/llm-streaming.md).
## Event-sourced sessions (dsh-session)
A `Session` is an append-only log of typed `SessionEvent`s — the single source of truth. The LLM message history is *derived* (`deriveMessages()`): user/assistant messages, tool results, and envelope-tagged context/steering messages come from their events in chronological order (raw `assistant/chunk` events are replay/UI data, skipped; the per-event mapping is in [session.md](core-data-structures/session.md)). Replay/fork = `ctx.sessions.create(id, { seed })`; trace/telemetry = listen to `session/event` ([event-sourcing RFC](rfc/implemented/architecture/2026-06-11-event-sourced-sessions.md)).
**Durability**: `session/event` is a synchronous notification; persistence backends buffer write-behind and drain at the awaited `session/flush` checkpoint at every turn end. The abstract `SessionPersistence` seam defines create/append/load/list over `SessionEvent` (no parallel persisted type); metadata travels as `SessionHeader`; crash recovery preserves an interrupted turn by closing it with a synthetic `turn/end {interrupted}`. Two backends (JSONL, SQLite) pass one shared contract suite ([persistence RFC](rfc/implemented/architecture/2026-06-14-session-persistence.md), [write coordinator RFC](rfc/implemented/architecture/2026-06-18-shared-persistence-write-coordinator.md)). Resume = `ctx.agents.resume({ resumeSessionId })`.
## Prompt assembly (dsh-system-prompt)
Plugins contribute `PromptSection`s (named, ordered, static or computed) and tool-schema providers; `assemble()` returns `PromptAssembly { sections, tools }` through the `system-prompt/assemble` waterfall. Tool schemas are deliberately part of the assembly — "what the model is told it can do" is one coherent thing — though adapters transmit them as the wire-level `tools` field ([RFC](rfc/implemented/architecture/2026-06-11-tool-schemas-in-prompt-assembly.md)).
## Tool pipeline (dsh-tools)
`ToolRegistry.register()` takes schema + `execute()`; schemas flow into the assembly automatically. `execute()` runs through a two-waterfall pipeline — `tools/pre-execute` (a `PreToolDecision`: allow/deny/ask) → core dispatch → `tools/post-execute` (a `PostToolDecision`: accept/block, replace content, attach context) — the seams where sandbox, permission, hook, and plan-mode plugins live. A thrown tool still reaches `post-execute` as an `isError` result.
## Agents (dsh-agent) and the loop (dsh-agent-loop)
`Agent` is the handle every plugin programs against: `send()` (queued), `steer()` (mid-turn injection, drained between steps), `inject()` (in-session context; a one-shot `injection` turn when idle), `cancel()` (the single public stop primitive: clears queued + steering work, aborts the in-flight step, drops a turn about to start), `whenIdle()` (quiescence observation, not teardown), plus `session`/`status`/`options`. A lifecycle owner tears down via `await AgentHandle.dispose()` — stop, await exit, unregister. Full semantics: [core.md](core-data-structures/core.md), [lifecycle RFC](rfc/implemented/architecture/2026-06-18-agent-lifecycle-and-ownership-seams.md).
**Subagents** are a seam, not a method on `Agent`: `ctx.subagents` is a named-provider registry (`spawn` starts fresh, `fork` seeds the child with the parent's completed-turn prefix, ACP drives an out-of-process child); children are ordinary `Agent`s. See [subagent.md](core-data-structures/subagent.md), [subagent RFC](rfc/implemented/feature/2026-06-21-subagent-capability-seam.md).
### Loop lifecycle (session / turn / step)
- **Session**: the whole event log of one agent.
- **Turn**: ≥1 queued message; steps run until the model stops requesting tools and no plugin requests continuation.
- **Step**: one model request + its tool executions.
```
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
create agent → emit agent/session-start(source) ⟵ once, before turn 1 (startup|resume)
forever:
wait for queued messages (idle)
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
emit agent/status(running)
TURN (error-contained — a throwing plugin ends the turn, never the loop):
Merge worktree-hooks-b-bash-seam into worktree-hooks-c-interception Bring the interception-seams branch onto current master (via A→B). The substantive reconciliation is master's compaction `agent/pre-step` serial seam meeting C's interception seams: - types.ts: keep BOTH master's `agent/pre-step` AND C's new interception events (`agent/prompt-submit`, `agent/session-start`, `agent/turn-continuation`→ `ContinuationDecision`); drop the turn-mirror declarations (removed on A). - loop.ts: the merged per-turn order is `turn/start` → per queued msg `agent/prompt-submit` (rewrite/inject/block) → (fully-blocked ⇒ zero-step `rejected`) → per step: drain steering → assemble system prompt → `agent/pre-step` (compaction, OUTSIDE the step) → `step/start` → single `deriveMessages()` → model → tools/pre-execute·dispatch·post-execute. No turn-mirror emits; `closeTurn()` is the A-simplified single-call form. - Docs (architecture, core.md, agent/agent-loop READMEs, catalog) reconciled to show C's interception seams alongside `agent/pre-step`, no turn/step mirrors. - rfc/README: dropped the stale `proposed/` compaction row (master moved that RFC to implemented/); kept C's new `pre-tool-input-rewrite` proposed row. - interception.spec.ts: migrated its two `agent/turn-end` reason collectors to the `turn/end` session event, and ADDED a cross-test proving a `prompt-submit` rewrite + additionalContext is VISIBLE to an `agent/pre-step` listener on the same turn — pinning the merged seam ordering (compaction sees the post-prompt-submit surface, not stale history).
2026-07-02 04:50:14 +08:00
'turn/start' ⟵ durable turn boundary (no agent/* mirror)
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
each queued msg: waterfall agent/prompt-submit ⟵ allow (rewrite/+context) | block
allow → session('user/message'…); inject additionalContext
every prompt blocked → 'turn/end'(rejected), 0 steps ⟵ zero-step turn, model never called
STEP loop:
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
drain steering (late steering from previous step's listeners)
assembly = ctx.systemPrompt.assemble() ⟵ waterfall system-prompt/assemble
await ctx.serial('agent/pre-step') ⟵ surface mutation (compaction) OUTSIDE the step
session('step/start') ⟵ durable step boundary (no agent/* mirror)
req = {model, system, tools, messages: session.deriveMessages(), signal}
refactor(compact): turn-agnostic retention + dedicated agent/pre-request seam Reform the compaction blueprint so a runaway turn survives and the design stops drifting across review rounds: - Drop in-flight-turn protection ("layer 2"). Retention is a uniform tail→head whole-unit walk; the only structural guard is step-alignment. A single turn that alone exceeds the window now compacts its own early closed steps instead of being retained verbatim (the failure mode that motivated this). - Move auto-compaction off the agent/request waterfall onto a new awaited agent/pre-request loop seam, fired before history derivation. Compaction mutates the surface; the loop derives once from the result — no double-derive, and a listener structurally cannot act on not-yet-derived messages. - Tighten compactIfNeeded to required (session, system, model, signal). - Enforce a single-pass convergence invariant in resolveConfig: reject configs where summarizationMaxTokens + retainTokens exceeds the threshold, so a compaction can never immediately re-trigger. - Document the crash vs recoverable failure taxonomy; core session repair stays compaction-agnostic (a log-only orphaned compact/start is inert). - Wire dsh-compact-basic into examples/coding-agent and add a with-key compaction e2e (compaction's first real-world exercise + runaway net). - Rewrite the RFC to encode the blueprint and move it to implemented/. The runaway-turn snapshot is a named deferred follow-up: dsh-llm-replay cannot yet serve the interleaved summarization model call.
2026-06-26 08:59:33 +08:00
req = waterfall agent/request ⟵ hooks, model switch
stream ctx.llm.stream(req) ⟵ waterfall llm/stream (raw chunks)
refactor(events): remove the agent/stream-chunk mirror of assistant/chunk The loop recorded every model token delta as a durable `assistant/chunk` session event AND emitted an identical live `agent/stream-chunk` Cordis event one line later. Same StreamChunk, same turn/step; the emit added only the live Agent handle, which the sole consumer discarded. This is the boundary-mirror duplication the event-domain work removed for turn/step boundaries, applied to the token stream — a follow-up the boundary RFC explicitly deferred. The premise is settled: chunk persistence is authoritative (the proposal to stop persisting chunks was rejected — replay/snapshots depend on it), so `assistant/chunk` on `session/event` is the load-bearing token stream and `agent/stream-chunk` is pure redundancy. - Remove the `agent/stream-chunk` declaration + emit; drop the now-unused StreamChunk import from dsh-agent's types. - Migrate `dsh-ui-stdio` (the only live consumer; ACP already reads assistant/chunk off session/event) to render assistant/chunk in its existing session/event listener. Consolidating to one listener also makes the inReasoning dim-SGR flag deterministic across chunk/boundary events (they no longer race across two listeners). - Repoint the agent-loop tests (cancel/loop) and ui-stdio tests to the session/event assistant/chunk feed. - New RFC (implemented/simplification/2026-07-02-remove-stream-chunk-mirror); amend the boundary RFC's retained-list entry to cross-link; update architecture, cookbook, event-domain-semantics, the ACP proposal, and the regenerated cordis catalog. Snapshot goldens unchanged (ACP never used the mirror), confirming no editor-facing transcript change.
2026-07-02 23:42:16 +08:00
session('assistant/chunk')
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
if assembler.finish is error/aborted: throw ⟵ adapter's in-band error path →
step error (turn ends error/aborted,
not a normal completed message)
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
msg = waterfall agent/step-result ⟵ runs BEFORE the log append, so the
simplify(session): fold trace-only usage/error events into load-bearing events The session event vocabulary carried two standalone trace-only events that were not load-bearing as separate records. Fold their facts into nearby load-bearing events and delete the standalone variants. - Token usage now rides on `assistant/message` as an optional `usage` field — the assembled model output and its accounting travel together. The loop folds `assembler.usage` onto the append instead of emitting a separate `usage` event. - The max-tokens path is the no-data-loss host: a step cut off with usage but EMPTY content (e.g. only a dropped tool call) previously emitted a standalone `usage`; it now records an empty-content `assistant/message { content: [], usage }`. `deriveMessages()` skips empty-content assistant messages, so the usage host never injects a spurious content-less assistant turn into the provider transcript. A step with neither content nor usage appends nothing. - An operational error's step number now rides on `turn/end.reason` for `kind: 'error'` (`{ kind: 'error', step, message, code? }`) — the durable turn outcome ACP and resume already consume. `failTurn` sets the reason directly (no separate session `error` event). `agent/error` + logging are unchanged for live diagnostics. - No format-version bump: pre-release, no persisted data, so per the format policy there is nothing to migrate or reject (the RFC's "refresh the format version" criterion over-reached). `version` stays 1. - ACP fixtures + goldens re-recorded (keyless replay): dropped standalone usage/error lines, usage folded onto assistant/message, error step on turn/end.reason. RFC moved proposed -> implemented with an implementation note recording the two scope refinements.
2026-06-21 10:00:06 +08:00
session('assistant/message' {content, usage?}) log records what tool dispatch uses
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
each tool-call (sequential, abort-checked between calls):
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
session('tool/call'); ctx.tools.execute() ⟵ waterfall tools/pre-execute (allow/
deny/ask gate) → dispatch → tools/post-execute (accept/block, replace, +context)
tool execution may append tool-owned session events, e.g. `todo/write`
session('tool/result')
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
append buffered post-execute additionalContext → session('context/message')(s)
⟵ after ALL tool/results (adjacency)
2026-07-04 15:36:40 +08:00
drain steering → session('steering/message')
session('step/end') ⟵ durable step boundary (no agent/* mirror)
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
cont = waterfall agent/turn-continuation(default = {action: hadToolCalls||steered
? 'continue' : 'stop'}) → ContinuationDecision
a continue's reason is recorded as next-step steering (same turn); steering pending
also forces continue (continuation OR step/end listeners — the /goal pattern)
if action==stop: break
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
session('turn/end') ⟵ durable turn boundary (no agent/* mirror)
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
await ctx.parallel('session/flush', session) ⟵ durability checkpoint (failure
reported via agent/error, not fatal)
leftover steering re-enqueued as queued messages ⟵ steering is never stranded
emit agent/status(idle) unless more queued
```
Error containment: a throwing listener or broken step ends the **turn** (`turn/end { reason: { kind: 'error', step, … } }`), never the driver loop; live diagnostics fire via `agent/error`; an adapter's in-band error/aborted finish chunk becomes a step error. `cancel()` is honored mid-stream and between tool calls; disposal mid-turn ends the turn `disposed`. A post-`turn/end` failure (a rejecting `session/flush`) is reported via `agent/error` only — the turn stays balanced, the backend keeps its buffer.
A turn ends with one `TurnEndReason` — `completed`, `aborted`, `error`, `disposed`, `max-tokens`, `rejected`, or `interrupted`; per-variant semantics (and the max-tokens-wins rule) are in [session.md § TurnEndReasonMap](core-data-structures/session.md#why-a-turn-ended-turnendreasonmap).
**Turn-enclosure invariant**: every session event lives inside a turn, making the turn the single durability/replay boundary — anything after the last `turn/end` is an interrupted-crash tail. `dsh-invariants` enforces it in dev ([invariant RFC](rfc/implemented/architecture/2026-06-15-turn-enclosure-invariant.md)).
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
### Event taxonomy
The `agent/*` events are declared in `dsh-agent` (so nothing depends on the loop package); each other service declares its own (`tools/*`, `llm/*`, `system-prompt/*`, `session/*`). The full catalog — signatures, dispatch modes, prose — is generated from source and freshness-gated: [cordis-catalog/events.md](cordis-catalog/events.md). Domain semantics (session = the fact log, agent = the live surface): [the event-domain RFC](rfc/implemented/architecture/2026-06-30-event-domain-semantics.md).
### Cordis waterfall semantics (important)
`ctx.waterfall` is **around-middleware**, not a value reducer. Each listener receives `(...args, next)`:
- call `next()` to delegate to later listeners (and ultimately the core behavior), possibly wrapping it;
- return a value **without** calling `next()` to short-circuit (veto);
- listeners run in registration order; `prepend: true` jumps the queue.
Composition caveat: values propagate through `next()`'s **return value** — a listener that returns a *new* object makes earlier listeners' mutations invisible downstream. Prefer mutate-then-`next()` for cooperative middleware; return a replacement only to take over the result.
## Extension guide
Plugin skeletons (tool, hook/permission gate, UI, protocol bridge) and the feature→mechanism map — which extension seam implements each product feature — live in [the extension cookbook](cookbook/extension-cookbook.md); step-by-step guides: [adding a package](cookbook/adding-a-package.md), [a tool](cookbook/adding-a-tool.md), [an LLM adapter](cookbook/adding-an-llm-adapter.md), [a vendored package](cookbook/adding-a-vendored-package.md).
## Deferred work (TODO)
Designed-for but not implemented: inter-agent channels beyond delegation (shared state, streaming output); the model-facing `/compact` consumer tool over `ctx.compact` ([compaction RFC](rfc/implemented/feature/2026-06-18-compaction-capability-seam.md)); parallel tool execution (concurrency-safety hints on `ToolDefinition`); session branching/tree if seed-based forking proves insufficient.