deepseek-harness/packages/core/agent-loop/src/loop.ts

935 lines
49 KiB
TypeScript
Raw Normal View History

Document the codebase thoroughly and tighten type safety Docs: per-folder README.md for packages/ (family overview + one per package: service, events, API, extension points, TODOs), examples/, and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md symlinks) for packages/ and vendor/; module-level doc comments in every packages/*/src file; richer JSDoc on all exported API (event side effects, disposal contracts, error behavior). Root AGENTS.md gains a "Type Safety and Documentation" policy section: the codebase aims to be very type-safe and well documented; type gymnastics are acceptable in core packages when they improve plugin-author DX; verbose docs are fine as long as they stay strictly in sync with the code. Type safety: removed the upstream-inherited "noImplicitAny": false from tsconfig.base.json — packages/* now compile under full strict mode; vendor/loader and vendor/include set it locally (vendor/cordis already did). Eliminated every `: any` / `as any` from packages and examples (catch clauses use unknown + a CodedError narrowing type; event data access uses discriminated-union narrowing). Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL — SchemaSpec with per-property `required: true` booleans, type-level InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and defineTool() so first-party tools get typed execute(args) with zero casts (raw JSON Schema still accepted for MCP interop; chosen over schemastery because it targets JSON Schema generation directly). echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
/**
* The agent loop driver: one `runLoop()` invocation drives one agent for its
* whole lifetime. Error-contained at the turn level — a throwing plugin ends
* the turn, never kills the loop. See the JSDoc on `runLoop()` for the full
* lifecycle pseudo-code.
*
* @module dsh-agent-loop/loop
*/
import type { Context } from 'cordis'
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
import type { FinishReason, GenerateOptions, LlmCallConfig, Message } from '@deepseek-ai/dsh-llm'
import { BlockAssembler, HarnessError, deepFreeze } from '@deepseek-ai/dsh-llm'
import type { ContinuationDecision, HookContext, PromptDecision } from '@deepseek-ai/dsh-agent'
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
import { canonicalHeader } from '@deepseek-ai/dsh-session'
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
import type { Session, TurnEndReason, TurnTrigger } from '@deepseek-ai/dsh-session'
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
import { createTransmissionLog, recordRequestHeader } from './request-log.ts'
import type { TransmissionLog } from './request-log.ts'
import { renderPrompt } from '@deepseek-ai/dsh-system-prompt'
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
import type { PromptAssembly } from '@deepseek-ai/dsh-system-prompt'
import type {} from '@deepseek-ai/dsh-tools'
import type { ReactLoopAgent } from './agent.ts'
Document the codebase thoroughly and tighten type safety Docs: per-folder README.md for packages/ (family overview + one per package: service, events, API, extension points, TODOs), examples/, and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md symlinks) for packages/ and vendor/; module-level doc comments in every packages/*/src file; richer JSDoc on all exported API (event side effects, disposal contracts, error behavior). Root AGENTS.md gains a "Type Safety and Documentation" policy section: the codebase aims to be very type-safe and well documented; type gymnastics are acceptable in core packages when they improve plugin-author DX; verbose docs are fine as long as they stay strictly in sync with the code. Type safety: removed the upstream-inherited "noImplicitAny": false from tsconfig.base.json — packages/* now compile under full strict mode; vendor/loader and vendor/include set it locally (vendor/cordis already did). Eliminated every `: any` / `as any` from packages and examples (catch clauses use unknown + a CodedError narrowing type; event data access uses discriminated-union narrowing). Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL — SchemaSpec with per-property `required: true` booleans, type-level InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and defineTool() so first-party tools get typed execute(args) with zero casts (raw JSON Schema still accepted for MCP interop; chosen over schemastery because it targets JSON Schema generation directly). echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
/** An Error with an optional machine-readable code (e.g., from LlmError or a throwing plugin). */
type CodedError = Error & { code?: string }
/**
* Normalize an arbitrary thrown value into a coded Error. A real Error passes
* through (its `code`, if any, is preserved by {@link errorData}); a non-Error
* throw is wrapped in a {@link HarnessError} with code `UNKNOWN` and the
* original value chained as `cause`, so a bad throw still carries a routable
* code instead of degrading to a bare message.
*/
function toError(error: unknown): CodedError {
return error instanceof Error ? error : new HarnessError(String(error), 'UNKNOWN', { cause: error })
}
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
/**
* Map a model-call {@link FinishReason} to the step error it should raise, or
* `undefined` when the step completed normally.
*
* Adapters report provider/transport failures one of two sanctioned ways (see
* the StreamChunk contract in dsh-llm): throw from `stream()` (handled by the
* caller's try/catch), OR end the stream with a finish-error/aborted chunk
* (the only option for adapters that can't throw mid-stream, e.g.
* library-backed ones). This translates the latter into a thrown step error
* so the turn ends error/aborted (the failure recorded on `turn/end.reason`),
* never as a normal `completed` assistant message.
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
*
* `FinishReason` is merge-extensible (plugins/adapters can add `kind`s), so
* the switch handles the known terminal-failure kinds and treats every other
* kind — `stop`, `tool-calls`, `max-tokens`, future additions — as success.
*/
function finishError(finish: FinishReason): CodedError | undefined {
switch (finish.kind) {
case 'error': {
const error: CodedError = new Error(finish.message)
if (finish.code !== undefined) error.code = finish.code
return error
}
case 'aborted': {
const error: CodedError = new Error('model stream aborted')
error.code = 'ABORTED'
return error
}
// stop / tool-calls / max-tokens / plugin-added kinds → not a failure.
default:
return undefined
}
}
/**
* Build the `{ message, code? }` part of an error payload, omitting the
* `code` key entirely when absent (exactOptionalPropertyTypes-correct).
*/
function errorData(err: CodedError): { message: string; code?: string } {
return { message: err.message, ...typeof err.code === 'string' ? { code: err.code } : {} }
}
/**
* The turn-end contribution of a step's *successful* finish, or `undefined`
* when the step finished ordinarily (a plain `completed`).
*
* {@link finishError} has already converted `error`/`aborted` finishes into
* thrown step errors, so the finishes that reach here are `stop`,
* `tool-calls`, `max-tokens`, or a future merge-extensible kind. Only
* `max-tokens` carries forward as a distinct {@link TurnEndReason}: a step that
* hit the output-token ceiling ended the turn cut-short rather than by the
* model's choice. `stop`/`tool-calls`/unknown kinds contribute nothing beyond
* the default `completed`. {@link runTurn} applies this with the rule "any
* `max-tokens` step in the turn makes the turn end `max-tokens`".
*/
function stepFinishReason(finish: FinishReason): TurnEndReason | undefined {
switch (finish.kind) {
case 'max-tokens':
return { kind: 'max-tokens' }
// stop / tool-calls / plugin-added kinds → no turn-end contribution
// beyond the default `completed`. FinishReason is merge-extensible, so a
// default (not assertNever) handles unknown kinds as ordinary success.
default:
return undefined
}
}
Document the codebase thoroughly and tighten type safety Docs: per-folder README.md for packages/ (family overview + one per package: service, events, API, extension points, TODOs), examples/, and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md symlinks) for packages/ and vendor/; module-level doc comments in every packages/*/src file; richer JSDoc on all exported API (event side effects, disposal contracts, error behavior). Root AGENTS.md gains a "Type Safety and Documentation" policy section: the codebase aims to be very type-safe and well documented; type gymnastics are acceptable in core packages when they improve plugin-author DX; verbose docs are fine as long as they stay strictly in sync with the code. Type safety: removed the upstream-inherited "noImplicitAny": false from tsconfig.base.json — packages/* now compile under full strict mode; vendor/loader and vendor/include set it locally (vendor/cordis already did). Eliminated every `: any` / `as any` from packages and examples (catch clauses use unknown + a CodedError narrowing type; event data access uses discriminated-union narrowing). Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL — SchemaSpec with per-property `required: true` booleans, type-level InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and defineTool() so first-party tools get typed execute(args) with zero casts (raw JSON Schema still accepted for MCP interop; chosen over schemastery because it targets JSON Schema generation directly). echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
/**
* Ambient handles the loop driver receives from the agent. Decouples the
* pure function `runLoop` from the mutable ReactLoopAgent fields, making the
Document the codebase thoroughly and tighten type safety Docs: per-folder README.md for packages/ (family overview + one per package: service, events, API, extension points, TODOs), examples/, and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md symlinks) for packages/ and vendor/; module-level doc comments in every packages/*/src file; richer JSDoc on all exported API (event side effects, disposal contracts, error behavior). Root AGENTS.md gains a "Type Safety and Documentation" policy section: the codebase aims to be very type-safe and well documented; type gymnastics are acceptable in core packages when they improve plugin-author DX; verbose docs are fine as long as they stay strictly in sync with the code. Type safety: removed the upstream-inherited "noImplicitAny": false from tsconfig.base.json — packages/* now compile under full strict mode; vendor/loader and vendor/include set it locally (vendor/cordis already did). Eliminated every `: any` / `as any` from packages and examples (catch clauses use unknown + a CodedError narrowing type; event data access uses discriminated-union narrowing). Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL — SchemaSpec with per-property `required: true` booleans, type-level InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and defineTool() so first-party tools get typed execute(args) with zero casts (raw JSON Schema still accepted for MCP interop; chosen over schemastery because it targets JSON Schema generation directly). echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
* loop testable without a real agent.
*/
export interface LoopHandle {
setStatus(status: 'idle' | 'running'): void
setAbort(controller: AbortController | undefined): void
/** Resolves when the agent is disposed — unblocks the idle wait. */
disposed: Promise<void>
isDisposed(): boolean
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
/**
* Whether a `cancel()` is pending for the current turn. The driver checks this
* at every decision point where a turn could start or continue (right after
* the idle wait, after the `running` flip, before each step, and at the
* continuation gate) and drops the about-to-run / continuing turn. Reset once
* per loop iteration via {@link clearCancel} after the turn returns, so the
* marker governs exactly one cancellation and never leaks to a later prompt.
*/
isCancelled(): boolean
/**
* The resolved reason for the pending cancel (`reason ?? 'cancelled'`), read
* by the marker branches (pre-step / continuation) so a turn dropped where no
* `AbortController` carries the reason still records the caller's
* `cancel(reason)` value — matching the mid-step abort path. Only meaningful
* when {@link isCancelled} is true.
*/
cancelReason(): string
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
/** Clear the cancel marker (called once per iteration after the turn returns). */
clearCancel(): void
/**
* Settle pending `whenIdle()` waiters WITHOUT a status transition. Used by the
* pre-step cancel-skip path: it drops the about-to-run turn and re-parks at the
* idle wait, so no `running→idle` transition fires to settle a `whenIdle()`
* waiter that was registered in the pre-step window — this settles it directly
* (it emits no `agent/status`, so an ACP `agent/status` listener never sees a
* spurious idle that would resolve a freshly-queued prompt as cancelled).
*/
settleIdle(): void
}
/**
* The agent loop. One invocation drives one agent for its whole lifetime:
*
* ```
fix(events): address Codex review — protect post-execute result, purge stale tools/execute refs Codex's PR-C review found two (A) blockers: - tools/post-execute could corrupt the protected outcome. postExecute passed the mutable `result` to listeners and then read result.callId / spread result on the return paths, so a listener mutating the reference (flipping isError, rewriting callId, injecting an error) escaped the decision channel. Now the authoritative callId/isError/error are SNAPSHOT before the waterfall and the return value is rebuilt from the snapshot + the typed PostToolDecision — the decision is the only sanctioned way to change the outcome, and callId is always exec.callId. Added a regression test that mutates the result reference and asserts it has no effect; proven to fail red on the unfixed code. - Public docs/JSDoc still advertised the removed `tools/execute` waterfall after the split. Swept every current-state reference to tools/pre-execute + tools/post-execute: the ToolRegistry class JSDoc (and the regenerated catalog), loop.ts's ASCII flow (also added the prompt-submit/session-start steps it was missing), the package-map READMEs (packages, core, agent-core), core-data-structures core.md/tools.md, the bash + acp + invariants src/READMEs (the deferred permission gate is the tools/pre-execute deny/ask seam now), the cookbook, and the implemented RFCs whose factual seam catalog drifted. codec.ts's totality prose now lists `rejected`. Proposed-RFC references are left as-is (frozen proposals, validated when built).
2026-06-30 20:25:30 +08:00
* create agent → emit agent/session-start(source) ⟵ once, before turn 1
* forever:
* wait for queued messages (idle)
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
* TURN (error-contained — a throwing plugin ends the turn, never the loop):
Merge worktree-hooks-b-bash-seam into worktree-hooks-c-interception Bring the interception-seams branch onto current master (via A→B). The substantive reconciliation is master's compaction `agent/pre-step` serial seam meeting C's interception seams: - types.ts: keep BOTH master's `agent/pre-step` AND C's new interception events (`agent/prompt-submit`, `agent/session-start`, `agent/turn-continuation`→ `ContinuationDecision`); drop the turn-mirror declarations (removed on A). - loop.ts: the merged per-turn order is `turn/start` → per queued msg `agent/prompt-submit` (rewrite/inject/block) → (fully-blocked ⇒ zero-step `rejected`) → per step: drain steering → assemble system prompt → `agent/pre-step` (compaction, OUTSIDE the step) → `step/start` → single `deriveMessages()` → model → tools/pre-execute·dispatch·post-execute. No turn-mirror emits; `closeTurn()` is the A-simplified single-call form. - Docs (architecture, core.md, agent/agent-loop READMEs, catalog) reconciled to show C's interception seams alongside `agent/pre-step`, no turn/step mirrors. - rfc/README: dropped the stale `proposed/` compaction row (master moved that RFC to implemented/); kept C's new `pre-tool-input-rewrite` proposed row. - interception.spec.ts: migrated its two `agent/turn-end` reason collectors to the `turn/end` session event, and ADDED a cross-test proving a `prompt-submit` rewrite + additionalContext is VISIBLE to an `agent/pre-step` listener on the same turn — pinning the merged seam ordering (compaction sees the post-prompt-submit surface, not stale history).
2026-07-02 04:50:14 +08:00
* 'turn/start'; each queued msg: waterfall agent/prompt-submit ⟵ durable turn boundary (no agent/* mirror)
fix(events): address Codex review — protect post-execute result, purge stale tools/execute refs Codex's PR-C review found two (A) blockers: - tools/post-execute could corrupt the protected outcome. postExecute passed the mutable `result` to listeners and then read result.callId / spread result on the return paths, so a listener mutating the reference (flipping isError, rewriting callId, injecting an error) escaped the decision channel. Now the authoritative callId/isError/error are SNAPSHOT before the waterfall and the return value is rebuilt from the snapshot + the typed PostToolDecision — the decision is the only sanctioned way to change the outcome, and callId is always exec.callId. Added a regression test that mutates the result reference and asserts it has no effect; proven to fail red on the unfixed code. - Public docs/JSDoc still advertised the removed `tools/execute` waterfall after the split. Swept every current-state reference to tools/pre-execute + tools/post-execute: the ToolRegistry class JSDoc (and the regenerated catalog), loop.ts's ASCII flow (also added the prompt-submit/session-start steps it was missing), the package-map READMEs (packages, core, agent-core), core-data-structures core.md/tools.md, the bash + acp + invariants src/READMEs (the deferred permission gate is the tools/pre-execute deny/ask seam now), the cookbook, and the implemented RFCs whose factual seam catalog drifted. codec.ts's totality prose now lists `rejected`. Proposed-RFC references are left as-is (frozen proposals, validated when built).
2026-06-30 20:25:30 +08:00
* allow → session('user/message'…) (+ inject additionalContext) | block → drop
Merge worktree-hooks-b-bash-seam into worktree-hooks-c-interception Bring the interception-seams branch onto current master (via A→B). The substantive reconciliation is master's compaction `agent/pre-step` serial seam meeting C's interception seams: - types.ts: keep BOTH master's `agent/pre-step` AND C's new interception events (`agent/prompt-submit`, `agent/session-start`, `agent/turn-continuation`→ `ContinuationDecision`); drop the turn-mirror declarations (removed on A). - loop.ts: the merged per-turn order is `turn/start` → per queued msg `agent/prompt-submit` (rewrite/inject/block) → (fully-blocked ⇒ zero-step `rejected`) → per step: drain steering → assemble system prompt → `agent/pre-step` (compaction, OUTSIDE the step) → `step/start` → single `deriveMessages()` → model → tools/pre-execute·dispatch·post-execute. No turn-mirror emits; `closeTurn()` is the A-simplified single-call form. - Docs (architecture, core.md, agent/agent-loop READMEs, catalog) reconciled to show C's interception seams alongside `agent/pre-step`, no turn/step mirrors. - rfc/README: dropped the stale `proposed/` compaction row (master moved that RFC to implemented/); kept C's new `pre-tool-input-rewrite` proposed row. - interception.spec.ts: migrated its two `agent/turn-end` reason collectors to the `turn/end` session event, and ADDED a cross-test proving a `prompt-submit` rewrite + additionalContext is VISIBLE to an `agent/pre-step` listener on the same turn — pinning the merged seam ordering (compaction sees the post-prompt-submit surface, not stale history).
2026-07-02 04:50:14 +08:00
* every prompt blocked → 'turn/end'(rejected), 0 steps
* STEP loop:
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
* drain steering → session('steering/message') ⟵ catches late steering
* assembly = ctx.systemPrompt.assemble({agent}) ⟵ waterfall system-prompt/assemble; renderPrompt
* (persona section + {{variables}}) IS the full prompt
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
* await ctx.serial('agent/pre-step') ⟵ surface mutation (compaction) OUTSIDE the step
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
* boundary = session.deriveMessages() ⟵ the reconstruction boundary: snapshot in the
* session('step/start') same sync frame, strictly before step/start
* config = waterfall agent/request(config) ⟵ frozen seed; a returned replacement switches
* prefix ??= waterfall agent/session-prefix ⟵ once per loop instance (first request):
* frozen session prefix; logged on the header,
* never session history
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
* session('request/header'|'request/header-delta') ⟵ the header event this request owes the
* log (initial/resume anchor, delta, fallback)
* req = freeze({header..., messages: prefix+boundary, sessionId, signal})
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
* stream ctx.llm.stream(req) ⟵ waterfall llm/stream (raw chunks, frozen req)
refactor(events): remove the agent/stream-chunk mirror of assistant/chunk The loop recorded every model token delta as a durable `assistant/chunk` session event AND emitted an identical live `agent/stream-chunk` Cordis event one line later. Same StreamChunk, same turn/step; the emit added only the live Agent handle, which the sole consumer discarded. This is the boundary-mirror duplication the event-domain work removed for turn/step boundaries, applied to the token stream — a follow-up the boundary RFC explicitly deferred. The premise is settled: chunk persistence is authoritative (the proposal to stop persisting chunks was rejected — replay/snapshots depend on it), so `assistant/chunk` on `session/event` is the load-bearing token stream and `agent/stream-chunk` is pure redundancy. - Remove the `agent/stream-chunk` declaration + emit; drop the now-unused StreamChunk import from dsh-agent's types. - Migrate `dsh-ui-stdio` (the only live consumer; ACP already reads assistant/chunk off session/event) to render assistant/chunk in its existing session/event listener. Consolidating to one listener also makes the inReasoning dim-SGR flag deterministic across chunk/boundary events (they no longer race across two listeners). - Repoint the agent-loop tests (cancel/loop) and ui-stdio tests to the session/event assistant/chunk feed. - New RFC (implemented/simplification/2026-07-02-remove-stream-chunk-mirror); amend the boundary RFC's retained-list entry to cross-link; update architecture, cookbook, event-domain-semantics, the ACP proposal, and the regenerated cordis catalog. Snapshot goldens unchanged (ACP never used the mirror), confirming no editor-facing transcript change.
2026-07-02 23:42:16 +08:00
* session('assistant/chunk')
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
* msg = waterfall agent/step-result ⟵ BEFORE the log append, so the
simplify(session): fold trace-only usage/error events into load-bearing events The session event vocabulary carried two standalone trace-only events that were not load-bearing as separate records. Fold their facts into nearby load-bearing events and delete the standalone variants. - Token usage now rides on `assistant/message` as an optional `usage` field — the assembled model output and its accounting travel together. The loop folds `assembler.usage` onto the append instead of emitting a separate `usage` event. - The max-tokens path is the no-data-loss host: a step cut off with usage but EMPTY content (e.g. only a dropped tool call) previously emitted a standalone `usage`; it now records an empty-content `assistant/message { content: [], usage }`. `deriveMessages()` skips empty-content assistant messages, so the usage host never injects a spurious content-less assistant turn into the provider transcript. A step with neither content nor usage appends nothing. - An operational error's step number now rides on `turn/end.reason` for `kind: 'error'` (`{ kind: 'error', step, message, code? }`) — the durable turn outcome ACP and resume already consume. `failTurn` sets the reason directly (no separate session `error` event). `agent/error` + logging are unchanged for live diagnostics. - No format-version bump: pre-release, no persisted data, so per the format policy there is nothing to migrate or reject (the RFC's "refresh the format version" criterion over-reached). `version` stays 1. - ACP fixtures + goldens re-recorded (keyless replay): dropped standalone usage/error lines, usage folded onto assistant/message, error step on turn/end.reason. RFC moved proposed -> implemented with an implementation note recording the two scope refinements.
2026-06-21 10:00:06 +08:00
* session('assistant/message' {content, usage?}) session records what actually ran
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
* each tool-call in msg (sequential, abort-checked):
fix(events): address Codex review — protect post-execute result, purge stale tools/execute refs Codex's PR-C review found two (A) blockers: - tools/post-execute could corrupt the protected outcome. postExecute passed the mutable `result` to listeners and then read result.callId / spread result on the return paths, so a listener mutating the reference (flipping isError, rewriting callId, injecting an error) escaped the decision channel. Now the authoritative callId/isError/error are SNAPSHOT before the waterfall and the return value is rebuilt from the snapshot + the typed PostToolDecision — the decision is the only sanctioned way to change the outcome, and callId is always exec.callId. Added a regression test that mutates the result reference and asserts it has no effect; proven to fail red on the unfixed code. - Public docs/JSDoc still advertised the removed `tools/execute` waterfall after the split. Swept every current-state reference to tools/pre-execute + tools/post-execute: the ToolRegistry class JSDoc (and the regenerated catalog), loop.ts's ASCII flow (also added the prompt-submit/session-start steps it was missing), the package-map READMEs (packages, core, agent-core), core-data-structures core.md/tools.md, the bash + acp + invariants src/READMEs (the deferred permission gate is the tools/pre-execute deny/ask seam now), the cookbook, and the implemented RFCs whose factual seam catalog drifted. codec.ts's totality prose now lists `rejected`. Proposed-RFC references are left as-is (frozen proposals, validated when built).
2026-06-30 20:25:30 +08:00
* session('tool/call'); ctx.tools.execute() ⟵ tools/pre-execute (allow/deny/ask)
* → dispatch → tools/post-execute
* session('tool/result')
fix(events): address Codex review — protect post-execute result, purge stale tools/execute refs Codex's PR-C review found two (A) blockers: - tools/post-execute could corrupt the protected outcome. postExecute passed the mutable `result` to listeners and then read result.callId / spread result on the return paths, so a listener mutating the reference (flipping isError, rewriting callId, injecting an error) escaped the decision channel. Now the authoritative callId/isError/error are SNAPSHOT before the waterfall and the return value is rebuilt from the snapshot + the typed PostToolDecision — the decision is the only sanctioned way to change the outcome, and callId is always exec.callId. Added a regression test that mutates the result reference and asserts it has no effect; proven to fail red on the unfixed code. - Public docs/JSDoc still advertised the removed `tools/execute` waterfall after the split. Swept every current-state reference to tools/pre-execute + tools/post-execute: the ToolRegistry class JSDoc (and the regenerated catalog), loop.ts's ASCII flow (also added the prompt-submit/session-start steps it was missing), the package-map READMEs (packages, core, agent-core), core-data-structures core.md/tools.md, the bash + acp + invariants src/READMEs (the deferred permission gate is the tools/pre-execute deny/ask seam now), the cookbook, and the implemented RFCs whose factual seam catalog drifted. codec.ts's totality prose now lists `rejected`. Proposed-RFC references are left as-is (frozen proposals, validated when built).
2026-06-30 20:25:30 +08:00
* append buffered post-execute additionalContext → session('context/message')(s)
2026-07-04 15:36:40 +08:00
* drain steering → session('steering/message')
* session('step/end') ⟵ durable step boundary (no agent/* mirror)
fix(events): address Codex review — protect post-execute result, purge stale tools/execute refs Codex's PR-C review found two (A) blockers: - tools/post-execute could corrupt the protected outcome. postExecute passed the mutable `result` to listeners and then read result.callId / spread result on the return paths, so a listener mutating the reference (flipping isError, rewriting callId, injecting an error) escaped the decision channel. Now the authoritative callId/isError/error are SNAPSHOT before the waterfall and the return value is rebuilt from the snapshot + the typed PostToolDecision — the decision is the only sanctioned way to change the outcome, and callId is always exec.callId. Added a regression test that mutates the result reference and asserts it has no effect; proven to fail red on the unfixed code. - Public docs/JSDoc still advertised the removed `tools/execute` waterfall after the split. Swept every current-state reference to tools/pre-execute + tools/post-execute: the ToolRegistry class JSDoc (and the regenerated catalog), loop.ts's ASCII flow (also added the prompt-submit/session-start steps it was missing), the package-map READMEs (packages, core, agent-core), core-data-structures core.md/tools.md, the bash + acp + invariants src/READMEs (the deferred permission gate is the tools/pre-execute deny/ask seam now), the cookbook, and the implemented RFCs whose factual seam catalog drifted. codec.ts's totality prose now lists `rejected`. Proposed-RFC references are left as-is (frozen proposals, validated when built).
2026-06-30 20:25:30 +08:00
* cont = waterfall agent/turn-continuation ⟵ ContinuationDecision; default
* {action: hadToolCalls||steered ? 'continue':'stop'}; a continue.reason is
* recorded as next-step steering
* if action==stop && steering arrived (step/end/continuation listeners): continue anyway
* if action==stop: break
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
* session('turn/end') ⟵ durable turn boundary (no agent/* mirror)
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
* await ctx.parallel('session/flush', session) ⟵ durability checkpoint
* re-enqueue leftover steering as queued ⟵ steering is never stranded
* idle (emit agent/status) unless more queued
* ```
* @param ctx - the plugin context the loop reaches events (agent/…, session/flush) and services (systemPrompt, llm, tools) through.
* @param agent - the agent this invocation drives for its whole lifetime (its inbox, session, and options).
* @param handle - the bridge to the agent's mutable state: status/abort setters plus the disposal and cancel-marker reads.
*/
export async function runLoop(ctx: Context, agent: ReactLoopAgent, handle: LoopHandle): Promise<void> {
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
// Per-instance transmission bookkeeping: whether THIS loop instance has
// anchored the log's header fold yet (its first request logs a
// 'initial'/'resume' request/header snapshot). Everything else the request
// needs is read from the session log itself — the loop holds no
// conversation state (the reconstructability RFC).
const transmission = createTransmissionLog()
const { session } = agent
while (!handle.isDisposed()) {
await agent.inbox.waitForQueued(handle.disposed)
if (handle.isDisposed()) break
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
// Pre-step cancel (window 1): a `cancel()` landed after a `send()` woke the
// idle wait but before we flip to `running`. The cancelled queued/steering
// work is already cleared by `cancel()`. Clear the marker, then:
// - if NOTHING new is queued, drop the about-to-run turn and re-park,
// settling any `whenIdle()` waiter DIRECTLY (no running→idle transition
// fires here to settle it) and WITHOUT emitting `agent/status` (an ACP
// listener must not see a spurious idle that resolves a freshly-queued
// prompt as cancelled);
// - if a NEW prompt was queued AFTER the cancel (a send() that raced in
// before the loop resumed), the marker was for the cancelled work only —
// fall through and run the new prompt's turn. Do NOT settle waiters here:
// a whenIdle() waiter must wait for that new turn's running→idle, not
// resolve before it runs (the quiescence contract).
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
if (handle.isCancelled()) {
handle.clearCancel()
if (!agent.inbox.hasQueued) {
handle.settleIdle()
continue
}
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
}
handle.setStatus('running')
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
// Pre-step cancel (window 2): `setStatus('running')` emits `agent/status`
// SYNCHRONOUSLY, so a `running` listener can `cancel()` in the gap between the
// check above and `runTurn`. Mirror window 1: clear the marker, then
// - if NOTHING new is queued, drop the about-to-run turn and transition
// back to `idle` (`running` was already emitted, so a real idle
// transition balances the status AND settles `whenIdle()` waiters);
// - if a NEW prompt was queued AFTER the cancel (a `running` listener that
// cancels then sends), the marker was for the cancelled work only — fall
// through and run the new prompt's turn (status is already `running`), so
// a `whenIdle()` waiter resolves on THAT turn's running→idle, not before
// it runs. Settling here would resolve quiescence while the replacement
// is still queued and unrun (the same early-resolve race window 1 fixes).
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
if (handle.isCancelled()) {
handle.clearCancel()
if (!agent.inbox.hasQueued) {
handle.setStatus('idle')
continue
}
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
}
// Re-derive the turn number from the log each iteration (do NOT keep a local
// counter): an idle `agent.inject()` can append its own one-shot turn while
// the loop waits above, so the next real turn must continue from whatever
// turn number is actually last in the log — a stale counter would collide.
const turn = lastTurnNumber(session) + 1
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
try {
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
await runTurn(ctx, agent, handle, turn, transmission)
Document the codebase thoroughly and tighten type safety Docs: per-folder README.md for packages/ (family overview + one per package: service, events, API, extension points, TODOs), examples/, and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md symlinks) for packages/ and vendor/; module-level doc comments in every packages/*/src file; richer JSDoc on all exported API (event side effects, disposal contracts, error behavior). Root AGENTS.md gains a "Type Safety and Documentation" policy section: the codebase aims to be very type-safe and well documented; type gymnastics are acceptable in core packages when they improve plugin-author DX; verbose docs are fine as long as they stay strictly in sync with the code. Type safety: removed the upstream-inherited "noImplicitAny": false from tsconfig.base.json — packages/* now compile under full strict mode; vendor/loader and vendor/include set it locally (vendor/cordis already did). Eliminated every `: any` / `as any` from packages and examples (catch clauses use unknown + a CodedError narrowing type; event data access uses discriminated-union narrowing). Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL — SchemaSpec with per-property `required: true` booleans, type-level InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and defineTool() so first-party tools get typed execute(args) with zero casts (raw JSON Schema still accepted for MCP interop; chosen over schemastery because it targets JSON Schema generation directly). echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
} catch (error: unknown) {
// Backstop: runTurn rethrows only a PRE-turn throw (the invariant guard
// before turn/start) — no turn/start was appended, so no turn is open and
// none is owed. A session `error` here would land outside any turn (after
// the previous turn/end), where the persistence backend drops it as a
// crash tail (the turn-enclosure RFC). Report via agent/error + the logger only; the
// driver survives and moves on.
const err = toError(error)
ctx.logger.warn(`agent "${agent.id}": turn ${turn} failed before it started: ${err.message}`)
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
try {
ctx.emit('agent/error', agent, turn, 0, err)
} catch { /* contained: a throwing agent/error listener must not kill the driver */ }
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
}
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
// Reset the cancel marker UNCONDITIONALLY here, after the turn returns and
// before the next iteration's idle wait. NOT gated on the idle transition
// below: a `send()` that lands during the cancelled turn's flush window makes
// `hasQueued` true at the `setStatus('idle')` guard, so an idle-gated reset
// would never fire and the stale marker would wrongly drop that next prompt's
// turn. Resetting per iteration scopes the marker to exactly the turn that was
// cancelled.
handle.clearCancel()
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
// Steering that arrived too late to join this turn (turn-end listeners,
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
// flush) becomes a queued message — it must never be stranded. (A cancelled
// turn already cleared its steering, so there is nothing to re-enqueue.)
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
for (const message of agent.inbox.drainSteering()) {
agent.inbox.enqueue(message)
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
if (!agent.inbox.hasQueued) handle.setStatus('idle')
}
}
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
async function runTurn(
ctx: Context, agent: ReactLoopAgent, handle: LoopHandle, turn: number, transmission: TransmissionLog,
): Promise<void> {
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
const { session } = agent
// --- Pre-turn. A throw here (the invariant guard) is owed NO turn/end —
// turn/start has not been appended — so it propagates to runLoop's backstop
// untouched. The queued messages are drained here but appended AFTER
// turn/start (below), so every event in the log lives inside a turn.
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
const queued = agent.inbox.drainQueued()
const first = queued[0]
/* v8 ignore next 3 -- invariant guard: runLoop only calls runTurn when hasQueued */
if (!first) throw new Error('runTurn invariant violated: no queued message at turn start')
const trigger: TurnTrigger = { kind: 'message', source: first.source }
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
let reason: TurnEndReason = { kind: 'completed' }
let step = 0
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
let stepOpen = false
let errorReported = false
// Close the open step exactly once (idempotent via stepOpen). Step boundaries
// are durable session events only — there is no agent/* step emit to mirror
// them (see the agent event-domain rule). A throwing step/end session-event
// listener must not abort finalization and strand the turn open (turn/end
// balance > notifying one bad listener); it is contained and surfaced as a
// turn error below.
const closeStep = (): boolean => {
if (!stepOpen) return false
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
stepOpen = false
fix(agent-loop): contain finalizer append-listener throws (review #32 round 2) The prior fix handled a throwing session/event listener on the turn/start append, but the SAME push-before-notify hazard remained on the three FINALIZER appends. Session.append pushes the event before notifying, so a throwing listener on a finalizer event left the event logged but aborted the rest of finalization — stranding the turn open. - failTurn(): set `reason` BEFORE appending the `error` event, and contain a throwing session/event listener on it (the event is already logged either way). Otherwise reason stayed unset, agent/error was skipped, and the caller's closeTurn(false) never ran → open turn. - closeStep(): the try/catch wrapped only the agent/step-end EMIT, not the step/end APPEND. A throwing session/event listener on step/end escaped — fatal when closeStep runs from the outer catch during finalization (turn/start + step/end but no turn/end). Now both the append and the emit are contained and surface as a turn error via failTurn. - closeTurn(): contain a throwing session/event listener on the turn/end append (it would propagate to the runLoop backstop from closeTurn(false), or skip the turn-end emit from closeTurn(true)). turn/end is logged either way, so the turn stays balanced. Regressions: a throwing session/event listener on the error event, on step/end during finalization (driven by a throwing agent/step-start), and on turn/end — each leaves a balanced turn and the loop survives.
2026-06-16 00:37:18 +08:00
// Session.append pushes step/end BEFORE notifying session/event listeners,
// so a throwing listener leaves step/end in the log (balance holds) but
// would otherwise abort finalization. Contain it and surface it as a turn
// error below.
fix(agent-loop): contain finalizer append-listener throws (review #32 round 2) The prior fix handled a throwing session/event listener on the turn/start append, but the SAME push-before-notify hazard remained on the three FINALIZER appends. Session.append pushes the event before notifying, so a throwing listener on a finalizer event left the event logged but aborted the rest of finalization — stranding the turn open. - failTurn(): set `reason` BEFORE appending the `error` event, and contain a throwing session/event listener on it (the event is already logged either way). Otherwise reason stayed unset, agent/error was skipped, and the caller's closeTurn(false) never ran → open turn. - closeStep(): the try/catch wrapped only the agent/step-end EMIT, not the step/end APPEND. A throwing session/event listener on step/end escaped — fatal when closeStep runs from the outer catch during finalization (turn/start + step/end but no turn/end). Now both the append and the emit are contained and surface as a turn error via failTurn. - closeTurn(): contain a throwing session/event listener on the turn/end append (it would propagate to the runLoop backstop from closeTurn(false), or skip the turn-end emit from closeTurn(true)). turn/end is logged either way, so the turn stays balanced. Regressions: a throwing session/event listener on the error event, on step/end during finalization (driven by a throwing agent/step-start), and on turn/end — each leaves a balanced turn and the loop survives.
2026-06-16 00:37:18 +08:00
let failure: unknown
try {
session.append('step/end', { turn, step })
} catch (error: unknown) {
failure = error
}
// A throwing step/end session-event listener surfaces as a turn error via
// failTurn (idempotent). This prevents a throwing listener from producing a
// silent "completed" turn when the step itself succeeded, AND keeps
// finalization going when closeStep runs from the outer catch.
if (failure !== undefined) {
failTurn(toError(failure))
return true
}
return false
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
simplify(session): fold trace-only usage/error events into load-bearing events The session event vocabulary carried two standalone trace-only events that were not load-bearing as separate records. Fold their facts into nearby load-bearing events and delete the standalone variants. - Token usage now rides on `assistant/message` as an optional `usage` field — the assembled model output and its accounting travel together. The loop folds `assembler.usage` onto the append instead of emitting a separate `usage` event. - The max-tokens path is the no-data-loss host: a step cut off with usage but EMPTY content (e.g. only a dropped tool call) previously emitted a standalone `usage`; it now records an empty-content `assistant/message { content: [], usage }`. `deriveMessages()` skips empty-content assistant messages, so the usage host never injects a spurious content-less assistant turn into the provider transcript. A step with neither content nor usage appends nothing. - An operational error's step number now rides on `turn/end.reason` for `kind: 'error'` (`{ kind: 'error', step, message, code? }`) — the durable turn outcome ACP and resume already consume. `failTurn` sets the reason directly (no separate session `error` event). `agent/error` + logging are unchanged for live diagnostics. - No format-version bump: pre-release, no persisted data, so per the format policy there is nothing to migrate or reject (the RFC's "refresh the format version" criterion over-reached). `version` stays 1. - ACP fixtures + goldens re-recorded (keyless replay): dropped standalone usage/error lines, usage folded onto assistant/message, error step on turn/end.reason. RFC moved proposed -> implemented with an implementation note recording the two scope refinements.
2026-06-21 10:00:06 +08:00
// Record a step/turn failure exactly once: set the error reason (carrying the
// failing `step` — the durable failure lives entirely on turn/end.reason, there
// is no separate session error event) and emit agent/error (contained — trap: a
// throwing agent/error listener must not re-escape and strand the turn).
// Disposal and abort set `reason` directly without calling this (they are not
// failures).
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
const failTurn = (err: CodedError): void => {
if (errorReported) return
errorReported = true
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
// The turn is always still open here: the only failure that can reach
// failTurn once turn/end is appended would be a throwing turn-boundary
// listener, and turn boundaries are durable session events with no agent/*
// mirror to throw. A throwing `turn/end` session-event listener is already
// contained inside closeTurn (append pushes before notifying, so the
// boundary is durable). So set the error reason for closeTurn to append.
reason = { kind: 'error', step, ...errorData(err) }
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
try {
ctx.emit('agent/error', agent, turn, step, err)
} catch {
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
// contained: the error is already captured on `reason`; a throwing
// agent/error listener must not prevent the turn from closing.
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
}
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
// Close the turn. Called exactly once per turn — the normal loop exit and the
// outer catch are mutually exclusive paths, and this never throws (the append
// is contained below), so there is no re-entry to guard against (unlike
// closeStep, which the cancel branches and the outer catch can both reach).
// Turn boundaries are durable session events only — there is no agent/* turn
// emit to mirror them (see the agent event-domain rule).
const closeTurn = (): void => {
fix(agent-loop): contain finalizer append-listener throws (review #32 round 2) The prior fix handled a throwing session/event listener on the turn/start append, but the SAME push-before-notify hazard remained on the three FINALIZER appends. Session.append pushes the event before notifying, so a throwing listener on a finalizer event left the event logged but aborted the rest of finalization — stranding the turn open. - failTurn(): set `reason` BEFORE appending the `error` event, and contain a throwing session/event listener on it (the event is already logged either way). Otherwise reason stayed unset, agent/error was skipped, and the caller's closeTurn(false) never ran → open turn. - closeStep(): the try/catch wrapped only the agent/step-end EMIT, not the step/end APPEND. A throwing session/event listener on step/end escaped — fatal when closeStep runs from the outer catch during finalization (turn/start + step/end but no turn/end). Now both the append and the emit are contained and surface as a turn error via failTurn. - closeTurn(): contain a throwing session/event listener on the turn/end append (it would propagate to the runLoop backstop from closeTurn(false), or skip the turn-end emit from closeTurn(true)). turn/end is logged either way, so the turn stays balanced. Regressions: a throwing session/event listener on the error event, on step/end during finalization (driven by a throwing agent/step-start), and on turn/end — each leaves a balanced turn and the loop survives.
2026-06-16 00:37:18 +08:00
// Session.append pushes turn/end BEFORE notifying session/event listeners,
// so a throwing listener leaves turn/end in the log (the turn is balanced)
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
// but would otherwise escape — from the outer catch it would propagate to
// the runLoop backstop. Contain it: the boundary is durable either way, and
// finalization must not abort on a bad listener.
fix(agent-loop): contain finalizer append-listener throws (review #32 round 2) The prior fix handled a throwing session/event listener on the turn/start append, but the SAME push-before-notify hazard remained on the three FINALIZER appends. Session.append pushes the event before notifying, so a throwing listener on a finalizer event left the event logged but aborted the rest of finalization — stranding the turn open. - failTurn(): set `reason` BEFORE appending the `error` event, and contain a throwing session/event listener on it (the event is already logged either way). Otherwise reason stayed unset, agent/error was skipped, and the caller's closeTurn(false) never ran → open turn. - closeStep(): the try/catch wrapped only the agent/step-end EMIT, not the step/end APPEND. A throwing session/event listener on step/end escaped — fatal when closeStep runs from the outer catch during finalization (turn/start + step/end but no turn/end). Now both the append and the emit are contained and surface as a turn error via failTurn. - closeTurn(): contain a throwing session/event listener on the turn/end append (it would propagate to the runLoop backstop from closeTurn(false), or skip the turn-end emit from closeTurn(true)). turn/end is logged either way, so the turn stays balanced. Regressions: a throwing session/event listener on the error event, on step/end during finalization (driven by a throwing agent/step-start), and on turn/end — each leaves a balanced turn and the loop survives.
2026-06-16 00:37:18 +08:00
try {
session.append('turn/end', { turn, reason })
} catch (error: unknown) {
ctx.logger.warn(`agent "${agent.id}": session/event listener threw on turn/end at turn ${turn}: ${toError(error).message}`)
}
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
try {
// --- Turn boundary. Once turn/start is appended, a turn/end is owed no
fix(agent-loop): decide turn balance + idle-injection flush from the log (review #32) Session.append pushes the event BEFORE notifying session/event listeners, so a throwing listener leaves the event in the log while the line after the append (a boolean flag) never runs. Both turn-balance decisions were gated on such flags, so a throwing listener could strand an open turn or skip a durability checkpoint. - loop.ts: the outer catch decided "turn/end owed" from `turnStarted`. A throwing listener on the turn/start append left turn/start logged but the flag false → catch rethrew and skipped turn/end → permanently open turn (violating ADR 0017). Now decided from the log (this turn's turn/start present), so the turn is always balanced; only a genuine pre-push failure (non-serializable trigger — turn/start never logged) is rethrown to the runLoop backstop. Removed the now-dead `turnStarted`. - agent.ts inject(): the idle one-shot-turn flush was gated on a `turnRecorded` flag set after append('turn/end'); a throwing turn/end listener skipped the flush, losing the balanced in-memory injection turn on crash. Now the flush decision is read from the log, the synthetic turn/end append contains a throwing listener (turn stays balanced), and a failing idle flush is reported via agent/error (step 0 convention) AND the logger — mirroring the loop's post-turn/end flush path — with a throwing agent/error listener contained. Rewrote the test that encoded the old (buggy) "turn/start listener throw is rethrown, no turn/end" semantics to assert the balanced-turn contract, and added regressions for the throwing-turn/end-listener flush and the agent/error report. Updated Agent.inject JSDoc.
2026-06-15 23:44:54 +08:00
// matter what throws below; the catch + closeTurn guarantee it (the catch
// decides "owed" from the log via isTurnOpen, so even a throwing turn/start
// listener — append pushes before notifying — still gets its turn/end).
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
session.append('turn/start', { turn, trigger })
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
// Each drained queued message runs the `agent/prompt-submit` waterfall before
// it becomes a `user/message` — a hook can rewrite the prompt or block it.
// Recorded INSIDE the turn (after turn/start) so every event is turn-enclosed;
// turn/end is now owed, so a throwing prompt-submit listener (the waterfall
// throws) is caught below and the turn still closes.
let anyAllowed = false
// Seeded with a floor (only observable if the batch were empty, which
// runTurn never allows — it is called with ≥1 queued message); each `block`
// decision carries a required `reason` and overwrites it, so a fully-blocked
// batch always reports the last vetoing reason.
let lastBlockReason = 'prompt blocked by hook'
for (const message of queued) {
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
const decision = await ctx.waterfall(
'agent/prompt-submit', agent, message.content, message.source,
() => Promise.resolve<PromptDecision>({ kind: 'allow' }),
)
if (decision.kind === 'block') {
lastBlockReason = decision.reason
// Record the veto durably: `PromptDecision.reason` is the durable record
// of why a prompt was blocked, but a fully-blocked batch's `rejected`
// turn/end only preserves the LAST reason, and a MIXED batch (this prompt
// blocked, another allowed) does not end `rejected` at all — so without
// this append a blocked prompt would vanish from the log whenever any
// sibling prompt is allowed. `prompt/blocked` sits in the open turn in
// place of the `user/message` this prompt would have become.
session.append('prompt/blocked', { content: message.content, source: message.source, reason: decision.reason })
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
continue
}
anyAllowed = true
// `allow.content` REPLACES the prompt bytes (a rewrite); absent keeps them.
const content = decision.content ?? message.content
session.append('user/message', { content, source: message.source }, { surfaceOp: 'append' })
// `allow.additionalContext` is a SEPARATE context/message the next request
// also sees. The turn is open, so inject() appends it into THIS turn.
if (decision.additionalContext) {
agent.inject(decision.additionalContext.content, { source: decision.additionalContext.source })
}
}
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
while (true) {
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
// A fully-blocked batch (every prompt vetoed by prompt-submit) opens a
// zero-step turn that ends `rejected`: break BEFORE the first step so the
// boundary stays balanced (turn/start → turn/end) and the block is a
// durable in-turn fact. `anyAllowed` never changes inside the loop, so this
// only ever fires on the first iteration.
if (!anyAllowed) {
reason = { kind: 'rejected', reason: lastBlockReason }
break
}
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
step += 1
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
// Steering from the previous round's continuation listeners joins before
// the request.
2026-07-04 15:36:40 +08:00
drainSteering(agent, turn)
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
// The step's AbortController exists BEFORE any async pre-step work so a
// dispose() or cancel() — in a synchronous turn-start listener or an
// async listener whose effect fires before we block — always has an armed
// abort to cancel against. isDisposed below covers disposal, which does
// NOT set the cancel marker. Cleared on every exit path below.
const abort = new AbortController()
handle.setAbort(abort)
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
// Assemble the system prompt for this step. Done HERE (before step/start)
// because the pre-step seam needs it: compaction measures token pressure
// against the system prompt (it counts toward the budget). runStep reuses
// this same assembly for the request, so the prompt is assembled once per
// step. renderPrompt IS the full prompt — the persona is the order-0
// section (registered by the AgentLoop plugin) and `{{variable}}`
// interpolation happens in the render, so there is no separate join.
const assembly = await ctx.systemPrompt.assemble({ agent })
const fullSystemPrompt = renderPrompt(assembly)
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
// Interruption landing after assembly: dispose() or cancel() in a
// turn-start listener (or a listener whose promise resolved before the
// await above) arms either handle.isDisposed() or handle.isCancelled().
// The Abort was created first, so any concurrent abort also lands on it.
// Drop the about-to-start step WITHOUT running the seam — no step is open
// yet, so end the turn accordingly (disposed wins for an unambiguous
// reason).
if (handle.isCancelled() || handle.isDisposed()) {
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
handle.setAbort(undefined)
reason = handle.isDisposed() ? { kind: 'disposed' } : { kind: 'aborted', reason: handle.cancelReason() }
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
break
}
// Pre-step surface-mutation checkpoint (compaction), fired OUTSIDE the
// step: after `turn/start` (and the prior step's close) but before
// `step/start`, so a compaction's log-only `compact/*` records and its
// replacement node land cleanly outside any step (honest structure that
// crash-safety relies on — a dangling `compact/start` sits before the
// synthetic `turn/end` repair appends). Serial (awaited, in order, no
// veto): each listener completes its surface mutation before the next, so
// concurrent listeners cannot interleave their `session.append`s. A
// throwing listener escapes to the outer catch, which closes the (not-yet-
// open) step as a no-op and ends the turn via failTurn — a broken
// pre-step plugin ends the turn, not the loop.
await ctx.serial('agent/pre-step', agent, turn, step, fullSystemPrompt, abort.signal)
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
// Interruption landing during the pre-step seam: do not open an empty step.
if (handle.isCancelled() || handle.isDisposed()) {
handle.setAbort(undefined)
reason = handle.isDisposed() ? { kind: 'disposed' } : { kind: 'aborted', reason: handle.cancelReason() }
break
}
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
// The reconstruction boundary (the reconstructability RFC): the request's
// messages are snapshotted HERE, in the same synchronous frame as the
// step/start append directly below — so the snapshot is exactly the
// derivation over the log prefix strictly before step/start's seq.
// Anything appended later — by a step/start session/event listener, an
// agent/request-window inject(), any concurrent task — lands after the
// boundary and joins the NEXT request. An external reconstructor
// recovers these exact messages by folding the surface over
// events[0..stepStartSeq).
const boundaryMessages = session.deriveMessages()
// Mark the step open BEFORE the append: Session.append pushes the event
// to the log before notifying session/event listeners, so a THROWING
// step/start listener leaves step/start in the log. Setting stepOpen first
// means the outer catch's closeStep() then appends the balancing step/end
// (turn stays enclosed) instead of stranding an open step under turn/end.
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
stepOpen = true
session.append('step/start', { turn, step })
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
// Cancel landing in the step-start window: a synchronous `session/event`
// step/start listener can cancel after the step is already open. Check
// AFTER the step/start append and before `runStep`: drop the step, end the
// turn accordingly. closeStep balances the already-appended step/start.
if (handle.isCancelled() || handle.isDisposed()) {
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
handle.setAbort(undefined)
reason = handle.isDisposed() ? { kind: 'disposed' } : { kind: 'aborted', reason: handle.cancelReason() }
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
closeStep()
break
}
let stepOutcome: { hadToolCalls: boolean; finish: FinishReason } | { error: Error }
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
try {
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
stepOutcome = await runStep(ctx, agent, turn, step, assembly, fullSystemPrompt, boundaryMessages, transmission, abort.signal)
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
} catch (error: unknown) {
stepOutcome = { error: toError(error) }
} finally {
handle.setAbort(undefined)
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
if ('error' in stepOutcome) {
// Steering that arrived during the failed step stays in the inbox —
// runLoop re-enqueues it as a queued message, so an abort-then-steer
// starts a fresh turn instead of being silently consumed.
closeStep()
const { error } = stepOutcome
if (handle.isDisposed()) {
reason = { kind: 'disposed' }
} else if (abort.signal.aborted) {
simplify(agent): drop the unused public Agent.abort(), keep whenIdle() The public Agent handle exposed abort() (step-only) and cancel() (queue-aware). No production caller used abort() — ACP maps session/cancel to cancel(), and lifecycle owners tear down via AgentHandle.dispose(); the loop's own stop paths abort their per-step AbortController directly. So abort() is latent generality that keeps a private loop mechanic public. RFC-premise correction: the public-agent-stop-surface RFC proposed removing whenIdle() too. Implementation found whenIdle() load-bearing — a real quiescence primitive with a deliberate loop contract (settle-without-transition, the replacement-turn race) and ACP test consumers; its proposed replacement ("observe the running->idle transition") is exactly the async-state race AGENTS.md warns against. So only abort() is removed; whenIdle() stays. The RFC is amended on the way to implemented/ to record the narrowed scope, and the new AGENTS.md "RFCs are proposals, not golden truth" principle (PR1) gets its worked example. - Remove Agent.abort() from the interface + the ReactLoopAgent impl; the no-arg 'aborted' default goes with it (cancel() keeps its 'cancelled' default). - Migrate tests: empty-queue abort() -> cancel(reason); the two review-fixes tests whose subject is the in-flight step's AbortController drive that controller directly via the private currentAbort field (cancel() would clear the inbox and destroy the queued steering one of them proves survives a step abort). The no-arg-default test is dropped (cancel()'s default is already covered in cancel.spec.ts). - Resulting public stop surface: cancel() + whenIdle(). Update agent/agent-loop READMEs, architecture.md, core.md type-equiv, the extension cookbook, the lifecycle RFC (short note), and the proposed ACP RFC. Implements docs/rfc/implemented/simplification/2026-06-20-public-agent-stop-surface.md
2026-06-21 05:50:39 +08:00
/* v8 ignore next -- signal.reason always set: cancel()/disposal provide a default */
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
reason = { kind: 'aborted', reason: String(abort.signal.reason ?? 'aborted') }
} else {
failTurn(error)
}
break
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
// The successful step's finish reason carries forward: a `max-tokens`
// step makes the whole turn end `max-tokens` (the ACP RFC's rule "any
// max-tokens step surfaces as max-tokens"). `stepFinishReason` returns
// `max-tokens` or `undefined`, so a later ordinary step never resets a
// max-tokens turn back to completed, and a never-truncated turn keeps the
// default `completed`. The disposal/abort/error branches above and the
// continuation-window disposal check below override this — they win.
const stepReason = stepFinishReason(stepOutcome.finish)
if (stepReason) reason = stepReason
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
// Steering that arrived during streaming/tool execution.
2026-07-04 15:36:40 +08:00
const steered = drainSteering(agent, turn)
if (closeStep()) break
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
const defaultDecision: ContinuationDecision = { action: stepOutcome.hadToolCalls || steered ? 'continue' : 'stop' }
let decision: ContinuationDecision
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
try {
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
decision = await ctx.waterfall(
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
'agent/turn-continuation', agent, turn, defaultDecision,
() => Promise.resolve(defaultDecision),
)
} catch (error: unknown) {
// A broken continuation plugin ends the turn, not the loop.
failTurn(toError(error))
break
}
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
// A forced `continue` may carry model-facing context: record it as
// next-STEP steering (the steering channel), so the continued turn's next
// iteration drains it before its request — the typed twin of the /goal
// step/end-steer pattern.
if (decision.action === 'continue' && decision.reason) {
agent.inbox.steer({ content: decision.reason.content, source: decision.reason.source })
}
let shouldContinue = decision.action === 'continue'
// Steering from step/end session-event or continuation listeners (the
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
// /goal pattern) demands the model see it — it overrides a stop decision;
// the next iteration's drain records it.
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
if (!shouldContinue && agent.inbox.hasSteering) shouldContinue = true
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
// A cancel that landed during the continuation window — after the step's
// AbortController was cleared (setAbort(undefined)) but before the next
// step starts — has no controller to observe it, so the turn-scoped marker
// ends the turn here. cancel() also cleared the steering FIFO, so the
// override above did not re-arm continuation.
if (handle.isCancelled()) {
reason = { kind: 'aborted', reason: handle.cancelReason() }
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
break
}
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
if (!shouldContinue || handle.isDisposed()) {
/* v8 ignore next -- disposal during continuation-decision window is a narrow race; error-path disposal is covered elsewhere */
if (handle.isDisposed()) reason = { kind: 'disposed' }
break
}
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
// Normal / inline-error loop exit: close the turn.
closeTurn()
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
} catch (error: unknown) {
fix(agent-loop): decide turn balance + idle-injection flush from the log (review #32) Session.append pushes the event BEFORE notifying session/event listeners, so a throwing listener leaves the event in the log while the line after the append (a boolean flag) never runs. Both turn-balance decisions were gated on such flags, so a throwing listener could strand an open turn or skip a durability checkpoint. - loop.ts: the outer catch decided "turn/end owed" from `turnStarted`. A throwing listener on the turn/start append left turn/start logged but the flag false → catch rethrew and skipped turn/end → permanently open turn (violating ADR 0017). Now decided from the log (this turn's turn/start present), so the turn is always balanced; only a genuine pre-push failure (non-serializable trigger — turn/start never logged) is rethrown to the runLoop backstop. Removed the now-dead `turnStarted`. - agent.ts inject(): the idle one-shot-turn flush was gated on a `turnRecorded` flag set after append('turn/end'); a throwing turn/end listener skipped the flush, losing the balanced in-memory injection turn on crash. Now the flush decision is read from the log, the synthetic turn/end append contains a throwing listener (turn stays balanced), and a failing idle flush is reported via agent/error (step 0 convention) AND the logger — mirroring the loop's post-turn/end flush path — with a throwing agent/error listener contained. Rewrote the test that encoded the old (buggy) "turn/start listener throw is rethrown, no turn/end" semantics to assert the balanced-turn contract, and added regressions for the throwing-turn/end-listener flush and the agent/error report. Updated Agent.inject JSDoc.
2026-06-15 23:44:54 +08:00
// Decide whether this turn was ever opened from the LOG, not a flag.
// Session.append pushes the event BEFORE notifying session/event listeners,
// so a throwing listener on the `turn/start` append leaves turn/start in the
// log even though execution never reached the lines after that append.
// Gating on a "turn started" boolean would skip turn/end and leave a
// permanently OPEN turn that poisons the next turn/replay (the turn-enclosure RFC). We
fix(agent-loop): decide turn balance + idle-injection flush from the log (review #32) Session.append pushes the event BEFORE notifying session/event listeners, so a throwing listener leaves the event in the log while the line after the append (a boolean flag) never runs. Both turn-balance decisions were gated on such flags, so a throwing listener could strand an open turn or skip a durability checkpoint. - loop.ts: the outer catch decided "turn/end owed" from `turnStarted`. A throwing listener on the turn/start append left turn/start logged but the flag false → catch rethrew and skipped turn/end → permanently open turn (violating ADR 0017). Now decided from the log (this turn's turn/start present), so the turn is always balanced; only a genuine pre-push failure (non-serializable trigger — turn/start never logged) is rethrown to the runLoop backstop. Removed the now-dead `turnStarted`. - agent.ts inject(): the idle one-shot-turn flush was gated on a `turnRecorded` flag set after append('turn/end'); a throwing turn/end listener skipped the flush, losing the balanced in-memory injection turn on crash. Now the flush decision is read from the log, the synthetic turn/end append contains a throwing listener (turn stays balanced), and a failing idle flush is reported via agent/error (step 0 convention) AND the logger — mirroring the loop's post-turn/end flush path — with a throwing agent/error listener contained. Rewrote the test that encoded the old (buggy) "turn/start listener throw is rethrown, no turn/end" semantics to assert the balanced-turn contract, and added regressions for the throwing-turn/end-listener flush and the agent/error report. Updated Agent.inject JSDoc.
2026-06-15 23:44:54 +08:00
// check the log for THIS turn's turn/start: present means a turn/end is owed
// and the normal-exit `closeTurn()` did NOT run (we are here because a throw
// preceded it — the two `closeTurn()` sites are on mutually exclusive paths),
// so this catch appends turn/end with the disposed/error reason chosen below.
// `closeStep()` IS idempotent (guarded by `stepOpen`) — it may have run
// already in a step branch, so running it again is a safe no-op. Absent
// turn/start means the append threw BEFORE its push (a non-serializable
// trigger — impossible for our fixed trigger); nothing was opened, so rethrow
// to the runLoop backstop.
fix(agent-loop): decide turn balance + idle-injection flush from the log (review #32) Session.append pushes the event BEFORE notifying session/event listeners, so a throwing listener leaves the event in the log while the line after the append (a boolean flag) never runs. Both turn-balance decisions were gated on such flags, so a throwing listener could strand an open turn or skip a durability checkpoint. - loop.ts: the outer catch decided "turn/end owed" from `turnStarted`. A throwing listener on the turn/start append left turn/start logged but the flag false → catch rethrew and skipped turn/end → permanently open turn (violating ADR 0017). Now decided from the log (this turn's turn/start present), so the turn is always balanced; only a genuine pre-push failure (non-serializable trigger — turn/start never logged) is rethrown to the runLoop backstop. Removed the now-dead `turnStarted`. - agent.ts inject(): the idle one-shot-turn flush was gated on a `turnRecorded` flag set after append('turn/end'); a throwing turn/end listener skipped the flush, losing the balanced in-memory injection turn on crash. Now the flush decision is read from the log, the synthetic turn/end append contains a throwing listener (turn stays balanced), and a failing idle flush is reported via agent/error (step 0 convention) AND the logger — mirroring the loop's post-turn/end flush path — with a throwing agent/error listener contained. Rewrote the test that encoded the old (buggy) "turn/start listener throw is rethrown, no turn/end" semantics to assert the balanced-turn contract, and added regressions for the throwing-turn/end-listener flush and the agent/error report. Updated Agent.inject JSDoc.
2026-06-15 23:44:54 +08:00
const turnStartLogged = session.events.some(e => e.type === 'turn/start' && e.data.turn === turn)
if (!turnStartLogged) throw error
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
closeStep()
// Choose the close reason. Disposal wins only if no error was already
// reported: a turn disposed mid-step sets reason=disposed in the step-error
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
// branch (without reporting an error), so preserve disposed rather than
// overwrite it. Otherwise a mid-step throw on a live agent is a real
// failure → failTurn. (errorReported is mutated only inside the failTurn
// closure, which the analyzer can't follow, hence the inline lint-disable.)
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
if (handle.isDisposed() && !errorReported) { // eslint-disable-line @typescript-eslint/no-unnecessary-condition
reason = { kind: 'disposed' }
} else {
failTurn(toError(error))
}
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
closeTurn()
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
// Durability checkpoint: persistence plugins drain write-behind buffers.
// A failing persistence plugin is reported but doesn't kill the agent.
try {
await ctx.parallel('session/flush', session)
Document the codebase thoroughly and tighten type safety Docs: per-folder README.md for packages/ (family overview + one per package: service, events, API, extension points, TODOs), examples/, and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md symlinks) for packages/ and vendor/; module-level doc comments in every packages/*/src file; richer JSDoc on all exported API (event side effects, disposal contracts, error behavior). Root AGENTS.md gains a "Type Safety and Documentation" policy section: the codebase aims to be very type-safe and well documented; type gymnastics are acceptable in core packages when they improve plugin-author DX; verbose docs are fine as long as they stay strictly in sync with the code. Type safety: removed the upstream-inherited "noImplicitAny": false from tsconfig.base.json — packages/* now compile under full strict mode; vendor/loader and vendor/include set it locally (vendor/cordis already did). Eliminated every `: any` / `as any` from packages and examples (catch clauses use unknown + a CodedError narrowing type; event data access uses discriminated-union narrowing). Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL — SchemaSpec with per-property `required: true` booleans, type-level InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and defineTool() so first-party tools get typed execute(args) with zero casts (raw JSON Schema still accepted for MCP interop; chosen over schemastery because it targets JSON Schema generation directly). echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
} catch (error: unknown) {
// The turn is already closed (turn/end appended above) and flush must run
// AFTER turn/end to be a checkpoint — so there is no in-turn position left
// for a session `error` event. Appending one here would land it after the
// last turn/end, where the persistence backend treats it as a crash tail
// and drops it on resume (the turn-enclosure RFC: every event is turn-enclosed). Report
// the failure via agent/error + the logger only; persistence keeps the
// buffered events for the next flush/dispose, so nothing is lost.
const err = toError(error)
ctx.logger.warn(`agent "${agent.id}": session/flush failed at turn ${turn}: ${err.message}`)
try {
ctx.emit('agent/error', agent, turn, step, err)
} catch {
// contained: a throwing agent/error listener must not escape the loop.
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
}
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
/** Drain the steering queue into the session. Returns whether any arrived. */
2026-07-04 15:36:40 +08:00
function drainSteering(agent: ReactLoopAgent, turn: number): boolean {
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
const messages = agent.inbox.drainSteering()
for (const message of messages) {
2026-06-17 19:25:29 +08:00
agent.session.append('steering/message', { turn, content: message.content, source: message.source }, { surfaceOp: 'append' })
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
return messages.length > 0
}
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
/** One step: build the request from the boundary snapshot + the step's
* header → compose the session prefix if this instance has none yet → log
* the header event the request owes → stream model → record → execute
* tools. The caller assembles the
feat(agent): add the agent/request-messages request-only message seam A new waterfall near request construction lets plugins contribute request-ONLY messages framing the derived history: RequestMessages { before, after } with a frozen empty seed, fired inside the open step after the agent/request config waterfall, so the step/start boundary snapshot and its same-sync-frame invariant are untouched. The request becomes messagePrefix + boundary snapshot + messageSuffix. Contributions never enter session history — deriveMessages() is unchanged — so the request header is their durable record: EpochHeader gains messagePrefix/messageSuffix (canonical absence for empty arrays), request/header-delta replaces either array whole with an empty array encoding the transition back to absence, and the dev-mode reconstruction cross-check now expects the folded header's framing around the boundary derivation. This is the seam for per-request advisory context that must be model-visible now without becoming durable history (a skills catalog, an environment reminder), keeping the base system prompt workspace-independent and provider prefix caches stable. The docs carry the channel cost model: session-frozen content belongs in before, low-frequency change notices belong in durable history via inject() (paid once, prefix-cached thereafter), and after is reserved for small frequently-refreshed state snapshots re-paid on every request they ride. No shipped producer yet, so ACP snapshot fixtures are byte-identical.
2026-07-07 19:42:30 +08:00
* system prompt, fires the `agent/pre-step` seam, snapshots the derivation,
* and opens the step BEFORE calling this, so `boundaryMessages` is exactly
* the surface prefix at step/start and already reflects any compaction. */
async function runStep(
ctx: Context,
agent: ReactLoopAgent,
turn: number,
step: number,
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
assembly: PromptAssembly,
system: string,
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
boundaryMessages: Message[],
transmission: TransmissionLog,
signal: AbortSignal,
): Promise<{ hadToolCalls: boolean; finish: FinishReason }> {
const { session, options } = agent
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
// Seed the call config: the first request of THIS loop instance seeds from
// current AgentOptions — explicit options always win over the logged
// baseline, which is what keeps fork model-overrides and resume-time
// reconfiguration correct. Later steps seed from the log's folded header,
// which by then is exactly what this instance last logged.
// One deep-cloned, frozen seed serves BOTH the listener chain and the
// no-listener fallback: structuredClone decouples it from the session's
// cached header fold (a raw reference would let a delegating listener
// mutate the fold in place and silently skip the delta log), and the freeze
// makes in-place shaping unrepresentable — a switch is a RETURNED
// replacement, which the header event below records.
const seedConfig: LlmCallConfig = deepFreeze(structuredClone(transmission.loggedHeader
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
// eslint-disable-next-line @typescript-eslint/no-non-null-assertion -- loggedHeader ⟹ a snapshot is in the log
? session.requestHeader()!.config
: { model: options.model ?? '' }))
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
// Shape the call config: listeners return a replacement to switch model or
// sampling (the seed is frozen — content shaping is not expressible here;
// model-visible content flows through the log channels). The header event
// below records whatever the request ACTUALLY uses, so a listener's switch
// is a logged, reconstructable fact, never silent drift.
const config = await ctx.waterfall('agent/request', agent, turn, step, seedConfig, () => Promise.resolve(seedConfig))
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
if (!config.model) {
throw new Error(`agent "${agent.id}" has no model: set AgentOptions.model or supply one via the agent/request waterfall`)
}
// Compose the session prefix ONCE per loop instance, lazily on its first
// request-building step: request-only messages placed in front of the
// ENTIRE derived history on every request this instance sends. The result
// is deep-cloned (decoupled from listener-held references), deep-frozen,
// and cached on the transmission bookkeeping, so reuse is structural — the
// prefix cannot change mid-session and the provider prefix cache holds by
// construction (resume = a new instance = a recompose, anchored by its
// 'resume' snapshot). The prefix is not session history — the header event
// below is its only durable record (EpochHeader.messagePrefix), which
// keeps the request a pure function of the log. The frozen empty seed
// serves both the listener chain and the no-listener fallback: a
// contribution is a RETURNED extension of `await next()`, never an
// in-place push.
if (transmission.sessionPrefix === undefined) {
const emptyPrefix: Message[] = deepFreeze([])
transmission.sessionPrefix = deepFreeze(structuredClone(await ctx.waterfall(
'agent/session-prefix', agent, emptyPrefix, signal,
() => Promise.resolve(emptyPrefix),
)))
}
const sessionPrefix = transmission.sessionPrefix
feat(agent): add the agent/request-messages request-only message seam A new waterfall near request construction lets plugins contribute request-ONLY messages framing the derived history: RequestMessages { before, after } with a frozen empty seed, fired inside the open step after the agent/request config waterfall, so the step/start boundary snapshot and its same-sync-frame invariant are untouched. The request becomes messagePrefix + boundary snapshot + messageSuffix. Contributions never enter session history — deriveMessages() is unchanged — so the request header is their durable record: EpochHeader gains messagePrefix/messageSuffix (canonical absence for empty arrays), request/header-delta replaces either array whole with an empty array encoding the transition back to absence, and the dev-mode reconstruction cross-check now expects the folded header's framing around the boundary derivation. This is the seam for per-request advisory context that must be model-visible now without becoming durable history (a skills catalog, an environment reminder), keeping the base system prompt workspace-independent and provider prefix caches stable. The docs carry the channel cost model: session-frozen content belongs in before, low-frequency change notices belong in durable history via inject() (paid once, prefix-cached thereafter), and after is reserved for small frequently-refreshed state snapshots re-paid on every request they ride. No shipped producer yet, so ACP snapshot fixtures are byte-identical.
2026-07-07 19:42:30 +08:00
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
// The request header (the log's request/header* vocabulary): canonical form,
feat(agent): add the agent/request-messages request-only message seam A new waterfall near request construction lets plugins contribute request-ONLY messages framing the derived history: RequestMessages { before, after } with a frozen empty seed, fired inside the open step after the agent/request config waterfall, so the step/start boundary snapshot and its same-sync-frame invariant are untouched. The request becomes messagePrefix + boundary snapshot + messageSuffix. Contributions never enter session history — deriveMessages() is unchanged — so the request header is their durable record: EpochHeader gains messagePrefix/messageSuffix (canonical absence for empty arrays), request/header-delta replaces either array whole with an empty array encoding the transition back to absence, and the dev-mode reconstruction cross-check now expects the folded header's framing around the boundary derivation. This is the seam for per-request advisory context that must be model-visible now without becoming durable history (a skills catalog, an environment reminder), keeping the base system prompt workspace-independent and provider prefix caches stable. The docs carry the channel cost model: session-frozen content belongs in before, low-frequency change notices belong in durable history via inject() (paid once, prefix-cached thereafter), and after is reserved for small frequently-refreshed state snapshots re-paid on every request they ride. No shipped producer yet, so ACP snapshot fixtures are byte-identical.
2026-07-07 19:42:30 +08:00
// recorded before dispatch so the log always explains the request —
// including the session prefix, which no other event carries.
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
const header = canonicalHeader({
config,
...system ? { system } : {},
...assembly.tools.length > 0 ? { tools: assembly.tools } : {},
...sessionPrefix.length > 0 ? { messagePrefix: sessionPrefix } : {},
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
})
recordRequestHeader(session, transmission, header)
// Build and freeze: the request is a pure function of (boundary snapshot,
// logged header) — llm/stream listeners and adapters read it, mutation
// throws. sessionId + frozen is the loop-built marker the dev invariant
// keys on. Message order: header.messagePrefix, then the boundary
// snapshot — the reconstruction equation the invariant recomputes.
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
const request: GenerateOptions = deepFreeze({
model: header.config.model,
messages: [...header.messagePrefix ?? [], ...boundaryMessages],
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
...header.system !== undefined ? { system: header.system } : {},
...header.tools !== undefined ? { tools: header.tools } : {},
...header.config.temperature !== undefined ? { temperature: header.config.temperature } : {},
...header.config.maxTokens !== undefined ? { maxTokens: header.config.maxTokens } : {},
...header.config.stop !== undefined ? { stop: header.config.stop } : {},
Add per-session snapshot replay for nested agents (PR2.5) The snapshot tier was built single-session: dsh-llm-replay served calls from one global positional cursor, and the harness harvested one session log. A subagent runs as a second agent with its own session, so a parent→child scenario could neither replay deterministically nor harvest the child's log. This resolves the TODO(subagent-snapshots) deferral from the subagent RFC. - Stamp the calling session id onto the model request: GenerateOptions.sessionId (typed Branded<'SessionId'> to avoid the dsh-llm↔dsh-session cycle), set by the agent loop from agent.session.id. Adapters ignore it; an llm/stream listener routes by it. - Key replay per session: dsh-llm-replay loads the parent log plus one per child (childFiles / $DSH_SNAPSHOT_CHILD_FILES), derives a script per recorded session, and binds each live (freshly-random) session to a recorded script by first-call order — parent first (earliest createdAt, first to stream). Keys by WHO calls, so it survives a future concurrent/backgrounded subagent; a global cursor would not. An unrecorded extra session fails loud. - Harvest every log: the harness collects all .jsonl across cwd buckets, ordered primary-first (top-level, then children by createdAt), and RunResult exposes the plural sessionLogs. The spec writes each back on record (session.jsonl + session.<n>.jsonl) and diffs each against its fixture on replay. - Wire the subagent seam + spawn + fork + tool into the acp-agent example (both cordis configs) and add two nested scenarios recorded against the real API: subagent-spawn (parent + 1 child) and subagent-multi (parent + 2 children, 3 sessions). Both replay keyless in the default gate. A new RFC documents the design (docs/rfc/implemented/testing/). Single-session replay is unchanged (a call with no sessionId is one anonymous primary session). TODO follow-up: a dedicated branded-ids package could own the SessionId brand and dissolve the cross-package cycle note; out of scope for this testing PR.
2026-06-22 08:39:36 +08:00
sessionId: session.id,
signal,
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
})
// --- Model call (streaming-first; raw chunks are the replay record) ---
const assembler = new BlockAssembler()
2026-06-17 19:25:29 +08:00
const chunkSeqs: number[] = []
for await (const chunk of ctx.llm.stream(request)) {
simplify(agent): drop the unused public Agent.abort(), keep whenIdle() The public Agent handle exposed abort() (step-only) and cancel() (queue-aware). No production caller used abort() — ACP maps session/cancel to cancel(), and lifecycle owners tear down via AgentHandle.dispose(); the loop's own stop paths abort their per-step AbortController directly. So abort() is latent generality that keeps a private loop mechanic public. RFC-premise correction: the public-agent-stop-surface RFC proposed removing whenIdle() too. Implementation found whenIdle() load-bearing — a real quiescence primitive with a deliberate loop contract (settle-without-transition, the replacement-turn race) and ACP test consumers; its proposed replacement ("observe the running->idle transition") is exactly the async-state race AGENTS.md warns against. So only abort() is removed; whenIdle() stays. The RFC is amended on the way to implemented/ to record the narrowed scope, and the new AGENTS.md "RFCs are proposals, not golden truth" principle (PR1) gets its worked example. - Remove Agent.abort() from the interface + the ReactLoopAgent impl; the no-arg 'aborted' default goes with it (cancel() keeps its 'cancelled' default). - Migrate tests: empty-queue abort() -> cancel(reason); the two review-fixes tests whose subject is the in-flight step's AbortController drive that controller directly via the private currentAbort field (cancel() would clear the inbox and destroy the queued steering one of them proves survives a step abort). The no-arg-default test is dropped (cancel()'s default is already covered in cancel.spec.ts). - Resulting public stop surface: cancel() + whenIdle(). Update agent/agent-loop READMEs, architecture.md, core.md type-equiv, the extension cookbook, the lifecycle RFC (short note), and the proposed ACP RFC. Implements docs/rfc/implemented/simplification/2026-06-20-public-agent-stop-surface.md
2026-06-21 05:50:39 +08:00
/* v8 ignore next -- signal.reason always set: cancel()/disposal provide a default */
if (signal.aborted) throw new Error(String(signal.reason ?? 'aborted'))
2026-06-17 19:25:29 +08:00
const chunkEvent = session.append('assistant/chunk', { turn, step, chunk })
chunkSeqs.push(chunkEvent.seq)
assembler.push(chunk)
}
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
// Adapters report provider/transport failures one of two sanctioned ways
// (see the StreamChunk contract in dsh-llm): throw from stream() — already
// handled by the caller's try/catch — OR end the stream with a
// finish-error/aborted chunk. finishError() maps the latter to the step
// error to raise (turn ends error/aborted, not a normal completed message).
const stepError = finishError(assembler.finish)
if (stepError) throw stepError
if (assembler.finish.kind === 'max-tokens') {
2026-06-19 00:37:12 +08:00
let message: Message = withoutToolCalls(assembler.message())
message = withoutToolCalls(await ctx.waterfall('agent/step-result', agent, turn, step, message, () => Promise.resolve(message)))
simplify(session): fold trace-only usage/error events into load-bearing events The session event vocabulary carried two standalone trace-only events that were not load-bearing as separate records. Fold their facts into nearby load-bearing events and delete the standalone variants. - Token usage now rides on `assistant/message` as an optional `usage` field — the assembled model output and its accounting travel together. The loop folds `assembler.usage` onto the append instead of emitting a separate `usage` event. - The max-tokens path is the no-data-loss host: a step cut off with usage but EMPTY content (e.g. only a dropped tool call) previously emitted a standalone `usage`; it now records an empty-content `assistant/message { content: [], usage }`. `deriveMessages()` skips empty-content assistant messages, so the usage host never injects a spurious content-less assistant turn into the provider transcript. A step with neither content nor usage appends nothing. - An operational error's step number now rides on `turn/end.reason` for `kind: 'error'` (`{ kind: 'error', step, message, code? }`) — the durable turn outcome ACP and resume already consume. `failTurn` sets the reason directly (no separate session `error` event). `agent/error` + logging are unchanged for live diagnostics. - No format-version bump: pre-release, no persisted data, so per the format policy there is nothing to migrate or reject (the RFC's "refresh the format version" criterion over-reached). `version` stays 1. - ACP fixtures + goldens re-recorded (keyless replay): dropped standalone usage/error lines, usage folded onto assistant/message, error step on turn/end.reason. RFC moved proposed -> implemented with an implementation note recording the two scope refinements.
2026-06-21 10:00:06 +08:00
// Fire the assistant/message when there is content OR usage: a max-tokens
// step can be cut off with empty content but still carry token accounting,
// and assistant/message is the only host for usage (there is no standalone
// usage event). An empty-content assistant/message is skipped by
// deriveMessages(), so hosting usage on it never injects a spurious assistant
// turn into derived history.
if (message.content.length > 0 || assembler.usage) {
// A max-tokens finish is itself a streamed `finish` chunk, so chunkSeqs is
// never empty here — pass the provenance unconditionally.
session.append(
'assistant/message',
{ turn, step, content: message.content, ...(assembler.usage ? { usage: assembler.usage } : {}) },
{ surfaceOp: 'append', sourceEventSeqs: chunkSeqs },
)
}
return { hadToolCalls: false, finish: assembler.finish }
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
// The step-result waterfall runs BEFORE the session append so the log (the
// source of truth for derived history and replay) records the message that
// tool dispatch actually uses.
let message: Message = assembler.message()
message = await ctx.waterfall('agent/step-result', agent, turn, step, message, () => Promise.resolve(message))
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
fix review findings: bump session format version + restore late turn-end warn Codex review of the trace-event fold found two merge-blockers. Blocker #1 — format version. Folding usage onto assistant/message and removing the standalone usage/error events changed the persisted SessionEventMap shape, which per the AGENTS.md "bump the version and reject — don't migrate" policy requires a backend to reject any non-current log. Centralize the version in an exported SESSION_FORMAT_VERSION constant (dsh-session), read by both write sites (Session constructor default, SessionStore.prepare header) and the coordinator's load-time assertVersion check. The constant is pinned at 0: while unreleased the on-disk format is unstable/pre-release, so breaking shape churn is absorbed at v0 (no monotonic bump until the first tagged release) and any non-0 log is rejected on load — no migration. Update every test/fixture/doc that stamps a currently-written header to the constant, bump the ACP snapshot fixture + golden headers to v0, and keep the version-rejection test meaningful by switching its bad value to a clearly non-current 99. AGENTS.md documents both the monotonic (SQLite SCHEMA_VERSION) and pinned-0 (session log) pre-release stances. Blocker #2 — restore the late turn-end warn. failTurn now sets the error reason only while the turn is still open; once turn/end is appended (a throwing agent/turn-end listener after closeTurn) the reason can no longer reach the durable log, so the late throw is logged via ctx.logger.warn instead of vanishing into a futile post-close assignment. A regression test asserts the warn fires. Also guard the normal-step assistant/message append with the same content-or-usage condition as the max-tokens branch (a content-less, usage-less step records no trace-only row), with a covering test.
2026-06-21 11:08:10 +08:00
// Same content-or-usage guard as the max-tokens branch: a step that finishes
// with neither assembled content nor usage (e.g. a bare `stop` finish that
// streamed nothing) records no assistant/message — an empty-content message
// exists only to host usage, and deriveMessages() skips it either way, so
// appending one with no usage would be a pure trace-only row.
//
// sourceEventSeqs records the assistant/chunk provenance, but is omitted when
// no chunks streamed (the surface invariant rejects an empty sourceEventSeqs).
fix review findings: bump session format version + restore late turn-end warn Codex review of the trace-event fold found two merge-blockers. Blocker #1 — format version. Folding usage onto assistant/message and removing the standalone usage/error events changed the persisted SessionEventMap shape, which per the AGENTS.md "bump the version and reject — don't migrate" policy requires a backend to reject any non-current log. Centralize the version in an exported SESSION_FORMAT_VERSION constant (dsh-session), read by both write sites (Session constructor default, SessionStore.prepare header) and the coordinator's load-time assertVersion check. The constant is pinned at 0: while unreleased the on-disk format is unstable/pre-release, so breaking shape churn is absorbed at v0 (no monotonic bump until the first tagged release) and any non-0 log is rejected on load — no migration. Update every test/fixture/doc that stamps a currently-written header to the constant, bump the ACP snapshot fixture + golden headers to v0, and keep the version-rejection test meaningful by switching its bad value to a clearly non-current 99. AGENTS.md documents both the monotonic (SQLite SCHEMA_VERSION) and pinned-0 (session log) pre-release stances. Blocker #2 — restore the late turn-end warn. failTurn now sets the error reason only while the turn is still open; once turn/end is appended (a throwing agent/turn-end listener after closeTurn) the reason can no longer reach the durable log, so the late throw is logged via ctx.logger.warn instead of vanishing into a futile post-close assignment. A regression test asserts the warn fires. Also guard the normal-step assistant/message append with the same content-or-usage condition as the max-tokens branch (a content-less, usage-less step records no trace-only row), with a covering test.
2026-06-21 11:08:10 +08:00
if (message.content.length > 0 || assembler.usage) {
session.append(
'assistant/message',
{ turn, step, content: message.content, ...(assembler.usage ? { usage: assembler.usage } : {}) },
{ surfaceOp: 'append', ...(chunkSeqs.length > 0 ? { sourceEventSeqs: chunkSeqs } : {}) },
)
}
// --- Tool execution (sequential; parallel execution is a TODO) ---
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
// ToolRegistry.execute converts tool failures (including aborts) into
// isError results, so abort is re-checked around every call here.
const toolCalls = message.content.filter(block => block.type === 'tool-call')
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
// Per-step buffer of `additionalContext` attached by tools/post-execute
// listeners. Appended as context/message(s) only AFTER every tool/result for
// the step, so a multi-call step keeps tool-call/result adjacency
// (interleaving context between a call's result and the next call's would
// break the pairing the next model request relies on).
const pendingContext: HookContext[] = []
for (const call of toolCalls) {
simplify(agent): drop the unused public Agent.abort(), keep whenIdle() The public Agent handle exposed abort() (step-only) and cancel() (queue-aware). No production caller used abort() — ACP maps session/cancel to cancel(), and lifecycle owners tear down via AgentHandle.dispose(); the loop's own stop paths abort their per-step AbortController directly. So abort() is latent generality that keeps a private loop mechanic public. RFC-premise correction: the public-agent-stop-surface RFC proposed removing whenIdle() too. Implementation found whenIdle() load-bearing — a real quiescence primitive with a deliberate loop contract (settle-without-transition, the replacement-turn race) and ACP test consumers; its proposed replacement ("observe the running->idle transition") is exactly the async-state race AGENTS.md warns against. So only abort() is removed; whenIdle() stays. The RFC is amended on the way to implemented/ to record the narrowed scope, and the new AGENTS.md "RFCs are proposals, not golden truth" principle (PR1) gets its worked example. - Remove Agent.abort() from the interface + the ReactLoopAgent impl; the no-arg 'aborted' default goes with it (cancel() keeps its 'cancelled' default). - Migrate tests: empty-queue abort() -> cancel(reason); the two review-fixes tests whose subject is the in-flight step's AbortController drive that controller directly via the private currentAbort field (cancel() would clear the inbox and destroy the queued steering one of them proves survives a step abort). The no-arg-default test is dropped (cancel()'s default is already covered in cancel.spec.ts). - Resulting public stop surface: cancel() + whenIdle(). Update agent/agent-loop READMEs, architecture.md, core.md type-equiv, the extension cookbook, the lifecycle RFC (short note), and the proposed ACP RFC. Implements docs/rfc/implemented/simplification/2026-06-20-public-agent-stop-surface.md
2026-06-21 05:50:39 +08:00
/* v8 ignore next -- signal.reason always set: cancel()/disposal provide a default */
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
if (signal.aborted) throw new Error(String(signal.reason ?? 'aborted'))
2026-06-17 19:25:29 +08:00
const callEvent = session.append('tool/call', { turn, step, callId: call.id, name: call.name, arguments: call.arguments })
let parsedArguments: unknown
try {
parsedArguments = call.arguments ? JSON.parse(call.arguments) : {}
} catch {
parsedArguments = call.arguments
}
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
// TODO(pre-tool-input-rewrite): tools/pre-execute deliberately cannot rewrite
// `arguments` — tool/call (the audit record) and assistant/message (the
// model-history source) are logged BEFORE execute, and live consumers (ACP,
// tool-bash presentation) read the pre-execution args, so an execution-only
// rewrite would desync the UI from what ran. Designing that consistently is
// its own proposed RFC (docs/rfc/proposed/feature/…-pre-tool-input-rewrite.md).
const result = await ctx.tools.execute({
callId: call.id,
name: call.name,
arguments: parsedArguments,
agent,
signal,
})
session.append('tool/result', {
turn, step,
// The correlation id MUST be the loop's authoritative call.id (the
// model-transcript id that deriveMessages turns into toolCallId), NOT
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
// result.callId — a post-execute waterfall listener returning a
// mismatched id would otherwise orphan the call↔result pairing in the
// next model request. A listener-internal id, if ever needed, belongs in
// a separate diagnostic field, never overloaded onto callId.
callId: call.id,
content: result.content,
isError: result.isError,
...result.error ? { error: result.error } : {},
// The tool's private presentation payload (e.g. a result-time diff),
// persisted so a UI bridge reproduces the card on replay.
...result.meta !== undefined ? { meta: result.meta } : {},
2026-06-17 19:25:29 +08:00
}, { surfaceOp: 'append', sourceEventSeqs: [callEvent.seq] })
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
// Buffer (don't append yet) any post-execute additionalContext for this call.
if (result.additionalContext) pendingContext.push(result.additionalContext)
// signal CAN flip during the await above (abort() inside a tool);
// the analyzer can't see through the await boundary.
simplify(agent): drop the unused public Agent.abort(), keep whenIdle() The public Agent handle exposed abort() (step-only) and cancel() (queue-aware). No production caller used abort() — ACP maps session/cancel to cancel(), and lifecycle owners tear down via AgentHandle.dispose(); the loop's own stop paths abort their per-step AbortController directly. So abort() is latent generality that keeps a private loop mechanic public. RFC-premise correction: the public-agent-stop-surface RFC proposed removing whenIdle() too. Implementation found whenIdle() load-bearing — a real quiescence primitive with a deliberate loop contract (settle-without-transition, the replacement-turn race) and ACP test consumers; its proposed replacement ("observe the running->idle transition") is exactly the async-state race AGENTS.md warns against. So only abort() is removed; whenIdle() stays. The RFC is amended on the way to implemented/ to record the narrowed scope, and the new AGENTS.md "RFCs are proposals, not golden truth" principle (PR1) gets its worked example. - Remove Agent.abort() from the interface + the ReactLoopAgent impl; the no-arg 'aborted' default goes with it (cancel() keeps its 'cancelled' default). - Migrate tests: empty-queue abort() -> cancel(reason); the two review-fixes tests whose subject is the in-flight step's AbortController drive that controller directly via the private currentAbort field (cancel() would clear the inbox and destroy the queued steering one of them proves survives a step abort). The no-arg-default test is dropped (cancel()'s default is already covered in cancel.spec.ts). - Resulting public stop surface: cancel() + whenIdle(). Update agent/agent-loop READMEs, architecture.md, core.md type-equiv, the extension cookbook, the lifecycle RFC (short note), and the proposed ACP RFC. Implements docs/rfc/implemented/simplification/2026-06-20-public-agent-stop-surface.md
2026-06-21 05:50:39 +08:00
/* v8 ignore start -- signal.reason default unreachable: cancel()/disposal always set it */
// eslint-disable-next-line @typescript-eslint/no-unnecessary-condition
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
if (signal.aborted) throw new Error(String(signal.reason ?? 'aborted'))
/* v8 ignore stop */
}
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
// Append buffered post-execute context AFTER every tool/result, preserving
// tool-call/result adjacency across the whole batch. inject() appends into the
// open turn (a context/message at its chronological position).
for (const context of pendingContext) {
agent.inject(context.content, { source: context.source })
}
return { hadToolCalls: toolCalls.length > 0, finish: assembler.finish }
}
2026-06-19 00:37:12 +08:00
function withoutToolCalls(message: Message): Message {
return { ...message, content: message.content.filter(block => block.type !== 'tool-call') }
}
/**
* The last turn number in a (possibly seeded) session log, or 0.
* @param session - the session whose log is scanned for the latest `turn/start`.
* @returns the latest `turn/start`'s turn number, or 0 when the log has none (the next turn is this plus one).
*/
export function lastTurnNumber(session: Session): number {
const lastStart = session.events.findLast(event => event.type === 'turn/start')
return lastStart?.data.turn ?? 0
}
/**
* Whether a turn is currently open in the session log (a `turn/start` with no
* matching later `turn/end`). Decided from the LOG, not agent status: status
* can be `running` while no turn is open (an `agent/status` listener firing
* before `turn/start`, or the post-`turn/end` flush window before status
* returns to idle), so status is not a reliable open-turn signal. Used by
* `inject()` to choose between appending into an open turn vs. wrapping the
* injection in its own one-shot turn (the turn-enclosure RFC).
* @param session - the session whose log is inspected.
* @returns true when the log's last turn boundary is a `turn/start` with no matching `turn/end` yet.
*/
export function isTurnOpen(session: Session): boolean {
const last = session.events.findLast(e => e.type === 'turn/start' || e.type === 'turn/end')
return last?.type === 'turn/start'
}