deepseek-harness/packages/core/agent-loop/src/loop.ts

741 lines
31 KiB
TypeScript
Raw Normal View History

/**
* Drives one agent across queued durable turns. Turn failures are contained so
* later work can run; the session log, not this driver, owns conversation state.
2026-07-19 22:50:49 +08:00
* See .agents/notes/implemented/architecture/2026-06-18-agent-lifecycle-and-ownership-seams.md.
* @module dsh-agent-loop/loop
*/
Document the codebase thoroughly and tighten type safety Docs: per-folder README.md for packages/ (family overview + one per package: service, events, API, extension points, TODOs), examples/, and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md symlinks) for packages/ and vendor/; module-level doc comments in every packages/*/src file; richer JSDoc on all exported API (event side effects, disposal contracts, error behavior). Root AGENTS.md gains a "Type Safety and Documentation" policy section: the codebase aims to be very type-safe and well documented; type gymnastics are acceptable in core packages when they improve plugin-author DX; verbose docs are fine as long as they stay strictly in sync with the code. Type safety: removed the upstream-inherited "noImplicitAny": false from tsconfig.base.json — packages/* now compile under full strict mode; vendor/loader and vendor/include set it locally (vendor/cordis already did). Eliminated every `: any` / `as any` from packages and examples (catch clauses use unknown + a CodedError narrowing type; event data access uses discriminated-union narrowing). Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL — SchemaSpec with per-property `required: true` booleans, type-level InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and defineTool() so first-party tools get typed execute(args) with zero casts (raw JSON Schema still accepted for MCP interop; chosen over schemastery because it targets JSON Schema generation directly). echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
import type { Context } from 'cordis'
import type { ContentBlock, FinishReason, GenerateOptions, LlmCallConfig, Message } from '@deepseek-ai/dsh-llm'
2026-07-14 21:57:52 +08:00
import { isDeepStrictEqual } from 'node:util'
import { BlockAssembler, HarnessError, assertNever, deepFreeze, errorChain, isLlmAdapterFailure } from '@deepseek-ai/dsh-llm'
import { agentEvents, assembleContextFor } from '@deepseek-ai/dsh-agent'
import type { AgentEventDispatch, ContinuationDecision, HookContext, PromptDecision, RequestError, RequestErrorDecision } from '@deepseek-ai/dsh-agent'
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
import { canonicalHeader } from '@deepseek-ai/dsh-session'
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
import type { Session, TurnEndReason, TurnTrigger } from '@deepseek-ai/dsh-session'
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
import { createTransmissionLog, recordRequestHeader } from './request-log.ts'
import type { TransmissionLog } from './request-log.ts'
import { renderPrompt } from '@deepseek-ai/dsh-system-prompt'
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
import type { PromptAssembly } from '@deepseek-ai/dsh-system-prompt'
import type {} from '@deepseek-ai/dsh-tools'
import { executeToolCalls } from './tool-calls.ts'
import type { Inbox } from './inbox.ts'
/** Normalize thrown values while preserving an existing error code. */
function toError(error: unknown): RequestError {
return error instanceof Error ? error : new HarnessError(String(error), 'UNKNOWN', { cause: error })
}
/** Distinguishes final model-request failures from failures in later step processing. */
class TerminalModelRequestFailure extends Error {
constructor(readonly requestError: RequestError) {
super(requestError.message, { cause: requestError })
this.name = 'TerminalModelRequestFailure'
}
}
/** Convert terminal failure finishes into step errors; unknown extensible finishes remain successful. */
function finishError(finish: FinishReason): RequestError | undefined {
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
switch (finish.kind) {
case 'error': {
const error: RequestError = new Error(finish.message)
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
if (finish.code !== undefined) error.code = finish.code
return error
}
case 'aborted': {
const error: RequestError = new Error('model stream aborted')
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
error.code = 'ABORTED'
return error
}
// stop / tool-calls / max-tokens / plugin-added kinds → not a failure.
default:
return undefined
}
}
/**
* Build the `{ message, code? }` part of an error payload, omitting the
* `code` key entirely when absent (exactOptionalPropertyTypes-correct).
* The durable message renders the full cause chain: `turn/end` is the single
* durable record of an in-turn failure, so a wrapper message alone (e.g.
* `fetch failed`) would lose the diagnosis the session log exists to keep.
*/
function errorData(err: RequestError): { message: string; code?: string } {
return { message: errorChain(err), ...typeof err.code === 'string' ? { code: err.code } : {} }
}
/** Map a successful max-token finish onto the turn reason; other successful finishes add nothing. */
function stepFinishReason(finish: FinishReason): TurnEndReason | undefined {
switch (finish.kind) {
case 'max-tokens':
return { kind: 'max-tokens' }
// stop / tool-calls / plugin-added kinds → no turn-end contribution
// beyond the default `completed`. FinishReason is merge-extensible, so a
// default (not assertNever) handles unknown kinds as ordinary success.
default:
return undefined
}
}
/** Mutable agent controls supplied to the loop driver. */
export interface LoopHandle {
/** Native-private agent inbox handed to the driver only at internal startup. */
readonly inbox: Inbox
2026-07-18 14:59:26 +08:00
/** Maximum parallel-safe calls allowed in one step. */
readonly maxParallelToolCalls: number
setStatus(status: 'idle' | 'running'): void
setAbort(controller: AbortController | undefined): void
/** Resolves when the agent is disposed — unblocks the idle wait. */
disposed: Promise<void>
isDisposed(): boolean
/** Whether cancellation is pending for the current loop iteration. */
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
isCancelled(): boolean
/** Resolved pending-cancellation reason; meaningful only while {@link isCancelled} is true. */
cancelReason(): string
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
/** Clear the cancel marker (called once per iteration after the turn returns). */
clearCancel(): void
/** Settle idle waiters before pre-running cancellation publishes idle. */
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
settleIdle(): void
2026-07-15 12:49:50 +08:00
/** Run an active tool-call batch, accepting post-tool context into the FIFO drained before settlement. */
readonly withToolBatch: <T>(run: (acceptContext: (context: HookContext) => void) => Promise<T>) => Promise<T>
}
/**
* Drive queued messages as independent durable turns until disposal. Plugin
* failures end the current turn without terminating the driver. The caller
* establishes the `ctx.agents.withInitiator()` boundary before entry; package-private
* orchestration recovers that exact Agent and captures its Session locally.
* @param ctx - the plugin context the loop reaches its initiating Agent,
* events (agent/…, session/flush), and services (systemPrompt, llm, tools)
* through.
* @param handle - the bridge to the agent's mutable state: status/abort setters plus the disposal and cancel-marker reads.
* @throws when no initiating Agent is active.
*/
export async function runLoop(ctx: Context, handle: LoopHandle): Promise<void> {
const agent = ctx.agents.requireInitiator()
// Per-instance prefix and request-header state; conversation history remains in the session log.
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
const transmission = createTransmissionLog()
const { session } = agent
// Fused subject and scope carrier for every agent event below.
const events = agentEvents(ctx, agent)
while (!handle.isDisposed()) {
// An idle listener can enqueue and cancel replacement work before the next
// wait is installed. Consume that empty marker before parking the driver.
if (handle.isCancelled()) {
handle.clearCancel()
if (!handle.inbox.hasQueued) {
handle.settleIdle()
handle.setStatus('idle')
continue
}
}
await handle.inbox.waitForQueued(handle.disposed)
if (handle.isDisposed()) break
// Cancellation between wake and `running` skips only the cancelled work;
// a replacement prompt still runs before the eventual idle transition.
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
if (handle.isCancelled()) {
handle.clearCancel()
if (!handle.inbox.hasQueued) {
// Settle before publishing idle: the already-idle path has no status
// transition, while an idle listener can register waiters for new work.
handle.settleIdle()
handle.setStatus('idle')
continue
}
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
}
handle.setStatus('running')
if (handle.isDisposed()) break
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
// A synchronous `running` listener can cancel before `runTurn`; balance the
// status only when no replacement prompt was queued by that listener.
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
if (handle.isCancelled()) {
handle.clearCancel()
if (!handle.inbox.hasQueued) {
handle.setStatus('idle')
continue
}
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
}
// Idle injection can add a turn, so derive the next number from the log.
const turn = lastTurnNumber(session) + 1
let terminalStopped = false
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
try {
terminalStopped = await runTurn(ctx, events, handle, turn, transmission)
Document the codebase thoroughly and tighten type safety Docs: per-folder README.md for packages/ (family overview + one per package: service, events, API, extension points, TODOs), examples/, and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md symlinks) for packages/ and vendor/; module-level doc comments in every packages/*/src file; richer JSDoc on all exported API (event side effects, disposal contracts, error behavior). Root AGENTS.md gains a "Type Safety and Documentation" policy section: the codebase aims to be very type-safe and well documented; type gymnastics are acceptable in core packages when they improve plugin-author DX; verbose docs are fine as long as they stay strictly in sync with the code. Type safety: removed the upstream-inherited "noImplicitAny": false from tsconfig.base.json — packages/* now compile under full strict mode; vendor/loader and vendor/include set it locally (vendor/cordis already did). Eliminated every `: any` / `as any` from packages and examples (catch clauses use unknown + a CodedError narrowing type; event data access uses discriminated-union narrowing). Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL — SchemaSpec with per-property `required: true` booleans, type-level InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and defineTool() so first-party tools get typed execute(args) with zero casts (raw JSON Schema still accepted for MCP interop; chosen over schemastery because it targets JSON Schema generation directly). echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
} catch (error: unknown) {
// Pre-turn failure has no durable boundary to close; report it without appending outside a turn.
const err = toError(error)
ctx.logger.warn(`agent "${agent.id}": turn ${turn} failed before it started: ${errorChain(err)}`)
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
try {
events.emit('agent/error', turn, 0, err)
} catch { /* contained: a throwing agent/error listener must not kill the driver */ }
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
}
// Reset per iteration, including when a prompt arrives during the flush window.
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
handle.clearCancel()
// Late steering becomes queued input unless terminal policy stopped the turn.
for (const message of handle.inbox.drainSteering()) {
if (!terminalStopped) handle.inbox.enqueue(message)
}
if (!handle.inbox.hasQueued) handle.setStatus('idle')
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
}
}
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
async function runTurn(
ctx: Context, events: AgentEventDispatch, handle: LoopHandle, turn: number, transmission: TransmissionLog,
): Promise<boolean> {
const agent = ctx.agents.requireInitiator()
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
const { session } = agent
const drainSteering = (): boolean => {
const messages = handle.inbox.drainSteering()
for (const message of messages) {
session.append('steering/message', { turn, content: message.content, source: message.source }, { surfaceOp: 'append' })
}
return messages.length > 0
}
// Claim one queued message before opening its turn, but append it only after `turn/start`.
const message = handle.inbox.dequeueQueued()
/* v8 ignore next 3 -- invariant guard: runLoop only calls runTurn when hasQueued */
if (!message) throw new Error('runTurn invariant violated: no queued message at turn start')
const trigger: TurnTrigger = { kind: 'message', source: message.source }
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
let reason: TurnEndReason = { kind: 'completed' }
let step = 0
let requestRetryAttempt = 0
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
let stepOpen = false
let errorReported = false
let terminalStopped = false
// Close the committed step once; pre-commit validation failure still escapes.
const closeStep = (): void => {
if (!stepOpen) return
session.append('step/end', { turn, step })
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
stepOpen = false
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
// Record the durable turn failure once and contain the live error notification.
const failTurn = (err: RequestError): void => {
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
if (errorReported) return
errorReported = true
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
reason = { kind: 'error', step, ...errorData(err) }
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
try {
events.emit('agent/error', turn, step, err)
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
} catch {
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
// contained: the error is already captured on `reason`; a throwing
// agent/error listener must not prevent the turn from closing.
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
}
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
// Pre-commit validation failure escapes rather than masquerading as a committed boundary.
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
const closeTurn = (): void => {
session.append('turn/end', { turn, reason })
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
try {
// --- Turn boundary. Once turn/start is appended, a turn/end is owed no
// matter what throws below; the catch + closeTurn guarantee it. A pre-commit
// veto leaves no turn/start in the log and therefore owes no turn/end.
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
session.append('turn/start', { turn, trigger })
// The claimed message runs the `agent/prompt-submit` waterfall before it
// becomes a `user/message` — a hook can rewrite the prompt or block it.
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
// Recorded INSIDE the turn (after turn/start) so every event is turn-enclosed;
// turn/end is now owed, so a throwing prompt-submit listener (the waterfall
// throws) is caught below and the turn still closes.
const promptDecision = await events.waterfall(
'agent/prompt-submit', message.content, message.source,
() => Promise.resolve<PromptDecision>({ kind: 'allow' }),
)
if (promptDecision.kind === 'block') {
session.append('prompt/blocked', { content: message.content, source: message.source, reason: promptDecision.reason })
reason = { kind: 'rejected', reason: promptDecision.reason }
} else {
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
// `allow.content` REPLACES the prompt bytes (a rewrite); absent keeps them.
const content = promptDecision.content ?? message.content
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
session.append('user/message', { content, source: message.source }, { surfaceOp: 'append' })
2026-07-13 16:31:03 +08:00
// Every `allow.additionalContexts` entry is a separate context/message the
// next request also sees. The turn is open, so inject() appends each one
// into THIS turn without flattening provenance or metadata.
for (const context of promptDecision.additionalContexts ?? []) {
2026-07-13 16:31:03 +08:00
agent.inject(context.content, {
source: context.source,
...context.meta !== undefined ? { meta: context.meta } : {},
})
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
}
}
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
while (true) {
// A blocked prompt closes its zero-step turn as rejected.
if (promptDecision.kind === 'block') break
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
step += 1
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
// Steering from the previous round's continuation listeners joins before
// the request.
drainSteering()
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
// The step's AbortController exists BEFORE any async pre-step work so a
// dispose() or cancel() — in a synchronous turn-start listener or an
// async listener whose effect fires before we block — always has an armed
// abort to cancel against. isDisposed below covers disposal, which does
// NOT set the cancel marker. Cleared on every exit path below.
const abort = new AbortController()
handle.setAbort(abort)
// Assemble once before pre-step so listener work and the request share one prompt value.
const assembly = await ctx.systemPrompt.assemble(assembleContextFor(agent))
const fullSystemPrompt = renderPrompt(assembly)
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
// Cancellation or disposal during assembly ends the turn before any step opens.
if (handle.isCancelled() || handle.isDisposed()) {
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
handle.setAbort(undefined)
reason = handle.isDisposed() ? { kind: 'disposed' } : { kind: 'aborted', reason: handle.cancelReason() }
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
break
}
// Compose the request-only prefix once per loop instance before the first
// request boundary. It precedes all derived history and is recorded only
// in the request header, not as session history.
if (transmission.sessionPrefix === undefined) {
const emptyPrefix: Message[] = deepFreeze([])
Merge origin/master: scope-aware fusion of the tools/execute seam, session-prefix, and tool-cordis Master brought 50 commits (the tool-cordis group, dsh-code-runtime + worker, the tools/execute around-dispatch seam + timeout-policy, repeat-tool-guard, agent/session-prefix, the ui reorganization). Beyond the ten textual conflicts, the merge reconciles master's new seams with this branch's scoped-registration world: - tools/execute (new waterfall around core dispatch): dispatched with the SAME exec.agent carrier as the pre/post waterfalls — an agent.ctx wrapper times/retries only its own agent's calls — and its base thunk resolves the tool through the caller's visible view (get(exec.name, exec.agent)), so a scoped/shadowed tool dispatches and a restricted-away global stays UNKNOWN_TOOL. Declared this: Scoped<ToolRegistry> with the scope-filtered doc sentence; invariants table + verify-scoped-dispatch pin it (21 events). - agent/session-prefix (new waterfall, once per loop instance): composed via the fused agentEvents dispatcher (scope-filtered like every agent-subject event), declared this: Scoped<Agent>, table-pinned. agent/pre-step keeps master's new sessionPrefix parameter with this branch's Scoped this. - timeout-policy reads the budget through the caller's visible view (get(exec.name, exec.agent)): a scoped tool's own timeoutMs governs its calls; a global name-twin's budget is never misapplied to a shadowing per-agent variant. - tool-cordis: cordis_inspect's tools section lists the CALLING agent's view (its description promises "what you can call"); the sandbox tool façade's reads resolve through the mount's own scope, mirroring where its register lands writes; sandboxRegisterTool's return type carries the exact-disposer union honestly. dsh-scope declared as peer+dev with the project reference. - doc-sync chain unions master's verify-cordis-api with this branch's verify-scoped-dispatch; the generated catalogs, event matrix (the zero-dispatcher guard passes over master's new events), module graph, and the cordis api-catalog are regenerated on the merged surface. Full gate sequence green on the merged tree: typecheck, lint, per-file 100% coverage (2668 tests), snapshots (38), doc-sync, module graph, build, hygiene, demo smoke.
2026-07-09 23:24:42 +08:00
const composed = await events.waterfall(
'agent/session-prefix', emptyPrefix, abort.signal,
() => Promise.resolve(emptyPrefix),
)
// Never cache an interrupted composition; the next turn recomposes it.
if (handle.isCancelled() || handle.isDisposed()) {
handle.setAbort(undefined)
reason = handle.isDisposed() ? { kind: 'disposed' } : { kind: 'aborted', reason: handle.cancelReason() }
break
}
transmission.sessionPrefix = deepFreeze(structuredClone(composed))
}
// Await surface mutations outside the step before snapshotting history.
await events.serial('agent/pre-step', turn, step, abort.signal)
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
// Interruption landing during the pre-step seam: do not open an empty step.
if (handle.isCancelled() || handle.isDisposed()) {
handle.setAbort(undefined)
reason = handle.isDisposed() ? { kind: 'disposed' } : { kind: 'aborted', reason: handle.cancelReason() }
break
}
// Snapshot the exact log prefix before step/start: the reconstruction
// boundary. Appends after this synchronous snapshot join the next request.
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
const boundaryMessages = session.deriveMessages()
session.append('step/start', { turn, step })
// Only a committed step/start creates a balancing obligation. A
// pre-commit veto throws before this assignment; post-commit observers
// are contained inside Session.append().
stepOpen = true
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
// Cancel landing in the step-start window: a synchronous `session/event`
// step/start listener can cancel after the step is already open. Check
// AFTER the step/start append and before `runStep`: drop the step, end the
// turn accordingly. closeStep balances the already-appended step/start.
if (handle.isCancelled() || handle.isDisposed()) {
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
handle.setAbort(undefined)
reason = handle.isDisposed() ? { kind: 'disposed' } : { kind: 'aborted', reason: handle.cancelReason() }
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
closeStep()
break
}
let stepOutcome:
| { hadToolCalls: boolean; finish: FinishReason }
| { requestError: RequestError }
| { error: RequestError }
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
try {
stepOutcome = await runStep(
ctx, events, handle, turn, step, assembly, fullSystemPrompt, boundaryMessages, transmission, abort.signal)
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
} catch (error: unknown) {
if (error instanceof TerminalModelRequestFailure) {
stepOutcome = { requestError: error.requestError }
} else {
stepOutcome = { error: toError(error) }
}
}
if ('requestError' in stepOutcome) {
// Recovery observes a balanced failed step and the original provider
// error while the failed step's signal remains the active owner.
closeStep()
if (handle.isDisposed() || abort.signal.aborted) {
handle.setAbort(undefined)
reason = handle.isDisposed()
? { kind: 'disposed' }
: { kind: 'aborted', reason: String(abort.signal.reason) }
break
}
const defaultDecision: RequestErrorDecision = { action: 'fail' }
let recoveryDecision: RequestErrorDecision = defaultDecision
try {
recoveryDecision = await events.waterfall(
'agent/request-error', turn, step, stepOutcome.requestError,
requestRetryAttempt, abort.signal,
() => Promise.resolve(defaultDecision),
)
} catch (recoveryError: unknown) {
ctx.logger.warn(
`agent "${agent.id}": request recovery failed at turn ${turn}, step ${step}: ${errorChain(recoveryError)}`,
)
}
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
handle.setAbort(undefined)
// Cancellation and disposal always win over either a recovery decision
// or a recovery-listener failure.
// eslint-disable-next-line @typescript-eslint/no-unnecessary-condition
if (handle.isDisposed() || abort.signal.aborted) {
reason = handle.isDisposed()
? { kind: 'disposed' }
: { kind: 'aborted', reason: String(abort.signal.reason) }
break
}
switch (recoveryDecision.action) {
case 'retry':
requestRetryAttempt += 1
continue
case 'fail':
failTurn(stepOutcome.requestError)
break
/* v8 ignore next -- closed-union exhaustiveness guard */
default:
assertNever(recoveryDecision, 'agent request-error decision')
}
break
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
if ('error' in stepOutcome) {
// Steering that arrived during the failed step stays in the inbox —
// runLoop re-enqueues it as a queued message, so an abort-then-steer
// starts a fresh turn instead of being silently consumed.
closeStep()
handle.setAbort(undefined)
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
const { error } = stepOutcome
/* v8 ignore next -- narrow race: disposal while non-request step work throws. */
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
if (handle.isDisposed()) {
reason = { kind: 'disposed' }
} else if (abort.signal.aborted) {
simplify(agent): drop the unused public Agent.abort(), keep whenIdle() The public Agent handle exposed abort() (step-only) and cancel() (queue-aware). No production caller used abort() — ACP maps session/cancel to cancel(), and lifecycle owners tear down via AgentHandle.dispose(); the loop's own stop paths abort their per-step AbortController directly. So abort() is latent generality that keeps a private loop mechanic public. RFC-premise correction: the public-agent-stop-surface RFC proposed removing whenIdle() too. Implementation found whenIdle() load-bearing — a real quiescence primitive with a deliberate loop contract (settle-without-transition, the replacement-turn race) and ACP test consumers; its proposed replacement ("observe the running->idle transition") is exactly the async-state race AGENTS.md warns against. So only abort() is removed; whenIdle() stays. The RFC is amended on the way to implemented/ to record the narrowed scope, and the new AGENTS.md "RFCs are proposals, not golden truth" principle (PR1) gets its worked example. - Remove Agent.abort() from the interface + the ReactLoopAgent impl; the no-arg 'aborted' default goes with it (cancel() keeps its 'cancelled' default). - Migrate tests: empty-queue abort() -> cancel(reason); the two review-fixes tests whose subject is the in-flight step's AbortController drive that controller directly via the private currentAbort field (cancel() would clear the inbox and destroy the queued steering one of them proves survives a step abort). The no-arg-default test is dropped (cancel()'s default is already covered in cancel.spec.ts). - Resulting public stop surface: cancel() + whenIdle(). Update agent/agent-loop READMEs, architecture.md, core.md type-equiv, the extension cookbook, the lifecycle RFC (short note), and the proposed ACP RFC. Implements docs/rfc/implemented/simplification/2026-06-20-public-agent-stop-surface.md
2026-06-21 05:50:39 +08:00
/* v8 ignore next -- signal.reason always set: cancel()/disposal provide a default */
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
reason = { kind: 'aborted', reason: String(abort.signal.reason ?? 'aborted') }
} else {
failTurn(error)
}
break
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
requestRetryAttempt = 0
// Preserve max-token completion unless a later disposal, abort, or error wins.
const stepReason = stepFinishReason(stepOutcome.finish)
if (stepReason) reason = stepReason
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
// Steering that arrived during streaming/tool execution.
const steered = drainSteering()
try {
await events.serial('agent/post-step', turn, step, abort.signal)
} catch (error: unknown) {
stepOutcome = { error: toError(error) }
}
if ('error' in stepOutcome) {
closeStep()
handle.setAbort(undefined)
/* v8 ignore next -- narrow race: disposal while a post-step listener throws. */
if (handle.isDisposed()) {
reason = { kind: 'disposed' }
} else if (abort.signal.aborted) {
/* v8 ignore next -- signal.reason always set by cancellation or disposal. */
reason = { kind: 'aborted', reason: String(abort.signal.reason ?? 'aborted') }
} else {
failTurn(stepOutcome.error)
}
break
}
if (handle.isDisposed() || abort.signal.aborted) {
reason = handle.isDisposed()
? { kind: 'disposed' }
: { kind: 'aborted', reason: String(abort.signal.reason) }
closeStep()
handle.setAbort(undefined)
break
}
closeStep()
handle.setAbort(undefined)
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
const defaultDecision: ContinuationDecision = { action: stepOutcome.hadToolCalls || steered ? 'continue' : 'stop' }
let decision: ContinuationDecision
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
try {
decision = await events.waterfall(
'agent/turn-continuation', turn, defaultDecision,
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
() => Promise.resolve(defaultDecision),
)
} catch (error: unknown) {
// A broken continuation plugin ends the turn, not the loop.
failTurn(toError(error))
break
}
// A continuation reason becomes next-step steering.
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
if (decision.action === 'continue' && decision.reason) {
handle.inbox.steer({ content: decision.reason.content, source: decision.reason.source })
feat(events): interception seams — the typed-Decision surface for hooks Reshape the agent's interception surface so every seam returns a small, typed Decision union, and the set covers the hook points a CC/Codex bridge (and a native plugin) needs. "Native hooks" are not a package — a native hook is just a cordis plugin on these canonical events; the bridges (a later PR) only translate an external protocol onto the same surface. dsh-agent: - NEW agent/session-start(agent, source) emit (once before turn 1; SessionStartSource startup|resume|clear|compact) — a pure notification, seeds context via inject(). - NEW agent/prompt-submit waterfall → PromptDecision (allow, optionally rewriting the prompt or attaching additionalContext, or block). - RESHAPE agent/turn-continuation boolean → ContinuationDecision ({action:'stop'} | {action:'continue', reason?}; a continue reason is recorded as next-step steering). - New HookContext envelope (required source — inject() would mislabel a missing one). dsh-tools: split the single tools/execute waterfall into tools/pre-execute (PreToolDecision allow/deny/ask gate) and tools/post-execute (PostToolDecision accept/block, optionally replacing content or attaching additionalContext). Core dispatch sits between as plain code; the tool body keeps its inner try/catch so a thrown tool still reaches post-execute as an isError. ToolExecutionResult gains additionalContext (ferried to the loop's per-step buffer). Input rewrite is deliberately NOT offered (a proposed RFC designs it consistently). dsh-session: new `rejected` TurnEndReason — a turn whose whole prompt batch was blocked by prompt-submit. agent-loop firing points: session-start emitted at create (source threaded — startup for create/fork, resume for resume()); prompt-submit per drained message with the always-open-turn rule (a fully-blocked batch is a zero-step rejected turn); the continuation reshape; post-tool additionalContext buffered and appended after all tool/results (adjacency). ACP codec maps rejected→cancelled. A worked native-plugin example (interception.spec.ts) proves all four seams compose end-to-end through the real loop with NO hook/* events (those belong to the bridge lib). All existing tools/execute + turn-continuation tests migrated. The tool-subagent abort test now aborts after a microtask so it still exercises the live onAbort bridge (execute() awaits pre-execute before the body runs). RFCs: implemented/feature/2026-06-30-interception-seams.md (the reshape) + proposed/feature/2026-06-30-pre-tool-input-rewrite.md (the deferred rewrite design).
2026-06-30 17:11:18 +08:00
}
let shouldContinue = decision.action === 'continue'
// Pending steering overrides an ordinary stop.
if (!shouldContinue && handle.inbox.hasSteering) shouldContinue = true
// Terminal policy is monotonic and runs after ordinary continuation folding.
let terminalStop = false
try {
const stop = await events.serial('agent/turn-stop', turn)
terminalStop = stop !== undefined
} catch (error: unknown) {
// A broken terminal policy is an ordinary continuation failure: fail
// this turn closed while leaving the driver alive for later turns.
failTurn(toError(error))
break
}
if (terminalStop) {
terminalStopped = true
// Terminal stop discards steering but preserves ordinary queued prompts.
handle.inbox.drainSteering()
shouldContinue = false
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
// The marker catches cancellation after the step controller was cleared.
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
if (handle.isCancelled()) {
reason = { kind: 'aborted', reason: handle.cancelReason() }
feat(agent): add queue-aware Agent.cancel() primitive abort() only kills the in-flight step, so a queued-but-not-yet-started prompt ran to completion after a cancel and a prompt accepted right after could be batched into the cancelled turn (the loop merges queued messages into one turn). This closes TODO(rfc010-cancel-prestep) with a distinct cancel() verb. cancel() clears the queued + steering FIFOs, aborts the in-flight step, and drives a turn-scoped marker on the LoopHandle that the driver checks at EVERY point a turn could start or continue: - right after the idle wait (window 1): drop the about-to-run turn and settle whenIdle() waiters directly (no running→idle transition fires, and no agent/status is emitted, so an ACP listener can't see a spurious idle that resolves a freshly-queued prompt as cancelled); - after the synchronous setStatus('running') emit (window 2): a running listener can cancel in the gap before runTurn; - in the step-start window (before runStep, after setAbort): a synchronous turn-start/step-start listener can cancel before any AbortController exists; - at the continuation gate: a cancel during the continuation waterfall (the finished step's controller already cleared) ends the turn aborted. The marker is ARMED only when there is something to cancel (running, an in-flight step, or queued/steering work) — an idle no-op cancel cannot leave it set to drop a later prompt — and RESET unconditionally once per loop iteration, so it governs exactly one turn and never leaks onto the next prompt (even when a send() lands in the cancelled turn's flush window). ACP session/cancel now maps to agent.cancel() (keeping the synchronous settlePrompt). Teardown/disconnect still use abort('disposed') until PR D, so the ACP README narrows the remaining best-effort window to teardown only. Tests (agent-loop/cancel.spec.ts) cover every window unit-level (the F1 hang guard: a whenIdle() waiter registered before a pre-step cancel resolves; the F2 leak guard: idle cancel then a prompt runs; mid-step, continuation, both pre-step windows, turn-start-listener, steering-cleared, marker-reset). ACP turns.spec.ts adds the through-bridge tests with NO intervening whenIdle (idle cancel→prompt runs; mid-stream cancel→immediate next prompt runs) and updates the stale pre-step test to the queue-aware guarantee. The existing cancel snapshot golden is byte-identical (it drives the new cancel() path end-to-end through the real subprocess), so no new golden is needed. 100% coverage.
2026-06-20 04:51:32 +08:00
break
}
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
if (!shouldContinue || handle.isDisposed()) {
/* v8 ignore next -- disposal during continuation-decision window is a narrow race; error-path disposal is covered elsewhere */
if (handle.isDisposed()) reason = { kind: 'disposed' }
break
}
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
// Normal / inline-error loop exit: close the turn.
closeTurn()
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
} catch (error: unknown) {
// Close only a turn whose start committed to the log.
fix(agent-loop): decide turn balance + idle-injection flush from the log (review #32) Session.append pushes the event BEFORE notifying session/event listeners, so a throwing listener leaves the event in the log while the line after the append (a boolean flag) never runs. Both turn-balance decisions were gated on such flags, so a throwing listener could strand an open turn or skip a durability checkpoint. - loop.ts: the outer catch decided "turn/end owed" from `turnStarted`. A throwing listener on the turn/start append left turn/start logged but the flag false → catch rethrew and skipped turn/end → permanently open turn (violating ADR 0017). Now decided from the log (this turn's turn/start present), so the turn is always balanced; only a genuine pre-push failure (non-serializable trigger — turn/start never logged) is rethrown to the runLoop backstop. Removed the now-dead `turnStarted`. - agent.ts inject(): the idle one-shot-turn flush was gated on a `turnRecorded` flag set after append('turn/end'); a throwing turn/end listener skipped the flush, losing the balanced in-memory injection turn on crash. Now the flush decision is read from the log, the synthetic turn/end append contains a throwing listener (turn stays balanced), and a failing idle flush is reported via agent/error (step 0 convention) AND the logger — mirroring the loop's post-turn/end flush path — with a throwing agent/error listener contained. Rewrote the test that encoded the old (buggy) "turn/start listener throw is rethrown, no turn/end" semantics to assert the balanced-turn contract, and added regressions for the throwing-turn/end-listener flush and the agent/error report. Updated Agent.inject JSDoc.
2026-06-15 23:44:54 +08:00
const turnStartLogged = session.events.some(e => e.type === 'turn/start' && e.data.turn === turn)
if (!turnStartLogged) throw error
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
closeStep()
// Preserve an established disposal reason; otherwise report the failure.
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
if (handle.isDisposed() && !errorReported) { // eslint-disable-line @typescript-eslint/no-unnecessary-condition
reason = { kind: 'disposed' }
} else {
failTurn(toError(error))
}
refactor(events): remove the turn boundary mirror events Complete the boundary-mirror removal begun with the step mirrors: drop `agent/turn-start` and `agent/turn-end` from the agent event taxonomy. Turn and step boundaries are now read exclusively off the durable `session/event` feed (`turn/start`/`turn/end`/`step/start`/`step/end`) — there is no `agent/*` mirror for any boundary. - loop.ts: delete both turn emits; `closeTurn` loses its `emit` parameter and its now-unreachable idempotency guard (it is called exactly once per turn, on mutually exclusive normal/catch paths); `failTurn` loses the dead post-close branch that only a throwing turn-end LISTENER could reach. - ui-stdio: render turn boundaries from `session/event`, recovering the short agent label from an `agent/created`→id map (the `turn/start` event carries only the turn number, and the session id is not reliably the agent id). ui-stdio is a disposable test REPL, so this migration retires the sole justification the event-domain-semantics RFC gave for KEEPING the turn mirrors. - Tests: reason/turn-number collectors and the boundary-ordering test now read `session/event`; the throwing-turn-boundary-LISTENER tests are deleted (that code path no longer exists). A new test covers the outer-catch disposed branch via a pre-step listener that disposes-then-throws (the surviving real path). - Docs: promote the "remove agent boundary mirror events" RFC to implemented (amended/narrowed — `agent/steering` is RETAINED, not a boundary mirror); update the event-domain-semantics + turn-enclosure RFCs, architecture.md, the cookbook, the ACP/agent/ui-stdio prose, and regenerate the cordis catalog. `agent/steering` and `agent/stream-chunk` are explicitly out of scope (not durable-boundary mirrors). ACP is unaffected — it already settles from the log's `turn/end` + `agent/status`; snapshot goldens are byte-unchanged.
2026-07-02 03:26:45 +08:00
closeTurn()
fix(agent-loop): always close a started turn and any open step on error (P1-5) After turn/start was appended, nothing guaranteed a matching turn/end: a throw from a boundary emit (agent/turn-start, agent/step-start, the normal-path agent/turn-end) escaped runTurn, and the outer runLoop backstop logged an error but never appended turn/end — leaving an unbalanced turn that replay, telemetry, and the invariants plugin all assume is impossible. runTurn is restructured around idempotent finalizers that satisfy the four traps a naive finally would hit: - closeStep()/closeTurn(emit) are guarded (stepOpen/turnEnded) so they run at most once; the agent/step-end and agent/error emits are contained so a throwing listener can't strand the turn open. - failTurn() records the single error event + reason and emits agent/error exactly once (errorReported guard) — no double-logging when the outer catch also runs (e.g. a step error followed by a throwing turn-end listener). - the catch closes an open step BEFORE turn/end (invariants reject turn/end while a step is open), and rethrows ONLY pre-turn throws (turnStarted false), where no turn/end is owed, so the backstop still nets them. - disposal precedence: reason stays disposed only when disposed AND no error was reported; otherwise the error reason wins. Tests (with the invariants plugin loaded as a balance oracle): throwing turn-start (one error, one turn/end, no step), throwing step-start (step/end before turn/end), throwing agent/error on a step-error path (balanced, loop survives), disposal mid-turn (reason disposed, no error event), a pre-turn turn/start-append throw (rethrown to the backstop, no turn/end owed), and a step error + throwing turn-end listener (error logged exactly once). Verified all six fail against a simulated finalizer bypass. dsh-invariants added as an agent-loop devDependency (test-only oracle; no package cycle).
2026-06-14 23:53:37 +08:00
}
// Flush through the store-owned durability checkpoint without killing the driver on failure.
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
try {
await ctx.sessions.flush(session)
Document the codebase thoroughly and tighten type safety Docs: per-folder README.md for packages/ (family overview + one per package: service, events, API, extension points, TODOs), examples/, and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md symlinks) for packages/ and vendor/; module-level doc comments in every packages/*/src file; richer JSDoc on all exported API (event side effects, disposal contracts, error behavior). Root AGENTS.md gains a "Type Safety and Documentation" policy section: the codebase aims to be very type-safe and well documented; type gymnastics are acceptable in core packages when they improve plugin-author DX; verbose docs are fine as long as they stay strictly in sync with the code. Type safety: removed the upstream-inherited "noImplicitAny": false from tsconfig.base.json — packages/* now compile under full strict mode; vendor/loader and vendor/include set it locally (vendor/cordis already did). Eliminated every `: any` / `as any` from packages and examples (catch clauses use unknown + a CodedError narrowing type; event data access uses discriminated-union narrowing). Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL — SchemaSpec with per-property `required: true` booleans, type-level InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and defineTool() so first-party tools get typed execute(args) with zero casts (raw JSON Schema still accepted for MCP interop; chosen over schemastery because it targets JSON Schema generation directly). echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
} catch (error: unknown) {
// The turn is closed, so report the failed flush live rather than append outside a turn.
const err = toError(error)
ctx.logger.warn(`agent "${agent.id}": session/flush failed at turn ${turn}: ${errorChain(err)}`)
try {
events.emit('agent/error', turn, step, err)
} catch {
// contained: a throwing agent/error listener must not escape the loop.
}
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
}
return terminalStopped
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
}
/**
* Run one committed step: transform call config, log the request header, build
* the request from the cached prefix plus the step-boundary snapshot, stream and
* record the response, then execute tools. The caller has already assembled the
* prompt, run `agent/pre-step`, snapshotted history, and opened the step.
*/
async function runStep(
ctx: Context,
events: AgentEventDispatch,
2026-07-15 12:22:39 +08:00
handle: LoopHandle,
turn: number,
step: number,
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001) Codex round 1 CBR-001: a head-anchored compaction checkpoint was mis-classified by the log-position step-alignment scan, so a second auto-compaction over a checkpoint-headed surface silently failed. Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a `replace` op lands a checkpoint at a high log seq whose SURFACE position is the head — its log neighbours (the open step's assistant/message) are not its surface neighbours, so the forward scan wrongly reported mid-step. Fix, per the agreed direction: - Replace the two log-position predicates with one surface-anchored helper `isToolPairingBalanced(nodes, events, beforeSeq)` in `dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is balanced when no unanswered tool-call precedes it on the surface; a region is collapsible iff both edges are balanced cuts. The open-tail and free-node cases fall out of the same counter. It also throws on a corrupt surface (a tool/result with no matching call). - Move compaction off the in-step seam to a new "pre-step" seam fired after turn/start and before step/start, so a compaction's log-only compact/* records and its replacement node land cleanly OUTSIDE any step (the honest structure crash-safety relies on). Renamed the event agent/pre-request → agent/pre-step and switched its dispatch from parallel → serial (listeners mutate the surface as a side effect; serial isolates them so concurrent appends can't interleave). Extended the catalog generator to accept @mode serial. Regression coverage: a real-loop test driving an auto-compaction asserts the landed checkpoint is a balanced cut on both sides; unit tests pin the checkpoint case, the mid-step injection case, multi-call steps, and the corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
assembly: PromptAssembly,
system: string,
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
boundaryMessages: Message[],
transmission: TransmissionLog,
signal: AbortSignal,
): Promise<{ hadToolCalls: boolean; finish: FinishReason }> {
const agent = ctx.agents.requireInitiator()
const { session, options } = agent
// Seed the first request from agent options and later requests from the logged header;
// detach and freeze so listeners must return an attributable replacement.
const seedConfig: LlmCallConfig = deepFreeze(structuredClone(transmission.loggedHeader
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
// eslint-disable-next-line @typescript-eslint/no-non-null-assertion -- loggedHeader ⟹ a snapshot is in the log
? session.requestHeader()!.config
2026-07-14 21:57:52 +08:00
: { provider: options.provider ?? '', model: options.model ?? '' }))
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
// Listener replacements are recorded in the request header before dispatch.
const config = await events.waterfall('agent/request', turn, step, seedConfig, () => Promise.resolve(seedConfig))
2026-07-14 21:57:52 +08:00
if (!config.provider || !config.model) {
throw new Error(`agent "${agent.id}" has no provider/model: set AgentOptions.provider and AgentOptions.model or supply both via the agent/request waterfall`)
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
}
// eslint-disable-next-line @typescript-eslint/no-non-null-assertion -- runTurn composes the prefix before every runStep call
const sessionPrefix = transmission.sessionPrefix!
feat(agent): add the agent/request-messages request-only message seam A new waterfall near request construction lets plugins contribute request-ONLY messages framing the derived history: RequestMessages { before, after } with a frozen empty seed, fired inside the open step after the agent/request config waterfall, so the step/start boundary snapshot and its same-sync-frame invariant are untouched. The request becomes messagePrefix + boundary snapshot + messageSuffix. Contributions never enter session history — deriveMessages() is unchanged — so the request header is their durable record: EpochHeader gains messagePrefix/messageSuffix (canonical absence for empty arrays), request/header-delta replaces either array whole with an empty array encoding the transition back to absence, and the dev-mode reconstruction cross-check now expects the folded header's framing around the boundary derivation. This is the seam for per-request advisory context that must be model-visible now without becoming durable history (a skills catalog, an environment reminder), keeping the base system prompt workspace-independent and provider prefix caches stable. The docs carry the channel cost model: session-frozen content belongs in before, low-frequency change notices belong in durable history via inject() (paid once, prefix-cached thereafter), and after is reserved for small frequently-refreshed state snapshots re-paid on every request they ride. No shipped producer yet, so ACP snapshot fixtures are byte-identical.
2026-07-07 19:42:30 +08:00
// Record the canonical header, including the otherwise-unlogged prefix, before dispatch.
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
const header = canonicalHeader({
config,
...system ? { system } : {},
...assembly.tools.length > 0 ? { tools: assembly.tools } : {},
...sessionPrefix.length > 0 ? { messagePrefix: sessionPrefix } : {},
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
})
recordRequestHeader(session, transmission, header)
// Freeze the logged header plus boundary snapshot; the prefix precedes derived history.
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
const request: GenerateOptions = deepFreeze({
2026-07-14 21:57:52 +08:00
provider: header.config.provider,
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
model: header.config.model,
messages: [...header.messagePrefix ?? [], ...boundaryMessages],
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
...header.system !== undefined ? { system: header.system } : {},
...header.tools !== undefined ? { tools: header.tools } : {},
...header.config.temperature !== undefined ? { temperature: header.config.temperature } : {},
...header.config.maxTokens !== undefined ? { maxTokens: header.config.maxTokens } : {},
...header.config.stop !== undefined ? { stop: header.config.stop } : {},
Add per-session snapshot replay for nested agents (PR2.5) The snapshot tier was built single-session: dsh-llm-replay served calls from one global positional cursor, and the harness harvested one session log. A subagent runs as a second agent with its own session, so a parent→child scenario could neither replay deterministically nor harvest the child's log. This resolves the TODO(subagent-snapshots) deferral from the subagent RFC. - Stamp the calling session id onto the model request: GenerateOptions.sessionId (typed Branded<'SessionId'> to avoid the dsh-llm↔dsh-session cycle), set by the agent loop from agent.session.id. Adapters ignore it; an llm/stream listener routes by it. - Key replay per session: dsh-llm-replay loads the parent log plus one per child (childFiles / $DSH_SNAPSHOT_CHILD_FILES), derives a script per recorded session, and binds each live (freshly-random) session to a recorded script by first-call order — parent first (earliest createdAt, first to stream). Keys by WHO calls, so it survives a future concurrent/backgrounded subagent; a global cursor would not. An unrecorded extra session fails loud. - Harvest every log: the harness collects all .jsonl across cwd buckets, ordered primary-first (top-level, then children by createdAt), and RunResult exposes the plural sessionLogs. The spec writes each back on record (session.jsonl + session.<n>.jsonl) and diffs each against its fixture on replay. - Wire the subagent seam + spawn + fork + tool into the acp-agent example (both cordis configs) and add two nested scenarios recorded against the real API: subagent-spawn (parent + 1 child) and subagent-multi (parent + 2 children, 3 sessions). Both replay keyless in the default gate. A new RFC documents the design (docs/rfc/implemented/testing/). Single-session replay is unchanged (a call with no sessionId is one anonymous primary session). TODO follow-up: a dedicated branded-ids package could own the SessionId brand and dissolve the cross-package cycle note; out of scope for this testing PR.
2026-06-22 08:39:36 +08:00
sessionId: session.id,
signal,
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall The loop is now transmission-stateless; a request is a pure function of (session log, this step's rendered assembly, current AgentOptions): - The reconstruction boundary is step/start: the messages snapshot is taken in the same synchronous frame immediately before the step/start append, so the request's messages are exactly the derivation over events[0..stepStartSeq) — an inject() from an agent/request listener (or any concurrent task) lands after the boundary and joins the NEXT request. This changes behavior for a synchronous step/start session/event listener that appends content (master derived after the append, so such a listener could reach the current request): agent/pre-step is the sanctioned seam for current-request content. - agent/request is re-typed to config-only: (agent, turn, step, config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes from AgentOptions on a loop instance's first request (explicit options beat the logged baseline — fork overrides and resume reconfiguration stay correct) and from the log's folded header afterwards; listeners return a replacement to switch. Content shaping through the request is no longer expressible — model-visible content flows through the log channels. - recordRequestHeader appends whatever header event the request owes the log before dispatch: an 'initial'/'resume' snapshot anchoring each loop instance, a round-trip-verified delta on change, a 'fallback' snapshot when the encoding cannot express it. Session.requestHeader() is the log's incrementally-folded baseline. - Requests are deep-frozen before dispatch (deepFreeze exempts the AbortSignal — freezing one breaks AbortController.abort() outright); frozen + sessionId is the loop-built marker the dev invariant keys on. Ported from #162 and re-anchored on the log: the append-extension / frozen-end-to-end / compaction-resend / prompt-change property tests, plus new specs for the boundary semantics, resume anchoring, and the end-to-end theorem (every recorded request rebuilds byte-equal from the log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against the real DeepSeek API. Snapshot goldens intentionally stale until the single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
})
// --- Model call (streaming-first; raw chunks are the replay record) ---
const assembler = new BlockAssembler()
2026-06-17 19:25:29 +08:00
const chunkSeqs: number[] = []
const stream = ctx.llm.stream(request)
try {
for await (const chunk of stream) {
/* v8 ignore next -- signal.reason always set: cancel()/disposal provide a default */
if (signal.aborted) throw new Error(String(signal.reason ?? 'aborted'))
const chunkEvent = session.append('assistant/chunk', { turn, step, chunk })
chunkSeqs.push(chunkEvent.seq)
assembler.push(chunk)
}
} catch (error: unknown) {
if (isLlmAdapterFailure(stream, error)) throw new TerminalModelRequestFailure(error)
throw error
}
// Normalize failure finish chunks into the same path as thrown stream errors.
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
const stepError = finishError(assembler.finish)
if (stepError) throw new TerminalModelRequestFailure(stepError)
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
const recordAssistantMessage = (
assembledContent: ContentBlock[],
message: Message,
preserveReplayState = true,
): void => {
session.append(
'assistant/message',
{
turn,
step,
content: message.content,
provenance: assistantProvenance(
header.config,
assembler.replayState,
preserveReplayState && isDeepStrictEqual(message.content, assembledContent),
),
...assembler.usage === undefined ? {} : { usage: assembler.usage },
},
{ surfaceOp: 'append', sourceEventSeqs: chunkSeqs },
)
}
// A rejected result still records the successful provider call without retaining rejected output.
const processStepResult = async (assembledContent: ContentBlock[], message: Message): Promise<Message> => {
try {
return await events.waterfall(
'agent/step-result', turn, step, message, () => Promise.resolve(message),
)
} catch (error: unknown) {
recordAssistantMessage(assembledContent, { ...message, content: [] }, false)
throw error
}
}
if (assembler.finish.kind === 'max-tokens') {
2026-07-14 21:57:52 +08:00
const assembled = assembler.message()
const assembledContent = structuredClone(assembled.content)
2026-07-14 21:57:52 +08:00
let message: Message = withoutToolCalls(assembled)
message = withoutToolCalls(await processStepResult(assembledContent, message))
// Preserve usage even when max-token truncation produced no content.
recordAssistantMessage(assembledContent, message)
return { hadToolCalls: false, finish: assembler.finish }
}
// Record the post-waterfall message that tool dispatch uses.
2026-07-14 21:57:52 +08:00
const assembled = assembler.message()
const assembledContent = structuredClone(assembled.content)
2026-07-14 21:57:52 +08:00
let message: Message = assembled
message = await processStepResult(assembledContent, message)
Fix architecture-review findings in the loop and service packages High (loop pipeline): agent/step-result now runs before the assistant/message append so the session log records what tool dispatch actually uses; abort is honored between tool calls, not just mid-stream; steering drains at step start, pending steering overrides a negative turn-continuation decision (/goal pattern), and leftover steering is re-enqueued as queued messages so it is never stranded; exceptions from turn-continuation listeners and session/flush are contained to the turn (error event + agent/error) instead of killing the driver loop. Medium: disposal emits agent/status('disposed') and mid-turn disposal records reason 'disposed'; duplicate LLM adapter registration throws (all-or-nothing); SessionEvent is a real discriminated union (casts removed); model-less agents fail with a clear actionable error unless agent/request supplies a model. Low: agent/queued and agent/steering carry the resolved MessageSource; streamBlocks() yields strictly in stream order and flushes delta-only blocks (matches generate()); BlockAssembler freezes blocks on block-end and ignores stragglers from malformed streams; turn numbering is a counter seeded from the log (fork-safe); LoopAgent's stop disposer is infallible (a throwing status listener cannot skip registry cleanup); AgentLoop.create uses a generator effect so stop and unregister are independent disposables; SessionStore wires onAppend inside its effect. 21 regression tests added (review-fixes.spec.ts), organized by finding. Docs updated: loop pseudocode (status emissions, ordering, error containment, steering guarantees) and waterfall composition caveat in docs/architecture.md; AGENTS.md notes that excessive tests are welcome.
2026-06-11 12:18:52 +08:00
// Every successful call records its completion anchor, including explicit
// empty chunk provenance for a contentless, usage-less provider response.
recordAssistantMessage(assembledContent, message)
// Dispatch may overlap; policy, durable results, and result context stay model-ordered.
const toolCalls = message.content.filter(block => block.type === 'tool-call')
2026-07-15 12:22:39 +08:00
if (toolCalls.length === 0) return { hadToolCalls: false, finish: assembler.finish }
2026-07-15 12:49:50 +08:00
return handle.withToolBatch(async (acceptContext) => {
await executeToolCalls(
ctx, turn, step, toolCalls, signal, handle.maxParallelToolCalls, acceptContext,
)
2026-07-15 12:22:39 +08:00
return { hadToolCalls: true, finish: assembler.finish }
})
}
2026-07-14 21:57:52 +08:00
/** Build durable assistant provenance, dropping replay state after any content rewrite. */
function assistantProvenance(config: LlmCallConfig, replayState: unknown, contentUnchanged: boolean): NonNullable<Message['provenance']> {
return {
provider: config.provider,
model: config.model,
...contentUnchanged && replayState !== undefined ? { replayState } : {},
}
}
2026-06-19 00:37:12 +08:00
function withoutToolCalls(message: Message): Message {
return { ...message, content: message.content.filter(block => block.type !== 'tool-call') }
}
/**
* The last turn number in a (possibly seeded) session log, or 0.
* @param session - the session whose log is scanned for the latest `turn/start`.
* @returns the latest `turn/start`'s turn number, or 0 when the log has none (the next turn is this plus one).
*/
export function lastTurnNumber(session: Session): number {
const lastStart = session.events.findLast(event => event.type === 'turn/start')
return lastStart?.data.turn ?? 0
}
/**
* Whether the session log has an unmatched `turn/start`. Agent status is not
* sufficient during pre-start and post-end windows.
* @param session - the session whose log is inspected.
* @returns true when the log's last turn boundary is a `turn/start` with no matching `turn/end` yet.
*/
export function isTurnOpen(session: Session): boolean {
const last = session.events.findLast(e => e.type === 'turn/start' || e.type === 'turn/end')
return last?.type === 'turn/start'
}