Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
import { describe, expect, it } from 'vitest'
|
|
|
|
|
import { Context } from 'cordis'
|
2026-06-17 21:25:47 +08:00
|
|
|
import LlmService, { CallId, StreamChunk } from '@deepseek-ai/dsh-llm'
|
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
|
|
|
import SessionStore, { SessionId, TurnEndReason } from '@deepseek-ai/dsh-session'
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
import SystemPrompt from '@deepseek-ai/dsh-system-prompt'
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
import ToolRegistry, { defineTool } from '@deepseek-ai/dsh-tools'
|
2026-07-14 02:32:35 +08:00
|
|
|
import AgentRegistry, { type Agent } from '@deepseek-ai/dsh-agent'
|
2026-07-14 01:59:21 +08:00
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
import AgentLoop from '@deepseek-ai/dsh-agent-loop'
|
2026-06-15 23:53:47 +08:00
|
|
|
import { MockAdapter, maxTokensResponse, textResponse, toolCallResponse } from './mock-adapter.ts'
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
function driverDone(agent: Agent): Promise<void> {
|
|
|
|
|
return (agent as Agent & { done: Promise<void> }).done
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-05 23:23:46 +08:00
|
|
|
async function harness(adapter: MockAdapter, persona = '') {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const ctx = new Context()
|
|
|
|
|
await ctx.plugin(LlmService)
|
|
|
|
|
await ctx.plugin(SessionStore)
|
2026-07-05 23:23:46 +08:00
|
|
|
await ctx.plugin(SystemPrompt, { persona })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
await ctx.plugin(ToolRegistry)
|
|
|
|
|
await ctx.plugin(AgentRegistry)
|
|
|
|
|
await ctx.plugin(AgentLoop, { agents: [] })
|
|
|
|
|
ctx.llm.registerAdapter(['mock'], adapter)
|
|
|
|
|
return ctx
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
/**
|
|
|
|
|
* Wait for the agent's NEXT transition to idle. Always event-based: callers
|
|
|
|
|
* invoke this right after send(), when the loop hasn't woken yet (status is
|
|
|
|
|
* still 'idle' synchronously), so polling the current status would lie.
|
|
|
|
|
*/
|
2026-07-14 02:32:35 +08:00
|
|
|
function waitForIdle(ctx: Context, agent: Agent): Promise<void> {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
return new Promise((resolve) => {
|
|
|
|
|
const dispose = ctx.on('agent/status', (subject, status) => {
|
|
|
|
|
if (subject === agent && status === 'idle') {
|
|
|
|
|
dispose()
|
|
|
|
|
resolve()
|
|
|
|
|
}
|
|
|
|
|
})
|
|
|
|
|
})
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
function send(agent: Agent, text: string) {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
agent.send([{ type: 'text', text }])
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
describe('agent loop', () => {
|
|
|
|
|
it('runs a simple turn: queued message → model → idle, with ordered events', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('hello there')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
2026-07-02 03:26:45 +08:00
|
|
|
// All boundaries — turn and step — are durable session events on the
|
|
|
|
|
// session/event feed (no agent/* mirror). Record them in fire order to
|
|
|
|
|
// assert the full boundary nesting.
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const order: string[] = []
|
2026-06-30 10:32:55 +08:00
|
|
|
ctx.on('session/event', (_session, event) => {
|
2026-07-02 03:26:45 +08:00
|
|
|
if (event.type === 'turn/start' || event.type === 'step/start' || event.type === 'step/end' || event.type === 'turn/end') {
|
|
|
|
|
order.push(event.type)
|
|
|
|
|
}
|
2026-06-30 10:32:55 +08:00
|
|
|
})
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-02 03:26:45 +08:00
|
|
|
expect(order).toEqual(['turn/start', 'step/start', 'step/end', 'turn/end'])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
const types = agent.session.events.map(e => e.type)
|
2026-06-15 20:56:17 +08:00
|
|
|
// turn/start opens the turn, THEN the queued user message is recorded inside
|
2026-06-21 10:00:06 +08:00
|
|
|
// it (every event is turn-enclosed), then the assembled message (carrying the
|
|
|
|
|
// step's usage).
|
2026-06-15 20:56:17 +08:00
|
|
|
expect(types[0]).toBe('turn/start')
|
|
|
|
|
expect(types[1]).toBe('user/message')
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(types).toContain('assistant/message')
|
2026-06-21 10:00:06 +08:00
|
|
|
const assistantMessage = agent.session.events.find(e => e.type === 'assistant/message')
|
|
|
|
|
expect(assistantMessage?.type === 'assistant/message' && assistantMessage.data.usage).toEqual({ inputTokens: 10, outputTokens: 'hello there'.length })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(types.at(-1)).toBe('turn/end')
|
|
|
|
|
|
|
|
|
|
// derived history: user + assistant
|
|
|
|
|
const messages = agent.session.deriveMessages()
|
|
|
|
|
expect(messages.map(m => m.role)).toEqual(['user', 'assistant'])
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
expect(messages[1]!.content).toEqual([{ type: 'text', text: 'hello there' }])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('round-trips tool calls: model requests tool → executes → result in next request', async () => {
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'echo', { text: 'ping' }, 'calling echo'),
|
|
|
|
|
textResponse('done'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
ctx.tools.register(defineTool({
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
name: 'echo',
|
|
|
|
|
description: 'echo back',
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
parameters: { text: { type: 'string' } },
|
|
|
|
|
async execute(args) {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
return [{ type: 'text', text: `echo: ${args.text}` }]
|
|
|
|
|
},
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'use the tool')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
// two model calls happened (tool-call step, then final step)
|
|
|
|
|
expect(adapter.requests).toHaveLength(2)
|
|
|
|
|
|
|
|
|
|
// the second request's derived history contains the tool result
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
const secondMessages = adapter.requests[1]!.messages
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const toolResultMessage = secondMessages.find(m =>
|
|
|
|
|
m.content.some(b => b.type === 'tool-result'))
|
|
|
|
|
expect(toolResultMessage).toBeDefined()
|
|
|
|
|
const block = toolResultMessage!.content.find(b => b.type === 'tool-result')!
|
|
|
|
|
expect(block).toMatchObject({ toolCallId: 'c1', isError: false })
|
Add ESLint: typescript-eslint strict-type-checked + stylistic formatting
Flat config with two layers. Correctness (type-checked): the headline
rules for this codebase are no-floating-promises / no-misused-promises
(a lost promise in the agent loop is our primary bug class),
switch-exhaustiveness-check (we switch over merge-extensible unions
everywhere), no-unnecessary-condition, require-await, and
no-explicit-any. Style (@stylistic): 2-space, no semicolons, single
quotes, trailing commas, max-len 140 — the existing house style, now
enforced instead of drifting between agents. vendor/ is excluded
(vendored source keeps upstream style); tests relax the rules that
fight test ergonomics (non-null assertions after expects, async mock
signatures, non-Error throws).
Code adjusted to pass: registry disposers wrap ctx.effect's
promise-returning disposer behind a sync () => void (our public API),
BlockAssembler gains an invariant-checking mustGet instead of non-null
assertions, lastTurnNumber uses findLast, waterfall tails return
Promise.resolve instead of async-without-await arrows, and the two
deliberate suppressions (non-exhaustive derivation switch, unbound
execute pass-through) carry justification comments.
yarn lint / yarn lint:fix added.
2026-06-11 14:17:58 +08:00
|
|
|
expect((block).content).toEqual([{ type: 'text', text: 'echo: ping' }])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
// session log records call + result
|
|
|
|
|
const types = agent.session.events.map(e => e.type)
|
|
|
|
|
expect(types).toContain('tool/call')
|
|
|
|
|
expect(types).toContain('tool/result')
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-03 17:12:00 +08:00
|
|
|
it('threads a tool-attached meta (execute object return) onto the tool/result event', async () => {
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'writer', { path: 'a.txt' }, 'writing'),
|
|
|
|
|
textResponse('done'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
// A tool that returns the { content, meta } object form: the loop must
|
|
|
|
|
// persist `meta` on the tool/result event so a UI reproduces the card on replay.
|
|
|
|
|
ctx.tools.register(defineTool({
|
|
|
|
|
name: 'writer',
|
|
|
|
|
description: 'writes a file',
|
|
|
|
|
parameters: { path: { type: 'string' } },
|
|
|
|
|
async execute() {
|
|
|
|
|
return { content: [{ type: 'text', text: 'ok' }], meta: { diffs: [{ path: 'a.txt', oldText: null, newText: 'x' }] } }
|
|
|
|
|
},
|
|
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-03 17:12:00 +08:00
|
|
|
|
|
|
|
|
send(agent, 'use the tool')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
const toolResult = agent.session.events.find(e => e.type === 'tool/result')
|
|
|
|
|
expect(toolResult?.type === 'tool/result' && toolResult.data.meta)
|
|
|
|
|
.toEqual({ diffs: [{ path: 'a.txt', oldText: null, newText: 'x' }] })
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-05 11:37:32 +08:00
|
|
|
it('renders harness identity, then the persona, then tool guidance — with {{variables}} resolved', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
2026-07-05 23:23:46 +08:00
|
|
|
// The persona is a TEMPLATE: {{model}} is the loop-registered variable
|
|
|
|
|
// projecting this agent's configured model, so the model knows its own name.
|
|
|
|
|
const ctx = await harness(adapter, 'You are a test agent on {{model}}.')
|
2026-07-05 01:54:46 +08:00
|
|
|
ctx.systemPrompt.section({ name: 'tool:noop', order: 100, text: 'Use the noop tool wisely.' })
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
ctx.tools.register(defineTool({
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
name: 'noop',
|
|
|
|
|
description: 'does nothing',
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
parameters: {},
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
async execute() {
|
|
|
|
|
return []
|
|
|
|
|
},
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
const request = adapter.requests[0]
|
2026-07-05 11:37:32 +08:00
|
|
|
expect(request!.system).toBe('You are an AI agent powered by the DeepSeek Harness SDK.\n\nYou are a test agent on mock.\n\nUse the noop tool wisely.')
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
expect(request!.tools?.map(t => t.name)).toEqual(['noop'])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-05 01:54:46 +08:00
|
|
|
it('resolves {{cwd}} from the agent session workspace (factory create with meta.cwd)', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
2026-07-05 23:23:46 +08:00
|
|
|
const ctx = await harness(adapter, 'Working in {{cwd}}.')
|
2026-07-11 22:55:26 +08:00
|
|
|
const handle = await ctx.agents.create({
|
2026-07-05 01:54:46 +08:00
|
|
|
sessionId: SessionId('s-cwd'),
|
|
|
|
|
meta: { cwd: '/work/space' },
|
2026-07-14 21:57:52 +08:00
|
|
|
agentOptions: { provider: 'mock', model: 'mock' },
|
2026-07-05 01:54:46 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
const agent = handle.agent
|
2026-07-05 01:54:46 +08:00
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-05 11:37:32 +08:00
|
|
|
expect(adapter.requests[0]!.system).toBe('You are an AI agent powered by the DeepSeek Harness SDK.\n\nWorking in /work/space.')
|
2026-07-05 01:54:46 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-05 02:42:48 +08:00
|
|
|
it('contains a strict-variable render failure: the turn errors, the loop keeps serving turns', async () => {
|
2026-07-12 03:36:43 +08:00
|
|
|
// A missing cwd variable must fail one turn without preventing a later valid turn.
|
2026-07-05 02:42:48 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok after rescue')])
|
2026-07-05 23:23:46 +08:00
|
|
|
const ctx = await harness(adapter, 'In {{cwd}}.')
|
2026-07-05 01:54:46 +08:00
|
|
|
const errors: Error[] = []
|
|
|
|
|
ctx.on('agent/error', (_agent, _turn, _step, error) => void errors.push(error))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-05 01:54:46 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests).toHaveLength(0) // the request was never sent
|
|
|
|
|
expect(errors.some(e => e.message.includes('no value for this assembly'))).toBe(true)
|
|
|
|
|
const turnEnd = agent.session.events.find(e => e.type === 'turn/end')
|
|
|
|
|
expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason.kind).toBe('error')
|
2026-07-05 02:42:48 +08:00
|
|
|
|
|
|
|
|
// The loop survived: a waterfall listener rescues {{cwd}} and the SAME
|
|
|
|
|
// agent completes a real model turn.
|
|
|
|
|
ctx.on('system-prompt/assemble', async (assembly, _context, next) => {
|
|
|
|
|
assembly.variables['cwd'] = '/rescued'
|
|
|
|
|
return next()
|
|
|
|
|
})
|
|
|
|
|
send(agent, 'again')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
2026-07-05 11:37:32 +08:00
|
|
|
expect(adapter.requests[0]!.system).toBe('You are an AI agent powered by the DeepSeek Harness SDK.\n\nIn /rescued.')
|
2026-07-05 02:42:48 +08:00
|
|
|
const turnEnds = agent.session.events.filter(e => e.type === 'turn/end')
|
|
|
|
|
expect(turnEnds).toHaveLength(2)
|
|
|
|
|
expect(turnEnds[1]?.type === 'turn/end' && turnEnds[1].data.reason.kind).toBe('completed')
|
2026-07-05 01:54:46 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-06 00:34:47 +08:00
|
|
|
it('supports the model-via-agent/request path with a {{model}} persona: the supplier states it via the assemble waterfall', async () => {
|
|
|
|
|
// AgentOptions.model unset: the model arrives in the agent/request
|
|
|
|
|
// waterfall (the loop's documented fallback — see runStep's no-model
|
|
|
|
|
// error). {{model}} renders BEFORE that waterfall, so the SAME plugin
|
|
|
|
|
// states the fact early on system-prompt/assemble — the owner of a
|
|
|
|
|
// late-bound fact owns stating it wherever it is claimed.
|
|
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter, 'You run on {{model}}.')
|
|
|
|
|
ctx.on('system-prompt/assemble', async (assembly, _context, next) => {
|
2026-07-14 21:57:52 +08:00
|
|
|
assembly.variables['provider'] = 'mock'
|
2026-07-06 00:34:47 +08:00
|
|
|
assembly.variables['model'] = 'mock'
|
|
|
|
|
return next()
|
|
|
|
|
})
|
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall
The loop is now transmission-stateless; a request is a pure function of
(session log, this step's rendered assembly, current AgentOptions):
- The reconstruction boundary is step/start: the messages snapshot is
taken in the same synchronous frame immediately before the step/start
append, so the request's messages are exactly the derivation over
events[0..stepStartSeq) — an inject() from an agent/request listener
(or any concurrent task) lands after the boundary and joins the NEXT
request. This changes behavior for a synchronous step/start
session/event listener that appends content (master derived after the
append, so such a listener could reach the current request):
agent/pre-step is the sanctioned seam for current-request content.
- agent/request is re-typed to config-only: (agent, turn, step,
config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes
from AgentOptions on a loop instance's first request (explicit options
beat the logged baseline — fork overrides and resume reconfiguration
stay correct) and from the log's folded header afterwards; listeners
return a replacement to switch. Content shaping through the request is
no longer expressible — model-visible content flows through the log
channels.
- recordRequestHeader appends whatever header event the request owes the
log before dispatch: an 'initial'/'resume' snapshot anchoring each
loop instance, a round-trip-verified delta on change, a 'fallback'
snapshot when the encoding cannot express it. Session.requestHeader()
is the log's incrementally-folded baseline.
- Requests are deep-frozen before dispatch (deepFreeze exempts the
AbortSignal — freezing one breaks AbortController.abort() outright);
frozen + sessionId is the loop-built marker the dev invariant keys on.
Ported from #162 and re-anchored on the log: the append-extension /
frozen-end-to-end / compaction-resend / prompt-change property tests,
plus new specs for the boundary semantics, resume anchoring, and the
end-to-end theorem (every recorded request rebuilds byte-equal from the
log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against
the real DeepSeek API. Snapshot goldens intentionally stale until the
single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
|
|
|
ctx.on('agent/request', async (_agent, _turn, _step, config, _next) => {
|
2026-07-14 21:57:52 +08:00
|
|
|
return { ...config, provider: 'mock', model: 'mock' }
|
2026-07-06 00:34:47 +08:00
|
|
|
})
|
2026-07-14 01:59:21 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a-late-model'), {})
|
2026-07-06 00:34:47 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
expect(adapter.requests[0]!.model).toBe('mock')
|
|
|
|
|
expect(adapter.requests[0]!.system).toBe('You are an AI agent powered by the DeepSeek Harness SDK.\n\nYou run on mock.')
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-11 22:55:26 +08:00
|
|
|
it.each([
|
|
|
|
|
['BigInt', { n: 1n }],
|
|
|
|
|
['Map', new Map([['key', 'value']])],
|
|
|
|
|
['class instance', new (class ResultMeta { x = 1 })()],
|
|
|
|
|
])('normalizes non-JSON tool meta (%s) before the durable result commit', async (_kind, meta) => {
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('bad-meta-call', 'bad-meta', {}, 'calling'),
|
|
|
|
|
textResponse('recovered'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
ctx.tools.register(defineTool({
|
|
|
|
|
name: 'bad-meta',
|
|
|
|
|
description: 'returns invalid durable metadata',
|
|
|
|
|
parameters: {},
|
|
|
|
|
execute: () => Promise.resolve({ content: [{ type: 'text' as const, text: 'apparent success' }], meta }),
|
|
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('bad-meta-agent'), { provider: 'mock', model: 'mock' })
|
2026-07-11 22:55:26 +08:00
|
|
|
|
|
|
|
|
send(agent, 'use the tool')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
const result = agent.session.events.find(event => event.type === 'tool/result')
|
|
|
|
|
expect(result?.type).toBe('tool/result')
|
|
|
|
|
if (result?.type === 'tool/result') {
|
|
|
|
|
expect(result.data.callId).toBe('bad-meta-call')
|
|
|
|
|
expect(result.data.isError).toBe(true)
|
|
|
|
|
expect(result.data.meta).toBeUndefined()
|
|
|
|
|
expect(result.data.content).toEqual([{
|
|
|
|
|
type: 'text',
|
2026-07-12 22:49:46 +08:00
|
|
|
text: 'Error: tool result must be losslessly JSON-serializable',
|
2026-07-11 22:55:26 +08:00
|
|
|
}])
|
|
|
|
|
}
|
|
|
|
|
// The normalized failure was durably logged and fed back to the model; the
|
|
|
|
|
// turn continued normally instead of failing after an apparent success.
|
|
|
|
|
expect(adapter.requests).toHaveLength(2)
|
|
|
|
|
expect(JSON.stringify(adapter.requests[1]!.messages)).toContain('losslessly JSON-serializable')
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-05 23:23:46 +08:00
|
|
|
it('omits the system field when a system-prompt/assemble veto empties the assembly', async () => {
|
|
|
|
|
// The documented escape valve: a deployment that must drop the harness
|
|
|
|
|
// openers short-circuits the assemble waterfall; the request then carries
|
|
|
|
|
// NO system field at all (not an empty string).
|
|
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
ctx.on('system-prompt/assemble', async () => ({ sections: [], tools: [], variables: {} }))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a-no-system'), { provider: 'mock', model: 'mock' })
|
2026-07-05 23:23:46 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
expect('system' in adapter.requests[0]!).toBe(false)
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-02 23:42:16 +08:00
|
|
|
it('records raw chunks for replay as assistant/chunk session events', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('abc')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
const chunkEvents = agent.session.events.filter(e => e.type === 'assistant/chunk')
|
|
|
|
|
// textResponse('abc') = block-start + 3 deltas + block-end + usage + finish = 7
|
|
|
|
|
expect(chunkEvents).toHaveLength(7)
|
|
|
|
|
// replay: chunk events alone re-assemble to the recorded assistant message
|
|
|
|
|
const deltaText = chunkEvents
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
.flatMap(e => e.type === 'assistant/chunk' ? [e.data.chunk] : [])
|
|
|
|
|
.filter((c: StreamChunk): c is Extract<StreamChunk, { type: 'text-delta' }> => c.type === 'text-delta')
|
|
|
|
|
.map(c => c.text)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
.join('')
|
|
|
|
|
expect(deltaText).toBe('abc')
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('injects steering between steps and continues the turn', async () => {
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'slow', {}),
|
|
|
|
|
textResponse('addressed the steering'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
ctx.tools.register(defineTool({
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
name: 'slow',
|
|
|
|
|
description: '',
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
parameters: {},
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
async execute() {
|
|
|
|
|
// steer while the turn is running (during tool execution)
|
|
|
|
|
agent.steer([{ type: 'text', text: 'change of plans' }])
|
|
|
|
|
return [{ type: 'text', text: 'tool done' }]
|
|
|
|
|
},
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
}))
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'start')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
const types = agent.session.events.map(e => e.type)
|
|
|
|
|
expect(types).toContain('steering/message')
|
|
|
|
|
// steering recorded before the second step's request derived its history
|
|
|
|
|
const steeringSeq = agent.session.events.find(e => e.type === 'steering/message')!.seq
|
|
|
|
|
const secondStepStart = agent.session.events.filter(e => e.type === 'step/start')[1]
|
|
|
|
|
expect(secondStepStart).toBeDefined()
|
|
|
|
|
expect(steeringSeq).toBeLessThan(secondStepStart!.seq)
|
|
|
|
|
|
|
|
|
|
// the second model request saw the steering content
|
|
|
|
|
const secondRequest = adapter.requests[1]
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
const flat = JSON.stringify(secondRequest!.messages)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(flat).toContain('change of plans')
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-17 17:27:53 +08:00
|
|
|
it('same-tick idle steering inherits one-send-one-turn FIFO behavior', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('first'), textResponse('second')])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
2026-07-17 17:27:53 +08:00
|
|
|
const idle = waitForIdle(ctx, agent)
|
|
|
|
|
agent.steer([{ type: 'text', text: 'first idle steer' }])
|
|
|
|
|
agent.steer([{ type: 'text', text: 'second idle steer' }])
|
|
|
|
|
await idle
|
|
|
|
|
|
|
|
|
|
expect(agent.session.events.filter(event => event.type === 'turn/start')).toHaveLength(2)
|
|
|
|
|
expect(agent.session.events
|
|
|
|
|
.filter(event => event.type === 'user/message')
|
|
|
|
|
.map(event => event.data.content)).toEqual([
|
|
|
|
|
[{ type: 'text', text: 'first idle steer' }],
|
|
|
|
|
[{ type: 'text', text: 'second idle steer' }],
|
|
|
|
|
])
|
|
|
|
|
expect(adapter.requests).toHaveLength(2)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
2026-06-15 20:56:17 +08:00
|
|
|
it('inject() while idle wraps context in a one-shot turn, visible to the next request', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
agent.inject([{ type: 'text', text: 'file changed: a.ts' }], { source: { kind: 'plugin', plugin: 'watcher' } })
|
2026-06-15 20:56:17 +08:00
|
|
|
// The idle inject records a self-contained turn (turn/start → context/message
|
|
|
|
|
// → turn/end) so the event stays turn-enclosed, but does NOT run the model.
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
await new Promise(r => setTimeout(r, 20))
|
|
|
|
|
expect(agent.status).toBe('idle')
|
|
|
|
|
expect(adapter.requests).toHaveLength(0)
|
2026-06-15 20:56:17 +08:00
|
|
|
const injectedTurn = agent.session.events.filter(e => e.type === 'turn/start')
|
|
|
|
|
expect(injectedTurn).toHaveLength(1)
|
|
|
|
|
const it0 = injectedTurn[0]!
|
|
|
|
|
expect(it0.type === 'turn/start' && it0.data.trigger.kind).toBe('injection')
|
|
|
|
|
expect(agent.session.events.at(-1)!.type).toBe('turn/end') // turn-enclosed
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
const flat = JSON.stringify(adapter.requests[0]!.messages)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(flat).toContain('file changed: a.ts')
|
2026-07-20 14:37:04 +08:00
|
|
|
expect(flat).not.toContain('<context source=')
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-20 14:37:04 +08:00
|
|
|
it('inject() persists structured context content verbatim with durable hidden meta', async () => {
|
2026-07-10 14:32:44 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('raw-context'), { provider: 'mock', model: 'mock' })
|
2026-07-10 14:32:44 +08:00
|
|
|
const text = '<system-reminder>Additional instructions from: pkg/AGENTS.md</system-reminder>'
|
|
|
|
|
const meta = {
|
|
|
|
|
kind: 'workspace-instructions',
|
|
|
|
|
version: 1,
|
|
|
|
|
changes: [{ action: 'set', scope: 'pkg', path: 'pkg/AGENTS.md', digest: 'abc123' }],
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
agent.inject([{ type: 'text', text }], {
|
|
|
|
|
source: { kind: 'plugin', plugin: 'workspace-context' },
|
|
|
|
|
meta,
|
|
|
|
|
})
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
const contextEvent = agent.session.events.find(event => event.type === 'context/message')
|
2026-07-20 14:37:04 +08:00
|
|
|
expect(contextEvent?.type === 'context/message' && contextEvent.data).toMatchObject({ meta })
|
2026-07-10 14:32:44 +08:00
|
|
|
const requestText = JSON.stringify(adapter.requests[0]!.messages)
|
|
|
|
|
expect(requestText).toContain('Additional instructions from: pkg/AGENTS.md')
|
|
|
|
|
expect(requestText).not.toContain('<context source=')
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-15 11:12:05 +08:00
|
|
|
it('defers inject() during tool execution until after the tool result', async () => {
|
2026-06-15 20:56:17 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'noticer', {}, 'calling'),
|
|
|
|
|
textResponse('done'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-15 11:12:05 +08:00
|
|
|
let visibleDuringTool = false
|
2026-07-18 13:39:12 +08:00
|
|
|
const meta = { kind: 'deferred-test', version: 1 }
|
2026-06-15 20:56:17 +08:00
|
|
|
ctx.tools.register(defineTool({
|
|
|
|
|
name: 'noticer',
|
|
|
|
|
description: 'injects a notice',
|
|
|
|
|
parameters: {},
|
|
|
|
|
async execute() {
|
2026-07-15 11:12:05 +08:00
|
|
|
await Promise.resolve()
|
|
|
|
|
const first = { type: 'text' as const, text: 'mid-turn notice' }
|
2026-07-18 13:39:12 +08:00
|
|
|
agent.inject([first], {
|
|
|
|
|
source: { kind: 'plugin', plugin: 'x' },
|
|
|
|
|
meta,
|
|
|
|
|
})
|
2026-07-15 11:12:05 +08:00
|
|
|
first.text = 'mutated after inject'
|
|
|
|
|
agent.inject([{ type: 'text', text: 'second notice' }], { source: { kind: 'plugin', plugin: 'x' } })
|
|
|
|
|
visibleDuringTool = agent.session.events.some(e => e.type === 'context/message')
|
2026-06-15 20:56:17 +08:00
|
|
|
return [{ type: 'text', text: 'ok' }]
|
|
|
|
|
},
|
|
|
|
|
}))
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-15 11:12:05 +08:00
|
|
|
expect(visibleDuringTool).toBe(false)
|
|
|
|
|
|
|
|
|
|
// The injection stays in the open turn, but its user-role context cannot
|
|
|
|
|
// split the assistant tool call from the provider's tool-result message.
|
2026-06-15 20:56:17 +08:00
|
|
|
const turnStarts = agent.session.events.filter(e => e.type === 'turn/start')
|
|
|
|
|
expect(turnStarts).toHaveLength(1)
|
|
|
|
|
const ts0 = turnStarts[0]!
|
|
|
|
|
expect(ts0.type === 'turn/start' && ts0.data.trigger.kind).toBe('message')
|
2026-07-15 11:12:05 +08:00
|
|
|
const result = agent.session.events.find(e => e.type === 'tool/result')!
|
|
|
|
|
const contexts = agent.session.events.filter(e => e.type === 'context/message')
|
|
|
|
|
expect(contexts).toHaveLength(2)
|
|
|
|
|
expect(result.seq).toBeLessThan(contexts[0]!.seq)
|
2026-07-18 13:39:12 +08:00
|
|
|
expect(contexts[0]?.type === 'context/message' && contexts[0].data).toMatchObject({
|
|
|
|
|
meta,
|
|
|
|
|
})
|
2026-07-15 11:12:05 +08:00
|
|
|
expect(contexts.flatMap(event => event.type === 'context/message' ? event.data.content : []))
|
|
|
|
|
.toEqual([
|
|
|
|
|
{ type: 'text', text: 'mid-turn notice' },
|
|
|
|
|
{ type: 'text', text: 'second notice' },
|
|
|
|
|
])
|
|
|
|
|
|
|
|
|
|
const secondRequest = adapter.requests[1]!.messages
|
|
|
|
|
const resultIndex = secondRequest.findIndex(message =>
|
|
|
|
|
message.content.some(block => block.type === 'tool-result'))
|
|
|
|
|
const contextIndexes = secondRequest.flatMap((message, index) =>
|
|
|
|
|
message.content.some(block => block.type === 'text'
|
|
|
|
|
&& (block.text.includes('mid-turn notice') || block.text.includes('second notice')))
|
|
|
|
|
? [index]
|
|
|
|
|
: [])
|
|
|
|
|
expect(resultIndex).toBeGreaterThanOrEqual(0)
|
|
|
|
|
expect(contextIndexes).toHaveLength(2)
|
|
|
|
|
expect(contextIndexes.every(index => index > resultIndex)).toBe(true)
|
2026-06-15 20:56:17 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-18 13:39:12 +08:00
|
|
|
it('rejects non-JSON context before it enters the active tool-batch FIFO', async () => {
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'invalid-injector', {}, 'calling'),
|
|
|
|
|
textResponse('done'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 14:50:55 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('invalid-context'), { provider: 'mock', model: 'mock' })
|
2026-07-18 13:39:12 +08:00
|
|
|
ctx.tools.register(defineTool({
|
|
|
|
|
name: 'invalid-injector',
|
|
|
|
|
description: 'attempts an invalid context injection',
|
|
|
|
|
parameters: {},
|
|
|
|
|
async execute() {
|
|
|
|
|
expect(() => {
|
|
|
|
|
agent.inject([{ type: 'text', text: 'invalid' }], {
|
|
|
|
|
source: { kind: 'plugin', plugin: 'test' },
|
|
|
|
|
meta: { bigint: 1n } as never,
|
|
|
|
|
})
|
|
|
|
|
}).toThrow('agent context must be losslessly JSON-serializable')
|
|
|
|
|
return [{ type: 'text', text: 'rejected invalid context' }]
|
|
|
|
|
},
|
|
|
|
|
}))
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(agent.session.events.some(event => event.type === 'context/message')).toBe(false)
|
2026-06-15 20:56:17 +08:00
|
|
|
})
|
|
|
|
|
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
it('agent/turn-continuation can force-continue (/loop pattern) and force-stop', async () => {
|
|
|
|
|
// force-continue: model never calls tools, but a plugin forces 3 steps
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
textResponse('step 1'),
|
|
|
|
|
textResponse('step 2'),
|
|
|
|
|
textResponse('step 3'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
let steps = 0
|
2026-06-30 10:32:55 +08:00
|
|
|
ctx.on('session/event', (_session, event) => { if (event.type === 'step/end') steps++ })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
ctx.on('agent/turn-continuation', async (_agent, _turn, _defaultDecision, next) => {
|
2026-06-30 17:11:18 +08:00
|
|
|
if (steps < 3) return { action: 'continue' as const }
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
return next()
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(steps).toBe(3)
|
|
|
|
|
expect(adapter.requests).toHaveLength(3)
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('agent/turn-continuation can veto continuation despite tool calls (budget-guard pattern)', async () => {
|
|
|
|
|
const adapter = new MockAdapter([toolCallResponse('c1', 'echo', { text: 'x' })])
|
|
|
|
|
const ctx = await harness(adapter)
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
ctx.tools.register(defineTool({
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
name: 'echo',
|
|
|
|
|
description: '',
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
parameters: { text: { type: 'string' } },
|
|
|
|
|
async execute(args) {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
return [{ type: 'text', text: String(args.text) }]
|
|
|
|
|
},
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
2026-06-30 17:11:18 +08:00
|
|
|
ctx.on('agent/turn-continuation', async () => ({ action: 'stop' }) as const)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
// only one model call despite the tool call requesting a follow-up
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
// tool still executed before the decision
|
|
|
|
|
expect(agent.session.events.some(e => e.type === 'tool/result')).toBe(true)
|
|
|
|
|
})
|
|
|
|
|
|
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall
The loop is now transmission-stateless; a request is a pure function of
(session log, this step's rendered assembly, current AgentOptions):
- The reconstruction boundary is step/start: the messages snapshot is
taken in the same synchronous frame immediately before the step/start
append, so the request's messages are exactly the derivation over
events[0..stepStartSeq) — an inject() from an agent/request listener
(or any concurrent task) lands after the boundary and joins the NEXT
request. This changes behavior for a synchronous step/start
session/event listener that appends content (master derived after the
append, so such a listener could reach the current request):
agent/pre-step is the sanctioned seam for current-request content.
- agent/request is re-typed to config-only: (agent, turn, step,
config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes
from AgentOptions on a loop instance's first request (explicit options
beat the logged baseline — fork overrides and resume reconfiguration
stay correct) and from the log's folded header afterwards; listeners
return a replacement to switch. Content shaping through the request is
no longer expressible — model-visible content flows through the log
channels.
- recordRequestHeader appends whatever header event the request owes the
log before dispatch: an 'initial'/'resume' snapshot anchoring each
loop instance, a round-trip-verified delta on change, a 'fallback'
snapshot when the encoding cannot express it. Session.requestHeader()
is the log's incrementally-folded baseline.
- Requests are deep-frozen before dispatch (deepFreeze exempts the
AbortSignal — freezing one breaks AbortController.abort() outright);
frozen + sessionId is the loop-built marker the dev invariant keys on.
Ported from #162 and re-anchored on the log: the append-extension /
frozen-end-to-end / compaction-resend / prompt-change property tests,
plus new specs for the boundary semantics, resume anchoring, and the
end-to-end theorem (every recorded request rebuilds byte-equal from the
log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against
the real DeepSeek API. Snapshot goldens intentionally stale until the
single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
|
|
|
it('agent/request waterfall switches models by returning a replacement config; the switch is logged', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall
The loop is now transmission-stateless; a request is a pure function of
(session log, this step's rendered assembly, current AgentOptions):
- The reconstruction boundary is step/start: the messages snapshot is
taken in the same synchronous frame immediately before the step/start
append, so the request's messages are exactly the derivation over
events[0..stepStartSeq) — an inject() from an agent/request listener
(or any concurrent task) lands after the boundary and joins the NEXT
request. This changes behavior for a synchronous step/start
session/event listener that appends content (master derived after the
append, so such a listener could reach the current request):
agent/pre-step is the sanctioned seam for current-request content.
- agent/request is re-typed to config-only: (agent, turn, step,
config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes
from AgentOptions on a loop instance's first request (explicit options
beat the logged baseline — fork overrides and resume reconfiguration
stay correct) and from the log's folded header afterwards; listeners
return a replacement to switch. Content shaping through the request is
no longer expressible — model-visible content flows through the log
channels.
- recordRequestHeader appends whatever header event the request owes the
log before dispatch: an 'initial'/'resume' snapshot anchoring each
loop instance, a round-trip-verified delta on change, a 'fallback'
snapshot when the encoding cannot express it. Session.requestHeader()
is the log's incrementally-folded baseline.
- Requests are deep-frozen before dispatch (deepFreeze exempts the
AbortSignal — freezing one breaks AbortController.abort() outright);
frozen + sessionId is the loop-built marker the dev invariant keys on.
Ported from #162 and re-anchored on the log: the append-extension /
frozen-end-to-end / compaction-resend / prompt-change property tests,
plus new specs for the boundary semantics, resume anchoring, and the
end-to-end theorem (every recorded request rebuilds byte-equal from the
log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against
the real DeepSeek API. Snapshot goldens intentionally stale until the
single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
|
|
|
ctx.on('agent/request', async (_agent, _turn, _step, config, _next) => {
|
|
|
|
|
// The seed is frozen — config is not a mutable per-call knob; a switch
|
|
|
|
|
// is proposed by returning a replacement, and the loop logs it.
|
|
|
|
|
expect(Object.isFrozen(config)).toBe(true)
|
|
|
|
|
expect(() => { (config as { model: string }).model = 'other-model' }).toThrow(TypeError)
|
|
|
|
|
return { ...config, model: 'other-model' }
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
expect(adapter.requests[0]!.model).toBe('other-model')
|
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall
The loop is now transmission-stateless; a request is a pure function of
(session log, this step's rendered assembly, current AgentOptions):
- The reconstruction boundary is step/start: the messages snapshot is
taken in the same synchronous frame immediately before the step/start
append, so the request's messages are exactly the derivation over
events[0..stepStartSeq) — an inject() from an agent/request listener
(or any concurrent task) lands after the boundary and joins the NEXT
request. This changes behavior for a synchronous step/start
session/event listener that appends content (master derived after the
append, so such a listener could reach the current request):
agent/pre-step is the sanctioned seam for current-request content.
- agent/request is re-typed to config-only: (agent, turn, step,
config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes
from AgentOptions on a loop instance's first request (explicit options
beat the logged baseline — fork overrides and resume reconfiguration
stay correct) and from the log's folded header afterwards; listeners
return a replacement to switch. Content shaping through the request is
no longer expressible — model-visible content flows through the log
channels.
- recordRequestHeader appends whatever header event the request owes the
log before dispatch: an 'initial'/'resume' snapshot anchoring each
loop instance, a round-trip-verified delta on change, a 'fallback'
snapshot when the encoding cannot express it. Session.requestHeader()
is the log's incrementally-folded baseline.
- Requests are deep-frozen before dispatch (deepFreeze exempts the
AbortSignal — freezing one breaks AbortController.abort() outright);
frozen + sessionId is the loop-built marker the dev invariant keys on.
Ported from #162 and re-anchored on the log: the append-extension /
frozen-end-to-end / compaction-resend / prompt-change property tests,
plus new specs for the boundary semantics, resume anchoring, and the
end-to-end theorem (every recorded request rebuilds byte-equal from the
log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against
the real DeepSeek API. Snapshot goldens intentionally stale until the
single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
|
|
|
// The header event records what the request ACTUALLY used — the switch is
|
|
|
|
|
// a reconstructable fact, not silent drift.
|
|
|
|
|
const headerEvent = agent.session.events.find(e => e.type === 'request/header')
|
|
|
|
|
expect(headerEvent?.type === 'request/header' && headerEvent.data.header.config.model).toBe('other-model')
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
it('agent/pre-step fires once per step before the step is opened', async () => {
|
2026-06-26 08:59:33 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'echo', {}, 'calling echo'),
|
|
|
|
|
textResponse('done'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
ctx.tools.register(defineTool({
|
|
|
|
|
name: 'echo', description: 'echo', parameters: {},
|
|
|
|
|
async execute() { return [{ type: 'text', text: 'echoed' }] },
|
|
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-26 08:59:33 +08:00
|
|
|
|
2026-07-15 16:50:44 +08:00
|
|
|
const fires: { turn: number; step: number; signal: AbortSignal }[] = []
|
|
|
|
|
ctx.on('agent/pre-step', (subject, turn, step, signal) => {
|
|
|
|
|
if (subject === agent) fires.push({ turn, step, signal })
|
2026-06-26 08:59:33 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-15 16:50:44 +08:00
|
|
|
expect(fires.map(({ turn, step }) => ({ turn, step }))).toEqual([
|
|
|
|
|
{ turn: 1, step: 1 },
|
|
|
|
|
{ turn: 1, step: 2 },
|
2026-06-26 08:59:33 +08:00
|
|
|
])
|
2026-07-15 16:50:44 +08:00
|
|
|
expect(fires.every(({ signal }) => signal instanceof AbortSignal)).toBe(true)
|
2026-06-26 08:59:33 +08:00
|
|
|
})
|
|
|
|
|
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
it('agent/pre-step fires BEFORE the step it precedes opens (events land outside the step)', async () => {
|
2026-07-13 23:27:00 +08:00
|
|
|
// The append lands before step/start, yet derive happens afterwards and the
|
|
|
|
|
// same step's request must include it.
|
2026-06-26 08:59:33 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-26 08:59:33 +08:00
|
|
|
|
|
|
|
|
let injected = false
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
ctx.on('agent/pre-step', (subject) => {
|
2026-06-26 08:59:33 +08:00
|
|
|
if (subject === agent && !injected) {
|
|
|
|
|
injected = true
|
|
|
|
|
subject.session.append('context/message', {
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
content: [{ type: 'text', text: 'INJECTED-IN-PRE-STEP' }],
|
2026-06-26 08:59:33 +08:00
|
|
|
source: { kind: 'plugin', plugin: 'test' },
|
|
|
|
|
}, { surfaceOp: 'append' })
|
|
|
|
|
}
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
// The adapter's request includes the node injected during pre-step (derive
|
|
|
|
|
// reflects it).
|
2026-06-26 08:59:33 +08:00
|
|
|
const text = JSON.stringify(adapter.requests[0]!.messages)
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
expect(text).toContain('INJECTED-IN-PRE-STEP')
|
|
|
|
|
|
|
|
|
|
// And the injected event sits BEFORE the first step/start in the log —
|
|
|
|
|
// the seam fired outside the step.
|
|
|
|
|
const events = agent.session.events
|
|
|
|
|
const injectedSeq = events.find(e => e.type === 'context/message')!.seq
|
|
|
|
|
const firstStepStartSeq = events.find(e => e.type === 'step/start')!.seq
|
|
|
|
|
expect(injectedSeq).toBeLessThan(firstStepStartSeq)
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('a throwing agent/pre-step listener ends the turn (error), not the loop', async () => {
|
2026-07-13 23:27:00 +08:00
|
|
|
// Before step/start, a pre-step throw reaches the turn catch: no step needs
|
|
|
|
|
// closing, the turn records error, and the loop remains available.
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('second turn ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
|
|
|
|
|
let throwOnce = true
|
|
|
|
|
ctx.on('agent/pre-step', () => {
|
|
|
|
|
if (throwOnce) { throwOnce = false; throw new Error('boom in pre-step') }
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
const errors: Error[] = []
|
|
|
|
|
ctx.on('agent/error', (_a, _t, _s, error) => void errors.push(error))
|
|
|
|
|
|
|
|
|
|
send(agent, 'first')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
// The first turn failed at step 1 (no model call happened), surfaced via
|
|
|
|
|
// agent/error, with the durable failure on turn/end.reason.
|
|
|
|
|
expect(errors).toHaveLength(1)
|
|
|
|
|
expect(errors[0]!.message).toContain('boom in pre-step')
|
|
|
|
|
expect(adapter.requests.length).toBe(0)
|
|
|
|
|
const firstTurnEnd = agent.session.events.find(e => e.type === 'turn/end')
|
|
|
|
|
expect(firstTurnEnd?.type === 'turn/end' && firstTurnEnd.data.reason).toMatchObject({ kind: 'error', step: 1 })
|
|
|
|
|
// The step opened-and-closed count stays balanced even though it never ran.
|
|
|
|
|
const types = agent.session.events.map(e => e.type)
|
|
|
|
|
expect(types.filter(t => t === 'step/start').length).toBe(types.filter(t => t === 'step/end').length)
|
|
|
|
|
|
|
|
|
|
// The loop survived: a second prompt runs a normal completed turn.
|
|
|
|
|
send(agent, 'second')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(adapter.requests.length).toBe(1)
|
|
|
|
|
const lastTurnEnd = agent.session.events.findLast(e => e.type === 'turn/end')
|
|
|
|
|
expect(lastTurnEnd?.type === 'turn/end' && lastTurnEnd.data.reason).toEqual({ kind: 'completed' })
|
2026-06-26 08:59:33 +08:00
|
|
|
})
|
|
|
|
|
|
simplify(agent): drop the unused public Agent.abort(), keep whenIdle()
The public Agent handle exposed abort() (step-only) and cancel() (queue-aware).
No production caller used abort() — ACP maps session/cancel to cancel(), and
lifecycle owners tear down via AgentHandle.dispose(); the loop's own stop paths
abort their per-step AbortController directly. So abort() is latent generality
that keeps a private loop mechanic public.
RFC-premise correction: the public-agent-stop-surface RFC proposed removing
whenIdle() too. Implementation found whenIdle() load-bearing — a real
quiescence primitive with a deliberate loop contract (settle-without-transition,
the replacement-turn race) and ACP test consumers; its proposed replacement
("observe the running->idle transition") is exactly the async-state race
AGENTS.md warns against. So only abort() is removed; whenIdle() stays. The RFC
is amended on the way to implemented/ to record the narrowed scope, and the new
AGENTS.md "RFCs are proposals, not golden truth" principle (PR1) gets its
worked example.
- Remove Agent.abort() from the interface + the ReactLoopAgent impl; the no-arg
'aborted' default goes with it (cancel() keeps its 'cancelled' default).
- Migrate tests: empty-queue abort() -> cancel(reason); the two review-fixes
tests whose subject is the in-flight step's AbortController drive that
controller directly via the private currentAbort field (cancel() would clear
the inbox and destroy the queued steering one of them proves survives a step
abort). The no-arg-default test is dropped (cancel()'s default is already
covered in cancel.spec.ts).
- Resulting public stop surface: cancel() + whenIdle(). Update agent/agent-loop
READMEs, architecture.md, core.md type-equiv, the extension cookbook, the
lifecycle RFC (short note), and the proposed ACP RFC.
Implements docs/rfc/implemented/simplification/2026-06-20-public-agent-stop-surface.md
2026-06-21 05:50:39 +08:00
|
|
|
it('cancel() mid-stream ends the turn with reason aborted', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter(['hang'])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
simplify(agent): drop the unused public Agent.abort(), keep whenIdle()
The public Agent handle exposed abort() (step-only) and cancel() (queue-aware).
No production caller used abort() — ACP maps session/cancel to cancel(), and
lifecycle owners tear down via AgentHandle.dispose(); the loop's own stop paths
abort their per-step AbortController directly. So abort() is latent generality
that keeps a private loop mechanic public.
RFC-premise correction: the public-agent-stop-surface RFC proposed removing
whenIdle() too. Implementation found whenIdle() load-bearing — a real
quiescence primitive with a deliberate loop contract (settle-without-transition,
the replacement-turn race) and ACP test consumers; its proposed replacement
("observe the running->idle transition") is exactly the async-state race
AGENTS.md warns against. So only abort() is removed; whenIdle() stays. The RFC
is amended on the way to implemented/ to record the narrowed scope, and the new
AGENTS.md "RFCs are proposals, not golden truth" principle (PR1) gets its
worked example.
- Remove Agent.abort() from the interface + the ReactLoopAgent impl; the no-arg
'aborted' default goes with it (cancel() keeps its 'cancelled' default).
- Migrate tests: empty-queue abort() -> cancel(reason); the two review-fixes
tests whose subject is the in-flight step's AbortController drive that
controller directly via the private currentAbort field (cancel() would clear
the inbox and destroy the queued steering one of them proves survives a step
abort). The no-arg-default test is dropped (cancel()'s default is already
covered in cancel.spec.ts).
- Resulting public stop surface: cancel() + whenIdle(). Update agent/agent-loop
READMEs, architecture.md, core.md type-equiv, the extension cookbook, the
lifecycle RFC (short note), and the proposed ACP RFC.
Implements docs/rfc/implemented/simplification/2026-06-20-public-agent-stop-surface.md
2026-06-21 05:50:39 +08:00
|
|
|
// wait until the stream is hanging, then cancel
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
await new Promise(r => setTimeout(r, 30))
|
|
|
|
|
expect(agent.status).toBe('running')
|
simplify(agent): drop the unused public Agent.abort(), keep whenIdle()
The public Agent handle exposed abort() (step-only) and cancel() (queue-aware).
No production caller used abort() — ACP maps session/cancel to cancel(), and
lifecycle owners tear down via AgentHandle.dispose(); the loop's own stop paths
abort their per-step AbortController directly. So abort() is latent generality
that keeps a private loop mechanic public.
RFC-premise correction: the public-agent-stop-surface RFC proposed removing
whenIdle() too. Implementation found whenIdle() load-bearing — a real
quiescence primitive with a deliberate loop contract (settle-without-transition,
the replacement-turn race) and ACP test consumers; its proposed replacement
("observe the running->idle transition") is exactly the async-state race
AGENTS.md warns against. So only abort() is removed; whenIdle() stays. The RFC
is amended on the way to implemented/ to record the narrowed scope, and the new
AGENTS.md "RFCs are proposals, not golden truth" principle (PR1) gets its
worked example.
- Remove Agent.abort() from the interface + the ReactLoopAgent impl; the no-arg
'aborted' default goes with it (cancel() keeps its 'cancelled' default).
- Migrate tests: empty-queue abort() -> cancel(reason); the two review-fixes
tests whose subject is the in-flight step's AbortController drive that
controller directly via the private currentAbort field (cancel() would clear
the inbox and destroy the queued steering one of them proves survives a step
abort). The no-arg-default test is dropped (cancel()'s default is already
covered in cancel.spec.ts).
- Resulting public stop surface: cancel() + whenIdle(). Update agent/agent-loop
READMEs, architecture.md, core.md type-equiv, the extension cookbook, the
lifecycle RFC (short note), and the proposed ACP RFC.
Implements docs/rfc/implemented/simplification/2026-06-20-public-agent-stop-surface.md
2026-06-21 05:50:39 +08:00
|
|
|
agent.cancel('user interrupt')
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(reasons).toEqual([{ kind: 'aborted', reason: 'user interrupt' }])
|
|
|
|
|
})
|
|
|
|
|
|
2026-06-15 23:53:47 +08:00
|
|
|
it('surfaces max-tokens as the turn-end reason when the last step is cut off', async () => {
|
|
|
|
|
// A single step that ends with a max-tokens finish (no tool calls): the
|
|
|
|
|
// turn stops by default and ends max-tokens, not completed.
|
|
|
|
|
const adapter = new MockAdapter([maxTokensResponse('truncat')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
expect(reasons).toEqual([{ kind: 'max-tokens' }])
|
2026-07-13 23:27:00 +08:00
|
|
|
// Assert the durable row, not only the live listener.
|
2026-06-15 23:53:47 +08:00
|
|
|
const turnEnd = agent.session.events.findLast(e => e.type === 'turn/end')
|
|
|
|
|
expect(turnEnd!.data.reason).toEqual({ kind: 'max-tokens' })
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('a max-tokens step earlier in a turn still surfaces as max-tokens after a later completed step', async () => {
|
2026-07-12 03:36:43 +08:00
|
|
|
// Step 1 is cut off (max-tokens, no tool calls → would stop by default), so continuation
|
|
|
|
|
// must be FORCED to reach step 2 which finishes normally (stop).
|
2026-06-15 23:53:47 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
maxTokensResponse('first half'),
|
|
|
|
|
textResponse('second half'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
let steps = 0
|
2026-06-30 10:32:55 +08:00
|
|
|
ctx.on('session/event', (_session, event) => { if (event.type === 'step/end') steps++ })
|
2026-06-15 23:53:47 +08:00
|
|
|
// Force exactly one continuation (step 1 → step 2), then defer to default
|
|
|
|
|
// (step 2 is a plain stop with no tool calls → stops).
|
|
|
|
|
ctx.on('agent/turn-continuation', async (_agent, _turn, _defaultDecision, next) => {
|
2026-06-30 17:11:18 +08:00
|
|
|
if (steps < 2) return { action: 'continue' as const }
|
2026-06-15 23:53:47 +08:00
|
|
|
return next()
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(steps).toBe(2)
|
|
|
|
|
expect(adapter.requests).toHaveLength(2)
|
2026-06-19 00:37:12 +08:00
|
|
|
expect(adapter.requests[1]!.messages).toEqual([
|
|
|
|
|
{ role: 'user', content: [{ type: 'text', text: 'go' }] },
|
2026-07-14 21:57:52 +08:00
|
|
|
{ role: 'assistant', content: [{ type: 'text', text: 'first half' }], provenance: { provider: 'mock', model: 'mock' } },
|
2026-06-19 00:37:12 +08:00
|
|
|
])
|
2026-06-15 23:53:47 +08:00
|
|
|
expect(reasons).toEqual([{ kind: 'max-tokens' }])
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('a completed step after no max-tokens keeps the turn completed (max-tokens does not leak across turns)', async () => {
|
|
|
|
|
// Two consecutive turns: turn 1 is cut off (max-tokens), turn 2 is a clean
|
|
|
|
|
// stop. The per-turn reason must be independent — turn 2 ends completed.
|
|
|
|
|
const adapter = new MockAdapter([maxTokensResponse('cut'), textResponse('clean')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
send(agent, 'first')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
send(agent, 'second')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(reasons).toEqual([{ kind: 'max-tokens' }, { kind: 'completed' }])
|
|
|
|
|
})
|
|
|
|
|
|
2026-06-17 21:25:47 +08:00
|
|
|
it('does not dispatch tool calls from a max-tokens-truncated step', async () => {
|
|
|
|
|
const callId = CallId('c1')
|
|
|
|
|
const adapter = new MockAdapter([[
|
|
|
|
|
{ type: 'block-start', index: 0, blockType: 'tool-call' },
|
|
|
|
|
{ type: 'tool-call-delta', index: 0, id: callId, name: 'echo', argumentsDelta: '{"text":"x"}' },
|
|
|
|
|
{ type: 'block-end', index: 0, block: { type: 'tool-call', id: callId, name: 'echo', arguments: '{"text":"x"}' } },
|
|
|
|
|
{ type: 'usage', usage: { inputTokens: 10, outputTokens: 5 } },
|
|
|
|
|
{ type: 'finish', reason: { kind: 'max-tokens' } },
|
|
|
|
|
]])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
let executions = 0
|
|
|
|
|
ctx.tools.register(defineTool({
|
|
|
|
|
name: 'echo',
|
|
|
|
|
description: '',
|
|
|
|
|
parameters: { text: { type: 'string' } },
|
|
|
|
|
async execute() {
|
|
|
|
|
executions += 1
|
|
|
|
|
return [{ type: 'text', text: 'should not run' }]
|
|
|
|
|
},
|
|
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-17 21:25:47 +08:00
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-17 21:25:47 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(executions).toBe(0)
|
|
|
|
|
expect(agent.session.events.some(e => e.type === 'tool/call')).toBe(false)
|
2026-06-19 00:37:12 +08:00
|
|
|
expect(agent.session.deriveMessages()).toEqual([{ role: 'user', content: [{ type: 'text', text: 'go' }] }])
|
|
|
|
|
expect(reasons).toEqual([{ kind: 'max-tokens' }])
|
2026-07-13 23:27:00 +08:00
|
|
|
// Empty content still needs an assistant/message to carry usage; derivation
|
|
|
|
|
// skips that host so it does not create a spurious assistant turn.
|
2026-06-21 10:00:06 +08:00
|
|
|
const assistantMessage = agent.session.events.find(e => e.type === 'assistant/message')
|
|
|
|
|
expect(assistantMessage?.type === 'assistant/message' && assistantMessage.data).toEqual({
|
2026-07-14 21:57:52 +08:00
|
|
|
turn: 1, step: 1, content: [], provenance: { provider: 'mock', model: 'mock' }, usage: { inputTokens: 10, outputTokens: 5 },
|
2026-06-21 10:00:06 +08:00
|
|
|
})
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-15 14:47:29 +08:00
|
|
|
it('appends an empty completion anchor for a max-tokens step with no usage', async () => {
|
|
|
|
|
// The truncated tool call is dropped from durable content, while the
|
|
|
|
|
// successful provider call still needs an exact replay anchor.
|
2026-06-21 10:00:06 +08:00
|
|
|
const callId = CallId('c1')
|
|
|
|
|
const adapter = new MockAdapter([[
|
|
|
|
|
{ type: 'block-start', index: 0, blockType: 'tool-call' },
|
|
|
|
|
{ type: 'tool-call-delta', index: 0, id: callId, name: 'echo', argumentsDelta: '{"text":"x"}' },
|
|
|
|
|
{ type: 'block-end', index: 0, block: { type: 'tool-call', id: callId, name: 'echo', arguments: '{"text":"x"}' } },
|
|
|
|
|
{ type: 'finish', reason: { kind: 'max-tokens' } },
|
|
|
|
|
]])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
ctx.tools.register(defineTool({
|
|
|
|
|
name: 'echo',
|
|
|
|
|
description: '',
|
|
|
|
|
parameters: { text: { type: 'string' } },
|
|
|
|
|
async execute() { return [{ type: 'text', text: 'should not run' }] },
|
|
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-21 10:00:06 +08:00
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-21 10:00:06 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(reasons).toEqual([{ kind: 'max-tokens' }])
|
2026-07-15 14:47:29 +08:00
|
|
|
const assistant = agent.session.events.find(e => e.type === 'assistant/message')!
|
|
|
|
|
expect(assistant.type === 'assistant/message' && assistant.data).toEqual({
|
|
|
|
|
turn: 1,
|
|
|
|
|
step: 1,
|
|
|
|
|
content: [],
|
2026-07-17 22:38:46 +08:00
|
|
|
provenance: { provider: 'mock', model: 'mock' },
|
2026-07-15 14:47:29 +08:00
|
|
|
})
|
|
|
|
|
expect(assistant.sourceEventSeqs?.length).toBeGreaterThan(0)
|
2026-06-21 10:00:06 +08:00
|
|
|
expect(agent.session.deriveMessages()).toEqual([{ role: 'user', content: [{ type: 'text', text: 'go' }] }])
|
2026-06-19 00:37:12 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-15 14:47:29 +08:00
|
|
|
it('appends an empty completion anchor for a normal stop with no usage', async () => {
|
|
|
|
|
// A clean content-less call stays absent from derived messages but remains
|
|
|
|
|
// a durable successful-call boundary for replay consumers.
|
2026-06-21 11:08:10 +08:00
|
|
|
const adapter = new MockAdapter([[{ type: 'finish', reason: { kind: 'stop' } }]])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-21 11:08:10 +08:00
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-21 11:08:10 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(reasons).toEqual([{ kind: 'completed' }])
|
2026-07-15 14:47:29 +08:00
|
|
|
const assistant = agent.session.events.find(e => e.type === 'assistant/message')!
|
|
|
|
|
expect(assistant.type === 'assistant/message' && assistant.data).toEqual({
|
|
|
|
|
turn: 1,
|
|
|
|
|
step: 1,
|
|
|
|
|
content: [],
|
2026-07-17 22:38:46 +08:00
|
|
|
provenance: { provider: 'mock', model: 'mock' },
|
2026-07-15 14:47:29 +08:00
|
|
|
})
|
|
|
|
|
expect(assistant.sourceEventSeqs?.length).toBe(1)
|
2026-06-21 11:08:10 +08:00
|
|
|
expect(agent.session.deriveMessages()).toEqual([{ role: 'user', content: [{ type: 'text', text: 'go' }] }])
|
|
|
|
|
})
|
|
|
|
|
|
2026-06-19 00:37:12 +08:00
|
|
|
it('keeps safe max-tokens assistant content while dropping truncated tool calls', async () => {
|
|
|
|
|
const callId = CallId('c1')
|
|
|
|
|
const adapter = new MockAdapter([[
|
|
|
|
|
{ type: 'block-start', index: 0, blockType: 'text' },
|
|
|
|
|
{ type: 'text-delta', index: 0, text: 'partial text' },
|
|
|
|
|
{ type: 'block-end', index: 0, block: { type: 'text', text: 'partial text' } },
|
|
|
|
|
{ type: 'block-start', index: 1, blockType: 'tool-call' },
|
|
|
|
|
{ type: 'tool-call-delta', index: 1, id: callId, name: 'echo', argumentsDelta: '{"text"' },
|
|
|
|
|
{ type: 'finish', reason: { kind: 'max-tokens' } },
|
|
|
|
|
]])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
let stepResults = 0
|
|
|
|
|
ctx.on('agent/step-result', async (_agent, _turn, _step, message, next) => {
|
|
|
|
|
stepResults += 1
|
|
|
|
|
expect(message.content).toEqual([{ type: 'text', text: 'partial text' }])
|
|
|
|
|
return next()
|
|
|
|
|
})
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-19 00:37:12 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(stepResults).toBe(1)
|
|
|
|
|
expect(agent.session.events.some(e => e.type === 'tool/call')).toBe(false)
|
2026-06-17 21:25:47 +08:00
|
|
|
expect(agent.session.deriveMessages()).toEqual([
|
|
|
|
|
{ role: 'user', content: [{ type: 'text', text: 'go' }] },
|
2026-07-14 21:57:52 +08:00
|
|
|
{ role: 'assistant', content: [{ type: 'text', text: 'partial text' }], provenance: { provider: 'mock', model: 'mock' } },
|
2026-06-17 21:25:47 +08:00
|
|
|
])
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-12 18:57:42 +08:00
|
|
|
it('contains a step/end observer failure without changing continuation', async () => {
|
2026-06-17 21:25:47 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'echo', { text: 'x' }),
|
2026-07-12 18:57:42 +08:00
|
|
|
textResponse('continued after tool call'),
|
2026-06-17 21:25:47 +08:00
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
ctx.tools.register(defineTool({
|
|
|
|
|
name: 'echo',
|
|
|
|
|
description: '',
|
|
|
|
|
parameters: { text: { type: 'string' } },
|
|
|
|
|
async execute(args) {
|
|
|
|
|
return [{ type: 'text', text: String(args.text) }]
|
|
|
|
|
},
|
|
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-17 21:25:47 +08:00
|
|
|
let threw = false
|
2026-07-12 18:57:42 +08:00
|
|
|
// Post-commit session observers cannot control the loop. The tool call still
|
|
|
|
|
// drives the second model request, and the turn completes normally.
|
2026-06-30 10:32:55 +08:00
|
|
|
ctx.on('session/event', (_session, event) => {
|
|
|
|
|
if (event.type === 'step/end' && !threw) { threw = true; throw new Error('bad step/end listener') }
|
2026-06-17 21:25:47 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-12 18:57:42 +08:00
|
|
|
expect(adapter.requests).toHaveLength(2)
|
2026-06-17 21:25:47 +08:00
|
|
|
const turnEnd = agent.session.events.findLast(e => e.type === 'turn/end')
|
2026-07-12 18:57:42 +08:00
|
|
|
expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason.kind).toBe('completed')
|
2026-06-17 21:25:47 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-17 17:14:52 +08:00
|
|
|
it('keeps same-tick sends in separate turns and checkpoints before the next starts', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('first answer'), textResponse('second answer')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-20 11:52:30 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-17 17:14:52 +08:00
|
|
|
|
|
|
|
|
const firstFlush = Promise.withResolvers<undefined>()
|
|
|
|
|
const releaseFirstFlush = Promise.withResolvers<undefined>()
|
|
|
|
|
let flushes = 0
|
|
|
|
|
ctx.on('session/flush', async (session) => {
|
|
|
|
|
if (session !== agent.session) return
|
|
|
|
|
flushes += 1
|
|
|
|
|
if (flushes === 1) {
|
|
|
|
|
firstFlush.resolve(undefined)
|
|
|
|
|
await releaseFirstFlush.promise
|
|
|
|
|
}
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
const turns: number[] = []
|
|
|
|
|
ctx.on('session/event', (session, event) => {
|
|
|
|
|
if (session === agent.session && event.type === 'turn/start') turns.push(event.data.turn)
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
const idle = waitForIdle(ctx, agent)
|
|
|
|
|
send(agent, 'first message')
|
|
|
|
|
send(agent, 'second message')
|
|
|
|
|
|
|
|
|
|
await firstFlush.promise
|
|
|
|
|
expect(turns).toEqual([1])
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
|
|
|
|
|
releaseFirstFlush.resolve(undefined)
|
|
|
|
|
await idle
|
|
|
|
|
|
|
|
|
|
expect(turns).toEqual([1, 2])
|
|
|
|
|
expect(flushes).toBe(2)
|
|
|
|
|
expect(adapter.requests).toHaveLength(2)
|
|
|
|
|
expect(JSON.stringify(adapter.requests[1]!.messages)).toContain('first answer')
|
|
|
|
|
expect(JSON.stringify(adapter.requests[1]!.messages)).toContain('second message')
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-17 17:27:53 +08:00
|
|
|
it('holds a turn-end listener send behind the closing turn checkpoint', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('first answer'), textResponse('second answer')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-20 11:52:30 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-17 17:27:53 +08:00
|
|
|
|
|
|
|
|
const firstFlush = Promise.withResolvers<undefined>()
|
|
|
|
|
const releaseFirstFlush = Promise.withResolvers<undefined>()
|
|
|
|
|
let flushes = 0
|
|
|
|
|
ctx.on('session/flush', async (session) => {
|
|
|
|
|
if (session !== agent.session) return
|
|
|
|
|
flushes += 1
|
|
|
|
|
if (flushes === 1) {
|
|
|
|
|
firstFlush.resolve(undefined)
|
|
|
|
|
await releaseFirstFlush.promise
|
|
|
|
|
}
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
const turns: number[] = []
|
|
|
|
|
const statuses: string[] = []
|
|
|
|
|
ctx.on('agent/status', (subject, status) => {
|
|
|
|
|
if (subject === agent) statuses.push(status)
|
|
|
|
|
})
|
|
|
|
|
ctx.on('session/event', (session, event) => {
|
|
|
|
|
if (session !== agent.session) return
|
|
|
|
|
if (event.type === 'turn/start') turns.push(event.data.turn)
|
|
|
|
|
if (event.type === 'turn/end' && event.data.turn === 1) send(agent, 'turn-end listener message')
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
const idle = waitForIdle(ctx, agent)
|
|
|
|
|
send(agent, 'first message')
|
|
|
|
|
await firstFlush.promise
|
|
|
|
|
|
|
|
|
|
expect(turns).toEqual([1])
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
|
|
|
|
|
releaseFirstFlush.resolve(undefined)
|
|
|
|
|
await idle
|
|
|
|
|
|
|
|
|
|
expect(turns).toEqual([1, 2])
|
|
|
|
|
expect(statuses).toEqual(['running', 'idle'])
|
|
|
|
|
expect(adapter.requests).toHaveLength(2)
|
|
|
|
|
expect(JSON.stringify(adapter.requests[1]!.messages)).toContain('first answer')
|
|
|
|
|
expect(JSON.stringify(adapter.requests[1]!.messages)).toContain('turn-end listener message')
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-17 17:14:52 +08:00
|
|
|
it('keeps a reentrant agent/queued send as the next independent turn', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('first'), textResponse('second')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-20 11:52:30 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-17 17:14:52 +08:00
|
|
|
|
|
|
|
|
let nested = false
|
|
|
|
|
ctx.on('agent/queued', (subject) => {
|
|
|
|
|
if (subject !== agent || nested) return
|
|
|
|
|
nested = true
|
|
|
|
|
send(agent, 'queued listener message')
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
const idle = waitForIdle(ctx, agent)
|
|
|
|
|
send(agent, 'outer message')
|
|
|
|
|
await idle
|
|
|
|
|
|
|
|
|
|
const turns = agent.session.events.filter(event => event.type === 'turn/start')
|
|
|
|
|
const messages = agent.session.events
|
|
|
|
|
.filter(event => event.type === 'user/message')
|
|
|
|
|
.map(event => event.data.content)
|
|
|
|
|
expect(turns).toHaveLength(2)
|
|
|
|
|
expect(messages).toEqual([
|
|
|
|
|
[{ type: 'text', text: 'outer message' }],
|
|
|
|
|
[{ type: 'text', text: 'queued listener message' }],
|
|
|
|
|
])
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('preserves independent turn sources across an adjacent microtask send', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('first'), textResponse('second')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-20 11:52:30 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-17 17:14:52 +08:00
|
|
|
|
|
|
|
|
const idle = waitForIdle(ctx, agent)
|
|
|
|
|
agent.send([{ type: 'text', text: 'user message' }])
|
|
|
|
|
await Promise.resolve()
|
|
|
|
|
agent.send(
|
|
|
|
|
[{ type: 'text', text: 'plugin message' }],
|
|
|
|
|
{ source: { kind: 'plugin', plugin: 'test' } },
|
|
|
|
|
)
|
|
|
|
|
await idle
|
|
|
|
|
|
|
|
|
|
const triggers = agent.session.events
|
|
|
|
|
.filter(event => event.type === 'turn/start')
|
|
|
|
|
.map(event => event.data.trigger)
|
|
|
|
|
const sources = agent.session.events
|
|
|
|
|
.filter(event => event.type === 'user/message')
|
|
|
|
|
.map(event => event.data.source)
|
|
|
|
|
expect(triggers).toEqual([
|
|
|
|
|
{ kind: 'message', source: { kind: 'user' } },
|
|
|
|
|
{ kind: 'message', source: { kind: 'plugin', plugin: 'test' } },
|
|
|
|
|
])
|
|
|
|
|
expect(sources).toEqual([
|
|
|
|
|
{ kind: 'user' },
|
|
|
|
|
{ kind: 'plugin', plugin: 'test' },
|
|
|
|
|
])
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('keeps a session-listener send after dequeue in the following turn', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('first'), textResponse('second')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
const turns: number[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/start') turns.push(event.data.turn) })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
// queue two messages while idle — first starts turn 1 immediately;
|
2026-07-02 23:42:16 +08:00
|
|
|
// queue the second during turn 1 when the first assistant chunk streams
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
let queued = false
|
2026-07-02 23:42:16 +08:00
|
|
|
ctx.on('session/event', (_s, event) => {
|
|
|
|
|
if (event.type === 'assistant/chunk' && !queued) {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
queued = true
|
|
|
|
|
send(agent, 'second message')
|
|
|
|
|
}
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'first message')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(turns).toEqual([1, 2])
|
|
|
|
|
expect(adapter.requests).toHaveLength(2)
|
2026-07-17 17:14:52 +08:00
|
|
|
expect(JSON.stringify(adapter.requests[1]!.messages)).toContain('first')
|
|
|
|
|
expect(JSON.stringify(adapter.requests[1]!.messages)).toContain('second message')
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('keeps a model-adapter callback send in the following turn', async () => {
|
2026-07-20 11:52:30 +08:00
|
|
|
const agentRef: { current?: Agent } = {}
|
2026-07-17 17:14:52 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
() => {
|
|
|
|
|
const agent = agentRef.current
|
|
|
|
|
if (agent === undefined) throw new Error('model callback ran before agent setup')
|
|
|
|
|
send(agent, 'model callback message')
|
|
|
|
|
return textResponse('first')
|
|
|
|
|
},
|
|
|
|
|
textResponse('second'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-20 11:52:30 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-17 17:14:52 +08:00
|
|
|
agentRef.current = agent
|
|
|
|
|
|
|
|
|
|
const idle = waitForIdle(ctx, agent)
|
|
|
|
|
send(agent, 'outer message')
|
|
|
|
|
await idle
|
|
|
|
|
|
|
|
|
|
const messages = agent.session.events
|
|
|
|
|
.filter(event => event.type === 'user/message')
|
|
|
|
|
.map(event => event.data.content)
|
|
|
|
|
expect(agent.session.events.filter(event => event.type === 'turn/start')).toHaveLength(2)
|
|
|
|
|
expect(messages).toEqual([
|
|
|
|
|
[{ type: 'text', text: 'outer message' }],
|
|
|
|
|
[{ type: 'text', text: 'model callback message' }],
|
|
|
|
|
])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('awaits session/flush at turn end (persistence checkpoint)', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
let flushed = 0
|
|
|
|
|
let flushedBeforeIdle = false
|
|
|
|
|
ctx.on('session/flush', async (session) => {
|
|
|
|
|
await new Promise(r => setTimeout(r, 10))
|
|
|
|
|
flushed++
|
|
|
|
|
flushedBeforeIdle = agent.status !== 'idle'
|
|
|
|
|
void session
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(flushed).toBe(1)
|
|
|
|
|
expect(flushedBeforeIdle).toBe(true)
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('errors from the model surface as agent/error and end the turn', async () => {
|
|
|
|
|
const adapter = new MockAdapter([]) // script exhausted → throws
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
const errors: Error[] = []
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
const reasons: TurnEndReason[] = []
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
ctx.on('agent/error', (_agent, _turn, _step, error) => void errors.push(error))
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(errors).toHaveLength(1)
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
expect(errors[0]!.message).toContain('script exhausted')
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(reasons[0]).toMatchObject({ kind: 'error' })
|
2026-06-21 10:00:06 +08:00
|
|
|
// The durable failure lives entirely on turn/end.reason (with the failing
|
|
|
|
|
// step), not a standalone error event.
|
|
|
|
|
const turnEnd = agent.session.events.find(e => e.type === 'turn/end')
|
|
|
|
|
expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason).toMatchObject({ kind: 'error', step: 1 })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('disposing the loop fiber mid-turn stops the loop (HMR safety)', async () => {
|
|
|
|
|
const adapter = new MockAdapter(['hang'])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
let agent!: Agent
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const fiber = await ctx.plugin(Object.assign((inner: Context) => {
|
2026-07-18 12:21:15 +08:00
|
|
|
agent = inner.agentLoop.create(SessionId('scoped'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
}, { inject: ['agentLoop'] }))
|
|
|
|
|
|
2026-07-14 01:59:21 +08:00
|
|
|
expect(ctx.agents.get(SessionId('scoped'))).toBe(agent)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
send(agent, 'go')
|
|
|
|
|
await new Promise(r => setTimeout(r, 30))
|
|
|
|
|
expect(agent.status).toBe('running')
|
|
|
|
|
|
|
|
|
|
await fiber.dispose()
|
2026-07-14 02:32:35 +08:00
|
|
|
await driverDone(agent)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
expect(agent.status).toBe('disposed')
|
2026-07-14 01:59:21 +08:00
|
|
|
expect(ctx.agents.get(SessionId('scoped'))).toBeUndefined()
|
Add ESLint: typescript-eslint strict-type-checked + stylistic formatting
Flat config with two layers. Correctness (type-checked): the headline
rules for this codebase are no-floating-promises / no-misused-promises
(a lost promise in the agent loop is our primary bug class),
switch-exhaustiveness-check (we switch over merge-extensible unions
everywhere), no-unnecessary-condition, require-await, and
no-explicit-any. Style (@stylistic): 2-space, no semicolons, single
quotes, trailing commas, max-len 140 — the existing house style, now
enforced instead of drifting between agents. vendor/ is excluded
(vendored source keeps upstream style); tests relax the rules that
fight test ergonomics (non-null assertions after expects, async mock
signatures, non-Error throws).
Code adjusted to pass: registry disposers wrap ctx.effect's
promise-returning disposer behind a sync () => void (our public API),
BlockAssembler gains an invariant-checking mustGet instead of non-null
assertions, lastTurnNumber uses findLast, waterfall tails return
Promise.resolve instead of async-without-await arrows, and the two
deliberate suppressions (non-exhaustive derivation switch, unbound
execute pass-through) carry justification comments.
yarn lint / yarn lint:fix added.
2026-06-11 14:17:58 +08:00
|
|
|
expect(() => { send(agent, 'too late') }).toThrow('disposed')
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
it('creates agents from config on startup', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('from config')])
|
|
|
|
|
const ctx = new Context()
|
|
|
|
|
await ctx.plugin(LlmService)
|
|
|
|
|
await ctx.plugin(SessionStore)
|
|
|
|
|
await ctx.plugin(SystemPrompt)
|
|
|
|
|
await ctx.plugin(ToolRegistry)
|
|
|
|
|
await ctx.plugin(AgentRegistry)
|
|
|
|
|
await ctx.plugin(AgentLoop, {
|
2026-07-18 12:21:15 +08:00
|
|
|
agents: [{ id: SessionId('config-agent'), provider: 'mock', model: 'mock' }],
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
})
|
|
|
|
|
ctx.llm.registerAdapter(['mock'], adapter)
|
|
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
const agent = ctx.agents.list()[0]!
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
expect(agent).toBeDefined()
|
2026-07-14 01:59:21 +08:00
|
|
|
expect(agent.id).toBe(agent.session.id)
|
|
|
|
|
expect(agent.id).toMatch(/^config-agent-session-/)
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
expect(agent.options.model).toBe('mock')
|
|
|
|
|
|
|
|
|
|
// the agent is alive: send triggers a turn
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
})
|
|
|
|
|
|
2026-06-25 23:35:13 +08:00
|
|
|
it('attaches config agent cwd to the fresh session header', async () => {
|
|
|
|
|
const ctx = new Context()
|
|
|
|
|
await ctx.plugin(LlmService)
|
|
|
|
|
await ctx.plugin(SessionStore)
|
|
|
|
|
await ctx.plugin(SystemPrompt)
|
|
|
|
|
await ctx.plugin(ToolRegistry)
|
|
|
|
|
await ctx.plugin(AgentRegistry)
|
|
|
|
|
await ctx.plugin(AgentLoop, {
|
2026-07-18 12:21:15 +08:00
|
|
|
agents: [{ id: SessionId('config-agent'), provider: 'mock', model: 'mock', cwd: '/work/project' }],
|
2026-06-25 23:35:13 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
const agent = ctx.agents.list()[0]!
|
2026-06-25 23:35:13 +08:00
|
|
|
expect(agent.session.header.cwd).toBe('/work/project')
|
|
|
|
|
})
|
|
|
|
|
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
it('replays a session log into an identical derived history', async () => {
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'echo', { text: 'x' }),
|
|
|
|
|
textResponse('done'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
ctx.tools.register(defineTool({
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
name: 'echo',
|
|
|
|
|
description: '',
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
parameters: { text: { type: 'string' } },
|
|
|
|
|
async execute(args) {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
return [{ type: 'text', text: String(args.text) }]
|
|
|
|
|
},
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
send(agent, 'run')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
|
|
|
const replayed = ctx.sessions.create(SessionId('replayed'), { seed: [...agent.session.events] })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(replayed.deriveMessages()).toEqual(agent.session.deriveMessages())
|
|
|
|
|
// event-by-event identity of types
|
|
|
|
|
expect(replayed.events.map(e => e.type)).toEqual(
|
Add ESLint: typescript-eslint strict-type-checked + stylistic formatting
Flat config with two layers. Correctness (type-checked): the headline
rules for this codebase are no-floating-promises / no-misused-promises
(a lost promise in the agent loop is our primary bug class),
switch-exhaustiveness-check (we switch over merge-extensible unions
everywhere), no-unnecessary-condition, require-await, and
no-explicit-any. Style (@stylistic): 2-space, no semicolons, single
quotes, trailing commas, max-len 140 — the existing house style, now
enforced instead of drifting between agents. vendor/ is excluded
(vendored source keeps upstream style); tests relax the rules that
fight test ergonomics (non-null assertions after expects, async mock
signatures, non-Error throws).
Code adjusted to pass: registry disposers wrap ctx.effect's
promise-returning disposer behind a sync () => void (our public API),
BlockAssembler gains an invariant-checking mustGet instead of non-null
assertions, lastTurnNumber uses findLast, waterfall tails return
Promise.resolve instead of async-without-await arrows, and the two
deliberate suppressions (non-exhaustive derivation switch, unbound
execute pass-through) carry justification comments.
yarn lint / yarn lint:fix added.
2026-06-11 14:17:58 +08:00
|
|
|
agent.session.events.map(e => e.type))
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
})
|