Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
import { describe, expect, it } from 'vitest'
|
build(vendor): rescope the vendored Cordis packages into @deepseek-ai
Machine-produced by `pnpm run rescope-vendor --apply` plus the regeneration it
prints: `pnpm install` for the lockfile, `pnpm run gen-third-party-notices`,
`verify-translation-pairing --write` for the touched bilingual pairs,
`gen-doc-graphs`, and one typert snapshot whose ids embed character offsets.
`pnpm run rescope-vendor --check` verifies the result.
Renames nine vendored packages (cordis, cosmokit, schemastery and the six
@cordisjs plugins) and every reference that resolves them: manifest names and
dependency keys, module specifiers including declare-module merges, cordis.yml
plugin names, tsconfig paths, every Markdown fence, and `docs/` prose.
Directory names, upstream versions, and dependency ranges are unchanged, so
vendor/README.md still reads as an upstream snapshot; its manifest table gains
an upstream-name column so THIRD_PARTY_NOTICES keeps MIT attribution pointed
at each fork's origin.
The tutorial tier follows the rename end to end: its yaml fences named plugins
the Loader can no longer resolve, its `ts ignore-check` fences disagreed with
the compiled fences beside them, and its prose quoted both. The contracts that
told readers to keep upstream names — the root convention and the vendoring
cookbook's tree comment and manifest invariant — now say to rescope instead.
Two rules read `@deepseek-ai/` as "another workspace plugin": the client bundle
purity gate now names the vendored libraries a browser bundle inlines, and the
files where a bare `cordis` is an agent-preset id keep that product data.
2026-08-10 22:04:06 +08:00
|
|
|
import { Context } from '@deepseek-ai/cordis'
|
2026-08-24 18:23:42 +08:00
|
|
|
import LlmRuntime, { createUserMessage, CallId, LlmError, ReasoningEffortId, StreamChunk } from '@deepseek-ai/dsh-llm'
|
2026-07-24 11:46:06 +08:00
|
|
|
import SessionStore, { SessionId, TurnEndReason } from '@deepseek-ai/dsh-session'
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
import SystemPrompt from '@deepseek-ai/dsh-system-prompt'
|
2026-08-13 00:36:22 +08:00
|
|
|
import ToolRuntime, { defineContentToolFixture } from '@deepseek-ai/dsh-tools'
|
2026-07-14 02:32:35 +08:00
|
|
|
import AgentRegistry, { type Agent } from '@deepseek-ai/dsh-agent'
|
2026-07-14 01:59:21 +08:00
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
import AgentLoop from '@deepseek-ai/dsh-agent-loop'
|
2026-06-15 23:53:47 +08:00
|
|
|
import { MockAdapter, maxTokensResponse, textResponse, toolCallResponse } from './mock-adapter.ts'
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
function driverDone(agent: Agent): Promise<void> {
|
|
|
|
|
return (agent as Agent & { done: Promise<void> }).done
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-05 23:23:46 +08:00
|
|
|
async function harness(adapter: MockAdapter, persona = '') {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const ctx = new Context()
|
2026-08-13 00:36:22 +08:00
|
|
|
await ctx.plugin(LlmRuntime)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
await ctx.plugin(SessionStore)
|
2026-07-05 23:23:46 +08:00
|
|
|
await ctx.plugin(SystemPrompt, { persona })
|
2026-08-13 00:36:22 +08:00
|
|
|
await ctx.plugin(ToolRuntime)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
await ctx.plugin(AgentRegistry)
|
|
|
|
|
await ctx.plugin(AgentLoop, { agents: [] })
|
|
|
|
|
ctx.llm.registerAdapter(['mock'], adapter)
|
|
|
|
|
return ctx
|
|
|
|
|
}
|
|
|
|
|
|
2026-08-04 21:02:28 +08:00
|
|
|
/** Wait for the agent's next transition to idle after a waking send. */
|
2026-07-14 02:32:35 +08:00
|
|
|
function waitForIdle(ctx: Context, agent: Agent): Promise<void> {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
return new Promise((resolve) => {
|
2026-08-06 12:13:14 +08:00
|
|
|
const dispose = ctx.on('agent/status', ({ agent: subject, status }) => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
if (subject === agent && status === 'idle') {
|
|
|
|
|
dispose()
|
|
|
|
|
resolve()
|
|
|
|
|
}
|
|
|
|
|
})
|
|
|
|
|
})
|
|
|
|
|
}
|
|
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
function send(agent: Agent, text: string) {
|
2026-07-28 13:55:59 +08:00
|
|
|
agent.followup(createUserMessage({ content: [{ type: 'text', text }], source: { kind: 'user' } }))
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
}
|
|
|
|
|
|
2026-08-07 14:54:31 +08:00
|
|
|
/** All user-message texts recorded in the log (to assert what actually ran). */
|
|
|
|
|
function userTexts(agent: Agent): string[] {
|
|
|
|
|
return agent.session.events
|
|
|
|
|
.filter(e => e.type === 'user/message')
|
|
|
|
|
.flatMap(e => e.type === 'user/message' ? e.data.content : [])
|
|
|
|
|
.flatMap(b => b.type === 'text' ? [b.text] : [])
|
|
|
|
|
}
|
|
|
|
|
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
describe('agent loop', () => {
|
2026-07-28 17:36:44 +08:00
|
|
|
it.each([0, -1, 1.5, Number.NaN, Number.MAX_SAFE_INTEGER + 1])(
|
|
|
|
|
'rejects invalid AgentOptions.maxTokens %s before publication',
|
|
|
|
|
async (maxTokens) => {
|
|
|
|
|
const ctx = await harness(new MockAdapter([]))
|
|
|
|
|
expect(() => ctx.agentLoop.create(
|
|
|
|
|
SessionId('invalid-max-tokens'),
|
|
|
|
|
{ provider: 'mock', model: 'mock', maxTokens },
|
|
|
|
|
)).toThrow('agent maxTokens must be a positive safe integer')
|
|
|
|
|
expect(ctx.agents.list()).toEqual([])
|
|
|
|
|
expect(ctx.sessions.list()).toEqual([])
|
|
|
|
|
},
|
|
|
|
|
)
|
|
|
|
|
|
2026-07-30 00:55:02 +08:00
|
|
|
it('seeds a valid AgentOptions.maxTokens into the first model request', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('bounded')])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
const agent = ctx.agentLoop.create(
|
|
|
|
|
SessionId('valid-max-tokens'),
|
|
|
|
|
{ provider: 'mock', model: 'mock', maxTokens: 256 },
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
send(agent, 'use the configured output limit')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests[0]?.maxTokens).toBe(256)
|
|
|
|
|
})
|
|
|
|
|
|
2026-08-24 18:23:42 +08:00
|
|
|
it('seeds an AgentOptions reasoning effort into the first model request', async () => {
|
|
|
|
|
const effort = ReasoningEffortId('high')
|
|
|
|
|
const adapter = new MockAdapter([textResponse('reasoned')], {
|
|
|
|
|
efforts: [{ id: effort, name: 'High' }],
|
|
|
|
|
defaultEffort: effort,
|
|
|
|
|
})
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
const agent = ctx.agentLoop.create(
|
|
|
|
|
SessionId('configured-reasoning-effort'),
|
|
|
|
|
{ provider: 'mock', model: 'mock', reasoningEffort: effort },
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
send(agent, 'use the configured reasoning effort')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests[0]?.reasoningEffort).toBe(effort)
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('validates reasoning effort in declarative agent config', () => {
|
|
|
|
|
const effort = ReasoningEffortId('high')
|
|
|
|
|
expect(AgentLoop.Config({
|
|
|
|
|
agents: [{ id: 'configured-agent', reasoningEffort: effort }],
|
|
|
|
|
}).agents[0]?.reasoningEffort).toBe(effort)
|
|
|
|
|
expect(() => AgentLoop.Config({
|
|
|
|
|
agents: [{ id: 'configured-agent', reasoningEffort: ReasoningEffortId('') }],
|
|
|
|
|
})).toThrow()
|
|
|
|
|
})
|
|
|
|
|
|
2026-08-03 12:25:33 +08:00
|
|
|
it('cancels queued wakeup work together with an active maintenance task', async () => {
|
2026-08-07 14:54:31 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('park reply')])
|
2026-08-03 12:25:33 +08:00
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
const agent = ctx.agentLoop.create(SessionId('cancel-maintenance-wakeup'), {
|
|
|
|
|
provider: 'mock',
|
|
|
|
|
model: 'mock',
|
|
|
|
|
})
|
|
|
|
|
const started = Promise.withResolvers<undefined>()
|
|
|
|
|
const maintenance = agent.runMaintenance(async (signal) => {
|
|
|
|
|
started.resolve(undefined)
|
|
|
|
|
await new Promise<void>((_resolve, reject) => {
|
|
|
|
|
signal.addEventListener('abort', () => {
|
|
|
|
|
reject(new Error('maintenance aborted', { cause: signal.reason }))
|
|
|
|
|
}, { once: true })
|
|
|
|
|
})
|
|
|
|
|
})
|
|
|
|
|
await started.promise
|
|
|
|
|
|
2026-08-07 14:54:31 +08:00
|
|
|
send(agent, 'discard this wakeup') // latched behind the live maintenance task
|
|
|
|
|
agent.cancel({ kind: 'user' }) // drops the queue and the latch, aborts maintenance
|
|
|
|
|
send(agent, 'park after cancellation') // newer intent: re-latched, replays at convergence
|
2026-08-03 12:25:33 +08:00
|
|
|
|
|
|
|
|
await expect(maintenance).rejects.toThrow('maintenance aborted')
|
|
|
|
|
await agent.whenIdle()
|
2026-08-07 14:54:31 +08:00
|
|
|
|
|
|
|
|
// The pre-cancel wakeup is gone; the post-cancel wake replays at convergence.
|
|
|
|
|
expect(userTexts(agent)).toEqual(['park after cancellation'])
|
|
|
|
|
expect(agent.inbox.nextTurn).toHaveLength(0)
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('replays a wake latched behind maintenance at convergence', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('wake reply')])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
const agent = ctx.agentLoop.create(SessionId('maintenance-wake-replay'), {
|
|
|
|
|
provider: 'mock',
|
|
|
|
|
model: 'mock',
|
|
|
|
|
})
|
|
|
|
|
const started = Promise.withResolvers<undefined>()
|
|
|
|
|
const finish = Promise.withResolvers<undefined>()
|
|
|
|
|
const maintenance = agent.runMaintenance(async () => {
|
|
|
|
|
started.resolve(undefined)
|
|
|
|
|
await finish.promise
|
|
|
|
|
})
|
|
|
|
|
await started.promise
|
|
|
|
|
|
|
|
|
|
send(agent, 'wake behind maintenance')
|
|
|
|
|
finish.resolve(undefined)
|
|
|
|
|
await maintenance
|
|
|
|
|
await agent.whenIdle()
|
|
|
|
|
|
|
|
|
|
expect(userTexts(agent)).toEqual(['wake behind maintenance'])
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('suppresses the replay when a latched maintenance wake is removed', async () => {
|
|
|
|
|
const adapter = new MockAdapter([])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
const agent = ctx.agentLoop.create(SessionId('maintenance-wake-removed'), {
|
|
|
|
|
provider: 'mock',
|
|
|
|
|
model: 'mock',
|
|
|
|
|
})
|
|
|
|
|
const started = Promise.withResolvers<undefined>()
|
|
|
|
|
const finish = Promise.withResolvers<undefined>()
|
|
|
|
|
const maintenance = agent.runMaintenance(async () => {
|
|
|
|
|
started.resolve(undefined)
|
|
|
|
|
await finish.promise
|
|
|
|
|
})
|
|
|
|
|
await started.promise
|
|
|
|
|
|
|
|
|
|
const wake = createUserMessage({ content: [{ type: 'text', text: 'removed wake' }], source: { kind: 'user' } })
|
|
|
|
|
agent.followup(wake)
|
|
|
|
|
agent.inbox.remove(wake.id)
|
|
|
|
|
finish.resolve(undefined)
|
|
|
|
|
await maintenance
|
|
|
|
|
await agent.whenIdle()
|
|
|
|
|
|
|
|
|
|
expect(userTexts(agent)).toEqual([])
|
2026-08-03 12:25:33 +08:00
|
|
|
expect(adapter.requests).toEqual([])
|
2026-08-07 14:54:31 +08:00
|
|
|
expect(agent.session.events.filter(e => e.type === 'turn/start')).toHaveLength(0)
|
2026-08-03 12:25:33 +08:00
|
|
|
})
|
|
|
|
|
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
it('runs a simple turn: queued message → model → idle, with ordered events', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('hello there')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
2026-07-02 03:26:45 +08:00
|
|
|
// All boundaries — turn and step — are durable session events on the
|
|
|
|
|
// session/event feed (no agent/* mirror). Record them in fire order to
|
|
|
|
|
// assert the full boundary nesting.
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const order: string[] = []
|
2026-06-30 10:32:55 +08:00
|
|
|
ctx.on('session/event', (_session, event) => {
|
2026-07-02 03:26:45 +08:00
|
|
|
if (event.type === 'turn/start' || event.type === 'step/start' || event.type === 'step/end' || event.type === 'turn/end') {
|
|
|
|
|
order.push(event.type)
|
|
|
|
|
}
|
2026-06-30 10:32:55 +08:00
|
|
|
})
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-02 03:26:45 +08:00
|
|
|
expect(order).toEqual(['turn/start', 'step/start', 'step/end', 'turn/end'])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
const types = agent.session.events.map(e => e.type)
|
2026-07-31 19:21:16 +08:00
|
|
|
// Durable inbox receipt precedes the turn-owned transcript.
|
2026-07-30 17:28:03 +08:00
|
|
|
expect(types[0]).toBe('agent/inbox/spliced')
|
|
|
|
|
expect(types).toContain('turn/start')
|
|
|
|
|
expect(types).toContain('user/message')
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(types).toContain('assistant/message')
|
2026-06-21 10:00:06 +08:00
|
|
|
const assistantMessage = agent.session.events.find(e => e.type === 'assistant/message')
|
|
|
|
|
expect(assistantMessage?.type === 'assistant/message' && assistantMessage.data.usage).toEqual({ inputTokens: 10, outputTokens: 'hello there'.length })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(types.at(-1)).toBe('turn/end')
|
|
|
|
|
|
|
|
|
|
// derived history: user + assistant
|
|
|
|
|
const messages = agent.session.deriveMessages()
|
|
|
|
|
expect(messages.map(m => m.role)).toEqual(['user', 'assistant'])
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
expect(messages[1]!.content).toEqual([{ type: 'text', text: 'hello there' }])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('round-trips tool calls: model requests tool → executes → result in next request', async () => {
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'echo', { text: 'ping' }, 'calling echo'),
|
|
|
|
|
textResponse('done'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-24 21:18:48 +08:00
|
|
|
ctx.tools.register(defineContentToolFixture({
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
name: 'echo',
|
|
|
|
|
description: 'echo back',
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
parameters: { text: { type: 'string' } },
|
|
|
|
|
async execute(args) {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
return [{ type: 'text', text: `echo: ${args.text}` }]
|
|
|
|
|
},
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'use the tool')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
// two model calls happened (tool-call step, then final step)
|
|
|
|
|
expect(adapter.requests).toHaveLength(2)
|
|
|
|
|
|
|
|
|
|
// the second request's derived history contains the tool result
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
const secondMessages = adapter.requests[1]!.messages
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const toolResultMessage = secondMessages.find(m =>
|
|
|
|
|
m.content.some(b => b.type === 'tool-result'))
|
|
|
|
|
expect(toolResultMessage).toBeDefined()
|
|
|
|
|
const block = toolResultMessage!.content.find(b => b.type === 'tool-result')!
|
|
|
|
|
expect(block).toMatchObject({ toolCallId: 'c1', isError: false })
|
Add ESLint: typescript-eslint strict-type-checked + stylistic formatting
Flat config with two layers. Correctness (type-checked): the headline
rules for this codebase are no-floating-promises / no-misused-promises
(a lost promise in the agent loop is our primary bug class),
switch-exhaustiveness-check (we switch over merge-extensible unions
everywhere), no-unnecessary-condition, require-await, and
no-explicit-any. Style (@stylistic): 2-space, no semicolons, single
quotes, trailing commas, max-len 140 — the existing house style, now
enforced instead of drifting between agents. vendor/ is excluded
(vendored source keeps upstream style); tests relax the rules that
fight test ergonomics (non-null assertions after expects, async mock
signatures, non-Error throws).
Code adjusted to pass: registry disposers wrap ctx.effect's
promise-returning disposer behind a sync () => void (our public API),
BlockAssembler gains an invariant-checking mustGet instead of non-null
assertions, lastTurnNumber uses findLast, waterfall tails return
Promise.resolve instead of async-without-await arrows, and the two
deliberate suppressions (non-exhaustive derivation switch, unbound
execute pass-through) carry justification comments.
yarn lint / yarn lint:fix added.
2026-06-11 14:17:58 +08:00
|
|
|
expect((block).content).toEqual([{ type: 'text', text: 'echo: ping' }])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
// session log records call + result
|
|
|
|
|
const types = agent.session.events.map(e => e.type)
|
|
|
|
|
expect(types).toContain('tool/call')
|
|
|
|
|
expect(types).toContain('tool/result')
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-05 11:37:32 +08:00
|
|
|
it('renders harness identity, then the persona, then tool guidance — with {{variables}} resolved', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
2026-07-05 23:23:46 +08:00
|
|
|
// The persona is a TEMPLATE: {{model}} is the loop-registered variable
|
|
|
|
|
// projecting this agent's configured model, so the model knows its own name.
|
|
|
|
|
const ctx = await harness(adapter, 'You are a test agent on {{model}}.')
|
2026-07-05 01:54:46 +08:00
|
|
|
ctx.systemPrompt.section({ name: 'tool:noop', order: 100, text: 'Use the noop tool wisely.' })
|
2026-07-24 21:18:48 +08:00
|
|
|
ctx.tools.register(defineContentToolFixture({
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
name: 'noop',
|
|
|
|
|
description: 'does nothing',
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
parameters: {},
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
async execute() {
|
|
|
|
|
return []
|
|
|
|
|
},
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
const request = adapter.requests[0]
|
2026-08-13 00:36:22 +08:00
|
|
|
expect(request!.system).toBe('You are an AI agent powered by DeepSeek Harness.\n\nYou are a test agent on mock.\n\nUse the noop tool wisely.')
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
expect(request!.tools?.map(t => t.name)).toEqual(['noop'])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-05 01:54:46 +08:00
|
|
|
it('resolves {{cwd}} from the agent session workspace (factory create with meta.cwd)', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
2026-07-05 23:23:46 +08:00
|
|
|
const ctx = await harness(adapter, 'Working in {{cwd}}.')
|
2026-07-11 22:55:26 +08:00
|
|
|
const handle = await ctx.agents.create({
|
2026-07-05 01:54:46 +08:00
|
|
|
sessionId: SessionId('s-cwd'),
|
|
|
|
|
meta: { cwd: '/work/space' },
|
2026-07-14 21:57:52 +08:00
|
|
|
agentOptions: { provider: 'mock', model: 'mock' },
|
2026-07-05 01:54:46 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
const agent = handle.agent
|
2026-07-05 01:54:46 +08:00
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-08-13 00:36:22 +08:00
|
|
|
expect(adapter.requests[0]!.system).toBe('You are an AI agent powered by DeepSeek Harness.\n\nWorking in /work/space.')
|
2026-07-05 01:54:46 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-05 02:42:48 +08:00
|
|
|
it('contains a strict-variable render failure: the turn errors, the loop keeps serving turns', async () => {
|
2026-07-12 03:36:43 +08:00
|
|
|
// A missing cwd variable must fail one turn without preventing a later valid turn.
|
2026-07-05 02:42:48 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok after rescue')])
|
2026-07-05 23:23:46 +08:00
|
|
|
const ctx = await harness(adapter, 'In {{cwd}}.')
|
2026-07-05 01:54:46 +08:00
|
|
|
const errors: Error[] = []
|
2026-08-06 12:13:14 +08:00
|
|
|
ctx.on('agent/error', ({ error }) => {
|
2026-07-24 21:18:48 +08:00
|
|
|
if (error instanceof Error) errors.push(error)
|
|
|
|
|
})
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-05 01:54:46 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests).toHaveLength(0) // the request was never sent
|
2026-07-31 13:54:36 +08:00
|
|
|
expect(errors.map(error => error.message)).toEqual([
|
|
|
|
|
'prompt variable "{{cwd}}" has no value for this assembly (section "deployment:persona")',
|
|
|
|
|
])
|
2026-07-05 01:54:46 +08:00
|
|
|
const turnEnd = agent.session.events.find(e => e.type === 'turn/end')
|
|
|
|
|
expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason.kind).toBe('error')
|
2026-07-30 17:28:03 +08:00
|
|
|
expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason.kind === 'error'
|
2026-08-03 16:33:23 +08:00
|
|
|
? turnEnd.data.reason.error.message
|
2026-07-30 17:28:03 +08:00
|
|
|
: '').toContain('no value for this assembly')
|
2026-07-05 02:42:48 +08:00
|
|
|
|
|
|
|
|
// The loop survived: a waterfall listener rescues {{cwd}} and the SAME
|
|
|
|
|
// agent completes a real model turn.
|
|
|
|
|
ctx.on('system-prompt/assemble', async (assembly, _context, next) => {
|
|
|
|
|
assembly.variables['cwd'] = '/rescued'
|
|
|
|
|
return next()
|
|
|
|
|
})
|
|
|
|
|
send(agent, 'again')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
2026-08-13 00:36:22 +08:00
|
|
|
expect(adapter.requests[0]!.system).toBe('You are an AI agent powered by DeepSeek Harness.\n\nIn /rescued.')
|
2026-07-05 02:42:48 +08:00
|
|
|
const turnEnds = agent.session.events.filter(e => e.type === 'turn/end')
|
|
|
|
|
expect(turnEnds).toHaveLength(2)
|
|
|
|
|
expect(turnEnds[1]?.type === 'turn/end' && turnEnds[1].data.reason.kind).toBe('completed')
|
2026-07-05 01:54:46 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-06 00:34:47 +08:00
|
|
|
it('supports the model-via-agent/request path with a {{model}} persona: the supplier states it via the assemble waterfall', async () => {
|
|
|
|
|
// AgentOptions.model unset: the model arrives in the agent/request
|
|
|
|
|
// waterfall (the loop's documented fallback — see runStep's no-model
|
|
|
|
|
// error). {{model}} renders BEFORE that waterfall, so the SAME plugin
|
|
|
|
|
// states the fact early on system-prompt/assemble — the owner of a
|
|
|
|
|
// late-bound fact owns stating it wherever it is claimed.
|
|
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter, 'You run on {{model}}.')
|
|
|
|
|
ctx.on('system-prompt/assemble', async (assembly, _context, next) => {
|
2026-07-14 21:57:52 +08:00
|
|
|
assembly.variables['provider'] = 'mock'
|
2026-07-06 00:34:47 +08:00
|
|
|
assembly.variables['model'] = 'mock'
|
|
|
|
|
return next()
|
|
|
|
|
})
|
2026-08-06 12:13:14 +08:00
|
|
|
ctx.on('agent/request', async (_payload, next) => {
|
2026-07-24 21:18:48 +08:00
|
|
|
const config = await next()
|
2026-07-14 21:57:52 +08:00
|
|
|
return { ...config, provider: 'mock', model: 'mock' }
|
2026-07-06 00:34:47 +08:00
|
|
|
})
|
2026-07-14 01:59:21 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a-late-model'), {})
|
2026-07-06 00:34:47 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
expect(adapter.requests[0]!.model).toBe('mock')
|
2026-08-13 00:36:22 +08:00
|
|
|
expect(adapter.requests[0]!.system).toBe('You are an AI agent powered by DeepSeek Harness.\n\nYou run on mock.')
|
2026-07-06 00:34:47 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-22 18:02:26 +08:00
|
|
|
it('omits the system field when system-prompt/assemble short-circuits with an empty assembly', async () => {
|
2026-07-05 23:23:46 +08:00
|
|
|
// The documented escape valve: a deployment that must drop the harness
|
|
|
|
|
// openers short-circuits the assemble waterfall; the request then carries
|
|
|
|
|
// NO system field at all (not an empty string).
|
|
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-30 22:09:15 +08:00
|
|
|
ctx.on('system-prompt/assemble', async () => ({ sections: [], contexts: [], tools: [], variables: {} }))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a-no-system'), { provider: 'mock', model: 'mock' })
|
2026-07-05 23:23:46 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
expect('system' in adapter.requests[0]!).toBe(false)
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-30 22:09:15 +08:00
|
|
|
it('materializes changed runtime context at the history tail without rewriting the system header', async () => {
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
textResponse('one'),
|
|
|
|
|
textResponse('two'),
|
|
|
|
|
textResponse('three'),
|
|
|
|
|
textResponse('four'),
|
|
|
|
|
textResponse('five'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
let mode = 'read-only'
|
|
|
|
|
const dispose = ctx.systemPrompt.context({ name: 'policy', order: 0, text: () => `Mode: ${mode}.` })
|
|
|
|
|
const agent = ctx.agentLoop.create(SessionId('a-runtime-context'), { provider: 'mock', model: 'mock' })
|
|
|
|
|
const contextEvents = () => agent.session.events.flatMap(event =>
|
|
|
|
|
event.type === 'user/message'
|
|
|
|
|
&& event.data.source.kind === 'plugin'
|
|
|
|
|
&& event.data.source.plugin === '@deepseek-ai/dsh-system-prompt'
|
|
|
|
|
? [event]
|
|
|
|
|
: [])
|
|
|
|
|
|
|
|
|
|
send(agent, 'first')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(contextEvents()).toHaveLength(1)
|
|
|
|
|
expect(contextEvents()[0]?.data.content).toEqual([{
|
|
|
|
|
type: 'text',
|
|
|
|
|
text: 'Current runtime context. This snapshot supersedes earlier runtime-context snapshots.\n\nMode: read-only.',
|
|
|
|
|
}])
|
|
|
|
|
|
|
|
|
|
send(agent, 'unchanged')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(contextEvents()).toHaveLength(1)
|
|
|
|
|
|
|
|
|
|
mode = 'danger-full-access'
|
|
|
|
|
send(agent, 'changed')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(contextEvents()).toHaveLength(2)
|
|
|
|
|
const changedBlock = contextEvents()[1]?.data.content[0]
|
|
|
|
|
expect(changedBlock?.type).toBe('text')
|
|
|
|
|
if (changedBlock?.type !== 'text') throw new Error('changed runtime context is not text')
|
|
|
|
|
expect(changedBlock.text).toContain('danger-full-access')
|
|
|
|
|
|
|
|
|
|
dispose()
|
|
|
|
|
send(agent, 'cleared')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(contextEvents()).toHaveLength(3)
|
|
|
|
|
expect(contextEvents()[2]?.data.content).toEqual([{
|
|
|
|
|
type: 'text',
|
|
|
|
|
text: 'Current runtime context: none. Earlier runtime-context snapshots no longer apply.',
|
|
|
|
|
}])
|
|
|
|
|
|
|
|
|
|
send(agent, 'still clear')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(contextEvents()).toHaveLength(3)
|
|
|
|
|
expect(adapter.requests.map(request => request.system)).toEqual(Array(5).fill(adapter.requests[0]?.system))
|
|
|
|
|
expect(agent.session.events.filter(event => event.type === 'request/header')).toHaveLength(1)
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('re-emits unchanged runtime context when a surface replacement removed the retained snapshot', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('one'), textResponse('two')])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
ctx.systemPrompt.context({ name: 'policy', order: 0, text: 'Mode: read-only.' })
|
|
|
|
|
const agent = ctx.agentLoop.create(SessionId('a-runtime-context-compacted'), { provider: 'mock', model: 'mock' })
|
|
|
|
|
|
|
|
|
|
send(agent, 'first')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
const contextEvent = agent.session.events.find(event =>
|
|
|
|
|
event.type === 'user/message'
|
|
|
|
|
&& event.data.source.kind === 'plugin'
|
|
|
|
|
&& event.data.source.plugin === '@deepseek-ai/dsh-system-prompt')
|
|
|
|
|
if (contextEvent?.type !== 'user/message') throw new Error('first turn did not materialize runtime context')
|
|
|
|
|
agent.session.append('user/message', createUserMessage({
|
|
|
|
|
content: [{ type: 'text', text: 'compacted summary' }],
|
|
|
|
|
source: { kind: 'plugin', plugin: 'test-compaction' },
|
|
|
|
|
}), {
|
|
|
|
|
surfaceOp: { op: 'replace', start: contextEvent.seq, end: contextEvent.seq },
|
|
|
|
|
sourceEventSeqs: [contextEvent.seq],
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'after compaction')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
const runtimeContexts = agent.session.events.flatMap(event =>
|
|
|
|
|
event.type === 'user/message'
|
|
|
|
|
&& event.data.source.kind === 'plugin'
|
|
|
|
|
&& event.data.source.plugin === '@deepseek-ai/dsh-system-prompt'
|
|
|
|
|
? [event]
|
|
|
|
|
: [])
|
|
|
|
|
expect(runtimeContexts).toHaveLength(2)
|
|
|
|
|
expect(adapter.requests[1]?.messages.some(message =>
|
|
|
|
|
message.source.kind === 'plugin'
|
|
|
|
|
&& message.source.plugin === '@deepseek-ai/dsh-system-prompt')).toBe(true)
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-31 13:08:45 +08:00
|
|
|
it('clears compacted runtime context after the active set becomes empty', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('one'), textResponse('two')])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
const dispose = ctx.systemPrompt.context({ name: 'policy', order: 0, text: 'Mode: read-only.' })
|
|
|
|
|
const agent = ctx.agentLoop.create(SessionId('a-runtime-context-compacted-clear'), { provider: 'mock', model: 'mock' })
|
|
|
|
|
|
|
|
|
|
send(agent, 'first')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
const contextEvent = agent.session.events.find(event =>
|
|
|
|
|
event.type === 'user/message'
|
|
|
|
|
&& event.data.source.kind === 'plugin'
|
|
|
|
|
&& event.data.source.plugin === '@deepseek-ai/dsh-system-prompt')
|
|
|
|
|
if (contextEvent?.type !== 'user/message') throw new Error('first turn did not materialize runtime context')
|
|
|
|
|
agent.session.append('user/message', createUserMessage({
|
|
|
|
|
content: [{ type: 'text', text: 'summary retaining old mode: read-only' }],
|
|
|
|
|
source: { kind: 'plugin', plugin: 'test-compaction' },
|
|
|
|
|
}), {
|
|
|
|
|
surfaceOp: { op: 'replace', start: contextEvent.seq, end: contextEvent.seq },
|
|
|
|
|
sourceEventSeqs: [contextEvent.seq],
|
|
|
|
|
})
|
|
|
|
|
dispose()
|
|
|
|
|
|
|
|
|
|
send(agent, 'after compaction')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
const clearing = adapter.requests[1]?.messages.find(message =>
|
|
|
|
|
message.source.kind === 'plugin'
|
|
|
|
|
&& message.source.plugin === '@deepseek-ai/dsh-system-prompt')
|
|
|
|
|
expect(clearing?.content).toEqual([{
|
|
|
|
|
type: 'text',
|
|
|
|
|
text: 'Current runtime context: none. Earlier runtime-context snapshots no longer apply.',
|
|
|
|
|
}])
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('does not clear runtime context after an unrelated replacement', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
const agent = ctx.agentLoop.create(SessionId('a-runtime-context-unrelated-compaction'), { provider: 'mock', model: 'mock' })
|
|
|
|
|
const original = agent.session.append('user/message', createUserMessage({
|
|
|
|
|
content: [{ type: 'text', text: 'old context' }],
|
|
|
|
|
source: { kind: 'plugin', plugin: 'test-context' },
|
|
|
|
|
}), { surfaceOp: 'append' })
|
|
|
|
|
agent.session.append('user/message', createUserMessage({
|
|
|
|
|
content: [{ type: 'text', text: 'compacted summary' }],
|
|
|
|
|
source: { kind: 'plugin', plugin: 'test-compaction' },
|
|
|
|
|
}), {
|
|
|
|
|
surfaceOp: { op: 'replace', start: original.seq, end: original.seq },
|
|
|
|
|
sourceEventSeqs: [original.seq],
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'after compaction')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(adapter.requests[0]?.messages.some(message =>
|
|
|
|
|
message.source.kind === 'plugin'
|
|
|
|
|
&& message.source.plugin === '@deepseek-ai/dsh-system-prompt')).toBe(false)
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-30 22:09:15 +08:00
|
|
|
it('replaces a malformed retained runtime-context message with the current complete snapshot', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
ctx.systemPrompt.context({ name: 'policy', order: 0, text: 'Mode: read-only.' })
|
|
|
|
|
const agent = ctx.agentLoop.create(SessionId('a-runtime-context-malformed'), { provider: 'mock', model: 'mock' })
|
|
|
|
|
agent.session.append('user/message', createUserMessage({
|
|
|
|
|
content: [{ type: 'text', text: 'broken' }, { type: 'text', text: 'snapshot' }],
|
|
|
|
|
source: { kind: 'plugin', plugin: '@deepseek-ai/dsh-system-prompt' },
|
|
|
|
|
}), { surfaceOp: 'append' })
|
|
|
|
|
|
|
|
|
|
send(agent, 'repair context')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
const runtimeContexts = agent.session.events.flatMap(event =>
|
|
|
|
|
event.type === 'user/message'
|
|
|
|
|
&& event.data.source.kind === 'plugin'
|
|
|
|
|
&& event.data.source.plugin === '@deepseek-ai/dsh-system-prompt'
|
|
|
|
|
? [event]
|
|
|
|
|
: [])
|
|
|
|
|
expect(runtimeContexts).toHaveLength(2)
|
|
|
|
|
expect(runtimeContexts[1]?.data.content).toEqual([{
|
|
|
|
|
type: 'text',
|
|
|
|
|
text: 'Current runtime context. This snapshot supersedes earlier runtime-context snapshots.\n\nMode: read-only.',
|
|
|
|
|
}])
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-02 23:42:16 +08:00
|
|
|
it('records raw chunks for replay as assistant/chunk session events', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('abc')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
const chunkEvents = agent.session.events.filter(e => e.type === 'assistant/chunk')
|
|
|
|
|
// textResponse('abc') = block-start + 3 deltas + block-end + usage + finish = 7
|
|
|
|
|
expect(chunkEvents).toHaveLength(7)
|
|
|
|
|
// replay: chunk events alone re-assemble to the recorded assistant message
|
|
|
|
|
const deltaText = chunkEvents
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
.flatMap(e => e.type === 'assistant/chunk' ? [e.data.chunk] : [])
|
|
|
|
|
.filter((c: StreamChunk): c is Extract<StreamChunk, { type: 'text-delta' }> => c.type === 'text-delta')
|
|
|
|
|
.map(c => c.text)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
.join('')
|
|
|
|
|
expect(deltaText).toBe('abc')
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('injects steering between steps and continues the turn', async () => {
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'slow', {}),
|
|
|
|
|
textResponse('addressed the steering'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-24 21:18:48 +08:00
|
|
|
ctx.tools.register(defineContentToolFixture({
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
name: 'slow',
|
|
|
|
|
description: '',
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
parameters: {},
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
async execute() {
|
|
|
|
|
// steer while the turn is running (during tool execution)
|
2026-07-28 13:55:59 +08:00
|
|
|
agent.steer(createUserMessage({ content: [{ type: 'text', text: 'change of plans' }], source: { kind: 'user' } }))
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
return [{ type: 'text', text: 'tool done' }]
|
|
|
|
|
},
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
}))
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'start')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-30 17:28:03 +08:00
|
|
|
const steering = agent.session.events.find(e =>
|
|
|
|
|
e.type === 'user/message' && JSON.stringify(e.data.content).includes('change of plans'))
|
|
|
|
|
expect(steering).toBeDefined()
|
2026-07-31 19:40:59 +08:00
|
|
|
// The entered batch is appended after the second step opens and before its
|
|
|
|
|
// request derives history.
|
2026-07-30 17:28:03 +08:00
|
|
|
const steeringSeq = steering!.seq
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const secondStepStart = agent.session.events.filter(e => e.type === 'step/start')[1]
|
|
|
|
|
expect(secondStepStart).toBeDefined()
|
2026-07-31 19:40:59 +08:00
|
|
|
expect(steeringSeq).toBeGreaterThan(secondStepStart!.seq)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
// the second model request saw the steering content
|
|
|
|
|
const secondRequest = adapter.requests[1]
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
const flat = JSON.stringify(secondRequest!.messages)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(flat).toContain('change of plans')
|
|
|
|
|
})
|
|
|
|
|
|
2026-08-04 21:02:28 +08:00
|
|
|
it('starts idle steering synchronously and enters later steering at the next step', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('first'), textResponse('second')])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
2026-07-17 17:27:53 +08:00
|
|
|
const idle = waitForIdle(ctx, agent)
|
2026-07-28 13:55:59 +08:00
|
|
|
agent.steer(createUserMessage({ content: [{ type: 'text', text: 'first idle steer' }], source: { kind: 'user' } }))
|
2026-08-04 21:02:28 +08:00
|
|
|
expect(agent.status).toBe('running')
|
|
|
|
|
expect(agent.session.events.filter(event => event.type === 'turn/start')).toHaveLength(1)
|
2026-07-28 13:55:59 +08:00
|
|
|
agent.steer(createUserMessage({ content: [{ type: 'text', text: 'second idle steer' }], source: { kind: 'user' } }))
|
2026-07-17 17:27:53 +08:00
|
|
|
await idle
|
|
|
|
|
|
2026-07-30 17:28:03 +08:00
|
|
|
expect(agent.session.events.filter(event => event.type === 'turn/start')).toHaveLength(1)
|
2026-07-17 17:27:53 +08:00
|
|
|
expect(agent.session.events
|
|
|
|
|
.filter(event => event.type === 'user/message')
|
|
|
|
|
.map(event => event.data.content)).toEqual([
|
|
|
|
|
[{ type: 'text', text: 'first idle steer' }],
|
|
|
|
|
[{ type: 'text', text: 'second idle steer' }],
|
|
|
|
|
])
|
2026-08-04 21:02:28 +08:00
|
|
|
expect(adapter.requests).toHaveLength(2)
|
2026-07-27 18:58:27 +08:00
|
|
|
expect(JSON.stringify(adapter.requests[0]?.messages)).toContain('first idle steer')
|
2026-08-04 21:02:28 +08:00
|
|
|
expect(JSON.stringify(adapter.requests[0]?.messages)).not.toContain('second idle steer')
|
|
|
|
|
expect(JSON.stringify(adapter.requests[1]?.messages)).toContain('second idle steer')
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-31 19:21:16 +08:00
|
|
|
it('stops after a throwing pre-step listener and retains later steering until a wakeup', async () => {
|
2026-07-24 16:05:52 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('recovered')])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
const agent = ctx.agentLoop.create(SessionId('failed-steering'), { provider: 'mock', model: 'mock' })
|
|
|
|
|
let fail = true
|
2026-08-06 12:13:14 +08:00
|
|
|
ctx.on('agent/pre-step', ({ agent: subject }, next) => {
|
2026-07-31 19:21:16 +08:00
|
|
|
if (subject !== agent || !fail) return next()
|
2026-07-24 16:05:52 +08:00
|
|
|
fail = false
|
2026-07-28 13:55:59 +08:00
|
|
|
subject.steer(createUserMessage({ content: [{ type: 'text', text: 'pending steering' }], source: { kind: 'user' } }))
|
2026-07-31 19:21:16 +08:00
|
|
|
throw new Error('pre-step failed')
|
2026-07-24 16:05:52 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'prompt')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-31 13:54:36 +08:00
|
|
|
expect(adapter.requests).toHaveLength(0)
|
2026-08-04 14:09:52 +08:00
|
|
|
expect(agent.session.events.filter(event => event.type === 'turn/start')).toHaveLength(1)
|
|
|
|
|
expect(agent.session.events.filter(event => event.type === 'turn/end')).toHaveLength(1)
|
2026-07-31 13:54:36 +08:00
|
|
|
expect(agent.inbox.nextStep).toHaveLength(1)
|
|
|
|
|
|
|
|
|
|
send(agent, 'resume')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-24 16:05:52 +08:00
|
|
|
expect(adapter.requests).toHaveLength(1)
|
2026-08-04 14:09:52 +08:00
|
|
|
expect(agent.session.events.filter(event => event.type === 'turn/start')).toHaveLength(2)
|
2026-07-24 16:05:52 +08:00
|
|
|
expect(JSON.stringify(adapter.requests[0]?.messages)).toContain('pending steering')
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-30 17:28:03 +08:00
|
|
|
it('inject() while idle durably stages context without opening a turn', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
2026-07-28 13:55:59 +08:00
|
|
|
agent.inject(createUserMessage({ content: [{ type: 'text', text: 'file changed: a.ts' }], source: { kind: 'plugin', plugin: 'watcher' } }))
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(agent.status).toBe('idle')
|
|
|
|
|
expect(adapter.requests).toHaveLength(0)
|
2026-07-24 16:05:52 +08:00
|
|
|
expect(agent.session.events.filter(event => event.type === 'turn/start')).toHaveLength(0)
|
|
|
|
|
expect(agent.session.events.at(-1)).toMatchObject({
|
2026-07-30 17:28:03 +08:00
|
|
|
type: 'agent/inbox/spliced',
|
2026-07-28 13:55:59 +08:00
|
|
|
data: {
|
2026-07-30 17:28:03 +08:00
|
|
|
target: 'next-step',
|
|
|
|
|
inserted: [{
|
|
|
|
|
role: 'user',
|
|
|
|
|
content: [{ type: 'text', text: 'file changed: a.ts' }],
|
|
|
|
|
source: { kind: 'plugin', plugin: 'watcher' },
|
|
|
|
|
}],
|
2026-07-28 13:55:59 +08:00
|
|
|
},
|
2026-07-24 16:05:52 +08:00
|
|
|
})
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
2026-07-24 16:05:52 +08:00
|
|
|
expect(agent.session.events.filter(event => event.type === 'turn/start')).toHaveLength(1)
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
const flat = JSON.stringify(adapter.requests[0]!.messages)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(flat).toContain('file changed: a.ts')
|
2026-07-20 14:37:04 +08:00
|
|
|
expect(flat).not.toContain('<context source=')
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-24 14:05:33 +08:00
|
|
|
it('inject() persists structured context content verbatim with durable source', async () => {
|
2026-07-10 14:32:44 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('raw-context'), { provider: 'mock', model: 'mock' })
|
2026-07-10 14:32:44 +08:00
|
|
|
const text = '<system-reminder>Additional instructions from: pkg/AGENTS.md</system-reminder>'
|
2026-08-13 00:36:22 +08:00
|
|
|
agent.inject(createUserMessage({ content: [{ type: 'text', text }], source: { kind: 'plugin', plugin: 'agent-instructions' } }))
|
2026-07-10 14:32:44 +08:00
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-23 19:15:45 +08:00
|
|
|
const contextEvent = agent.session.events.find(event => event.type === 'user/message' && event.data.source.kind === 'plugin')
|
2026-07-24 14:05:33 +08:00
|
|
|
expect(contextEvent?.type === 'user/message' && contextEvent.data.source)
|
2026-08-13 00:36:22 +08:00
|
|
|
.toEqual({ kind: 'plugin', plugin: 'agent-instructions' })
|
2026-07-10 14:32:44 +08:00
|
|
|
const requestText = JSON.stringify(adapter.requests[0]!.messages)
|
|
|
|
|
expect(requestText).toContain('Additional instructions from: pkg/AGENTS.md')
|
|
|
|
|
expect(requestText).not.toContain('<context source=')
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-15 11:12:05 +08:00
|
|
|
it('defers inject() during tool execution until after the tool result', async () => {
|
2026-06-15 20:56:17 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'noticer', {}, 'calling'),
|
|
|
|
|
textResponse('done'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-15 11:12:05 +08:00
|
|
|
let visibleDuringTool = false
|
2026-07-24 21:18:48 +08:00
|
|
|
ctx.tools.register(defineContentToolFixture({
|
2026-06-15 20:56:17 +08:00
|
|
|
name: 'noticer',
|
|
|
|
|
description: 'injects a notice',
|
|
|
|
|
parameters: {},
|
|
|
|
|
async execute() {
|
2026-07-15 11:12:05 +08:00
|
|
|
await Promise.resolve()
|
|
|
|
|
const first = { type: 'text' as const, text: 'mid-turn notice' }
|
2026-07-28 13:55:59 +08:00
|
|
|
agent.inject(createUserMessage({ content: [first], source: { kind: 'plugin', plugin: 'x' } }))
|
2026-07-15 11:12:05 +08:00
|
|
|
first.text = 'mutated after inject'
|
2026-07-28 13:55:59 +08:00
|
|
|
agent.inject(createUserMessage({ content: [{ type: 'text', text: 'second notice' }], source: { kind: 'plugin', plugin: 'x' } }))
|
2026-07-23 19:15:45 +08:00
|
|
|
visibleDuringTool = agent.session.events.some(e => e.type === 'user/message' && e.data.source.kind === 'plugin')
|
2026-06-15 20:56:17 +08:00
|
|
|
return [{ type: 'text', text: 'ok' }]
|
|
|
|
|
},
|
|
|
|
|
}))
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-15 11:12:05 +08:00
|
|
|
expect(visibleDuringTool).toBe(false)
|
|
|
|
|
|
|
|
|
|
// The injection stays in the open turn, but its user-role context cannot
|
|
|
|
|
// split the assistant tool call from the provider's tool-result message.
|
2026-06-15 20:56:17 +08:00
|
|
|
const turnStarts = agent.session.events.filter(e => e.type === 'turn/start')
|
|
|
|
|
expect(turnStarts).toHaveLength(1)
|
2026-07-15 11:12:05 +08:00
|
|
|
const result = agent.session.events.find(e => e.type === 'tool/result')!
|
2026-07-23 19:15:45 +08:00
|
|
|
const contexts = agent.session.events.filter(e => e.type === 'user/message' && e.data.source.kind === 'plugin')
|
2026-07-15 11:12:05 +08:00
|
|
|
expect(contexts).toHaveLength(2)
|
|
|
|
|
expect(result.seq).toBeLessThan(contexts[0]!.seq)
|
2026-07-23 19:15:45 +08:00
|
|
|
expect(contexts.flatMap(event => event.type === 'user/message' ? event.data.content : []))
|
2026-07-15 11:12:05 +08:00
|
|
|
.toEqual([
|
2026-07-27 18:58:27 +08:00
|
|
|
{ type: 'text', text: 'mid-turn notice' },
|
2026-07-15 11:12:05 +08:00
|
|
|
{ type: 'text', text: 'second notice' },
|
|
|
|
|
])
|
|
|
|
|
|
|
|
|
|
const secondRequest = adapter.requests[1]!.messages
|
|
|
|
|
const resultIndex = secondRequest.findIndex(message =>
|
|
|
|
|
message.content.some(block => block.type === 'tool-result'))
|
|
|
|
|
const contextIndexes = secondRequest.flatMap((message, index) =>
|
|
|
|
|
message.content.some(block => block.type === 'text'
|
|
|
|
|
&& (block.text.includes('mid-turn notice') || block.text.includes('second notice')))
|
|
|
|
|
? [index]
|
|
|
|
|
: [])
|
|
|
|
|
expect(resultIndex).toBeGreaterThanOrEqual(0)
|
2026-07-27 18:58:27 +08:00
|
|
|
expect(contextIndexes).toHaveLength(2)
|
2026-07-15 11:12:05 +08:00
|
|
|
expect(contextIndexes.every(index => index > resultIndex)).toBe(true)
|
2026-06-15 20:56:17 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-18 13:39:12 +08:00
|
|
|
it('rejects non-JSON context before it enters the active tool-batch FIFO', async () => {
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'invalid-injector', {}, 'calling'),
|
|
|
|
|
textResponse('done'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 14:50:55 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('invalid-context'), { provider: 'mock', model: 'mock' })
|
2026-07-24 21:18:48 +08:00
|
|
|
ctx.tools.register(defineContentToolFixture({
|
2026-07-18 13:39:12 +08:00
|
|
|
name: 'invalid-injector',
|
|
|
|
|
description: 'attempts an invalid context injection',
|
|
|
|
|
parameters: {},
|
|
|
|
|
async execute() {
|
|
|
|
|
expect(() => {
|
2026-07-28 13:55:59 +08:00
|
|
|
agent.inject(createUserMessage({ content: [{ type: 'text', text: 'invalid' }], source: { kind: 'plugin', plugin: 'test', bigint: 1n } as never }))
|
2026-07-18 13:39:12 +08:00
|
|
|
}).toThrow('agent context must be losslessly JSON-serializable')
|
|
|
|
|
return [{ type: 'text', text: 'rejected invalid context' }]
|
|
|
|
|
},
|
|
|
|
|
}))
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-23 19:15:45 +08:00
|
|
|
expect(agent.session.events.some(event => event.type === 'user/message' && event.data.source.kind === 'plugin')).toBe(false)
|
2026-06-15 20:56:17 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-27 17:38:42 +08:00
|
|
|
it('agent/turn-stopping can steer another step (/loop pattern)', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
textResponse('step 1'),
|
|
|
|
|
textResponse('step 2'),
|
|
|
|
|
textResponse('step 3'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
let steps = 0
|
2026-06-30 10:32:55 +08:00
|
|
|
ctx.on('session/event', (_session, event) => { if (event.type === 'step/end') steps++ })
|
2026-08-06 12:13:14 +08:00
|
|
|
ctx.on('agent/turn-stopping', ({ agent: subject }) => {
|
2026-07-24 21:18:48 +08:00
|
|
|
if (steps < 3) {
|
2026-07-28 13:55:59 +08:00
|
|
|
subject.steer(createUserMessage({ content: [{ type: 'text', text: 'continue' }], source: { kind: 'plugin', plugin: 'loop-test' } }))
|
2026-07-24 21:18:48 +08:00
|
|
|
}
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(steps).toBe(3)
|
|
|
|
|
expect(adapter.requests).toHaveLength(3)
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-24 21:18:48 +08:00
|
|
|
it('a tool can conclude the turn despite owing a follow-up request', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([toolCallResponse('c1', 'echo', { text: 'x' })])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-24 21:18:48 +08:00
|
|
|
ctx.tools.register(defineContentToolFixture({
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
name: 'echo',
|
|
|
|
|
description: '',
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
parameters: { text: { type: 'string' } },
|
2026-07-24 21:18:48 +08:00
|
|
|
async execute(args, exec) {
|
|
|
|
|
exec.concludeTurn()
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
return [{ type: 'text', text: String(args.text) }]
|
|
|
|
|
},
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
// only one model call despite the tool call requesting a follow-up
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
2026-07-24 21:18:48 +08:00
|
|
|
// The tool still executes and durably records its result.
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(agent.session.events.some(e => e.type === 'tool/result')).toBe(true)
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-30 17:28:03 +08:00
|
|
|
it('continues for steering that arrived during a concluding tool step', async () => {
|
2026-07-26 11:47:25 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'finalize', {}),
|
|
|
|
|
textResponse('next turn reply'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
|
|
|
|
ctx.tools.register(defineContentToolFixture({
|
|
|
|
|
name: 'finalize',
|
|
|
|
|
description: '',
|
|
|
|
|
parameters: {},
|
|
|
|
|
async execute(_args, exec) {
|
|
|
|
|
// Steering lands while the concluding tool is still executing.
|
2026-07-28 13:55:59 +08:00
|
|
|
agent.steer(createUserMessage({ content: [{ type: 'text', text: 'late steering' }], source: { kind: 'user' } }))
|
2026-07-26 11:47:25 +08:00
|
|
|
exec.concludeTurn()
|
|
|
|
|
return [{ type: 'text', text: 'final' }]
|
|
|
|
|
},
|
|
|
|
|
}))
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-30 17:28:03 +08:00
|
|
|
expect(adapter.requests).toHaveLength(2)
|
2026-07-26 11:47:25 +08:00
|
|
|
const events = agent.session.events.map(event => event.type)
|
|
|
|
|
expect(events.filter(type => type === 'turn/end')).toHaveLength(1)
|
2026-07-30 17:28:03 +08:00
|
|
|
expect(JSON.stringify(adapter.requests[1]?.messages)).toContain('late steering')
|
2026-07-26 11:47:25 +08:00
|
|
|
const texts = adapter.requests[1]!.messages
|
|
|
|
|
.flatMap(message => message.content)
|
|
|
|
|
.filter(block => block.type === 'text')
|
|
|
|
|
.map(block => block.text)
|
|
|
|
|
expect(texts).toContain('late steering')
|
|
|
|
|
})
|
|
|
|
|
|
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall
The loop is now transmission-stateless; a request is a pure function of
(session log, this step's rendered assembly, current AgentOptions):
- The reconstruction boundary is step/start: the messages snapshot is
taken in the same synchronous frame immediately before the step/start
append, so the request's messages are exactly the derivation over
events[0..stepStartSeq) — an inject() from an agent/request listener
(or any concurrent task) lands after the boundary and joins the NEXT
request. This changes behavior for a synchronous step/start
session/event listener that appends content (master derived after the
append, so such a listener could reach the current request):
agent/pre-step is the sanctioned seam for current-request content.
- agent/request is re-typed to config-only: (agent, turn, step,
config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes
from AgentOptions on a loop instance's first request (explicit options
beat the logged baseline — fork overrides and resume reconfiguration
stay correct) and from the log's folded header afterwards; listeners
return a replacement to switch. Content shaping through the request is
no longer expressible — model-visible content flows through the log
channels.
- recordRequestHeader appends whatever header event the request owes the
log before dispatch: an 'initial'/'resume' snapshot anchoring each
loop instance, a round-trip-verified delta on change, a 'fallback'
snapshot when the encoding cannot express it. Session.requestHeader()
is the log's incrementally-folded baseline.
- Requests are deep-frozen before dispatch (deepFreeze exempts the
AbortSignal — freezing one breaks AbortController.abort() outright);
frozen + sessionId is the loop-built marker the dev invariant keys on.
Ported from #162 and re-anchored on the log: the append-extension /
frozen-end-to-end / compaction-resend / prompt-change property tests,
plus new specs for the boundary semantics, resume anchoring, and the
end-to-end theorem (every recorded request rebuilds byte-equal from the
log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against
the real DeepSeek API. Snapshot goldens intentionally stale until the
single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
|
|
|
it('agent/request waterfall switches models by returning a replacement config; the switch is logged', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
2026-08-06 12:13:14 +08:00
|
|
|
ctx.on('agent/request', async (_payload, next) => {
|
2026-07-24 21:18:48 +08:00
|
|
|
const config = await next()
|
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall
The loop is now transmission-stateless; a request is a pure function of
(session log, this step's rendered assembly, current AgentOptions):
- The reconstruction boundary is step/start: the messages snapshot is
taken in the same synchronous frame immediately before the step/start
append, so the request's messages are exactly the derivation over
events[0..stepStartSeq) — an inject() from an agent/request listener
(or any concurrent task) lands after the boundary and joins the NEXT
request. This changes behavior for a synchronous step/start
session/event listener that appends content (master derived after the
append, so such a listener could reach the current request):
agent/pre-step is the sanctioned seam for current-request content.
- agent/request is re-typed to config-only: (agent, turn, step,
config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes
from AgentOptions on a loop instance's first request (explicit options
beat the logged baseline — fork overrides and resume reconfiguration
stay correct) and from the log's folded header afterwards; listeners
return a replacement to switch. Content shaping through the request is
no longer expressible — model-visible content flows through the log
channels.
- recordRequestHeader appends whatever header event the request owes the
log before dispatch: an 'initial'/'resume' snapshot anchoring each
loop instance, a round-trip-verified delta on change, a 'fallback'
snapshot when the encoding cannot express it. Session.requestHeader()
is the log's incrementally-folded baseline.
- Requests are deep-frozen before dispatch (deepFreeze exempts the
AbortSignal — freezing one breaks AbortController.abort() outright);
frozen + sessionId is the loop-built marker the dev invariant keys on.
Ported from #162 and re-anchored on the log: the append-extension /
frozen-end-to-end / compaction-resend / prompt-change property tests,
plus new specs for the boundary semantics, resume anchoring, and the
end-to-end theorem (every recorded request rebuilds byte-equal from the
log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against
the real DeepSeek API. Snapshot goldens intentionally stale until the
single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
|
|
|
// The seed is frozen — config is not a mutable per-call knob; a switch
|
|
|
|
|
// is proposed by returning a replacement, and the loop logs it.
|
|
|
|
|
expect(Object.isFrozen(config)).toBe(true)
|
|
|
|
|
expect(() => { (config as { model: string }).model = 'other-model' }).toThrow(TypeError)
|
|
|
|
|
return { ...config, model: 'other-model' }
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
Enable maximum-strict TypeScript across our packages
tsconfig.base.json adds noUncheckedIndexedAccess,
exactOptionalPropertyTypes, noImplicitOverride,
noFallthroughCasesInSwitch, noUnusedLocals, and noUnusedParameters on
top of strict. Vendored packages opt out of the new flags locally
(their tsconfigs are ours to regenerate; their source is not), keeping
upstream-sync friendliness.
Our code fixed accordingly: index accesses acknowledge undefined
(assembler flush cursors, lastTurnNumber); optional properties are
omitted instead of set-to-undefined (GenerateResult.usage,
ToolDefinition.strict, GenerateOptions.system/tools, error payloads
via an errorData helper); Session.onAppend is explicitly
`(…) => void | undefined`; tests and examples updated for unused
parameters and indexed access.
2026-06-11 14:02:47 +08:00
|
|
|
expect(adapter.requests[0]!.model).toBe('other-model')
|
loop: every request is built from the log — boundary snapshot, header events, config-only waterfall
The loop is now transmission-stateless; a request is a pure function of
(session log, this step's rendered assembly, current AgentOptions):
- The reconstruction boundary is step/start: the messages snapshot is
taken in the same synchronous frame immediately before the step/start
append, so the request's messages are exactly the derivation over
events[0..stepStartSeq) — an inject() from an agent/request listener
(or any concurrent task) lands after the boundary and joins the NEXT
request. This changes behavior for a synchronous step/start
session/event listener that appends content (master derived after the
append, so such a listener could reach the current request):
agent/pre-step is the sanctioned seam for current-request content.
- agent/request is re-typed to config-only: (agent, turn, step,
config: LlmCallConfig, next) → LlmCallConfig. The frozen seed comes
from AgentOptions on a loop instance's first request (explicit options
beat the logged baseline — fork overrides and resume reconfiguration
stay correct) and from the log's folded header afterwards; listeners
return a replacement to switch. Content shaping through the request is
no longer expressible — model-visible content flows through the log
channels.
- recordRequestHeader appends whatever header event the request owes the
log before dispatch: an 'initial'/'resume' snapshot anchoring each
loop instance, a round-trip-verified delta on change, a 'fallback'
snapshot when the encoding cannot express it. Session.requestHeader()
is the log's incrementally-folded baseline.
- Requests are deep-frozen before dispatch (deepFreeze exempts the
AbortSignal — freezing one breaks AbortController.abort() outright);
frozen + sessionId is the loop-built marker the dev invariant keys on.
Ported from #162 and re-anchored on the log: the append-extension /
frozen-end-to-end / compaction-resend / prompt-change property tests,
plus new specs for the boundary semantics, resume anchoring, and the
end-to-end theorem (every recorded request rebuilds byte-equal from the
log alone). Live cache-hit e2e (request-cache.e2e.ts) verified against
the real DeepSeek API. Snapshot goldens intentionally stale until the
single re-record after the compact/summary envelope lands.
2026-07-06 03:07:34 +08:00
|
|
|
// The header event records what the request ACTUALLY used — the switch is
|
|
|
|
|
// a reconstructable fact, not silent drift.
|
|
|
|
|
const headerEvent = agent.session.events.find(e => e.type === 'request/header')
|
|
|
|
|
expect(headerEvent?.type === 'request/header' && headerEvent.data.header.config.model).toBe('other-model')
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-31 19:21:16 +08:00
|
|
|
it('agent/pre-step fires once per proposed step before the step is opened', async () => {
|
2026-06-26 08:59:33 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'echo', {}, 'calling echo'),
|
|
|
|
|
textResponse('done'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-24 21:18:48 +08:00
|
|
|
ctx.tools.register(defineContentToolFixture({
|
2026-06-26 08:59:33 +08:00
|
|
|
name: 'echo', description: 'echo', parameters: {},
|
|
|
|
|
async execute() { return [{ type: 'text', text: 'echoed' }] },
|
|
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-26 08:59:33 +08:00
|
|
|
|
2026-07-15 16:50:44 +08:00
|
|
|
const fires: { turn: number; step: number; signal: AbortSignal }[] = []
|
2026-08-06 12:13:14 +08:00
|
|
|
ctx.on('agent/pre-step', ({ agent: subject, turn, step, signal }, next) => {
|
2026-07-15 16:50:44 +08:00
|
|
|
if (subject === agent) fires.push({ turn, step, signal })
|
2026-07-31 19:21:16 +08:00
|
|
|
return next()
|
2026-06-26 08:59:33 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-15 16:50:44 +08:00
|
|
|
expect(fires.map(({ turn, step }) => ({ turn, step }))).toEqual([
|
|
|
|
|
{ turn: 1, step: 1 },
|
|
|
|
|
{ turn: 1, step: 2 },
|
2026-06-26 08:59:33 +08:00
|
|
|
])
|
2026-07-15 16:50:44 +08:00
|
|
|
expect(fires.every(({ signal }) => signal instanceof AbortSignal)).toBe(true)
|
2026-06-26 08:59:33 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-31 19:21:16 +08:00
|
|
|
it('agent/pre-step fires before its step boundary opens and before the request', async () => {
|
2026-06-26 08:59:33 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-26 08:59:33 +08:00
|
|
|
|
2026-07-31 19:21:16 +08:00
|
|
|
let boundaryOpen = true
|
2026-08-06 12:13:14 +08:00
|
|
|
ctx.on('agent/pre-step', ({ agent: subject }, next) => {
|
2026-07-30 17:28:03 +08:00
|
|
|
if (subject === agent) boundaryOpen = subject.session.events.at(-1)?.type === 'step/start'
|
2026-07-31 19:21:16 +08:00
|
|
|
return next()
|
2026-06-26 08:59:33 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-31 19:21:16 +08:00
|
|
|
expect(boundaryOpen).toBe(false)
|
2026-07-30 17:28:03 +08:00
|
|
|
expect(adapter.requests).toHaveLength(1)
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-31 19:21:16 +08:00
|
|
|
it('a throwing agent/pre-step listener fails the proposal, not the loop', async () => {
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('second turn ok')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
|
|
|
|
|
let throwOnce = true
|
2026-08-06 12:13:14 +08:00
|
|
|
ctx.on('agent/pre-step', (_payload, next) => {
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
if (throwOnce) { throwOnce = false; throw new Error('boom in pre-step') }
|
2026-07-31 19:21:16 +08:00
|
|
|
return next()
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
const errors: Error[] = []
|
2026-08-06 12:13:14 +08:00
|
|
|
ctx.on('agent/error', ({ error }) => {
|
2026-07-24 21:18:48 +08:00
|
|
|
if (error instanceof Error) errors.push(error)
|
|
|
|
|
})
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
|
|
|
|
|
send(agent, 'first')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
2026-08-04 14:09:52 +08:00
|
|
|
// The first proposal failed inside a balanced turn without calling the model.
|
2026-07-31 13:54:36 +08:00
|
|
|
expect(errors.map(error => error.message)).toEqual(['boom in pre-step'])
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
expect(adapter.requests.length).toBe(0)
|
2026-08-04 14:09:52 +08:00
|
|
|
expect(agent.session.events.some(event => event.type === 'turn/start')).toBe(true)
|
|
|
|
|
expect(agent.session.events.some(event => event.type === 'turn/end')).toBe(true)
|
fix(compact): decide step-alignment from surface tool-pairing, fire compaction pre-step (CBR-001)
Codex round 1 CBR-001: a head-anchored compaction checkpoint was
mis-classified by the log-position step-alignment scan, so a second
auto-compaction over a checkpoint-headed surface silently failed.
Root cause: `isStepAlignedStart/End` scanned the LOG by seq, but a
`replace` op lands a checkpoint at a high log seq whose SURFACE position
is the head — its log neighbours (the open step's assistant/message) are
not its surface neighbours, so the forward scan wrongly reported mid-step.
Fix, per the agreed direction:
- Replace the two log-position predicates with one surface-anchored
helper `isToolPairingBalanced(nodes, events, beforeSeq)` in
`dsh-session` (renamed step-boundary.ts → tool-pairing.ts). A cut is
balanced when no unanswered tool-call precedes it on the surface; a
region is collapsible iff both edges are balanced cuts. The open-tail
and free-node cases fall out of the same counter. It also throws on a
corrupt surface (a tool/result with no matching call).
- Move compaction off the in-step seam to a new "pre-step" seam fired
after turn/start and before step/start, so a compaction's log-only
compact/* records and its replacement node land cleanly OUTSIDE any
step (the honest structure crash-safety relies on). Renamed the event
agent/pre-request → agent/pre-step and switched its dispatch from
parallel → serial (listeners mutate the surface as a side effect;
serial isolates them so concurrent appends can't interleave). Extended
the catalog generator to accept @mode serial.
Regression coverage: a real-loop test driving an auto-compaction asserts
the landed checkpoint is a balanced cut on both sides; unit tests pin the
checkpoint case, the mid-step injection case, multi-call steps, and the
corrupt-surface guard. Proven red on the old log-position logic.
2026-06-26 13:51:01 +08:00
|
|
|
|
|
|
|
|
// The loop survived: a second prompt runs a normal completed turn.
|
|
|
|
|
send(agent, 'second')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(adapter.requests.length).toBe(1)
|
|
|
|
|
const lastTurnEnd = agent.session.events.findLast(e => e.type === 'turn/end')
|
|
|
|
|
expect(lastTurnEnd?.type === 'turn/end' && lastTurnEnd.data.reason).toEqual({ kind: 'completed' })
|
2026-06-26 08:59:33 +08:00
|
|
|
})
|
|
|
|
|
|
simplify(agent): drop the unused public Agent.abort(), keep whenIdle()
The public Agent handle exposed abort() (step-only) and cancel() (queue-aware).
No production caller used abort() — ACP maps session/cancel to cancel(), and
lifecycle owners tear down via AgentHandle.dispose(); the loop's own stop paths
abort their per-step AbortController directly. So abort() is latent generality
that keeps a private loop mechanic public.
RFC-premise correction: the public-agent-stop-surface RFC proposed removing
whenIdle() too. Implementation found whenIdle() load-bearing — a real
quiescence primitive with a deliberate loop contract (settle-without-transition,
the replacement-turn race) and ACP test consumers; its proposed replacement
("observe the running->idle transition") is exactly the async-state race
AGENTS.md warns against. So only abort() is removed; whenIdle() stays. The RFC
is amended on the way to implemented/ to record the narrowed scope, and the new
AGENTS.md "RFCs are proposals, not golden truth" principle (PR1) gets its
worked example.
- Remove Agent.abort() from the interface + the ReactLoopAgent impl; the no-arg
'aborted' default goes with it (cancel() keeps its 'cancelled' default).
- Migrate tests: empty-queue abort() -> cancel(reason); the two review-fixes
tests whose subject is the in-flight step's AbortController drive that
controller directly via the private currentAbort field (cancel() would clear
the inbox and destroy the queued steering one of them proves survives a step
abort). The no-arg-default test is dropped (cancel()'s default is already
covered in cancel.spec.ts).
- Resulting public stop surface: cancel() + whenIdle(). Update agent/agent-loop
READMEs, architecture.md, core.md type-equiv, the extension cookbook, the
lifecycle RFC (short note), and the proposed ACP RFC.
Implements docs/rfc/implemented/simplification/2026-06-20-public-agent-stop-surface.md
2026-06-21 05:50:39 +08:00
|
|
|
it('cancel() mid-stream ends the turn with reason aborted', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter(['hang'])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
simplify(agent): drop the unused public Agent.abort(), keep whenIdle()
The public Agent handle exposed abort() (step-only) and cancel() (queue-aware).
No production caller used abort() — ACP maps session/cancel to cancel(), and
lifecycle owners tear down via AgentHandle.dispose(); the loop's own stop paths
abort their per-step AbortController directly. So abort() is latent generality
that keeps a private loop mechanic public.
RFC-premise correction: the public-agent-stop-surface RFC proposed removing
whenIdle() too. Implementation found whenIdle() load-bearing — a real
quiescence primitive with a deliberate loop contract (settle-without-transition,
the replacement-turn race) and ACP test consumers; its proposed replacement
("observe the running->idle transition") is exactly the async-state race
AGENTS.md warns against. So only abort() is removed; whenIdle() stays. The RFC
is amended on the way to implemented/ to record the narrowed scope, and the new
AGENTS.md "RFCs are proposals, not golden truth" principle (PR1) gets its
worked example.
- Remove Agent.abort() from the interface + the ReactLoopAgent impl; the no-arg
'aborted' default goes with it (cancel() keeps its 'cancelled' default).
- Migrate tests: empty-queue abort() -> cancel(reason); the two review-fixes
tests whose subject is the in-flight step's AbortController drive that
controller directly via the private currentAbort field (cancel() would clear
the inbox and destroy the queued steering one of them proves survives a step
abort). The no-arg-default test is dropped (cancel()'s default is already
covered in cancel.spec.ts).
- Resulting public stop surface: cancel() + whenIdle(). Update agent/agent-loop
READMEs, architecture.md, core.md type-equiv, the extension cookbook, the
lifecycle RFC (short note), and the proposed ACP RFC.
Implements docs/rfc/implemented/simplification/2026-06-20-public-agent-stop-surface.md
2026-06-21 05:50:39 +08:00
|
|
|
// wait until the stream is hanging, then cancel
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
await new Promise(r => setTimeout(r, 30))
|
|
|
|
|
expect(agent.status).toBe('running')
|
2026-07-16 18:12:34 +08:00
|
|
|
agent.cancel({ kind: 'user' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-30 17:28:03 +08:00
|
|
|
expect(reasons).toEqual([{ kind: 'aborted', reason: { kind: 'user' } }])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
2026-06-15 23:53:47 +08:00
|
|
|
it('surfaces max-tokens as the turn-end reason when the last step is cut off', async () => {
|
|
|
|
|
// A single step that ends with a max-tokens finish (no tool calls): the
|
|
|
|
|
// turn stops by default and ends max-tokens, not completed.
|
|
|
|
|
const adapter = new MockAdapter([maxTokensResponse('truncat')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
|
|
|
|
expect(reasons).toEqual([{ kind: 'max-tokens' }])
|
2026-07-13 23:27:00 +08:00
|
|
|
// Assert the durable row, not only the live listener.
|
2026-06-15 23:53:47 +08:00
|
|
|
const turnEnd = agent.session.events.findLast(e => e.type === 'turn/end')
|
|
|
|
|
expect(turnEnd!.data.reason).toEqual({ kind: 'max-tokens' })
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('a max-tokens step earlier in a turn still surfaces as max-tokens after a later completed step', async () => {
|
2026-07-12 03:36:43 +08:00
|
|
|
// Step 1 is cut off (max-tokens, no tool calls → would stop by default), so continuation
|
|
|
|
|
// must be FORCED to reach step 2 which finishes normally (stop).
|
2026-06-15 23:53:47 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
maxTokensResponse('first half'),
|
|
|
|
|
textResponse('second half'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
let steps = 0
|
2026-06-30 10:32:55 +08:00
|
|
|
ctx.on('session/event', (_session, event) => { if (event.type === 'step/end') steps++ })
|
2026-06-15 23:53:47 +08:00
|
|
|
// Force exactly one continuation (step 1 → step 2), then defer to default
|
|
|
|
|
// (step 2 is a plain stop with no tool calls → stops).
|
2026-08-06 12:13:14 +08:00
|
|
|
ctx.on('agent/turn-stopping', ({ agent: subject }) => {
|
2026-07-24 21:18:48 +08:00
|
|
|
if (steps < 2) {
|
2026-07-28 13:55:59 +08:00
|
|
|
subject.steer(createUserMessage({ content: [{ type: 'text', text: 'continue after truncation' }], source: { kind: 'plugin', plugin: 'max-tokens-test' } }))
|
2026-07-24 21:18:48 +08:00
|
|
|
}
|
2026-06-15 23:53:47 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(steps).toBe(2)
|
|
|
|
|
expect(adapter.requests).toHaveLength(2)
|
2026-06-19 00:37:12 +08:00
|
|
|
expect(adapter.requests[1]!.messages).toEqual([
|
2026-07-28 14:15:23 +08:00
|
|
|
{
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'user',
|
|
|
|
|
content: [{ type: 'text', text: 'go' }],
|
|
|
|
|
source: { kind: 'user' },
|
|
|
|
|
},
|
|
|
|
|
{
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'assistant',
|
|
|
|
|
content: [{ type: 'text', text: 'first half' }],
|
|
|
|
|
source: { kind: 'model', provider: 'mock', model: 'mock' },
|
|
|
|
|
},
|
|
|
|
|
{
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'user',
|
|
|
|
|
content: [{ type: 'text', text: 'continue after truncation' }],
|
|
|
|
|
source: { kind: 'plugin', plugin: 'max-tokens-test' },
|
|
|
|
|
},
|
2026-06-19 00:37:12 +08:00
|
|
|
])
|
2026-08-03 20:36:26 +08:00
|
|
|
// A max-token step is sticky: the later completed step must not
|
|
|
|
|
// downgrade the turn outcome.
|
|
|
|
|
expect(reasons).toEqual([{ kind: 'max-tokens' }])
|
2026-06-15 23:53:47 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('a completed step after no max-tokens keeps the turn completed (max-tokens does not leak across turns)', async () => {
|
|
|
|
|
// Two consecutive turns: turn 1 is cut off (max-tokens), turn 2 is a clean
|
|
|
|
|
// stop. The per-turn reason must be independent — turn 2 ends completed.
|
|
|
|
|
const adapter = new MockAdapter([maxTokensResponse('cut'), textResponse('clean')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-15 23:53:47 +08:00
|
|
|
|
|
|
|
|
send(agent, 'first')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
send(agent, 'second')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(reasons).toEqual([{ kind: 'max-tokens' }, { kind: 'completed' }])
|
|
|
|
|
})
|
|
|
|
|
|
2026-06-17 21:25:47 +08:00
|
|
|
it('does not dispatch tool calls from a max-tokens-truncated step', async () => {
|
|
|
|
|
const callId = CallId('c1')
|
|
|
|
|
const adapter = new MockAdapter([[
|
|
|
|
|
{ type: 'block-start', index: 0, blockType: 'tool-call' },
|
|
|
|
|
{ type: 'tool-call-delta', index: 0, id: callId, name: 'echo', argumentsDelta: '{"text":"x"}' },
|
|
|
|
|
{ type: 'block-end', index: 0, block: { type: 'tool-call', id: callId, name: 'echo', arguments: '{"text":"x"}' } },
|
|
|
|
|
{ type: 'usage', usage: { inputTokens: 10, outputTokens: 5 } },
|
|
|
|
|
{ type: 'finish', reason: { kind: 'max-tokens' } },
|
|
|
|
|
]])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
let executions = 0
|
2026-07-24 21:18:48 +08:00
|
|
|
ctx.tools.register(defineContentToolFixture({
|
2026-06-17 21:25:47 +08:00
|
|
|
name: 'echo',
|
|
|
|
|
description: '',
|
|
|
|
|
parameters: { text: { type: 'string' } },
|
|
|
|
|
async execute() {
|
|
|
|
|
executions += 1
|
|
|
|
|
return [{ type: 'text', text: 'should not run' }]
|
|
|
|
|
},
|
|
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-17 21:25:47 +08:00
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-17 21:25:47 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(executions).toBe(0)
|
|
|
|
|
expect(agent.session.events.some(e => e.type === 'tool/call')).toBe(false)
|
2026-07-28 14:15:23 +08:00
|
|
|
expect(agent.session.deriveMessages()).toEqual([{
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'user',
|
|
|
|
|
content: [{ type: 'text', text: 'go' }],
|
|
|
|
|
source: { kind: 'user' },
|
|
|
|
|
}])
|
2026-06-19 00:37:12 +08:00
|
|
|
expect(reasons).toEqual([{ kind: 'max-tokens' }])
|
2026-07-13 23:27:00 +08:00
|
|
|
// Empty content still needs an assistant/message to carry usage; derivation
|
|
|
|
|
// skips that host so it does not create a spurious assistant turn.
|
2026-06-21 10:00:06 +08:00
|
|
|
const assistantMessage = agent.session.events.find(e => e.type === 'assistant/message')
|
|
|
|
|
expect(assistantMessage?.type === 'assistant/message' && assistantMessage.data).toEqual({
|
2026-07-28 14:15:23 +08:00
|
|
|
turn: 1,
|
|
|
|
|
step: 1,
|
|
|
|
|
message: {
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'assistant',
|
|
|
|
|
content: [],
|
|
|
|
|
source: { kind: 'model', provider: 'mock', model: 'mock' },
|
|
|
|
|
},
|
|
|
|
|
usage: { inputTokens: 10, outputTokens: 5 },
|
2026-06-21 10:00:06 +08:00
|
|
|
})
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-15 14:47:29 +08:00
|
|
|
it('appends an empty completion anchor for a max-tokens step with no usage', async () => {
|
|
|
|
|
// The truncated tool call is dropped from durable content, while the
|
|
|
|
|
// successful provider call still needs an exact replay anchor.
|
2026-06-21 10:00:06 +08:00
|
|
|
const callId = CallId('c1')
|
|
|
|
|
const adapter = new MockAdapter([[
|
|
|
|
|
{ type: 'block-start', index: 0, blockType: 'tool-call' },
|
|
|
|
|
{ type: 'tool-call-delta', index: 0, id: callId, name: 'echo', argumentsDelta: '{"text":"x"}' },
|
|
|
|
|
{ type: 'block-end', index: 0, block: { type: 'tool-call', id: callId, name: 'echo', arguments: '{"text":"x"}' } },
|
|
|
|
|
{ type: 'finish', reason: { kind: 'max-tokens' } },
|
|
|
|
|
]])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-24 21:18:48 +08:00
|
|
|
ctx.tools.register(defineContentToolFixture({
|
2026-06-21 10:00:06 +08:00
|
|
|
name: 'echo',
|
|
|
|
|
description: '',
|
|
|
|
|
parameters: { text: { type: 'string' } },
|
|
|
|
|
async execute() { return [{ type: 'text', text: 'should not run' }] },
|
|
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-21 10:00:06 +08:00
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-21 10:00:06 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(reasons).toEqual([{ kind: 'max-tokens' }])
|
2026-07-15 14:47:29 +08:00
|
|
|
const assistant = agent.session.events.find(e => e.type === 'assistant/message')!
|
|
|
|
|
expect(assistant.type === 'assistant/message' && assistant.data).toEqual({
|
|
|
|
|
turn: 1,
|
|
|
|
|
step: 1,
|
2026-07-28 14:15:23 +08:00
|
|
|
message: {
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'assistant',
|
|
|
|
|
content: [],
|
|
|
|
|
source: { kind: 'model', provider: 'mock', model: 'mock' },
|
|
|
|
|
},
|
2026-07-15 14:47:29 +08:00
|
|
|
})
|
|
|
|
|
expect(assistant.sourceEventSeqs?.length).toBeGreaterThan(0)
|
2026-07-28 14:15:23 +08:00
|
|
|
expect(agent.session.deriveMessages()).toEqual([{
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'user',
|
|
|
|
|
content: [{ type: 'text', text: 'go' }],
|
|
|
|
|
source: { kind: 'user' },
|
|
|
|
|
}])
|
2026-06-19 00:37:12 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-15 14:47:29 +08:00
|
|
|
it('appends an empty completion anchor for a normal stop with no usage', async () => {
|
|
|
|
|
// A clean content-less call stays absent from derived messages but remains
|
|
|
|
|
// a durable successful-call boundary for replay consumers.
|
2026-06-21 11:08:10 +08:00
|
|
|
const adapter = new MockAdapter([[{ type: 'finish', reason: { kind: 'stop' } }]])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-21 11:08:10 +08:00
|
|
|
|
|
|
|
|
const reasons: TurnEndReason[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
2026-06-21 11:08:10 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(reasons).toEqual([{ kind: 'completed' }])
|
2026-07-15 14:47:29 +08:00
|
|
|
const assistant = agent.session.events.find(e => e.type === 'assistant/message')!
|
|
|
|
|
expect(assistant.type === 'assistant/message' && assistant.data).toEqual({
|
|
|
|
|
turn: 1,
|
|
|
|
|
step: 1,
|
2026-07-28 14:15:23 +08:00
|
|
|
message: {
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'assistant',
|
|
|
|
|
content: [],
|
|
|
|
|
source: { kind: 'model', provider: 'mock', model: 'mock' },
|
|
|
|
|
},
|
2026-07-15 14:47:29 +08:00
|
|
|
})
|
|
|
|
|
expect(assistant.sourceEventSeqs?.length).toBe(1)
|
2026-07-28 14:15:23 +08:00
|
|
|
expect(agent.session.deriveMessages()).toEqual([{
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'user',
|
|
|
|
|
content: [{ type: 'text', text: 'go' }],
|
|
|
|
|
source: { kind: 'user' },
|
|
|
|
|
}])
|
2026-06-21 11:08:10 +08:00
|
|
|
})
|
|
|
|
|
|
2026-06-19 00:37:12 +08:00
|
|
|
it('keeps safe max-tokens assistant content while dropping truncated tool calls', async () => {
|
|
|
|
|
const callId = CallId('c1')
|
|
|
|
|
const adapter = new MockAdapter([[
|
|
|
|
|
{ type: 'block-start', index: 0, blockType: 'text' },
|
|
|
|
|
{ type: 'text-delta', index: 0, text: 'partial text' },
|
|
|
|
|
{ type: 'block-end', index: 0, block: { type: 'text', text: 'partial text' } },
|
|
|
|
|
{ type: 'block-start', index: 1, blockType: 'tool-call' },
|
|
|
|
|
{ type: 'tool-call-delta', index: 1, id: callId, name: 'echo', argumentsDelta: '{"text"' },
|
2026-08-15 16:07:30 +08:00
|
|
|
{
|
|
|
|
|
type: 'finish',
|
|
|
|
|
reason: { kind: 'max-tokens' },
|
|
|
|
|
replayState: { response: { responseId: 'resp-1' }, blocks: ['text-meta', 'tool-meta'] },
|
|
|
|
|
},
|
|
|
|
|
], textResponse('continued')])
|
2026-06-19 00:37:12 +08:00
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-19 00:37:12 +08:00
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
2026-08-15 16:07:30 +08:00
|
|
|
send(agent, 'continue')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
2026-06-19 00:37:12 +08:00
|
|
|
|
|
|
|
|
expect(agent.session.events.some(e => e.type === 'tool/call')).toBe(false)
|
2026-08-15 16:07:30 +08:00
|
|
|
// The follow-up request replays the truncated message with its replay
|
|
|
|
|
// metadata pruned in step with the dropped tool call.
|
|
|
|
|
expect(adapter.requests[1]?.messages[1]?.source).toEqual({
|
|
|
|
|
kind: 'model',
|
|
|
|
|
provider: 'mock',
|
|
|
|
|
model: 'mock',
|
|
|
|
|
replayState: { response: { responseId: 'resp-1' }, blocks: ['text-meta'] },
|
|
|
|
|
})
|
2026-06-17 21:25:47 +08:00
|
|
|
expect(agent.session.deriveMessages()).toEqual([
|
2026-07-28 14:15:23 +08:00
|
|
|
{
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'user',
|
|
|
|
|
content: [{ type: 'text', text: 'go' }],
|
|
|
|
|
source: { kind: 'user' },
|
|
|
|
|
},
|
|
|
|
|
{
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'assistant',
|
|
|
|
|
content: [{ type: 'text', text: 'partial text' }],
|
2026-08-15 16:07:30 +08:00
|
|
|
source: {
|
|
|
|
|
kind: 'model',
|
|
|
|
|
provider: 'mock',
|
|
|
|
|
model: 'mock',
|
|
|
|
|
replayState: { response: { responseId: 'resp-1' }, blocks: ['text-meta'] },
|
|
|
|
|
},
|
|
|
|
|
},
|
|
|
|
|
{
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'user',
|
|
|
|
|
content: [{ type: 'text', text: 'continue' }],
|
|
|
|
|
source: { kind: 'user' },
|
|
|
|
|
},
|
|
|
|
|
{
|
|
|
|
|
id: expect.any(String) as unknown,
|
|
|
|
|
role: 'assistant',
|
|
|
|
|
content: [{ type: 'text', text: 'continued' }],
|
2026-07-28 14:15:23 +08:00
|
|
|
source: { kind: 'model', provider: 'mock', model: 'mock' },
|
|
|
|
|
},
|
2026-06-17 21:25:47 +08:00
|
|
|
])
|
|
|
|
|
})
|
|
|
|
|
|
2026-07-12 18:57:42 +08:00
|
|
|
it('contains a step/end observer failure without changing continuation', async () => {
|
2026-06-17 21:25:47 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'echo', { text: 'x' }),
|
2026-07-12 18:57:42 +08:00
|
|
|
textResponse('continued after tool call'),
|
2026-06-17 21:25:47 +08:00
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-24 21:18:48 +08:00
|
|
|
ctx.tools.register(defineContentToolFixture({
|
2026-06-17 21:25:47 +08:00
|
|
|
name: 'echo',
|
|
|
|
|
description: '',
|
|
|
|
|
parameters: { text: { type: 'string' } },
|
|
|
|
|
async execute(args) {
|
|
|
|
|
return [{ type: 'text', text: String(args.text) }]
|
|
|
|
|
},
|
|
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-06-17 21:25:47 +08:00
|
|
|
let threw = false
|
2026-07-12 18:57:42 +08:00
|
|
|
// Post-commit session observers cannot control the loop. The tool call still
|
|
|
|
|
// drives the second model request, and the turn completes normally.
|
2026-06-30 10:32:55 +08:00
|
|
|
ctx.on('session/event', (_session, event) => {
|
|
|
|
|
if (event.type === 'step/end' && !threw) { threw = true; throw new Error('bad step/end listener') }
|
2026-06-17 21:25:47 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'go')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-12 18:57:42 +08:00
|
|
|
expect(adapter.requests).toHaveLength(2)
|
2026-06-17 21:25:47 +08:00
|
|
|
const turnEnd = agent.session.events.findLast(e => e.type === 'turn/end')
|
2026-07-12 18:57:42 +08:00
|
|
|
expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason.kind).toBe('completed')
|
2026-06-17 21:25:47 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-30 17:28:03 +08:00
|
|
|
it('contains a reentrant send attempted during durable inbox publication', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('first')])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const ctx = await harness(adapter)
|
2026-07-20 11:52:30 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-17 17:14:52 +08:00
|
|
|
|
|
|
|
|
let nested = false
|
2026-07-30 17:28:03 +08:00
|
|
|
ctx.on('session/event', (session, event) => {
|
|
|
|
|
if (session !== agent.session || event.type !== 'agent/inbox/spliced'
|
|
|
|
|
|| event.data.inserted.length === 0 || nested) return
|
2026-07-17 17:14:52 +08:00
|
|
|
nested = true
|
|
|
|
|
send(agent, 'queued listener message')
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
const idle = waitForIdle(ctx, agent)
|
|
|
|
|
send(agent, 'outer message')
|
|
|
|
|
await idle
|
|
|
|
|
|
|
|
|
|
const turns = agent.session.events.filter(event => event.type === 'turn/start')
|
|
|
|
|
const messages = agent.session.events
|
|
|
|
|
.filter(event => event.type === 'user/message')
|
|
|
|
|
.map(event => event.data.content)
|
2026-07-30 17:28:03 +08:00
|
|
|
expect(turns).toHaveLength(1)
|
|
|
|
|
expect(messages).toEqual([[{ type: 'text', text: 'outer message' }]])
|
2026-07-17 17:14:52 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('preserves independent turn sources across an adjacent microtask send', async () => {
|
|
|
|
|
const adapter = new MockAdapter([textResponse('first'), textResponse('second')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-20 11:52:30 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-17 17:14:52 +08:00
|
|
|
|
|
|
|
|
const idle = waitForIdle(ctx, agent)
|
2026-07-28 13:55:59 +08:00
|
|
|
agent.followup(createUserMessage({ content: [{ type: 'text', text: 'user message' }], source: { kind: 'user' } }))
|
2026-07-17 17:14:52 +08:00
|
|
|
await Promise.resolve()
|
2026-07-28 13:55:59 +08:00
|
|
|
agent.followup(createUserMessage({ content: [{ type: 'text', text: 'plugin message' }], source: { kind: 'plugin', plugin: 'test' } }))
|
2026-07-17 17:14:52 +08:00
|
|
|
await idle
|
|
|
|
|
|
2026-07-30 13:49:57 +08:00
|
|
|
const turns = agent.session.events.filter(event => event.type === 'turn/start')
|
2026-07-17 17:14:52 +08:00
|
|
|
const sources = agent.session.events
|
|
|
|
|
.filter(event => event.type === 'user/message')
|
|
|
|
|
.map(event => event.data.source)
|
2026-07-30 13:49:57 +08:00
|
|
|
expect(turns).toHaveLength(2)
|
2026-07-17 17:14:52 +08:00
|
|
|
expect(sources).toEqual([
|
|
|
|
|
{ kind: 'user' },
|
|
|
|
|
{ kind: 'plugin', plugin: 'test' },
|
|
|
|
|
])
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('keeps a session-listener send after dequeue in the following turn', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([textResponse('first'), textResponse('second')])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
const turns: number[] = []
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/start') turns.push(event.data.turn) })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
// queue two messages while idle — first starts turn 1 immediately;
|
2026-07-02 23:42:16 +08:00
|
|
|
// queue the second during turn 1 when the first assistant chunk streams
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
let queued = false
|
2026-07-02 23:42:16 +08:00
|
|
|
ctx.on('session/event', (_s, event) => {
|
|
|
|
|
if (event.type === 'assistant/chunk' && !queued) {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
queued = true
|
2026-07-30 17:28:03 +08:00
|
|
|
queueMicrotask(() => { send(agent, 'second message') })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
}
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
send(agent, 'first message')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
|
|
|
|
expect(turns).toEqual([1, 2])
|
|
|
|
|
expect(adapter.requests).toHaveLength(2)
|
2026-07-17 17:14:52 +08:00
|
|
|
expect(JSON.stringify(adapter.requests[1]!.messages)).toContain('first')
|
|
|
|
|
expect(JSON.stringify(adapter.requests[1]!.messages)).toContain('second message')
|
|
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('keeps a model-adapter callback send in the following turn', async () => {
|
2026-07-20 11:52:30 +08:00
|
|
|
const agentRef: { current?: Agent } = {}
|
2026-07-17 17:14:52 +08:00
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
() => {
|
|
|
|
|
const agent = agentRef.current
|
|
|
|
|
if (agent === undefined) throw new Error('model callback ran before agent setup')
|
|
|
|
|
send(agent, 'model callback message')
|
|
|
|
|
return textResponse('first')
|
|
|
|
|
},
|
|
|
|
|
textResponse('second'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-20 11:52:30 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
2026-07-17 17:14:52 +08:00
|
|
|
agentRef.current = agent
|
|
|
|
|
|
|
|
|
|
const idle = waitForIdle(ctx, agent)
|
|
|
|
|
send(agent, 'outer message')
|
|
|
|
|
await idle
|
|
|
|
|
|
|
|
|
|
const messages = agent.session.events
|
|
|
|
|
.filter(event => event.type === 'user/message')
|
|
|
|
|
.map(event => event.data.content)
|
|
|
|
|
expect(agent.session.events.filter(event => event.type === 'turn/start')).toHaveLength(2)
|
|
|
|
|
expect(messages).toEqual([
|
|
|
|
|
[{ type: 'text', text: 'outer message' }],
|
|
|
|
|
[{ type: 'text', text: 'model callback message' }],
|
|
|
|
|
])
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-30 17:28:03 +08:00
|
|
|
it('records normalized model errors on the turn boundary', async () => {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const adapter = new MockAdapter([]) // script exhausted → throws
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
2026-07-31 13:54:36 +08:00
|
|
|
const errors: unknown[] = []
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
const reasons: TurnEndReason[] = []
|
2026-08-06 12:13:14 +08:00
|
|
|
ctx.on('agent/error', ({ error }) => {
|
2026-07-31 13:54:36 +08:00
|
|
|
errors.push(error)
|
2026-07-24 21:18:48 +08:00
|
|
|
})
|
2026-07-02 03:26:45 +08:00
|
|
|
ctx.on('session/event', (_s, event) => { if (event.type === 'turn/end') reasons.push(event.data.reason) })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
2026-07-31 13:54:36 +08:00
|
|
|
expect(errors).toHaveLength(1)
|
|
|
|
|
expect(errors[0]).toBeInstanceOf(LlmError)
|
|
|
|
|
expect((errors[0] as LlmError).failure).toEqual({
|
|
|
|
|
message: 'MockAdapter: script exhausted',
|
|
|
|
|
code: 'UNKNOWN',
|
|
|
|
|
})
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(reasons[0]).toMatchObject({ kind: 'error' })
|
2026-07-31 13:54:36 +08:00
|
|
|
// The durable failure and live relay describe the same failed turn.
|
2026-06-21 10:00:06 +08:00
|
|
|
const turnEnd = agent.session.events.find(e => e.type === 'turn/end')
|
2026-07-30 17:28:03 +08:00
|
|
|
expect(turnEnd?.type === 'turn/end' && turnEnd.data.reason).toMatchObject({ kind: 'error' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
|
|
|
|
it('disposing the loop fiber mid-turn stops the loop (HMR safety)', async () => {
|
|
|
|
|
const adapter = new MockAdapter(['hang'])
|
|
|
|
|
const ctx = await harness(adapter)
|
|
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
let agent!: Agent
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
const fiber = await ctx.plugin(Object.assign((inner: Context) => {
|
2026-07-18 12:21:15 +08:00
|
|
|
agent = inner.agentLoop.create(SessionId('scoped'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
}, { inject: ['agentLoop'] }))
|
|
|
|
|
|
2026-07-14 01:59:21 +08:00
|
|
|
expect(ctx.agents.get(SessionId('scoped'))).toBe(agent)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
send(agent, 'go')
|
|
|
|
|
await new Promise(r => setTimeout(r, 30))
|
|
|
|
|
expect(agent.status).toBe('running')
|
|
|
|
|
|
|
|
|
|
await fiber.dispose()
|
2026-07-14 02:32:35 +08:00
|
|
|
await driverDone(agent)
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
|
2026-07-24 21:58:07 +08:00
|
|
|
expect(ctx.agents.get(SessionId('scoped'))).toBeUndefined()
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
it('creates agents from config on startup', async () => {
|
2026-08-24 18:23:42 +08:00
|
|
|
const effort = ReasoningEffortId('high')
|
|
|
|
|
const adapter = new MockAdapter([textResponse('from config')], {
|
|
|
|
|
efforts: [{ id: effort, name: 'High' }],
|
|
|
|
|
defaultEffort: effort,
|
|
|
|
|
})
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
const ctx = new Context()
|
2026-08-13 00:36:22 +08:00
|
|
|
await ctx.plugin(LlmRuntime)
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
await ctx.plugin(SessionStore)
|
|
|
|
|
await ctx.plugin(SystemPrompt)
|
2026-08-13 00:36:22 +08:00
|
|
|
await ctx.plugin(ToolRuntime)
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
await ctx.plugin(AgentRegistry)
|
|
|
|
|
await ctx.plugin(AgentLoop, {
|
2026-08-24 18:23:42 +08:00
|
|
|
agents: [{ id: SessionId('config-agent'), provider: 'mock', model: 'mock', reasoningEffort: effort }],
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
})
|
|
|
|
|
ctx.llm.registerAdapter(['mock'], adapter)
|
|
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
const agent = ctx.agents.list()[0]!
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
expect(agent).toBeDefined()
|
2026-07-14 01:59:21 +08:00
|
|
|
expect(agent.id).toBe(agent.session.id)
|
|
|
|
|
expect(agent.id).toMatch(/^config-agent-session-/)
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
expect(agent.options.model).toBe('mock')
|
|
|
|
|
|
|
|
|
|
// the agent is alive: send triggers a turn
|
|
|
|
|
send(agent, 'hi')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
expect(adapter.requests).toHaveLength(1)
|
2026-08-24 18:23:42 +08:00
|
|
|
expect(adapter.requests[0]?.reasoningEffort).toBe(effort)
|
|
|
|
|
const header = agent.session.events.find(event => event.type === 'request/header')
|
|
|
|
|
expect(header?.type === 'request/header' && header.data.header.config.reasoningEffort).toBe(effort)
|
Enforce 100% per-file test coverage on packages/*/src
vitest coverage (v8 provider) with per-file 100% thresholds for
statements, branches, functions, and lines. Scope: our runtime source
only — types-only files, vendor/ (upstream code), and examples/
(exercised by the demo smoke test) are excluded. yarn test:coverage
runs the gate.
59 tests added to close every gap: llm generate-waterfall and adapter
disposal; assembler edge protocol (duplicate block-start, stragglers
after block-end, id fallback, usage omission, invariant violation);
the whole Inbox surface incl. the wakeup-overwrite race; LoopAgent
disposed-state throws and double-stop idempotence; config-driven agent
creation; loop backstop catches (throwing turn-start/turn-end
listeners, non-Error throws, non-JSON tool arguments); system-prompt
dynamic sections and disposer paths; tools errorMessage fallbacks and
the full schema-DSL emission matrix. Genuinely unreachable defensive
guards carry /* v8 ignore */ comments with stated reasons rather than
deletion (132 tests total).
2026-06-11 14:58:36 +08:00
|
|
|
})
|
|
|
|
|
|
2026-06-25 23:35:13 +08:00
|
|
|
it('attaches config agent cwd to the fresh session header', async () => {
|
|
|
|
|
const ctx = new Context()
|
2026-08-13 00:36:22 +08:00
|
|
|
await ctx.plugin(LlmRuntime)
|
2026-06-25 23:35:13 +08:00
|
|
|
await ctx.plugin(SessionStore)
|
|
|
|
|
await ctx.plugin(SystemPrompt)
|
2026-08-13 00:36:22 +08:00
|
|
|
await ctx.plugin(ToolRuntime)
|
2026-06-25 23:35:13 +08:00
|
|
|
await ctx.plugin(AgentRegistry)
|
|
|
|
|
await ctx.plugin(AgentLoop, {
|
2026-07-18 12:21:15 +08:00
|
|
|
agents: [{ id: SessionId('config-agent'), provider: 'mock', model: 'mock', cwd: '/work/project' }],
|
2026-06-25 23:35:13 +08:00
|
|
|
})
|
|
|
|
|
|
2026-07-14 02:32:35 +08:00
|
|
|
const agent = ctx.agents.list()[0]!
|
2026-06-25 23:35:13 +08:00
|
|
|
expect(agent.session.header.cwd).toBe('/work/project')
|
|
|
|
|
})
|
|
|
|
|
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
it('replays a session log into an identical derived history', async () => {
|
|
|
|
|
const adapter = new MockAdapter([
|
|
|
|
|
toolCallResponse('c1', 'echo', { text: 'x' }),
|
|
|
|
|
textResponse('done'),
|
|
|
|
|
])
|
|
|
|
|
const ctx = await harness(adapter)
|
2026-07-24 21:18:48 +08:00
|
|
|
ctx.tools.register(defineContentToolFixture({
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
name: 'echo',
|
|
|
|
|
description: '',
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
parameters: { text: { type: 'string' } },
|
|
|
|
|
async execute(args) {
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
return [{ type: 'text', text: String(args.text) }]
|
|
|
|
|
},
|
Document the codebase thoroughly and tighten type safety
Docs: per-folder README.md for packages/ (family overview + one per
package: service, events, API, extension points, TODOs), examples/,
and examples/echo-agent/; folder-level AGENTS.md (+ CLAUDE.md
symlinks) for packages/ and vendor/; module-level doc comments in
every packages/*/src file; richer JSDoc on all exported API
(event side effects, disposal contracts, error behavior). Root
AGENTS.md gains a "Type Safety and Documentation" policy section:
the codebase aims to be very type-safe and well documented; type
gymnastics are acceptable in core packages when they improve
plugin-author DX; verbose docs are fine as long as they stay strictly
in sync with the code.
Type safety: removed the upstream-inherited "noImplicitAny": false
from tsconfig.base.json — packages/* now compile under full strict
mode; vendor/loader and vendor/include set it locally (vendor/cordis
already did). Eliminated every `: any` / `as any` from packages and
examples (catch clauses use unknown + a CodedError narrowing type;
event data access uses discriminated-union narrowing).
Typed tool schemas: new @deepseek-ai/dsh-tools schema DSL —
SchemaSpec with per-property `required: true` booleans, type-level
InferArgs<S>, a runtime SchemaSpec → JSON Schema converter, and
defineTool() so first-party tools get typed execute(args) with zero
casts (raw JSON Schema still accepted for MCP interop; chosen over
schemastery because it targets JSON Schema generation directly).
echo-tool and all test tools migrated; +7 tests.
2026-06-11 12:39:27 +08:00
|
|
|
}))
|
2026-07-18 12:21:15 +08:00
|
|
|
const agent = ctx.agentLoop.create(SessionId('a1'), { provider: 'mock', model: 'mock' })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
send(agent, 'run')
|
|
|
|
|
await waitForIdle(ctx, agent)
|
|
|
|
|
|
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
|
|
|
const replayed = ctx.sessions.create(SessionId('replayed'), { seed: [...agent.session.events] })
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
expect(replayed.deriveMessages()).toEqual(agent.session.deriveMessages())
|
feat(session): project the inherited-history boundary into the log
A plugin owning a standalone open/close bracket cannot tell a dead marker
from a live one: an unmatched `compact/start` reads identically whether the
previous writer died mid-compaction or a compaction is running now.
`Session.firstLiveSeq` already holds that answer exactly, but only in memory.
Append the log-only `session/inherited` event at that seq from the seeded
constructor — the single waist all six seeded-start paths pass through
(resume, configured startup on a persisted id, `sessions.fork()`, a subagent
fork child, `adopt()`'s live prefix, and a bare seeded `create`). Read it
through the new `isInheritedSeq(events, seq)`.
The constructor placement means persistence needs no changes: the marker is
already in `events` when a backend captures the creation seed, so it rides
the ordinary seed path with no load-time write. It also covers fork, where
the inherited bracket's owner may still be running — the case a
persistence-layer boundary could not reach.
Activity ordering excludes the boundary through `lastActivityTime()`, since
lazy resume makes browsing a pickup and the three call sites would otherwise
float every opened session to the top of a picker or list.
2026-07-30 11:38:51 +08:00
|
|
|
// event-by-event identity of types over the inherited prefix
|
|
|
|
|
expect(replayed.events.slice(0, agent.session.seq).map(e => e.type)).toEqual(
|
Add ESLint: typescript-eslint strict-type-checked + stylistic formatting
Flat config with two layers. Correctness (type-checked): the headline
rules for this codebase are no-floating-promises / no-misused-promises
(a lost promise in the agent loop is our primary bug class),
switch-exhaustiveness-check (we switch over merge-extensible unions
everywhere), no-unnecessary-condition, require-await, and
no-explicit-any. Style (@stylistic): 2-space, no semicolons, single
quotes, trailing commas, max-len 140 — the existing house style, now
enforced instead of drifting between agents. vendor/ is excluded
(vendored source keeps upstream style); tests relax the rules that
fight test ergonomics (non-null assertions after expects, async mock
signatures, non-Error throws).
Code adjusted to pass: registry disposers wrap ctx.effect's
promise-returning disposer behind a sync () => void (our public API),
BlockAssembler gains an invariant-checking mustGet instead of non-null
assertions, lastTurnNumber uses findLast, waterfall tails return
Promise.resolve instead of async-without-await arrows, and the two
deliberate suppressions (non-exhaustive derivation switch, unbound
execute pass-through) carry justification comments.
yarn lint / yarn lint:fix added.
2026-06-11 14:17:58 +08:00
|
|
|
agent.session.events.map(e => e.type))
|
2026-07-30 15:32:06 +08:00
|
|
|
expect(replayed.events.at(-1)?.type).toBe('session/end-seed')
|
Implement the agent loop plugin
@deepseek-ai/dsh-agent-loop: LoopAgent (inbox with queued + steering
FIFOs, per-step AbortController) and the streaming-first
session/turn/step loop. Extension seams: agent/request,
agent/step-result, agent/turn-continuation waterfalls; raw chunks
logged for replay while BlockAssembler builds the assembled message;
steering drains between steps; session/flush awaited at turn end.
16 tests with a scripted mock adapter cover turn lifecycle ordering,
tool round-trips, steering, inject(), continuation override/veto,
mid-stream abort, queued turn chaining, replay equivalence, and
mid-turn fiber disposal (HMR safety).
2026-06-11 10:54:31 +08:00
|
|
|
})
|
|
|
|
|
})
|