Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
import { mkdtempSync } from 'node:fs'
import { tmpdir } from 'node:os'
import { join } from 'node:path'
import { describe , expect , it , vi } from 'vitest'
import { Context } from 'cordis'
import { CallId } from '@deepseek-ai/dsh-llm'
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
import { BashExecutor , BashTaskId } from '@deepseek-ai/dsh-bash'
import type { BashExecRequest , BashExecSpec , BashRunResult , BashTask , BashTaskRead , OwnerToken } from '@deepseek-ai/dsh-bash'
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
import SystemPrompt from '@deepseek-ai/dsh-system-prompt'
import ToolRegistry from '@deepseek-ai/dsh-tools'
2026-06-20 08:14:27 +08:00
import AgentRegistry from '@deepseek-ai/dsh-agent'
import type { Agent } from '@deepseek-ai/dsh-agent'
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
import { LocalBashExecutor } from '@deepseek-ai/dsh-bash-local'
import * as ToolBash from '@deepseek-ai/dsh-tool-bash'
import { renderResult } from '@deepseek-ai/dsh-tool-bash'
const spillDir = mkdtempSync ( join ( tmpdir ( ) , 'dsh-tool-bash-spec-' ) )
async function setup() {
const ctx = new Context ( )
await ctx . plugin ( SystemPrompt )
await ctx . plugin ( ToolRegistry )
2026-06-20 08:14:27 +08:00
await ctx . plugin ( AgentRegistry )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
await ctx . plugin ( LocalBashExecutor , { timeoutMs : 10_000 } )
; ( ctx . bash as LocalBashExecutor ) . internals = { spillDir , graceMs : 200 }
await ctx . plugin ( ToolBash )
return ctx
}
2026-06-20 08:14:27 +08:00
/ * *
* Build a fake { @link Agent } whose session token is ` sessionId ` , REGISTER it in
* ` ctx.agents ` ( the completion - notice path finds the owning agent by scanning
* the registry for a matching ` session.header.id ` ) , and return it . The returned
* agent is also passed to ` execute ` as ` exec.agent ` so it owns the spawned task .
* The registration disposer is tracked so { @link unregisterFakeAgents } can drop
* it ( simulating the owning session disconnecting before a task completes ) .
* /
const fakeAgentDisposers = new Map < Context , ( ( ) = > void ) [ ] > ( )
function registerFakeAgent ( ctx : Context , sessionId : string , inject : ( . . . args : unknown [ ] ) = > void ) : Agent {
2026-06-20 09:57:05 +08:00
// The registry KEY (agent.id) is deliberately DIFFERENT from the session
// token (session.header.id) — a config agent has `agentId !== sessionId`. The
// owner token IS the session id, so the notice path must find the agent by
// `session.header.id`, NOT the registry key. Using distinct values here makes
// the test fail if a regression matched on the wrong field (a same-value fake
// would pass either way — the "hits the line but not the scenario" trap).
2026-06-21 11:08:10 +08:00
const agent = { id : ` agent- ${ sessionId } ` , inject , session : { header : { version : 0 , id : sessionId , createdAt : 0 } } } as unknown as Agent
2026-06-20 08:14:27 +08:00
const dispose = ctx . agents . register ( agent )
const list = fakeAgentDisposers . get ( ctx ) ? ? [ ]
list . push ( dispose )
fakeAgentDisposers . set ( ctx , list )
return agent
}
/** Unregister every fake agent in this ctx (simulate the owning session disconnecting). */
function unregisterFakeAgents ( ctx : Context ) : void {
for ( const dispose of fakeAgentDisposers . get ( ctx ) ? ? [ ] ) dispose ( )
fakeAgentDisposers . delete ( ctx )
}
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
let callCounter = 0
function call ( ctx : Context , name : string , args : unknown ) {
return ctx . tools . execute ( { callId : CallId ( ` call- ${ ++ callCounter } ` ) , name , arguments : args } )
}
function text ( result : { content : { type : string ; text? : string } [ ] } ) : string {
return result . content . filter ( block = > block . type === 'text' ) . map ( block = > block . text ) . join ( '' )
}
2026-06-23 16:27:49 +08:00
async function callUntilText (
ctx : Context ,
name : string ,
args : unknown ,
expected : string ,
timeoutMs = 5 _000 ,
) : Promise < Awaited < ReturnType < typeof call > > > {
const deadline = Date . now ( ) + timeoutMs
let last : Awaited < ReturnType < typeof call > > | undefined
while ( Date . now ( ) < deadline ) {
last = await call ( ctx , name , args )
if ( text ( last ) . includes ( expected ) ) return last
await new Promise ( resolve = > setTimeout ( resolve , 20 ) )
}
throw new Error ( ` ${ name } output did not include ${ JSON . stringify ( expected ) } ; last text was ${ JSON . stringify ( last !== undefined ? text ( last ) : '' ) } ` )
}
2026-06-19 01:54:57 +08:00
class LossyReadBashExecutor extends BashExecutor {
private readonly task : BashTask = {
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
id : BashTaskId ( 'bash-lossy' ) ,
2026-06-19 01:54:57 +08:00
command : 'fake' ,
status : 'running' ,
exitCode : null ,
signal : null ,
done : Promise.resolve ( ) ,
}
resolve ( request : BashExecRequest ) : BashExecSpec {
return {
command : request.command ,
workdir : request.workdir ? ? process . cwd ( ) ,
timeoutMs : request.timeoutMs ? ? 0 ,
. . . request . signal ? { signal : request.signal } : { } ,
2026-06-20 08:14:27 +08:00
owner : request.owner ,
2026-06-19 01:54:57 +08:00
}
}
run ( ) : Promise < BashRunResult > {
return Promise . reject ( new Error ( 'not used' ) )
}
start ( ) : BashTask {
return this . task
}
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
get ( id : BashTaskId ) : BashTask | undefined {
2026-06-19 01:54:57 +08:00
return id === this . task . id ? this . task : undefined
}
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
ownerOf ( ) : OwnerToken | undefined {
2026-06-20 08:14:27 +08:00
return undefined
}
2026-06-19 01:54:57 +08:00
list ( ) : BashTask [ ] {
return [ this . task ]
}
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
readOutput ( id : BashTaskId ) : BashTaskRead {
2026-06-19 01:54:57 +08:00
if ( id !== this . task . id ) throw new Error ( ` unknown bash task " ${ id } " ` )
return { task : this.task , delta : 'tail' , lossy : true }
}
kill ( ) : boolean {
return false
}
}
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
describe ( 'bash tool' , ( ) = > {
it ( 'returns stdout for a successful command' , async ( ) = > {
const ctx = await setup ( )
const result = await call ( ctx , 'bash' , { command : 'echo hello' , description : 'test command' } )
expect ( result . isError ) . toBe ( false )
expect ( text ( result ) ) . toBe ( 'hello\n' )
} )
it ( 'reports (no output) for silent commands' , async ( ) = > {
const ctx = await setup ( )
const result = await call ( ctx , 'bash' , { command : 'true' , description : 'test command' } )
expect ( text ( result ) ) . toBe ( '(no output)' )
} )
it ( 'marks stderr sections' , async ( ) = > {
const ctx = await setup ( )
const result = await call ( ctx , 'bash' , { command : 'echo out; echo err >&2' , description : 'test command' } )
expect ( text ( result ) ) . toBe ( 'out\n[stderr]\nerr\n' )
expect ( result . isError ) . toBe ( false )
} )
it ( 'reports non-zero exits without isError' , async ( ) = > {
const ctx = await setup ( )
const result = await call ( ctx , 'bash' , { command : 'echo failing; exit 3' , description : 'test command' } )
expect ( result . isError ) . toBe ( false )
expect ( text ( result ) ) . toBe ( 'failing\n[exit code: 3]' )
} )
it ( 'reports timeout kills with both markers (timeout first)' , async ( ) = > {
const ctx = await setup ( )
const result = await call ( ctx , 'bash' , { command : 'sleep 60' , description : 'test command' , timeoutMs : 100 } )
expect ( result . isError ) . toBe ( false )
expect ( text ( result ) ) . toBe ( '(no output)\n[timed out after 100ms]\n[killed by signal: SIGTERM]' )
} )
it ( 'reports a timeout even when the command traps the signal and exits 0' , async ( ) = > {
// The signal-independent timeout marker: a trapped SIGTERM that exits 0
// after our timer fired must NOT look like a clean success. (bash may
// print "Terminated" to stderr for the killed sleep — environment
// dependent — so assert the marker, not the exact body.)
const ctx = await setup ( )
const result = await call ( ctx , 'bash' , { command : 'trap "exit 0" TERM; sleep 60' , description : 'test command' , timeoutMs : 100 } )
expect ( result . isError ) . toBe ( false )
expect ( text ( result ) ) . toContain ( '[timed out after 100ms]' )
expect ( text ( result ) ) . not . toContain ( '[exit code:' )
} )
it ( 'reports truncation with the spill path' , async ( ) = > {
const ctx = new Context ( )
await ctx . plugin ( SystemPrompt )
await ctx . plugin ( ToolRegistry )
await ctx . plugin ( LocalBashExecutor , { maxOutputBytes : 100 } )
; ( ctx . bash as LocalBashExecutor ) . internals = { spillDir , graceMs : 200 }
await ctx . plugin ( ToolBash )
const result = await call ( ctx , 'bash' , { command : 'for i in $(seq 1 100); do printf "line-%04d\\n" $i; done' , description : 'test command' } )
expect ( text ( result ) ) . toContain ( '[output truncated; full output: ' )
expect ( text ( result ) ) . toContain ( 'line-0100' )
} )
it ( 'honors workdir' , async ( ) = > {
const ctx = await setup ( )
const result = await call ( ctx , 'bash' , { command : 'pwd' , description : 'test command' , workdir : '/tmp' } )
expect ( text ( result ) . trim ( ) ) . toMatch ( /\/tmp$/ )
} )
it ( 'surfaces spawn failures as isError' , async ( ) = > {
const ctx = await setup ( )
const result = await call ( ctx , 'bash' , { command : 'true' , description : 'test command' , workdir : '/nonexistent-dsh' } )
expect ( result . isError ) . toBe ( true )
expect ( text ( result ) ) . toMatch ( /ENOENT/ )
} )
it ( 'surfaces aborts as isError' , async ( ) = > {
const ctx = await setup ( )
const controller = new AbortController ( )
const pending = ctx . tools . execute ( {
callId : CallId ( 'call-abort' ) ,
name : 'bash' ,
arguments : { command : 'sleep 60' , description : 'test command' } ,
signal : controller.signal ,
} )
setTimeout ( ( ) = > { controller . abort ( ) } , 50 )
const result = await pending
expect ( result . isError ) . toBe ( true )
expect ( text ( result ) ) . toMatch ( /aborted/ )
} )
2026-06-13 23:00:42 +08:00
// Type and required-key violations are now rejected by the harness
2026-06-18 02:18:24 +08:00
// (defineTool validates against the SchemaSpec — the arg-validation RFC) before execute.
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
it . each ( [
2026-06-13 23:00:42 +08:00
[ { } , /missing required property "command"/ ] ,
[ { command : 42 , description : 'd' } , /"command" must be a string/ ] ,
[ { command : 'x' } , /missing required property "description"/ ] ,
[ { command : 'x' , description : 7 } , /"description" must be a string/ ] ,
[ { command : 'x' , description : 'd' , timeoutMs : 'soon' } , /"timeoutMs" must be a number/ ] ,
[ { command : 'x' , description : 'd' , workdir : 7 } , /"workdir" must be a string/ ] ,
[ { command : 'x' , description : 'd' , run_in_background : 'yes' } , /"run_in_background" must be a boolean/ ] ,
] ) ( 'rejects schema-invalid args %j' , async ( args , pattern ) = > {
const ctx = await setup ( )
const result = await call ( ctx , 'bash' , args )
expect ( result . isError ) . toBe ( true )
expect ( text ( result ) ) . toMatch ( pattern )
} )
// Value constraints the SchemaSpec can't express stay in the tool body.
it . each ( [
[ { command : ' ' , description : 'd' } , /invalid command/ ] ,
[ { command : 'x' , description : ' ' } , /invalid description/ ] ,
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
[ { command : 'x' , description : 'd' , timeoutMs : - 1 } , /invalid timeoutMs/ ] ,
[ { command : 'x' , description : 'd' , timeoutMs : Number.NaN } , /invalid timeoutMs/ ] ,
2026-06-13 23:00:42 +08:00
] ) ( 'rejects value-invalid args %j' , async ( args , pattern ) = > {
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
const ctx = await setup ( )
const result = await call ( ctx , 'bash' , args )
expect ( result . isError ) . toBe ( true )
expect ( text ( result ) ) . toMatch ( pattern )
} )
it ( 'registers all three schemas in the system prompt assembly' , async ( ) = > {
const ctx = await setup ( )
const names = ctx . tools . schemas ( ) . map ( schema = > schema . name )
expect ( names ) . toEqual ( [ 'bash' , 'bash_output' , 'bash_kill' ] )
const bashSchema = ctx . tools . schemas ( ) [ 0 ] !
expect ( bashSchema . parameters ) . toMatchObject ( {
type : 'object' ,
required : [ 'command' , 'description' ] ,
} )
} )
it ( 'unregisters everything when the plugin fiber is disposed (HMR safety)' , async ( ) = > {
const ctx = new Context ( )
await ctx . plugin ( SystemPrompt )
await ctx . plugin ( ToolRegistry )
await ctx . plugin ( LocalBashExecutor , { } )
const fiber = await ctx . plugin ( ToolBash )
expect ( ctx . tools . schemas ( ) ) . toHaveLength ( 3 )
await fiber . dispose ( )
expect ( ctx . tools . schemas ( ) ) . toHaveLength ( 0 )
} )
it ( 'tools depend on the executor: no registration without ctx.bash' , async ( ) = > {
const ctx = new Context ( )
await ctx . plugin ( SystemPrompt )
await ctx . plugin ( ToolRegistry )
// inject: ['tools', 'bash'] keeps the plugin pending until bash exists.
await ctx . plugin ( ToolBash )
expect ( ctx . tools . schemas ( ) ) . toHaveLength ( 0 )
await ctx . plugin ( LocalBashExecutor , { } )
await new Promise ( resolve = > setTimeout ( resolve , 0 ) )
expect ( ctx . tools . schemas ( ) ) . toHaveLength ( 3 )
} )
} )
describe ( 'background tools' , ( ) = > {
it ( 'bash with run_in_background returns a task id immediately' , async ( ) = > {
const ctx = await setup ( )
const result = await call ( ctx , 'bash' , { command : 'sleep 0.2; echo bg-done' , description : 'test command' , run_in_background : true } )
expect ( result . isError ) . toBe ( false )
expect ( text ( result ) ) . toMatch ( /^started background task bash-\d+$/ )
} )
it ( 'bash_output polls incrementally and reports status' , async ( ) = > {
const ctx = await setup ( )
2026-06-23 16:27:49 +08:00
const started = await call ( ctx , 'bash' , { command : 'echo first; sleep 1; echo second' , description : 'test command' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
2026-06-23 16:27:49 +08:00
const first = await callUntilText ( ctx , 'bash_output' , { task_id : id } , 'first' )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
expect ( text ( first ) ) . toContain ( 'first' )
expect ( text ( first ) ) . toContain ( '[status: running]' )
await ctx . bash . get ( id ) ! . done
const second = await call ( ctx , 'bash_output' , { task_id : id } )
expect ( text ( second ) ) . toContain ( 'second' )
expect ( text ( second ) ) . not . toContain ( 'first' )
expect ( text ( second ) ) . toContain ( '[status: completed, exit code: 0]' )
const third = await call ( ctx , 'bash_output' , { task_id : id } )
expect ( text ( third ) ) . toContain ( '(no new output)' )
} )
it ( 'bash_output flags lossy reads with spill paths' , async ( ) = > {
const ctx = new Context ( )
await ctx . plugin ( SystemPrompt )
await ctx . plugin ( ToolRegistry )
await ctx . plugin ( LocalBashExecutor , { maxOutputBytes : 100 } )
; ( ctx . bash as LocalBashExecutor ) . internals = { spillDir , graceMs : 200 }
await ctx . plugin ( ToolBash )
const started = await call ( ctx , 'bash' , { command : 'for i in $(seq 1 200); do printf "line-%04d\\n" $i; done' , description : 'test command' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
await ctx . bash . get ( id ) ! . done
const read = await call ( ctx , 'bash_output' , { task_id : id } )
expect ( text ( read ) ) . toContain ( '[some output was dropped from memory; full output: ' )
} )
2026-06-19 01:54:57 +08:00
it ( 'bash_output reports unavailable when a lossy read has no safe spill path' , async ( ) = > {
const ctx = new Context ( )
await ctx . plugin ( SystemPrompt )
await ctx . plugin ( ToolRegistry )
await ctx . plugin ( LossyReadBashExecutor )
await ctx . plugin ( ToolBash )
const read = await call ( ctx , 'bash_output' , { task_id : 'bash-lossy' } )
expect ( text ( read ) ) . toBe ( 'tail\n[some output was dropped from memory; full output: (unavailable)]\n[status: running]' )
} )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
it ( 'bash_kill stops a running task; repeat reports already-finished' , async ( ) = > {
const ctx = await setup ( )
const started = await call ( ctx , 'bash' , { command : 'sleep 60' , description : 'test command' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
const killed = await call ( ctx , 'bash_kill' , { task_id : id } )
expect ( text ( killed ) ) . toBe ( ` killed background task ${ id } ` )
await ctx . bash . get ( id ) ! . done
const again = await call ( ctx , 'bash_kill' , { task_id : id } )
expect ( text ( again ) ) . toBe ( ` task ${ id } had already finished ` )
const status = await call ( ctx , 'bash_output' , { task_id : id } )
expect ( text ( status ) ) . toContain ( '[status: killed by SIGTERM]' )
} )
it ( 'unknown task ids are isError for both tools' , async ( ) = > {
const ctx = await setup ( )
const read = await call ( ctx , 'bash_output' , { task_id : 'bash-999' } )
expect ( read . isError ) . toBe ( true )
expect ( text ( read ) ) . toMatch ( /unknown bash task/ )
const kill = await call ( ctx , 'bash_kill' , { task_id : 'bash-999' } )
expect ( kill . isError ) . toBe ( true )
} )
it . each ( [
2026-06-13 23:00:42 +08:00
[ 'bash_output' , { } , /missing required property "task_id"/ ] ,
[ 'bash_output' , { task_id : 9 } , /"task_id" must be a string/ ] ,
[ 'bash_kill' , { task_id : '' } , /invalid task_id/ ] ,
] ) ( '%s rejects invalid task_id %j' , async ( tool , args , pattern ) = > {
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
const ctx = await setup ( )
const result = await call ( ctx , tool , args )
expect ( result . isError ) . toBe ( true )
2026-06-13 23:00:42 +08:00
expect ( text ( result ) ) . toMatch ( pattern )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
} )
2026-06-20 08:14:27 +08:00
it ( 'injects a completion notice into the owning agent (found via the registry by session token)' , async ( ) = > {
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
const ctx = await setup ( )
const inject = vi . fn ( )
2026-06-20 08:14:27 +08:00
// The notice path looks the agent up in ctx.agents by its session token, so
// the agent must be REGISTERED (not merely passed to execute). Mount a
// registry and register a fake whose session.header.id IS the owner token.
const agent = registerFakeAgent ( ctx , 'bg' , inject )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
const started = await ctx . tools . execute ( {
callId : CallId ( 'call-bg' ) ,
name : 'bash' ,
arguments : { command : 'true' , description : 'test command' , run_in_background : true } ,
agent ,
} )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
await ctx . bash . get ( id ) ! . done
expect ( inject ) . toHaveBeenCalledTimes ( 1 )
const [ content , options ] = inject . mock . calls [ 0 ] as [
{ type : string ; text : string } [ ] ,
{ source : { kind : string ; plugin : string } } ,
]
expect ( content [ 0 ] ! . text ) . toContain ( ` background bash task ${ id } finished ` )
expect ( content [ 0 ] ! . text ) . toContain ( 'bash_output' )
expect ( options . source ) . toEqual ( { kind : 'plugin' , plugin : 'tool-bash' } )
} )
it ( 'swallows ONLY the disposed-agent inject error' , async ( ) = > {
const ctx = await setup ( )
2026-06-20 08:14:27 +08:00
const agent = registerFakeAgent ( ctx , 'bg' , ( ) = > { throw new Error ( 'agent "x" is disposed' ) } )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
const started = await ctx . tools . execute ( {
callId : CallId ( 'call-bg2' ) ,
name : 'bash' ,
arguments : { command : 'true' , description : 'test command' , run_in_background : true } ,
agent ,
} )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
await expect ( ctx . bash . get ( id ) ! . done ) . resolves . toBeUndefined ( )
} )
it ( 'rethrows a non-disposed inject failure (not blindly swallowed)' , async ( ) = > {
const ctx = await setup ( )
// A real bug in inject (not the benign disposed race) must surface — the
// base-class notifier contains it (logs, does not reject task.done), but
// the listener itself must have thrown rather than silently eaten it.
const errorSpy = vi . spyOn ( console , 'error' ) . mockImplementation ( ( ) = > undefined )
try {
2026-06-20 08:14:27 +08:00
const agent = registerFakeAgent ( ctx , 'bg' , ( ) = > { throw new Error ( 'unexpected inject bug' ) } )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
const started = await ctx . tools . execute ( {
callId : CallId ( 'call-bg3' ) ,
name : 'bash' ,
arguments : { command : 'true' , description : 'test command' , run_in_background : true } ,
agent ,
} )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
await ctx . bash . get ( id ) ! . done
// notifyTaskDone caught and logged the rethrown error.
expect ( errorSpy ) . toHaveBeenCalled ( )
const logged = errorSpy . mock . calls . flat ( ) . some ( arg = > arg instanceof Error && arg . message === 'unexpected inject bug' )
expect ( logged ) . toBe ( true )
} finally {
errorSpy . mockRestore ( )
}
} )
2026-06-20 08:14:27 +08:00
it ( 'drops the notice cleanly when the owning agent is gone from the registry by completion' , async ( ) = > {
// A bash task (owned by the host-scoped bash-local fiber) can OUTLIVE its
// per-session agent — e.g. the ACP session disconnects and its AgentHandle
// disposes while the background task is still running. The owner token is
// still on the task, but no live agent carries it anymore, so the registry
// lookup finds nothing and the notice is dropped (no throw).
const ctx = await setup ( )
const inject = vi . fn ( )
const agent = registerFakeAgent ( ctx , 'bg' , inject )
const started = await ctx . tools . execute ( {
callId : CallId ( 'call-bg4' ) ,
name : 'bash' ,
arguments : { command : 'true' , description : 'test command' , run_in_background : true } ,
agent ,
} )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
2026-06-20 08:14:27 +08:00
// Unregister the agent BEFORE the task completes (simulate disconnect).
unregisterFakeAgents ( ctx )
2026-06-21 06:17:11 +08:00
await expect ( ctx . bash . get ( id ) ! . done ) . resolves . toBeUndefined ( )
2026-06-20 08:14:27 +08:00
expect ( inject ) . not . toHaveBeenCalled ( )
} )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
it ( 'does not notify when no agent owned the task' , async ( ) = > {
const ctx = await setup ( )
const started = await call ( ctx , 'bash' , { command : 'true' , description : 'test command' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
await expect ( ctx . bash . get ( id ) ! . done ) . resolves . toBeUndefined ( )
} )
} )
2026-06-16 19:23:21 +08:00
describe ( 'background task ownership (cross-session isolation)' , ( ) = > {
/** Run a tool on behalf of a specific agent (sets exec.agent). */
function callAs ( ctx : Context , agent : import ( '@deepseek-ai/dsh-agent' ) . Agent | undefined , name : string , args : unknown ) {
return ctx . tools . execute ( { callId : CallId ( ` own- ${ ++ callCounter } ` ) , name , arguments : args , . . . agent ? { agent } : { } } )
}
2026-06-20 08:14:27 +08:00
// Ownership is by TOKEN (session.header.id), NOT agent object identity — so
// each agent needs a DISTINCT session id, else every fake yields the same
// token and the isolation tests pass for the wrong reason (all tasks owned by
// the same token). The impl reads `session.header.id`, so the fakes MUST carry
// it.
const fakeAgent = ( sessionId : string ) = >
2026-06-21 11:08:10 +08:00
( { inject : ( ) = > undefined , session : { header : { version : 0 , id : sessionId , createdAt : 0 } } } ) as unknown as import ( '@deepseek-ai/dsh-agent' ) . Agent
2026-06-20 08:14:27 +08:00
it ( 'rejects bash_output/bash_kill for a task owned by a DIFFERENT session token' , async ( ) = > {
2026-06-16 19:23:21 +08:00
const ctx = await setup ( )
2026-06-20 08:14:27 +08:00
const a = fakeAgent ( 'sess-a' )
const b = fakeAgent ( 'sess-b' )
2026-06-16 19:23:21 +08:00
// Agent A starts a long-running background task.
const started = await callAs ( ctx , a , 'bash' , { command : 'sleep 60' , description : 'bg' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
2026-06-16 19:23:21 +08:00
2026-06-20 08:14:27 +08:00
// Agent B (a different session token) cannot read or kill A's task.
2026-06-16 19:23:21 +08:00
const readByB = await callAs ( ctx , b , 'bash_output' , { task_id : id } )
expect ( readByB . isError ) . toBe ( true )
expect ( text ( readByB ) ) . toMatch ( /belongs to another session/ )
const killByB = await callAs ( ctx , b , 'bash_kill' , { task_id : id } )
expect ( killByB . isError ) . toBe ( true )
expect ( text ( killByB ) ) . toMatch ( /belongs to another session/ )
// The task is still running (B's kill did nothing) — A can still kill it.
const killByA = await callAs ( ctx , a , 'bash_kill' , { task_id : id } )
expect ( killByA . isError ) . toBe ( false )
expect ( text ( killByA ) ) . toBe ( ` killed background task ${ id } ` )
} )
2026-06-20 13:52:02 +08:00
it ( 'a DIFFERENT Agent object with the SAME session token may access the task (ownership is by token, not object identity)' , async ( ) = > {
// Ownership fences by session.header.id, NOT Agent object identity. Two
// distinct Agent objects sharing one session token (e.g. an agent re-created
// on the same session) are the SAME owner.
2026-06-20 08:14:27 +08:00
const ctx = await setup ( )
const a1 = fakeAgent ( 'sess-shared' )
const a2 = fakeAgent ( 'sess-shared' ) // distinct object, same token
const started = await callAs ( ctx , a1 , 'bash' , { command : 'sleep 60' , description : 'bg' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
2026-06-20 08:14:27 +08:00
const readByA2 = await callAs ( ctx , a2 , 'bash_output' , { task_id : id } )
expect ( readByA2 . isError ) . toBe ( false )
await callAs ( ctx , a1 , 'bash_kill' , { task_id : id } ) // cleanup
} )
2026-06-16 19:23:21 +08:00
it ( 'the no-agent (non-loop) caller cannot access an owned task' , async ( ) = > {
const ctx = await setup ( )
2026-06-20 08:14:27 +08:00
const a = fakeAgent ( 'sess-a' )
2026-06-16 19:23:21 +08:00
const started = await callAs ( ctx , a , 'bash' , { command : 'sleep 60' , description : 'bg' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
2026-06-20 08:14:27 +08:00
// A call with no exec.agent has no token → cannot prove ownership of an owned task.
2026-06-16 19:23:21 +08:00
const read = await callAs ( ctx , undefined , 'bash_output' , { task_id : id } )
expect ( read . isError ) . toBe ( true )
expect ( text ( read ) ) . toMatch ( /belongs to another session/ )
await callAs ( ctx , a , 'bash_kill' , { task_id : id } ) // cleanup
} )
it ( 'an UNOWNED task (started with no agent) is accessible to anyone' , async ( ) = > {
const ctx = await setup ( )
2026-06-20 08:14:27 +08:00
// Started by a non-loop caller (no exec.agent) → no owner token recorded.
2026-06-16 19:23:21 +08:00
const started = await callAs ( ctx , undefined , 'bash' , { command : 'sleep 60' , description : 'bg' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
2026-06-16 19:23:21 +08:00
// Any agent (and the no-agent caller) may read/kill it.
2026-06-20 08:14:27 +08:00
const read = await callAs ( ctx , fakeAgent ( 'sess-x' ) , 'bash_output' , { task_id : id } )
2026-06-16 19:23:21 +08:00
expect ( read . isError ) . toBe ( false )
const killed = await callAs ( ctx , undefined , 'bash_kill' , { task_id : id } )
expect ( killed . isError ) . toBe ( false )
} )
2026-06-20 08:14:27 +08:00
it ( 'the owner can still access its task AFTER it completes (owner token persists on the task)' , async ( ) = > {
2026-06-16 19:23:21 +08:00
const ctx = await setup ( )
2026-06-20 08:14:27 +08:00
const a = fakeAgent ( 'sess-a' )
const b = fakeAgent ( 'sess-b' )
2026-06-16 19:23:21 +08:00
const started = await callAs ( ctx , a , 'bash' , { command : 'echo done' , description : 'bg' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
2026-06-16 19:23:21 +08:00
await ctx . bash . get ( id ) ! . done
// Completion does NOT clear ownership: B is still rejected, A still allowed.
const readByB = await callAs ( ctx , b , 'bash_output' , { task_id : id } )
expect ( readByB . isError ) . toBe ( true )
expect ( text ( readByB ) ) . toMatch ( /belongs to another session/ )
const readByA = await callAs ( ctx , a , 'bash_output' , { task_id : id } )
expect ( readByA . isError ) . toBe ( false )
} )
2026-06-20 08:14:27 +08:00
it ( 'ownership SURVIVES an independent tool-bash HMR reload (token lives on the executor)' , async ( ) = > {
// The owner token lives on the TASK inside the executor (dsh-bash fiber), NOT
// in a tool-bash plugin-local map. So reloading ONLY tool-bash (executor +
2026-06-20 13:52:02 +08:00
// task survive) preserves ownership. This is the regression guard: a
// plugin-local map would make B accessible after reload, and this test would
// catch it.
2026-06-16 19:23:21 +08:00
const ctx = new Context ( )
await ctx . plugin ( SystemPrompt )
await ctx . plugin ( ToolRegistry )
await ctx . plugin ( LocalBashExecutor , { timeoutMs : 10_000 } )
; ( ctx . bash as LocalBashExecutor ) . internals = { spillDir , graceMs : 200 }
const fiber = await ctx . plugin ( ToolBash )
2026-06-20 08:14:27 +08:00
const a = fakeAgent ( 'sess-a' )
const b = fakeAgent ( 'sess-b' )
2026-06-16 19:23:21 +08:00
const started = await callAs ( ctx , a , 'bash' , { command : 'sleep 60' , description : 'bg' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
2026-06-16 19:23:21 +08:00
// Before reload: B is rejected (A owns it).
expect ( ( await callAs ( ctx , b , 'bash_output' , { task_id : id } ) ) . isError ) . toBe ( true )
2026-06-20 08:14:27 +08:00
// Reload ONLY tool-bash; the executor and its running task (with its owner
// token) survive.
2026-06-16 19:23:21 +08:00
await fiber . dispose ( )
await ctx . plugin ( ToolBash )
expect ( ctx . bash . get ( id ) ? . status ) . toBe ( 'running' )
2026-06-20 08:14:27 +08:00
expect ( ctx . bash . ownerOf ( id ) ) . toBe ( 'sess-a' )
2026-06-16 19:23:21 +08:00
2026-06-20 08:14:27 +08:00
// After reload, ownership is INTACT → B is STILL rejected.
expect ( ( await callAs ( ctx , b , 'bash_output' , { task_id : id } ) ) . isError ) . toBe ( true )
await callAs ( ctx , a , 'bash_kill' , { task_id : id } ) // cleanup
2026-06-16 19:23:21 +08:00
} )
} )
2026-06-17 10:01:18 +08:00
describe ( 'session-cwd routing (per-session workdir)' , ( ) = > {
function callAs ( ctx : Context , agent : import ( '@deepseek-ai/dsh-agent' ) . Agent | undefined , args : unknown ) {
return ctx . tools . execute ( { callId : CallId ( ` cwd- ${ ++ callCounter } ` ) , name : 'bash' , arguments : args , . . . agent ? { agent } : { } } )
}
// An agent whose session header carries a cwd (what session/new records).
const agentInCwd = ( cwd : string ) = >
2026-06-21 11:08:10 +08:00
( { inject : ( ) = > undefined , session : { header : { version : 0 , id : 'c' , createdAt : 0 , cwd } } } ) as unknown as import ( '@deepseek-ai/dsh-agent' ) . Agent
2026-06-17 10:01:18 +08:00
it ( 'defaults bash to the agent\'s session cwd (not the server launch dir)' , async ( ) = > {
const ctx = await setup ( )
const result = await callAs ( ctx , agentInCwd ( '/tmp' ) , { command : 'pwd' , description : 'pwd' } )
expect ( text ( result ) . trim ( ) ) . toMatch ( /\/tmp$/ )
} )
it ( 'an explicit absolute workdir overrides the session cwd' , async ( ) = > {
const ctx = await setup ( )
const result = await callAs ( ctx , agentInCwd ( '/' ) , { command : 'pwd' , description : 'pwd' , workdir : '/tmp' } )
expect ( text ( result ) . trim ( ) ) . toMatch ( /\/tmp$/ )
} )
it ( 'a relative workdir is resolved against the session cwd' , async ( ) = > {
const ctx = await setup ( )
// session cwd /usr + relative 'bin' → /usr/bin
const result = await callAs ( ctx , agentInCwd ( '/usr' ) , { command : 'pwd' , description : 'pwd' , workdir : 'bin' } )
expect ( text ( result ) . trim ( ) ) . toMatch ( /\/usr\/bin$/ )
} )
it ( 'two sessions with different cwds each run bash in their own dir' , async ( ) = > {
const ctx = await setup ( )
const inUsr = await callAs ( ctx , agentInCwd ( '/usr' ) , { command : 'pwd' , description : 'pwd' } )
const inTmp = await callAs ( ctx , agentInCwd ( '/tmp' ) , { command : 'pwd' , description : 'pwd' } )
expect ( text ( inUsr ) . trim ( ) ) . toMatch ( /\/usr$/ )
expect ( text ( inTmp ) . trim ( ) ) . toMatch ( /\/tmp$/ )
} )
it ( 'falls back to the executor default when the agent has no session cwd' , async ( ) = > {
const ctx = await setup ( )
// No exec.agent at all → executor uses its config/process.cwd() default.
const result = await ctx . tools . execute ( { callId : CallId ( 'cwd-noagent' ) , name : 'bash' , arguments : { command : 'pwd' , description : 'pwd' } } )
expect ( result . isError ) . toBe ( false )
expect ( text ( result ) . trim ( ) . length ) . toBeGreaterThan ( 0 )
} )
} )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
describe ( 'renderResult' , ( ) = > {
const base = {
exitCode : 0 as number | null ,
signal : null as NodeJS . Signals | null ,
timedOut : false ,
aborted : false ,
timeoutMs : 1000 ,
stdout : { text : '' , truncated : false } ,
stderr : { text : '' , truncated : false } ,
}
it ( 'renders stderr-only output without a stdout prefix' , ( ) = > {
expect ( renderResult ( { . . . base , stderr : { text : 'err\n' , truncated : false } } ) )
. toBe ( '[stderr]\nerr\n' )
} )
it ( 'adds a separator when stdout does not end with a newline' , ( ) = > {
expect ( renderResult ( {
. . . base ,
stdout : { text : 'out' , truncated : false } ,
stderr : { text : 'err' , truncated : false } ,
} ) ) . toBe ( 'out\n[stderr]\nerr' )
} )
it ( 'appends exit-code markers after a newline for unterminated output' , ( ) = > {
expect ( renderResult ( { . . . base , exitCode : 7 , stdout : { text : 'x' , truncated : false } } ) )
. toBe ( 'x\n[exit code: 7]' )
} )
it ( 'renders signal kills without the timeout marker when not timed out' , ( ) = > {
expect ( renderResult ( { . . . base , exitCode : null , signal : 'SIGKILL' } ) )
. toBe ( '(no output)\n[killed by signal: SIGKILL]' )
} )
it ( 'reports a timeout that exited 0 (trapped signal) without a kill marker' , ( ) = > {
expect ( renderResult ( { . . . base , exitCode : 0 , signal : null , timedOut : true } ) )
. toBe ( '(no output)\n[timed out after 1000ms]' )
} )
it ( 'orders the timeout marker before a kill marker' , ( ) = > {
expect ( renderResult ( { . . . base , exitCode : null , signal : 'SIGTERM' , timedOut : true } ) )
. toBe ( '(no output)\n[timed out after 1000ms]\n[killed by signal: SIGTERM]' )
} )
it ( 'notes truncation with a fallback when the spill path is missing' , ( ) = > {
expect ( renderResult ( { . . . base , stdout : { text : 'tail' , truncated : true } } ) )
. toBe ( 'tail\n[output truncated; full output: (unavailable)]' )
} )
} )
describe ( 'status lines' , ( ) = > {
it ( 'reports kills without a recorded signal (executor raced process exit)' , async ( ) = > {
const ctx = await setup ( )
const started = await call ( ctx , 'bash' , { command : 'sleep 60' , description : 'test command' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
const task = ctx . bash . get ( id ) !
await call ( ctx , 'bash_kill' , { task_id : id } )
await task . done
// Simulate the variant where the close event carried no signal.
task . signal = null
const read = await call ( ctx , 'bash_output' , { task_id : id } )
expect ( text ( read ) ) . toContain ( '[status: killed]' )
} )
it ( 'reports completed tasks with a null exit code as exit 0' , async ( ) = > {
const ctx = await setup ( )
const started = await call ( ctx , 'bash' , { command : 'true' , description : 'test command' , run_in_background : true } )
feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.
- Extract the `Branded<B>` primitive into a new standalone type-only package
`@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
at the config boundary and the inner create()/resume casts disappear (only the
genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
Map keys and public params/exports (SessionStore, AgentRegistry + factory
options, the ACP session-id surface + ToolPresenter CallId map, the
persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
the Branded type-equiv at dsh-brand, fix stale param types in the session/
agent/bash READMEs, regenerate the cordis catalog + module graph.
Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:17:25 +08:00
const id = BashTaskId ( /task (bash-\d+)/ . exec ( text ( started ) ) ! [ 1 ] ! )
Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools
Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):
- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
runtime arg validation and background completion notices via
agent.inject(). Non-zero exits are reported, not errored.
Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
2026-06-12 23:28:44 +08:00
const task = ctx . bash . get ( id ) !
await task . done
// Defensive: completed tasks always carry an exit code in practice; the
// ?? 0 fallback covers task shapes from other executor implementations.
task . exitCode = null
const read = await call ( ctx , 'bash_output' , { task_id : id } )
expect ( text ( read ) ) . toContain ( '[status: completed, exit code: 0]' )
} )
} )
feat(acp): tool-owned tool-call UI presentation (title/command/output)
In Zed the tool-call card showed only "bash" — the bare tool name — instead
of what the command does. Fix it by letting each TOOL own how its calls render,
rather than the bridge special-casing names.
dsh-tools: add an optional two-state presentation seam to ToolDefinition /
defineTool — `presentCall(args)` (pending: title, kind, rawInput) and
`presentResult(args, result)` (completed: title?, content?). Provider-neutral
`ToolCallKind`/`ToolCallPresentation`/`ToolResultPresentation` vocabulary so
tools never depend on ACP. defineTool soft-validates args (display runs on log
replay, so a malformed/old shape returns undefined instead of throwing).
dsh-tool-bash: bash declares presentCall (model `description` → title, exact
`command` → rawInput, kind execute) and presentResult (wrap output in a fenced
```console block — a UI-only affordance kept out of the model-facing result);
bash_output/bash_kill present task-scoped titles.
dsh-acp: inject `tools`; a per-session `ToolPresenter` looks the tool up by name
and maps its neutral presentation to the ACP tool_call/tool_call_update wire
shape, with a generic fallback (title = name) for tools that declare nothing.
Because the `tool/result` event carries only {callId, content, isError}, the
presenter keeps a small bridge-local map of ONLY in-flight calls' (name, args),
keyed by callId and removed as each result is presented — no event-schema or
core change. Replay uses a throwaway presenter so loaded sessions render
identically to live ones.
Tests: dsh-tools defineTool presenters (typed args, soft-validate), tool-bash
bash/bash_output/bash_kill presenters, acp ToolPresenter (tool-owned mapping,
unknown-callId fallback, in-flight-only map), and an end-to-end turn through the
bridge. The key-gated e2e now asserts a real bash call's title is the model
description (not "bash") and rawInput is the command — verified against the real
DeepSeek model. The test harness derives its inject from the bridge's exported
`inject` so it can't drift again.
2026-06-18 09:01:36 +08:00
describe ( 'tool-owned UI presentation (presentCall / presentResult)' , ( ) = > {
2026-07-03 02:04:03 +08:00
it ( 'bash presentCall: a foreground run is a terminal card (command title, description, workdir → cwd absolute or relative)' , async ( ) = > {
feat(acp): tool-owned tool-call UI presentation (title/command/output)
In Zed the tool-call card showed only "bash" — the bare tool name — instead
of what the command does. Fix it by letting each TOOL own how its calls render,
rather than the bridge special-casing names.
dsh-tools: add an optional two-state presentation seam to ToolDefinition /
defineTool — `presentCall(args)` (pending: title, kind, rawInput) and
`presentResult(args, result)` (completed: title?, content?). Provider-neutral
`ToolCallKind`/`ToolCallPresentation`/`ToolResultPresentation` vocabulary so
tools never depend on ACP. defineTool soft-validates args (display runs on log
replay, so a malformed/old shape returns undefined instead of throwing).
dsh-tool-bash: bash declares presentCall (model `description` → title, exact
`command` → rawInput, kind execute) and presentResult (wrap output in a fenced
```console block — a UI-only affordance kept out of the model-facing result);
bash_output/bash_kill present task-scoped titles.
dsh-acp: inject `tools`; a per-session `ToolPresenter` looks the tool up by name
and maps its neutral presentation to the ACP tool_call/tool_call_update wire
shape, with a generic fallback (title = name) for tools that declare nothing.
Because the `tool/result` event carries only {callId, content, isError}, the
presenter keeps a small bridge-local map of ONLY in-flight calls' (name, args),
keyed by callId and removed as each result is presented — no event-schema or
core change. Replay uses a throwaway presenter so loaded sessions render
identically to live ones.
Tests: dsh-tools defineTool presenters (typed args, soft-validate), tool-bash
bash/bash_output/bash_kill presenters, acp ToolPresenter (tool-owned mapping,
unknown-callId fallback, in-flight-only map), and an end-to-end turn through the
bridge. The key-gated e2e now asserts a real bash call's title is the model
description (not "bash") and rawInput is the command — verified against the real
DeepSeek model. The test harness derives its inject from the bridge's exported
`inject` so it can't drift again.
2026-06-18 09:01:36 +08:00
const ctx = await setup ( )
2026-07-03 02:04:03 +08:00
// No explicit workdir → a terminal card with no cwd (the UI bridge fills the
// session cwd it owns; the pure presenter can't see it).
2026-06-18 17:25:09 +08:00
expect ( ctx . tools . get ( 'bash' ) ? . presentCall ? . ( { command : 'ls -la src' , description : 'List files in src' } ) )
2026-07-03 02:04:03 +08:00
. toEqual ( { card : 'terminal' , title : 'ls -la src' , description : 'List files in src' } )
2026-06-18 18:54:32 +08:00
// An ABSOLUTE workdir is surfaced verbatim as the terminal cwd header.
2026-06-18 17:25:09 +08:00
expect ( ctx . tools . get ( 'bash' ) ? . presentCall ? . ( { command : 'pwd' , description : 'Print dir' , workdir : '/tmp/x' } ) )
2026-07-03 02:04:03 +08:00
. toEqual ( { card : 'terminal' , title : 'pwd' , description : 'Print dir' , cwd : '/tmp/x' } )
2026-06-18 18:54:32 +08:00
// A RELATIVE workdir is passed through AS-IS (the bridge resolves it against
// the session cwd, matching where execution runs) — not dropped.
2026-06-18 17:25:09 +08:00
expect ( ctx . tools . get ( 'bash' ) ? . presentCall ? . ( { command : 'pwd' , description : 'Print dir' , workdir : 'sub' } ) )
2026-07-03 02:04:03 +08:00
. toEqual ( { card : 'terminal' , title : 'pwd' , description : 'Print dir' , cwd : 'sub' } )
feat(acp): tool-owned tool-call UI presentation (title/command/output)
In Zed the tool-call card showed only "bash" — the bare tool name — instead
of what the command does. Fix it by letting each TOOL own how its calls render,
rather than the bridge special-casing names.
dsh-tools: add an optional two-state presentation seam to ToolDefinition /
defineTool — `presentCall(args)` (pending: title, kind, rawInput) and
`presentResult(args, result)` (completed: title?, content?). Provider-neutral
`ToolCallKind`/`ToolCallPresentation`/`ToolResultPresentation` vocabulary so
tools never depend on ACP. defineTool soft-validates args (display runs on log
replay, so a malformed/old shape returns undefined instead of throwing).
dsh-tool-bash: bash declares presentCall (model `description` → title, exact
`command` → rawInput, kind execute) and presentResult (wrap output in a fenced
```console block — a UI-only affordance kept out of the model-facing result);
bash_output/bash_kill present task-scoped titles.
dsh-acp: inject `tools`; a per-session `ToolPresenter` looks the tool up by name
and maps its neutral presentation to the ACP tool_call/tool_call_update wire
shape, with a generic fallback (title = name) for tools that declare nothing.
Because the `tool/result` event carries only {callId, content, isError}, the
presenter keeps a small bridge-local map of ONLY in-flight calls' (name, args),
keyed by callId and removed as each result is presented — no event-schema or
core change. Replay uses a throwaway presenter so loaded sessions render
identically to live ones.
Tests: dsh-tools defineTool presenters (typed args, soft-validate), tool-bash
bash/bash_output/bash_kill presenters, acp ToolPresenter (tool-owned mapping,
unknown-callId fallback, in-flight-only map), and an end-to-end turn through the
bridge. The key-gated e2e now asserts a real bash call's title is the model
description (not "bash") and rawInput is the command — verified against the real
DeepSeek model. The test harness derives its inject from the bridge's exported
`inject` so it can't drift again.
2026-06-18 09:01:36 +08:00
} )
2026-07-03 02:04:03 +08:00
it ( 'bash presentResult: a terminal result carries RAW output (newlines intact) + parsed exit code' , async ( ) = > {
feat(acp): tool-owned tool-call UI presentation (title/command/output)
In Zed the tool-call card showed only "bash" — the bare tool name — instead
of what the command does. Fix it by letting each TOOL own how its calls render,
rather than the bridge special-casing names.
dsh-tools: add an optional two-state presentation seam to ToolDefinition /
defineTool — `presentCall(args)` (pending: title, kind, rawInput) and
`presentResult(args, result)` (completed: title?, content?). Provider-neutral
`ToolCallKind`/`ToolCallPresentation`/`ToolResultPresentation` vocabulary so
tools never depend on ACP. defineTool soft-validates args (display runs on log
replay, so a malformed/old shape returns undefined instead of throwing).
dsh-tool-bash: bash declares presentCall (model `description` → title, exact
`command` → rawInput, kind execute) and presentResult (wrap output in a fenced
```console block — a UI-only affordance kept out of the model-facing result);
bash_output/bash_kill present task-scoped titles.
dsh-acp: inject `tools`; a per-session `ToolPresenter` looks the tool up by name
and maps its neutral presentation to the ACP tool_call/tool_call_update wire
shape, with a generic fallback (title = name) for tools that declare nothing.
Because the `tool/result` event carries only {callId, content, isError}, the
presenter keeps a small bridge-local map of ONLY in-flight calls' (name, args),
keyed by callId and removed as each result is presented — no event-schema or
core change. Replay uses a throwaway presenter so loaded sessions render
identically to live ones.
Tests: dsh-tools defineTool presenters (typed args, soft-validate), tool-bash
bash/bash_output/bash_kill presenters, acp ToolPresenter (tool-owned mapping,
unknown-callId fallback, in-flight-only map), and an end-to-end turn through the
bridge. The key-gated e2e now asserts a real bash call's title is the model
description (not "bash") and rawInput is the command — verified against the real
DeepSeek model. The test harness derives its inject from the bridge's exported
`inject` so it can't drift again.
2026-06-18 09:01:36 +08:00
const ctx = await setup ( )
const present = ctx . tools . get ( 'bash' ) ! . presentResult ! (
{ command : 'echo hi' , description : 'echo' } ,
{ content : [ { type : 'text' , text : 'hi\n[exit code: 0]\n\n' } ] , isError : false } ,
)
2026-07-03 02:04:03 +08:00
// A terminal result keeps the RAW bytes (newlines intact) a terminal renderer
// needs; the bridge derives the fenced fallback. exitCode is parsed back from
// the [exit code: N] marker.
expect ( present ) . toEqual ( { card : 'terminal' , output : 'hi\n[exit code: 0]\n\n' , exitCode : 0 } )
feat(acp): tool-owned tool-call UI presentation (title/command/output)
In Zed the tool-call card showed only "bash" — the bare tool name — instead
of what the command does. Fix it by letting each TOOL own how its calls render,
rather than the bridge special-casing names.
dsh-tools: add an optional two-state presentation seam to ToolDefinition /
defineTool — `presentCall(args)` (pending: title, kind, rawInput) and
`presentResult(args, result)` (completed: title?, content?). Provider-neutral
`ToolCallKind`/`ToolCallPresentation`/`ToolResultPresentation` vocabulary so
tools never depend on ACP. defineTool soft-validates args (display runs on log
replay, so a malformed/old shape returns undefined instead of throwing).
dsh-tool-bash: bash declares presentCall (model `description` → title, exact
`command` → rawInput, kind execute) and presentResult (wrap output in a fenced
```console block — a UI-only affordance kept out of the model-facing result);
bash_output/bash_kill present task-scoped titles.
dsh-acp: inject `tools`; a per-session `ToolPresenter` looks the tool up by name
and maps its neutral presentation to the ACP tool_call/tool_call_update wire
shape, with a generic fallback (title = name) for tools that declare nothing.
Because the `tool/result` event carries only {callId, content, isError}, the
presenter keeps a small bridge-local map of ONLY in-flight calls' (name, args),
keyed by callId and removed as each result is presented — no event-schema or
core change. Replay uses a throwaway presenter so loaded sessions render
identically to live ones.
Tests: dsh-tools defineTool presenters (typed args, soft-validate), tool-bash
bash/bash_output/bash_kill presenters, acp ToolPresenter (tool-owned mapping,
unknown-callId fallback, in-flight-only map), and an end-to-end turn through the
bridge. The key-gated e2e now asserts a real bash call's title is the model
description (not "bash") and rawInput is the command — verified against the real
DeepSeek model. The test harness derives its inject from the bridge's exported
`inject` so it can't drift again.
2026-06-18 09:01:36 +08:00
} )
2026-06-18 18:54:32 +08:00
it ( 'bash presentResult: a non-zero exit and a signal kill parse into exitCode / signal' , async ( ) = > {
const ctx = await setup ( )
const args = { command : 'x' , description : 'x' }
const nonzero = ctx . tools . get ( 'bash' ) ! . presentResult ! ( args , { content : [ { type : 'text' , text : 'oops\n[exit code: 3]' } ] , isError : false } )
2026-07-03 02:04:03 +08:00
expect ( nonzero ) . toEqual ( { card : 'terminal' , output : 'oops\n[exit code: 3]' , exitCode : 3 } )
2026-06-18 18:54:32 +08:00
const killed = ctx . tools . get ( 'bash' ) ! . presentResult ! ( args , { content : [ { type : 'text' , text : 'gone\n[killed by signal: SIGKILL]' } ] , isError : false } )
2026-07-03 02:04:03 +08:00
expect ( killed ) . toEqual ( { card : 'terminal' , output : 'gone\n[killed by signal: SIGKILL]' , signal : 'SIGKILL' } )
2026-06-18 18:54:32 +08:00
} )
it ( 'bash presentResult exit parse is the inverse of renderResult markers (round-trip)' , async ( ) = > {
const ctx = await setup ( )
const present = ctx . tools . get ( 'bash' ) !
// For each renderResult outcome, the rendered text fed back through
// presentResult recovers the matching structured exit — the parse and the
// marker emission co-evolve in one file, so this pins the pair.
const base = {
aborted : false ,
timeoutMs : 1000 ,
stdout : { text : 'out' , truncated : false } ,
stderr : { text : '' , truncated : false } ,
}
const cases = [
{ result : { . . . base , exitCode : 0 , signal : null , timedOut : false } , expect : { exitCode : 0 } } ,
{ result : { . . . base , exitCode : 7 , signal : null , timedOut : false } , expect : { exitCode : 7 } } ,
{ result : { . . . base , exitCode : null , signal : 'SIGTERM' as const , timedOut : false } , expect : { signal : 'SIGTERM' } } ,
// A trapped-timeout run that exits 0 has no signal/exit marker → reads as exit 0 (it did exit 0).
{ result : { . . . base , exitCode : 0 , signal : null , timedOut : true } , expect : { exitCode : 0 } } ,
]
for ( const c of cases ) {
const rendered = renderResult ( c . result )
const out = present . presentResult ! ( { command : 'x' , description : 'x' } , { content : [ { type : 'text' , text : rendered } ] , isError : false } )
2026-07-03 02:04:03 +08:00
// Drop card + output; the remaining fields are the parsed exit.
const { card : _c , output : _o , . . . exit } = out as { card : string ; output? : string ; exitCode? : number ; signal? : string }
2026-06-18 18:54:32 +08:00
expect ( exit ) . toEqual ( c . expect )
}
} )
2026-06-18 19:35:15 +08:00
it ( 'bash presentResult: a clean exit-0 whose output ENDS in marker-like text is NOT read as a failure' , async ( ) = > {
const ctx = await setup ( )
const args = { command : 'printf "[exit code: 5]"' , description : 'print' }
// A successful command can print text that looks like a marker. renderResult
// for a clean exit 0 appends NOTHING (and no trailing newline), so the body's
// own tail is `[exit code: 5]`. The parse requires a LEADING newline before
// the marker (renderResult always inserts one before a REAL marker), so this
// no-trailing-newline body is NOT mistaken for a failure → exitCode 0.
const out = ctx . tools . get ( 'bash' ) ! . presentResult ! ( args , { content : [ { type : 'text' , text : '[exit code: 5]' } ] , isError : false } )
2026-07-03 02:04:03 +08:00
expect ( out ) . toEqual ( { card : 'terminal' , output : '[exit code: 5]' , exitCode : 0 } )
2026-06-18 19:35:15 +08:00
// Same for a fake signal marker with no leading newline.
const sig = ctx . tools . get ( 'bash' ) ! . presentResult ! ( args , { content : [ { type : 'text' , text : '[killed by signal: SIGKILL]' } ] , isError : false } )
2026-07-03 02:04:03 +08:00
expect ( sig ) . toEqual ( { card : 'terminal' , output : '[killed by signal: SIGKILL]' , exitCode : 0 } )
2026-06-18 19:35:15 +08:00
} )
2026-07-03 02:04:03 +08:00
it ( 'bash presentCall/presentResult: a run_in_background call is a generic card and its ack carries no exit pill' , async ( ) = > {
2026-06-18 19:35:15 +08:00
const ctx = await setup ( )
2026-07-03 02:04:03 +08:00
// The background start returns a task-id ack, not a streamed run — a generic
// execute card with the command as rawInput and the description as content.
2026-06-18 19:35:15 +08:00
const call = ctx . tools . get ( 'bash' ) ! . presentCall ! ( { command : 'sleep 100' , description : 'wait' , run_in_background : true } )
2026-07-03 02:04:03 +08:00
expect ( call ) . toEqual ( { card : 'generic' , title : 'sleep 100' , kind : 'execute' , rawInput : 'sleep 100' , content : [ { type : 'text' , text : 'wait' } ] } )
// The ack result is a generic fenced-text card — no terminal output / exit pill.
2026-06-18 19:35:15 +08:00
const result = ctx . tools . get ( 'bash' ) ! . presentResult ! (
{ command : 'sleep 100' , description : 'wait' , run_in_background : true } ,
{ content : [ { type : 'text' , text : 'started background task bash-1' } ] , isError : false } ,
)
2026-07-03 02:04:03 +08:00
expect ( result ) . toEqual ( { card : 'generic' , content : [ { type : 'text' , text : '```console\nstarted background task bash-1\n```' } ] } )
2026-06-18 19:35:15 +08:00
} )
2026-07-03 02:04:03 +08:00
it ( 'bash presentResult: an isError result is a generic card (no real process exit to report)' , async ( ) = > {
2026-06-18 19:35:15 +08:00
const ctx = await setup ( )
// A spawn failure / abort has no process exit — the body is an error message,
2026-07-03 02:04:03 +08:00
// not renderResult output, so a generic fenced card, no terminal output/exit.
2026-06-18 19:35:15 +08:00
const out = ctx . tools . get ( 'bash' ) ! . presentResult ! (
{ command : 'x' , description : 'x' } ,
{ content : [ { type : 'text' , text : 'command aborted' } ] , isError : true } ,
)
2026-07-03 02:04:03 +08:00
expect ( out ) . toEqual ( { card : 'generic' , content : [ { type : 'text' , text : '```console\ncommand aborted\n```' } ] } )
2026-06-18 19:35:15 +08:00
} )
feat(acp): tool-owned tool-call UI presentation (title/command/output)
In Zed the tool-call card showed only "bash" — the bare tool name — instead
of what the command does. Fix it by letting each TOOL own how its calls render,
rather than the bridge special-casing names.
dsh-tools: add an optional two-state presentation seam to ToolDefinition /
defineTool — `presentCall(args)` (pending: title, kind, rawInput) and
`presentResult(args, result)` (completed: title?, content?). Provider-neutral
`ToolCallKind`/`ToolCallPresentation`/`ToolResultPresentation` vocabulary so
tools never depend on ACP. defineTool soft-validates args (display runs on log
replay, so a malformed/old shape returns undefined instead of throwing).
dsh-tool-bash: bash declares presentCall (model `description` → title, exact
`command` → rawInput, kind execute) and presentResult (wrap output in a fenced
```console block — a UI-only affordance kept out of the model-facing result);
bash_output/bash_kill present task-scoped titles.
dsh-acp: inject `tools`; a per-session `ToolPresenter` looks the tool up by name
and maps its neutral presentation to the ACP tool_call/tool_call_update wire
shape, with a generic fallback (title = name) for tools that declare nothing.
Because the `tool/result` event carries only {callId, content, isError}, the
presenter keeps a small bridge-local map of ONLY in-flight calls' (name, args),
keyed by callId and removed as each result is presented — no event-schema or
core change. Replay uses a throwaway presenter so loaded sessions render
identically to live ones.
Tests: dsh-tools defineTool presenters (typed args, soft-validate), tool-bash
bash/bash_output/bash_kill presenters, acp ToolPresenter (tool-owned mapping,
unknown-callId fallback, in-flight-only map), and an end-to-end turn through the
bridge. The key-gated e2e now asserts a real bash call's title is the model
description (not "bash") and rawInput is the command — verified against the real
DeepSeek model. The test harness derives its inject from the bridge's exported
`inject` so it can't drift again.
2026-06-18 09:01:36 +08:00
it ( 'bash presentResult: leaves a non-text (unexpected) result untouched → undefined (UI keeps raw content)' , async ( ) = > {
const ctx = await setup ( )
const present = ctx . tools . get ( 'bash' ) ! . presentResult ! (
{ command : 'x' , description : 'x' } ,
{ content : [ { type : 'image' , url : 'https://x/y.png' } ] , isError : false } ,
)
expect ( present ) . toBeUndefined ( )
} )
it ( 'bash presentResult: a result that is not exactly one block → undefined (no single text to fence)' , async ( ) = > {
const ctx = await setup ( )
const args = { command : 'x' , description : 'x' }
// Empty content (no block) and multi-block content both fall through.
expect ( ctx . tools . get ( 'bash' ) ! . presentResult ! ( args , { content : [ ] , isError : false } ) ) . toBeUndefined ( )
expect ( ctx . tools . get ( 'bash' ) ! . presentResult ! ( args , {
content : [ { type : 'text' , text : 'a' } , { type : 'text' , text : 'b' } ] ,
isError : false ,
} ) ) . toBeUndefined ( )
} )
it ( 'bash_output / bash_kill presentCall: a readable task-scoped title, task id as rawInput' , async ( ) = > {
const ctx = await setup ( )
expect ( ctx . tools . get ( 'bash_output' ) ! . presentCall ! ( { task_id : 'bash-3' } ) )
2026-07-03 02:04:03 +08:00
. toEqual ( { card : 'generic' , title : 'Read output from background task bash-3' , kind : 'execute' , rawInput : 'bash-3' } )
feat(acp): tool-owned tool-call UI presentation (title/command/output)
In Zed the tool-call card showed only "bash" — the bare tool name — instead
of what the command does. Fix it by letting each TOOL own how its calls render,
rather than the bridge special-casing names.
dsh-tools: add an optional two-state presentation seam to ToolDefinition /
defineTool — `presentCall(args)` (pending: title, kind, rawInput) and
`presentResult(args, result)` (completed: title?, content?). Provider-neutral
`ToolCallKind`/`ToolCallPresentation`/`ToolResultPresentation` vocabulary so
tools never depend on ACP. defineTool soft-validates args (display runs on log
replay, so a malformed/old shape returns undefined instead of throwing).
dsh-tool-bash: bash declares presentCall (model `description` → title, exact
`command` → rawInput, kind execute) and presentResult (wrap output in a fenced
```console block — a UI-only affordance kept out of the model-facing result);
bash_output/bash_kill present task-scoped titles.
dsh-acp: inject `tools`; a per-session `ToolPresenter` looks the tool up by name
and maps its neutral presentation to the ACP tool_call/tool_call_update wire
shape, with a generic fallback (title = name) for tools that declare nothing.
Because the `tool/result` event carries only {callId, content, isError}, the
presenter keeps a small bridge-local map of ONLY in-flight calls' (name, args),
keyed by callId and removed as each result is presented — no event-schema or
core change. Replay uses a throwaway presenter so loaded sessions render
identically to live ones.
Tests: dsh-tools defineTool presenters (typed args, soft-validate), tool-bash
bash/bash_output/bash_kill presenters, acp ToolPresenter (tool-owned mapping,
unknown-callId fallback, in-flight-only map), and an end-to-end turn through the
bridge. The key-gated e2e now asserts a real bash call's title is the model
description (not "bash") and rawInput is the command — verified against the real
DeepSeek model. The test harness derives its inject from the bridge's exported
`inject` so it can't drift again.
2026-06-18 09:01:36 +08:00
expect ( ctx . tools . get ( 'bash_kill' ) ! . presentCall ! ( { task_id : 'bash-3' } ) )
2026-07-03 02:04:03 +08:00
. toEqual ( { card : 'generic' , title : 'Kill background task bash-3' , kind : 'execute' , rawInput : 'bash-3' } )
feat(acp): tool-owned tool-call UI presentation (title/command/output)
In Zed the tool-call card showed only "bash" — the bare tool name — instead
of what the command does. Fix it by letting each TOOL own how its calls render,
rather than the bridge special-casing names.
dsh-tools: add an optional two-state presentation seam to ToolDefinition /
defineTool — `presentCall(args)` (pending: title, kind, rawInput) and
`presentResult(args, result)` (completed: title?, content?). Provider-neutral
`ToolCallKind`/`ToolCallPresentation`/`ToolResultPresentation` vocabulary so
tools never depend on ACP. defineTool soft-validates args (display runs on log
replay, so a malformed/old shape returns undefined instead of throwing).
dsh-tool-bash: bash declares presentCall (model `description` → title, exact
`command` → rawInput, kind execute) and presentResult (wrap output in a fenced
```console block — a UI-only affordance kept out of the model-facing result);
bash_output/bash_kill present task-scoped titles.
dsh-acp: inject `tools`; a per-session `ToolPresenter` looks the tool up by name
and maps its neutral presentation to the ACP tool_call/tool_call_update wire
shape, with a generic fallback (title = name) for tools that declare nothing.
Because the `tool/result` event carries only {callId, content, isError}, the
presenter keeps a small bridge-local map of ONLY in-flight calls' (name, args),
keyed by callId and removed as each result is presented — no event-schema or
core change. Replay uses a throwaway presenter so loaded sessions render
identically to live ones.
Tests: dsh-tools defineTool presenters (typed args, soft-validate), tool-bash
bash/bash_output/bash_kill presenters, acp ToolPresenter (tool-owned mapping,
unknown-callId fallback, in-flight-only map), and an end-to-end turn through the
bridge. The key-gated e2e now asserts a real bash call's title is the model
description (not "bash") and rawInput is the command — verified against the real
DeepSeek model. The test harness derives its inject from the bridge's exported
`inject` so it can't drift again.
2026-06-18 09:01:36 +08:00
} )
it ( 'presentCall validates softly: malformed args (missing required description) return undefined, never throw' , async ( ) = > {
const ctx = await setup ( )
// defineTool wraps presentCall to soft-validate against the schema and fall
// back to undefined (a generic UI presentation) rather than throwing on the
// display path — it may run on replay of arbitrary logged args. The
// ToolDefinition.presentCall takes `unknown`, so a malformed shape needs no cast.
expect ( ctx . tools . get ( 'bash' ) ? . presentCall ? . ( { command : 'ls' } ) ) . toBeUndefined ( )
} )
} )
feat(bash): add stdin + extra env to the executor seam as a trusted-plugin surface
The hooks subsystem runs external hook commands the Claude Code / Codex way:
JSON payload on stdin, context in CLAUDE_PROJECT_DIR / CLAUDE_PLUGIN_ROOT env.
Reusing the ctx.bash seam for that needs two new inputs — but stdin and arbitrary
env are exactly what dsh-bash-local's credential scrub exists to keep away from
model-driven commands. So this adds them as a TRUSTED-PLUGIN surface:
- BashExecRequest + BashExecSpec gain optional `stdin` and `env`. They are plain
optionals on the resolved spec (not required-but-nullable like `owner`): a
missing one means "none", the safe default, not a security footgun.
- dsh-bash-local threads them through resolve/run/start. `env` merges AFTER the
credential scrub, so a trusted caller's explicit entry wins even on a
credential-shaped name — the scrub guards the harness's OWN ambient creds from
model-driven commands, not a trusted plugin. stdin is always a pipe, closed
immediately (with bytes when supplied, empty otherwise — EOF as before); an
EPIPE from a child that exits without reading is swallowed.
- The model-facing dsh-tool-bash NEVER forwards model input into stdin/env (its
request is command/workdir/timeoutMs/signal/owner only). A regression guard
drives the real tool with adversarial args and asserts the request carries
neither field — proven to go red if the consumer ever forwards them.
Configurable scrub (in an earlier sketch) is dropped as speculative: the explicit
`env` field already gives a trusted caller full control, and no caller needs to
broaden the ambient scrub. Documented in a new architecture RFC, the bash.md
type-equiv blocks, and the three bash READMEs.
2026-06-30 13:52:25 +08:00
docs(bash): reframe stdin/env — the scrub is the security control, not a trust boundary
Address review: the "trusted-plugin surface" framing overstated the security
story. A model driving the `bash` tool already has equivalent power to set env
vars and feed stdin through ordinary shell syntax (`FOO=bar cmd`, heredocs), so
the `env`/`stdin` seam fields grant it no new capability — and they cannot
exfiltrate the harness's ambient credentials, because the credential SCRUB in
dsh-bash-local (which strips *KEY*/*SECRET*/*TOKEN* from process.env before the
child sees it) is the actual control, and it works regardless of these fields
(tool-call args are static JSON, never shell-evaluated).
So drop the "dangerous / trusted-plugin boundary" language across the RFC, the
three bash-package READMEs, the bash/src/types.ts JSDoc, and docs/bash.md (both
the type-equiv blocks — kept 1:1 with source — and the prose). The reality that
remains: the `bash` tool doesn't EXPOSE env/stdin as parameters because they'd
be redundant with shell syntax; the fields exist for in-process plugins (the
hooks bridges) to pass a JSON payload + CLAUDE_* vars cleanly. The guard test is
kept but reframed: it catches a future `...args` spread that would silently
forward model input into the post-scrub env merge, NOT a trust wall. No code or
behavior change.
2026-07-02 04:08:24 +08:00
describe ( 'the model-facing bash tool builds its request from named args only (no {...args} forward)' , ( ) = > {
feat(bash): add stdin + extra env to the executor seam as a trusted-plugin surface
The hooks subsystem runs external hook commands the Claude Code / Codex way:
JSON payload on stdin, context in CLAUDE_PROJECT_DIR / CLAUDE_PLUGIN_ROOT env.
Reusing the ctx.bash seam for that needs two new inputs — but stdin and arbitrary
env are exactly what dsh-bash-local's credential scrub exists to keep away from
model-driven commands. So this adds them as a TRUSTED-PLUGIN surface:
- BashExecRequest + BashExecSpec gain optional `stdin` and `env`. They are plain
optionals on the resolved spec (not required-but-nullable like `owner`): a
missing one means "none", the safe default, not a security footgun.
- dsh-bash-local threads them through resolve/run/start. `env` merges AFTER the
credential scrub, so a trusted caller's explicit entry wins even on a
credential-shaped name — the scrub guards the harness's OWN ambient creds from
model-driven commands, not a trusted plugin. stdin is always a pipe, closed
immediately (with bytes when supplied, empty otherwise — EOF as before); an
EPIPE from a child that exits without reading is swallowed.
- The model-facing dsh-tool-bash NEVER forwards model input into stdin/env (its
request is command/workdir/timeoutMs/signal/owner only). A regression guard
drives the real tool with adversarial args and asserts the request carries
neither field — proven to go red if the consumer ever forwards them.
Configurable scrub (in an earlier sketch) is dropped as speculative: the explicit
`env` field already gives a trusted caller full control, and no caller needs to
broaden the ambient scrub. Documented in a new architecture RFC, the bash.md
type-equiv blocks, and the three bash READMEs.
2026-06-30 13:52:25 +08:00
/ * *
* Records every { @link BashExecRequest } the consumer hands to ` resolve() ` , so a
docs(bash): reframe stdin/env — the scrub is the security control, not a trust boundary
Address review: the "trusted-plugin surface" framing overstated the security
story. A model driving the `bash` tool already has equivalent power to set env
vars and feed stdin through ordinary shell syntax (`FOO=bar cmd`, heredocs), so
the `env`/`stdin` seam fields grant it no new capability — and they cannot
exfiltrate the harness's ambient credentials, because the credential SCRUB in
dsh-bash-local (which strips *KEY*/*SECRET*/*TOKEN* from process.env before the
child sees it) is the actual control, and it works regardless of these fields
(tool-call args are static JSON, never shell-evaluated).
So drop the "dangerous / trusted-plugin boundary" language across the RFC, the
three bash-package READMEs, the bash/src/types.ts JSDoc, and docs/bash.md (both
the type-equiv blocks — kept 1:1 with source — and the prose). The reality that
remains: the `bash` tool doesn't EXPOSE env/stdin as parameters because they'd
be redundant with shell syntax; the fields exist for in-process plugins (the
hooks bridges) to pass a JSON payload + CLAUDE_* vars cleanly. The guard test is
kept but reframed: it catches a future `...args` spread that would silently
forward model input into the post-scrub env merge, NOT a trust wall. No code or
behavior change.
2026-07-02 04:08:24 +08:00
* test can assert what the model - facing tool DID and DID NOT forward . The ` bash `
* tool does not expose ` stdin ` / ` env ` as parameters ( bash syntax already gives a
* model that power ) , so it must build its request from named args only and
* never spread unknown tool - call keys into it . This guard ' s job is to catch a
* future refactor that blindly forwards ` ...args ` — which would silently thread
* model input into the post - scrub ` env ` merge — NOT to defend a trust boundary
* ( the credential scrub in dsh - bash - local is the security control ; see the
* bash - stdin - env RFC ) . Foreground ` run() ` returns a canned result ; ` start() ` is
* unused here .
feat(bash): add stdin + extra env to the executor seam as a trusted-plugin surface
The hooks subsystem runs external hook commands the Claude Code / Codex way:
JSON payload on stdin, context in CLAUDE_PROJECT_DIR / CLAUDE_PLUGIN_ROOT env.
Reusing the ctx.bash seam for that needs two new inputs — but stdin and arbitrary
env are exactly what dsh-bash-local's credential scrub exists to keep away from
model-driven commands. So this adds them as a TRUSTED-PLUGIN surface:
- BashExecRequest + BashExecSpec gain optional `stdin` and `env`. They are plain
optionals on the resolved spec (not required-but-nullable like `owner`): a
missing one means "none", the safe default, not a security footgun.
- dsh-bash-local threads them through resolve/run/start. `env` merges AFTER the
credential scrub, so a trusted caller's explicit entry wins even on a
credential-shaped name — the scrub guards the harness's OWN ambient creds from
model-driven commands, not a trusted plugin. stdin is always a pipe, closed
immediately (with bytes when supplied, empty otherwise — EOF as before); an
EPIPE from a child that exits without reading is swallowed.
- The model-facing dsh-tool-bash NEVER forwards model input into stdin/env (its
request is command/workdir/timeoutMs/signal/owner only). A regression guard
drives the real tool with adversarial args and asserts the request carries
neither field — proven to go red if the consumer ever forwards them.
Configurable scrub (in an earlier sketch) is dropped as speculative: the explicit
`env` field already gives a trusted caller full control, and no caller needs to
broaden the ambient scrub. Documented in a new architecture RFC, the bash.md
type-equiv blocks, and the three bash READMEs.
2026-06-30 13:52:25 +08:00
* /
class RecordingBashExecutor extends BashExecutor {
readonly requests : BashExecRequest [ ] = [ ]
resolve ( request : BashExecRequest ) : BashExecSpec {
this . requests . push ( request )
return {
command : request.command ,
workdir : request.workdir ? ? process . cwd ( ) ,
timeoutMs : request.timeoutMs ? ? 0 ,
. . . request . signal ? { signal : request.signal } : { } ,
. . . request . stdin !== undefined ? { stdin : request.stdin } : { } ,
. . . request . env !== undefined ? { env : request.env } : { } ,
owner : request.owner ,
}
}
run ( ) : Promise < BashRunResult > {
return Promise . resolve ( {
exitCode : 0 , signal : null , timedOut : false , aborted : false , timeoutMs : 0 ,
stdout : { text : 'ok' , truncated : false } , stderr : { text : '' , truncated : false } ,
} )
}
start ( ) : BashTask { throw new Error ( 'unused' ) }
get ( ) : BashTask | undefined { return undefined }
ownerOf ( ) : OwnerToken | undefined { return undefined }
list ( ) : BashTask [ ] { return [ ] }
readOutput ( ) : BashTaskRead { throw new Error ( 'unused' ) }
kill ( ) : boolean { return false }
}
async function setupRecording() {
const ctx = new Context ( )
await ctx . plugin ( SystemPrompt )
await ctx . plugin ( ToolRegistry )
await ctx . plugin ( AgentRegistry )
await ctx . plugin ( RecordingBashExecutor )
await ctx . plugin ( ToolBash )
return { ctx , bash : ctx.bash as RecordingBashExecutor }
}
docs(bash): reframe stdin/env — the scrub is the security control, not a trust boundary
Address review: the "trusted-plugin surface" framing overstated the security
story. A model driving the `bash` tool already has equivalent power to set env
vars and feed stdin through ordinary shell syntax (`FOO=bar cmd`, heredocs), so
the `env`/`stdin` seam fields grant it no new capability — and they cannot
exfiltrate the harness's ambient credentials, because the credential SCRUB in
dsh-bash-local (which strips *KEY*/*SECRET*/*TOKEN* from process.env before the
child sees it) is the actual control, and it works regardless of these fields
(tool-call args are static JSON, never shell-evaluated).
So drop the "dangerous / trusted-plugin boundary" language across the RFC, the
three bash-package READMEs, the bash/src/types.ts JSDoc, and docs/bash.md (both
the type-equiv blocks — kept 1:1 with source — and the prose). The reality that
remains: the `bash` tool doesn't EXPOSE env/stdin as parameters because they'd
be redundant with shell syntax; the fields exist for in-process plugins (the
hooks bridges) to pass a JSON payload + CLAUDE_* vars cleanly. The guard test is
kept but reframed: it catches a future `...args` spread that would silently
forward model input into the post-scrub env merge, NOT a trust wall. No code or
behavior change.
2026-07-02 04:08:24 +08:00
it ( 'does not forward env/stdin even when the model includes them as extra arguments' , async ( ) = > {
feat(bash): add stdin + extra env to the executor seam as a trusted-plugin surface
The hooks subsystem runs external hook commands the Claude Code / Codex way:
JSON payload on stdin, context in CLAUDE_PROJECT_DIR / CLAUDE_PLUGIN_ROOT env.
Reusing the ctx.bash seam for that needs two new inputs — but stdin and arbitrary
env are exactly what dsh-bash-local's credential scrub exists to keep away from
model-driven commands. So this adds them as a TRUSTED-PLUGIN surface:
- BashExecRequest + BashExecSpec gain optional `stdin` and `env`. They are plain
optionals on the resolved spec (not required-but-nullable like `owner`): a
missing one means "none", the safe default, not a security footgun.
- dsh-bash-local threads them through resolve/run/start. `env` merges AFTER the
credential scrub, so a trusted caller's explicit entry wins even on a
credential-shaped name — the scrub guards the harness's OWN ambient creds from
model-driven commands, not a trusted plugin. stdin is always a pipe, closed
immediately (with bytes when supplied, empty otherwise — EOF as before); an
EPIPE from a child that exits without reading is swallowed.
- The model-facing dsh-tool-bash NEVER forwards model input into stdin/env (its
request is command/workdir/timeoutMs/signal/owner only). A regression guard
drives the real tool with adversarial args and asserts the request carries
neither field — proven to go red if the consumer ever forwards them.
Configurable scrub (in an earlier sketch) is dropped as speculative: the explicit
`env` field already gives a trusted caller full control, and no caller needs to
broaden the ambient scrub. Documented in a new architecture RFC, the bash.md
type-equiv blocks, and the three bash READMEs.
2026-06-30 13:52:25 +08:00
const { ctx , bash } = await setupRecording ( )
docs(bash): reframe stdin/env — the scrub is the security control, not a trust boundary
Address review: the "trusted-plugin surface" framing overstated the security
story. A model driving the `bash` tool already has equivalent power to set env
vars and feed stdin through ordinary shell syntax (`FOO=bar cmd`, heredocs), so
the `env`/`stdin` seam fields grant it no new capability — and they cannot
exfiltrate the harness's ambient credentials, because the credential SCRUB in
dsh-bash-local (which strips *KEY*/*SECRET*/*TOKEN* from process.env before the
child sees it) is the actual control, and it works regardless of these fields
(tool-call args are static JSON, never shell-evaluated).
So drop the "dangerous / trusted-plugin boundary" language across the RFC, the
three bash-package READMEs, the bash/src/types.ts JSDoc, and docs/bash.md (both
the type-equiv blocks — kept 1:1 with source — and the prose). The reality that
remains: the `bash` tool doesn't EXPOSE env/stdin as parameters because they'd
be redundant with shell syntax; the fields exist for in-process plugins (the
hooks bridges) to pass a JSON payload + CLAUDE_* vars cleanly. The guard test is
kept but reframed: it catches a future `...args` spread that would silently
forward model input into the post-scrub env merge, NOT a trust wall. No code or
behavior change.
2026-07-02 04:08:24 +08:00
// Extra args: the model includes `env` and `stdin` keys hoping they reach the
// executor. The bash tool's schema ignores unknown keys, and execute() builds
// the request from only command/workdir/timeoutMs/signal — so the recorded
// request carries NEITHER. (Not a security wall — the model could set an env
// var or feed stdin via shell syntax anyway; this just keeps the request
// shape honest so a future `...args` spread can't silently forward input.)
feat(bash): add stdin + extra env to the executor seam as a trusted-plugin surface
The hooks subsystem runs external hook commands the Claude Code / Codex way:
JSON payload on stdin, context in CLAUDE_PROJECT_DIR / CLAUDE_PLUGIN_ROOT env.
Reusing the ctx.bash seam for that needs two new inputs — but stdin and arbitrary
env are exactly what dsh-bash-local's credential scrub exists to keep away from
model-driven commands. So this adds them as a TRUSTED-PLUGIN surface:
- BashExecRequest + BashExecSpec gain optional `stdin` and `env`. They are plain
optionals on the resolved spec (not required-but-nullable like `owner`): a
missing one means "none", the safe default, not a security footgun.
- dsh-bash-local threads them through resolve/run/start. `env` merges AFTER the
credential scrub, so a trusted caller's explicit entry wins even on a
credential-shaped name — the scrub guards the harness's OWN ambient creds from
model-driven commands, not a trusted plugin. stdin is always a pipe, closed
immediately (with bytes when supplied, empty otherwise — EOF as before); an
EPIPE from a child that exits without reading is swallowed.
- The model-facing dsh-tool-bash NEVER forwards model input into stdin/env (its
request is command/workdir/timeoutMs/signal/owner only). A regression guard
drives the real tool with adversarial args and asserts the request carries
neither field — proven to go red if the consumer ever forwards them.
Configurable scrub (in an earlier sketch) is dropped as speculative: the explicit
`env` field already gives a trusted caller full control, and no caller needs to
broaden the ambient scrub. Documented in a new architecture RFC, the bash.md
type-equiv blocks, and the three bash READMEs.
2026-06-30 13:52:25 +08:00
await ctx . tools . execute ( {
2026-07-02 04:30:40 +08:00
callId : CallId ( 'no-forward-1' ) ,
feat(bash): add stdin + extra env to the executor seam as a trusted-plugin surface
The hooks subsystem runs external hook commands the Claude Code / Codex way:
JSON payload on stdin, context in CLAUDE_PROJECT_DIR / CLAUDE_PLUGIN_ROOT env.
Reusing the ctx.bash seam for that needs two new inputs — but stdin and arbitrary
env are exactly what dsh-bash-local's credential scrub exists to keep away from
model-driven commands. So this adds them as a TRUSTED-PLUGIN surface:
- BashExecRequest + BashExecSpec gain optional `stdin` and `env`. They are plain
optionals on the resolved spec (not required-but-nullable like `owner`): a
missing one means "none", the safe default, not a security footgun.
- dsh-bash-local threads them through resolve/run/start. `env` merges AFTER the
credential scrub, so a trusted caller's explicit entry wins even on a
credential-shaped name — the scrub guards the harness's OWN ambient creds from
model-driven commands, not a trusted plugin. stdin is always a pipe, closed
immediately (with bytes when supplied, empty otherwise — EOF as before); an
EPIPE from a child that exits without reading is swallowed.
- The model-facing dsh-tool-bash NEVER forwards model input into stdin/env (its
request is command/workdir/timeoutMs/signal/owner only). A regression guard
drives the real tool with adversarial args and asserts the request carries
neither field — proven to go red if the consumer ever forwards them.
Configurable scrub (in an earlier sketch) is dropped as speculative: the explicit
`env` field already gives a trusted caller full control, and no caller needs to
broaden the ambient scrub. Documented in a new architecture RFC, the bash.md
type-equiv blocks, and the three bash READMEs.
2026-06-30 13:52:25 +08:00
name : 'bash' ,
arguments : {
command : 'echo hi' ,
description : 'echo' ,
env : { SNEAKY_API_KEY : 'leak' } ,
stdin : 'malicious payload' ,
} ,
} )
expect ( bash . requests ) . toHaveLength ( 1 )
const request = bash . requests [ 0 ] !
expect ( request . command ) . toBe ( 'echo hi' )
expect ( 'env' in request ) . toBe ( false )
expect ( 'stdin' in request ) . toBe ( false )
} )
it ( 'a background bash call likewise carries no env/stdin' , async ( ) = > {
const { ctx , bash } = await setupRecording ( )
// start() throws in this recorder, but resolve() runs first and records the
2026-07-02 04:30:40 +08:00
// request — which is all this no-forward assertion needs.
feat(bash): add stdin + extra env to the executor seam as a trusted-plugin surface
The hooks subsystem runs external hook commands the Claude Code / Codex way:
JSON payload on stdin, context in CLAUDE_PROJECT_DIR / CLAUDE_PLUGIN_ROOT env.
Reusing the ctx.bash seam for that needs two new inputs — but stdin and arbitrary
env are exactly what dsh-bash-local's credential scrub exists to keep away from
model-driven commands. So this adds them as a TRUSTED-PLUGIN surface:
- BashExecRequest + BashExecSpec gain optional `stdin` and `env`. They are plain
optionals on the resolved spec (not required-but-nullable like `owner`): a
missing one means "none", the safe default, not a security footgun.
- dsh-bash-local threads them through resolve/run/start. `env` merges AFTER the
credential scrub, so a trusted caller's explicit entry wins even on a
credential-shaped name — the scrub guards the harness's OWN ambient creds from
model-driven commands, not a trusted plugin. stdin is always a pipe, closed
immediately (with bytes when supplied, empty otherwise — EOF as before); an
EPIPE from a child that exits without reading is swallowed.
- The model-facing dsh-tool-bash NEVER forwards model input into stdin/env (its
request is command/workdir/timeoutMs/signal/owner only). A regression guard
drives the real tool with adversarial args and asserts the request carries
neither field — proven to go red if the consumer ever forwards them.
Configurable scrub (in an earlier sketch) is dropped as speculative: the explicit
`env` field already gives a trusted caller full control, and no caller needs to
broaden the ambient scrub. Documented in a new architecture RFC, the bash.md
type-equiv blocks, and the three bash READMEs.
2026-06-30 13:52:25 +08:00
await ctx . tools . execute ( {
2026-07-02 04:30:40 +08:00
callId : CallId ( 'no-forward-2' ) ,
feat(bash): add stdin + extra env to the executor seam as a trusted-plugin surface
The hooks subsystem runs external hook commands the Claude Code / Codex way:
JSON payload on stdin, context in CLAUDE_PROJECT_DIR / CLAUDE_PLUGIN_ROOT env.
Reusing the ctx.bash seam for that needs two new inputs — but stdin and arbitrary
env are exactly what dsh-bash-local's credential scrub exists to keep away from
model-driven commands. So this adds them as a TRUSTED-PLUGIN surface:
- BashExecRequest + BashExecSpec gain optional `stdin` and `env`. They are plain
optionals on the resolved spec (not required-but-nullable like `owner`): a
missing one means "none", the safe default, not a security footgun.
- dsh-bash-local threads them through resolve/run/start. `env` merges AFTER the
credential scrub, so a trusted caller's explicit entry wins even on a
credential-shaped name — the scrub guards the harness's OWN ambient creds from
model-driven commands, not a trusted plugin. stdin is always a pipe, closed
immediately (with bytes when supplied, empty otherwise — EOF as before); an
EPIPE from a child that exits without reading is swallowed.
- The model-facing dsh-tool-bash NEVER forwards model input into stdin/env (its
request is command/workdir/timeoutMs/signal/owner only). A regression guard
drives the real tool with adversarial args and asserts the request carries
neither field — proven to go red if the consumer ever forwards them.
Configurable scrub (in an earlier sketch) is dropped as speculative: the explicit
`env` field already gives a trusted caller full control, and no caller needs to
broaden the ambient scrub. Documented in a new architecture RFC, the bash.md
type-equiv blocks, and the three bash READMEs.
2026-06-30 13:52:25 +08:00
name : 'bash' ,
arguments : {
command : 'sleep 1' ,
description : 'sleep' ,
run_in_background : true ,
env : { TOKEN : 'leak' } ,
stdin : 'x' ,
} ,
} )
expect ( bash . requests ) . toHaveLength ( 1 )
const request = bash . requests [ 0 ] !
expect ( 'env' in request ) . toBe ( false )
expect ( 'stdin' in request ) . toBe ( false )
// The owner token IS set on a background call (the isolation fence) — proving
// the recorder sees the real request the consumer built, so the absent
// env/stdin above is a real negative, not a recorder that drops everything.
expect ( 'owner' in request ) . toBe ( true )
} )
} )