The first real agent wiring: DeepSeek V4 + the bash tool suite + stdio chat + JSONL persistence, runnable via yarn demo:coding (reads the gitignored repo-root .env through process.loadEnvFile). - examples/coding-agent: cordis.yml wiring both real plugin families (llm-deepseek with !!js env secrets; bash-local + tool-bash), a bash-only coding system prompt, a max-steps-guard plugin (bounds runaway turns via the agent/turn-continuation waterfall — abort() from step-end is a no-op by then), and a stdio UI with dimmed reasoning and exit-on-idle for piped stdin. - e2e (yarn test:e2e, key-gated): full-loop.e2e.ts runs a real model against the real bash tool; coding-task.e2e.ts is the swebench-style smoke — the model fixes a buggy add.js in a temp dir and the test re-runs node add.test.js itself rather than trusting the agent. - docs/cookbook: adding-a-package (the verified checklist), adding-a-tool (execute() contract, background pattern, seams), adding-an-llm-adapter (protocol obligations, mock-server testing, e2e policy). AGENTS.md layout/commands/secrets sections updated; architecture.md points at both examples and the cookbook. - vitest.e2e.config.ts: serialize test files + retry twice — parallel e2e files trip the shared internal key's concurrency quota. - fix: the !js YAML tag spelling in docs/JSDoc is actually !!js (js-yaml resolves custom tags under tag:yaml.org,2002:js).
3.4 KiB
Cookbook: adding an LLM adapter
How to connect a new model provider. Reference implementations:
packages/llm-deepseek (hand-rolled HTTP/SSE) and packages/llm-pi-ai
(wrapping an LLM library). Read the StreamChunk doc in
packages/llm/src/types.ts first — it records the protocol conventions both
adapters were verified against.
The shape
class MyAdapter extends LlmAdapter {
async * stream(options: GenerateOptions): AsyncIterable<StreamChunk> { … }
}
export const name = 'llm-myprovider'
export const inject = ['llm']
export const Config: z<Config> = z.object({ apiKey: z.string(), … })
export function apply(ctx: Context, config: Config) {
ctx.llm.registerAdapter(['model-a', 'model-b'], new MyAdapter(…))
}
Registration is effect-based (HMR-safe); one adapter per model name —
duplicates throw. Secrets are cordis-native: schemastery Config with env
fallbacks, fed from cordis.yml via !!js process.env.MY_KEY. Never read
ad-hoc key files in code.
Protocol obligations (the contract two implementations verified)
- Emit
usageBEFOREfinish; emit NOTHING afterfinish. The robust way: buffer finish/usage until the provider's end-of-stream marker, then flush (handles providers that send trailing usage-only chunks). - Tool-call
argumentsare RAW JSON strings end-to-end; stream fragments asargumentsDelta. If your provider hands back parsed objects, re-stringify atblock-end. - Allocate block
indexes in first-seen stream order; reuse the index for every delta of the same block. - Errors have exactly two sanctioned paths: THROW from
stream()(transport and protocol failures — useLlmErrorwith a stable code), or end the stream withfinish {kind: 'error' | 'aborted'}(provider in-band failures). Consumers handle both; pick per failure class and document it. - Honor
options.signal(pass it to fetch / your SDK). prefilland other unsupportedGenerateOptionsfields: throwLlmError(..., 'UNSUPPORTED')rather than silently dropping.
Provider-specific request knobs (thinking modes, effort levels) belong in
the ADAPTER's Config, not in GenerateOptions — the core vocabulary stays
provider-neutral.
Structure that worked
Split the adapter into testable stages (llm-deepseek's layout): wire types
(types.ts, coverage-exempt) → request serializer → SSE/transport parser →
chunk-translation state machine → a thin adapter class wiring them. Each
stage gets its own unit suite.
Testing
- Unit: mock the provider, not the harness. A scripted
node:httpserver speaking the provider's wire format covers happy paths, every error status, malformed payloads, premature closes, and aborts — no network, and it drives the 100% per-file coverage gate. Works for SDK-backed adapters too (point the SDK's baseURL at the mock). - Hostile framing tests. Split stream payloads at arbitrary byte positions (including mid-UTF-8) — real networks do.
- E2E:
tests/*.e2e.tsunderyarn test:e2e, gated withdescribe.skipIf(!process.env.MY_KEY)so CI (no secrets) stays green. Cover each model × each provider mode you map (thinking on/off, effort levels), a tool-call round trip INCLUDING the follow-up turn with results in history, and loose assertions only (substring/structure, bounded maxTokens — real models are nondeterministic). - Register the e2e file pattern in
knip.json(per-workspaceentryoverride) or knip flags it unused.