deepseek-harness/vitest.e2e.config.ts

59 lines
2.6 KiB
TypeScript
Raw Normal View History

Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
import tsconfigPaths from 'vite-tsconfig-paths'
import { defineConfig } from 'vitest/config'
import { vitestExecArgv } from './vitest.shared.ts'
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
// Real-API suite, separate because it spends tokens. Each test self-skips without
2026-07-17 23:33:30 +08:00
// its provider credential for keyless CI; credentialed workflows preflight the
// secrets they require. Values may come from the environment or gitignored root
// `.env`, with provider-specific endpoint overrides where supported.
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
try {
// Node >= 21.7 native; throws when the file does not exist.
process.loadEnvFile(new URL('.env', import.meta.url).pathname)
} catch {
// No .env — fine, the environment may already carry the variables.
}
2026-07-06 01:10:27 +08:00
const DEFAULT_E2E_MAX_WORKERS = 4
function positiveIntFromEnv(name: string, fallback: number): number {
const raw = process.env[name]
if (raw === undefined || raw === '') return fallback
const value = Number(raw)
if (!Number.isInteger(value) || value < 1) {
throw new Error(`${name} must be a positive integer, got ${JSON.stringify(raw)}`)
}
return value
}
const e2eMaxWorkers = positiveIntFromEnv('DSH_E2E_MAX_WORKERS', DEFAULT_E2E_MAX_WORKERS)
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
export default defineConfig({
// Same resolution note as vitest.config.ts: bare workspace names resolve
// through the tsconfig.base.json paths facade (no include = match-all, so
// client-package sources get mapping too — dropping /client subpath imports
// onto package exports would load browser dist bundles into node).
// Built-artifact e2e suites are unaffected: their built-ness lives in
// subprocesses and createRequire lookups, which bypass vite resolution
// entirely.
plugins: [tsconfigPaths({ projects: ['./tsconfig.base.json'] })],
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
test: {
execArgv: vitestExecArgv,
setupFiles: ['./scripts/test-invariants.ts'],
// apps/cli only, not apps/*: apps/web/tests/*.e2e.ts needs the built
// frontend dist and runs under vitest.web.config.ts (the test:web job).
include: ['packages/*/*/tests/**/*.e2e.ts', 'apps/cli/tests/**/*.e2e.ts', 'examples/*/tests/**/*.e2e.ts'],
// Real model calls: generous timeouts, and retries for transient flakes
// (the shared internal key hits concurrency quotas). No coverage — the
// unit suites own the coverage gate.
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
testTimeout: 120_000,
hookTimeout: 30_000,
retry: 2,
2026-07-06 01:10:27 +08:00
// Run files in a bounded pool: enough lower-level parallelism to keep CI
// and local with-key runs moving, while leaving a resource knob for shared
// API quotas (`DSH_E2E_MAX_WORKERS=1` restores serial execution).
fileParallelism: e2eMaxWorkers > 1,
maxWorkers: e2eMaxWorkers,
Add two DeepSeek LLM adapters: dsh-llm-deepseek and dsh-llm-pi-ai The first real LlmAdapter implementations, shipped as a deliberate pair: same models and wire protocol, completely different internals, so the StreamChunk protocol is verified across independent implementations. - dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation state machine against the official chat-completions format (thinking mode via top-level thinking/reasoning_effort; the empty-string reasoning_content first chunk; usage attached to the finish chunk or trailing; reasoning_content passback on tool-call turns; disjoint cache-token accounting). - dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai, mapping its event vocabulary (parsed tool arguments, in-stream error events, folded reasoning tokens) onto the same chunks. The agent loop now honors the in-band error path: an adapter that ends its stream with finish {kind:error|aborted} (the only option for adapters that can't throw mid-stream, like pi-ai) is translated into a step error, so the turn ends error/aborted with a logged error event instead of a normal completed assistant message. This makes the StreamChunk error contract real for both adapters; docs/architecture.md and the StreamChunk doc are updated accordingly. New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated real-API matrices for both adapters across V4 Flash/Pro and all thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit suites run against local node:http mock SSE servers at 100% per-file coverage.
2026-06-13 00:28:29 +08:00
},
})