deepseek-harness/.agents/notes/implemented/process/2026-08-27-gate-runner-fail-fast.md

8 KiB

Agent Note: Gate-runner fail-fast

Status: implemented

English | 中文

Parallel pre-push gates owns the bounded gate scheduler in scripts/run-gates.ts; this note adds one scheduling option to that scheduler.

Problem

The gate scheduler in scripts/run-gates.ts runs every independent gate in an aggregate to completion and reports run-gates: N passed, M failed. A gate failure does not stop the remaining gates; only gates whose needs dependency failed are skipped. On an aggregate that is already red, the remaining gates keep consuming runner time and produce evidence that cannot change the verdict. The largest single cost is the instrumented coverage run in the ci-coverage aggregate, which has taken about 27 minutes; in ci-consumers, the Node compatibility smoke runs independently of the build, so it keeps running after a build failure that already settles the verdict.

GitHub Actions provides no native cross-job cancellation: fail-fast applies only inside a matrix, and the all checks passed aggregate settles only after every needed job finishes, so it cannot cancel siblings early. The only in-repository lever is the gate scheduler itself.

Decision

run-gates.ts accepts a fail-fast scheduling option. When enabled, the first blocking gate failure (a gate whose allowFailure is not true) aborts the aggregate: the shared AbortSignal terminates every running gate's process tree, and every not-yet-run gate is recorded as skipped with the error aborted by fail-fast: <label> failed (or aborted by fail-fast: host interruption when the host signal aborted the run). The exit status remains 1, including when a killed child traps the signal and exits zero: such a result carries the abort mark and is recorded skipped, never passed. A gate that settled before the abort took effect keeps its real result, so the summary stays truthful about what produced evidence.

Termination covers the whole tree, not just the direct pnpm wrapper: POSIX signals the detached child's process group (kill(-pid, SIGTERM), escalating unconditionally to SIGKILL after 5 seconds) and additionally signals every transitive descendant read from the process table, so the detached leaves of a nested run-gates (the check:node-compat and check:ci:lint:contracts-ready gates inside ci-consumers) are killed without relying on the inner scheduler's own escalation. The descendant list is primed at spawn and refreshed every 5 seconds while the child runs, each tick merging the fresh enumeration into the live-filtered cache (the enumeration is asynchronous — a slow WMI/CIM call is bounded by its own 10-second timeout and never blocks the gate's output draining or exit handling; an enumeration still in flight when the gate settles is cancelled rather than left holding stdio handles, and a snapshot that settles after the child exited is still merged, because the child may be gone while a grandchild keeps close pending — exactly when the abort needs the list) so a descendant reparented by an exited intermediate stays tracked across ticks; at abort the fresh enumeration is merged into the same list, then re-signalled on the escalation, because the group kill reaps the direct child and reparents its detached descendants, making them unreachable by parent id afterwards; settlement waits until the group and the captured descendants are gone. On the abort path only, a bounded pipe-drain timer force-closes the stdio streams 10 seconds after the abort, so a descendant holding the write ends (uninterruptible I/O included) cannot keep close pending to the job timeout; ordinary runs keep waiting rather than report passed over a live leak. Windows runs taskkill /PID <pid> /T /F immediately, because a taskkill without /F does not terminate console processes, which is what gate commands are; the same process-table enumeration as POSIX supplies a descendant list there, and each captured descendant is also taskkilled, because a taskkill /T rooted at a pid that already exited finds nothing. Windows never reparents, so an exited root's descendants keep it as their parent and remain reachable through the table; the same sampler cadence as POSIX keeps the cache crossing a vanished intermediate's table record (the enumeration is bounded by a 10-second PowerShell timeout so a hung WMI/CIM call cannot stall the abort path). Without this, Windows has no signal forwarding and a wrapper-only kill would orphan the script tree on the shared self-hosted pool. Children are detached into their own POSIX process group only when fail-fast is enabled; ordinary runs keep them in the host group so terminal Ctrl+C still reaches them. Host SIGINT/SIGTERM on a fail-fast run is forwarded to the abort path, so an interrupted or runner-cancelled run drains and kills its gate trees instead of orphaning them.

The option is enabled through DSH_GATE_FAIL_FAST (accepted values: 1 or unset; anything else fails loud through the existing flagEnabled contract) on every run-gates aggregate job in ci.yml: the three blocking Linux jobs (node-24 static, node-24-coverage, node-24-consumers), the Node compatibility matrix (node-compat), and the two native Windows lanes that drive aggregates (windows-build, windows-coverage). scripts/ci-workflow.spec.ts pins the flag on those jobs and pins its absence on windows-observational, so removing it fails the CI gate.

The windows-observational lane stays complete: it is continue-on-error by design and exists to collect as much Windows-native evidence per run as possible, so the first failure must not truncate the rest. The windows Wine lane and windows-native-tests run a single script or Vitest command rather than a run-gates aggregate, so the scheduler option does not apply to them. The master serial standby lanes (serial-linux-selfhosted, serial-windows) and the manual runner benchmarks do not set the flag: they are completeness drills that must execute the full aggregate to prove pool readiness.

Consequences

A red pull-request run ends sooner. The largest saving is in ci-coverage: a failing exempt-heavy gate aborts the multi-minute instrumented coverage gate instead of letting it run out.

The trade-off is diagnostic: one push returns only the first blocking failure instead of the full failure set, so resolving several independent failures may take more push-fix rounds. Killed gates are recorded as skipped with the fail-fast error, so the summary line N passed, M failed, K skipped remains truthful about what produced evidence and what did not. A gate that ignores SIGTERM is force-killed after the 5-second grace. A tree that survives both signals holds the aggregate only while its direct child's stdio stays open; once close fires, the group-liveness poll gives up after 8 seconds and the run settles with a loud gate tree not quiescent warning instead of reporting a clean tree.

Alternatives considered

Cross-job cancellation watchdog. A job that polls sibling conclusions and calls the run-cancel API would stop all lanes on the first failure. It is not native, adds a polling dependency and token surface, and discards the parallel evidence other jobs have already produced. Rejected; fail-fast at the scheduler is orthogonal to the job topology and carries none of that.

Consolidating the three Linux jobs into one check. A single job could fail fast natively, but it would lose the independent runner allocation whose queue-delay overlap is documented in the independent CI consumer build note, and it would make the coverage long tail the tail of the whole job. Rejected; fail-fast applies within the existing job split instead.

Signaling only the direct child. The first implementation sent SIGTERM to the pnpm wrapper and relied on pnpm forwarding it to the script child. A probe confirmed the forwarding on POSIX. Rejected: Windows has no signal forwarding, and a wrapper-only kill orphans the script tree on the shared self-hosted pool; the tree termination above covers both platforms.