deepseek-harness/.agents/notes/implemented/bug-fix/2026-08-29-code-runtime-python-reply-backlog-and-surrogate-count.md

34 lines
4 KiB
Markdown
Raw Normal View History

fix(code-runtime-python): bound reply and call backlogs, snapshot binding metadata, and compact the reply queue Review findings on the CPython backend: a child that never reads fd 3 leaves the reply pipe full forever, so the drain loop waits on 'drain' while every call frame it keeps sending resolves a binding and queues another reply — the backlog (and the binding results it pins) would grow until the wall clock. sendReply now caps the pending backlog at MAX_PENDING_REPLIES and settles the run as worker-exit past it, mirroring the frame cap; a child flooding calls against a binding that never settles would otherwise bypass that cap (pendingReplies grows only after the await), so the dispatcher counts in-flight binding calls before dispatch and releases the slot in the async body's finally, capping outstanding closures at the same bound. The drain also compacts its consumed prefix (replyQueue.splice(0, head)) once head reaches the bound, so a drain that stays alive without emptying cannot grow the backing store linearly with cumulative throughput. The completion-value meter counted lone surrogates with _SURROGATE.findall(folded), materializing one single-character string per surrogate: a surrogate-dense value near the budget (millions of surrogates, each serializing to six bytes) allocated millions of objects before the meter returned, defeating the meter's counting-without-building contract. The count is now the length difference between folded and the without string the meter already computes; a standalone equivalence check confirms it matches findall across lone-high, lone-low, paired, astral, and mixed cases. validateBindings read namespace.global/errorClass.name/memberNameProperty several times and retained the original errorClass object for the boot frame, whose JSON.stringify re-read it after validation: a stateful getter could throw or change between the two stages, turning the seam-misuse rejection into a worker-exit or injecting an unvalidated name. Each field is now read once into a plain value and the bindings map stores a plain { name, memberNameProperty } copy, so validation and the boot frame see identical values. Regression tests: a hostile child floods 5000 sequential valid calls without reading fd 3 and the run settles worker-exit with the reply-queue message before maxWallMs; a 3,000,000-surrogate completion succeeds at an 18,000,002-byte budget and reports output-limit one byte under; a 5000-call flood against a never-settling binding settles worker-exit with the call-backlog message; getter-backed namespace metadata that throws or changes on a second read boots and runs with each field read exactly once; a two-wave flood whose replies exceed the writable high-water mark drives the drain past the compaction bound mid-delivery and verifies all 1524 replies arrive. README Known Limitations gains the reply-backlog and call-backlog bounds (en/zh, pairing re-recorded); a new Agent Note registers the findings.
2026-08-29 23:51:17 +08:00
# Agent Note: Bound the reply backlog and count lone surrogates without a match list in the CPython backend
Status: implemented
English | [中文](2026-08-29-code-runtime-python-reply-backlog-and-surrogate-count.zh.md)
## Problem
A further review round on the CPython subprocess backend (packages/experimental/code-runtime-python) surfaced two unbounded-allocation findings. First, `replyQueue` had no bound: a child that never reads fd 3 keeps the reply pipe full forever, so the drain loop waits on `drain` while every call frame it keeps sending resolves a binding and queues another reply — the backlog (and the binding results it pins) grows until the wall clock. Second, `_json_str_cost` counted lone surrogates with `_SURROGATE.findall(folded)`, which materializes one single-character string per surrogate: a surrogate-dense completion value near the budget (each surrogate serializes to six bytes, so a budget-sized value holds millions of them) allocates millions of objects before the meter returns, defeating the meter's own contract of counting without building.
## Decision
### The reply backlog is capped at 1024 pending frames
`sendReply` now counts pending replies separately from the consumed slots the drain loop clears, and settles the run as a `worker-exit` with a reply-queue message before pushing when the backlog reaches `MAX_PENDING_REPLIES`. The counter is decremented as the drain writes each frame and reset when the drain finishes, so it measures only replies the host still holds. This mirrors the frame cap's treatment of an oversized inbound frame: a child that stops participating in the protocol fails the run early instead of growing host memory until the wall clock. It is a count bound, not a byte bound — binding results carry no seam-level byte cap, so the bound limits how many are retained, not how large any one is.
### Lone surrogates are counted by length difference, not by a match list
`_json_str_cost` computed `lone = len(_SURROGATE.findall(folded))`, building a list of one single-character string per lone surrogate. The count is now the length difference between `folded` and `without = _SURROGATE.sub("", folded)`: after pair-combining, every remaining surrogate is lone and exactly one code point, so the number removed is the count, and the `without` string is needed by the meter anyway. The meter returns the identical byte cost with no per-surrogate objects.
## Testing
- `tests/runtime.spec.ts` — a hostile child floods 5000 sequential valid call frames and never reads fd 3; the run settles as `worker-exit` with the reply-queue message long before `maxWallMs`, proving the backlog cap fires instead of a wall-clock timeout. A surrogate-dense completion of 3,000,000 lone surrogates pins the boundary at scale: 18,000,002 serialized bytes succeed at an 18,000,002 budget and report `output-limit` one byte under, proving the meter counts every surrogate exactly (the len-diff is verified equal to the old findall count across lone-high, lone-low, paired, astral, and mixed cases).
## Alternatives considered
**Pause the fd-3 read side while waiting for drain instead of capping the queue.** Rejected: pausing reads would also stall processing of `done` and `log` frames the child may send after its last call, changing settlement timing; a count cap is deterministic and matches the existing frame-cap pattern.
**Keep findall and rely on the character-count lower bound.** Rejected: the lower bound admits a string by CHARACTER count while each surrogate serializes to six bytes, so a budget-sized surrogate-dense string passes it and reaches the meter; the match list is exactly the allocation the meter exists to avoid.
## Consequences
A child that stops consuming its replies now fails the run as a `worker-exit` once 1024 replies are retained, bounding host memory without a wall-clock wait. The completion-value meter counts lone surrogates with no per-surrogate allocation, keeping its documented counting-without-building contract for surrogate-dense values.