Review findings on the CPython backend: a child that never reads fd 3 leaves
the reply pipe full forever, so the drain loop waits on 'drain' while every
call frame it keeps sending resolves a binding and queues another reply —
the backlog (and the binding results it pins) would grow until the wall
clock. sendReply now caps the pending backlog at MAX_PENDING_REPLIES and
settles the run as worker-exit past it, mirroring the frame cap; a child
flooding calls against a binding that never settles would otherwise bypass
that cap (pendingReplies grows only after the await), so the dispatcher
counts in-flight binding calls before dispatch and releases the slot in the
async body's finally, capping outstanding closures at the same bound. The
drain also compacts its consumed prefix (replyQueue.splice(0, head)) once
head reaches the bound, so a drain that stays alive without emptying cannot
grow the backing store linearly with cumulative throughput.
The completion-value meter counted lone surrogates with
_SURROGATE.findall(folded), materializing one single-character string per
surrogate: a surrogate-dense value near the budget (millions of surrogates,
each serializing to six bytes) allocated millions of objects before the meter
returned, defeating the meter's counting-without-building contract. The count
is now the length difference between folded and the without string the meter
already computes; a standalone equivalence check confirms it matches findall
across lone-high, lone-low, paired, astral, and mixed cases.
validateBindings read namespace.global/errorClass.name/memberNameProperty
several times and retained the original errorClass object for the boot
frame, whose JSON.stringify re-read it after validation: a stateful getter
could throw or change between the two stages, turning the seam-misuse
rejection into a worker-exit or injecting an unvalidated name. Each field is
now read once into a plain value and the bindings map stores a plain
{ name, memberNameProperty } copy, so validation and the boot frame see
identical values.
Regression tests: a hostile child floods 5000 sequential valid calls without
reading fd 3 and the run settles worker-exit with the reply-queue message
before maxWallMs; a 3,000,000-surrogate completion succeeds at an
18,000,002-byte budget and reports output-limit one byte under; a 5000-call
flood against a never-settling binding settles worker-exit with the
call-backlog message; getter-backed namespace metadata that throws or
changes on a second read boots and runs with each field read exactly once; a
two-wave flood whose replies exceed the writable high-water mark drives the
drain past the compaction bound mid-delivery and verifies all 1524 replies
arrive. README Known Limitations gains the reply-backlog and call-backlog
bounds (en/zh, pairing re-recorded); a new Agent Note registers the findings.
3.6 KiB
Agent Note: 在 CPython 后端限制回复积压并改用长度差计数孤立代理项
Status: implemented
English | 中文
Problem
对 CPython 子进程后端(packages/experimental/code-runtime-python)的又一轮评审浮出两项无界分配发现。其一,replyQueue 没有上限:从不读取 fd 3 的子进程让回复管道永远占满,排空循环只能等待 drain,而它持续发送的每个调用帧都会解析一个 binding 并入队一条回复——积压(连同其钉住的 binding 结果)一直增长到墙钟。其二,_json_str_cost 用 _SURROGATE.findall(folded) 计数孤立代理项,每个代理项物化一个单字符字符串:接近预算的代理项密集完成值(每个代理项序列化为六个字节,预算大小的值可容纳数百万个)会在计量返回前分配数百万个对象,违背计量器自身「计数而不构建」的契约。
Decision
回复积压限制为 1024 个待发帧
sendReply 现在把待发回复数与排空循环已清空的槽位分开计数,当积压达到 MAX_PENDING_REPLIES 时,在入队前以带回复队列消息的 worker-exit 结算运行。计数器在排空写入每帧时递减、排空结束时重置,因此只度量宿主仍持有的回复。这与帧上限对超大入站帧的处理一致:停止参与协议的子进程让运行提前失败,而不是让宿主内存增长到墙钟。这是计数上限而非字节上限——binding 结果在 seam 层没有字节上限,因此该上限限制保留的数量,而非单个结果的大小。
孤立代理项改用长度差计数,而非匹配列表
_json_str_cost 原先计算 lone = len(_SURROGATE.findall(folded)),为每个孤立代理项构建一个单字符字符串的列表。现在计数改为 folded 与 without = _SURROGATE.sub("", folded) 的长度差:配对合并后,剩余的每个代理项都是孤立且恰好一个码点,因此被移除的数量即计数,而 without 字符串本就是计量需要的。计量器返回完全相同的字节成本,且不产生任何按代理项计的对象。
Testing
tests/runtime.spec.ts——敌意子进程洪泛 5000 个连续合法调用帧且从不读取 fd 3;运行在远早于maxWallMs时以带回复队列消息的worker-exit结算,证明积压上限先于墙钟超时触发。3,000,000 个孤立代理项的代理项密集完成值在规模上钉住边界:18,000,002 个序列化字节在 18,000,002 预算下成功、少一个字节时报output-limit,证明计量器精确计数每个代理项(长度差在孤立高、孤立低、配对、星面和混合用例下与旧 findall 计数逐一相等,已实测验证)。
Alternatives considered
在等待 drain 时暂停 fd-3 读侧而非限制队列。 拒绝:暂停读取也会让子进程在最后一个调用后可能发送的 done 与 log 帧处理停滞,改变结算时机;计数上限是确定性的,且与既有帧上限模式一致。
保留 findall 并依赖字符计数下界。 拒绝:下界按字符数放行字符串,而每个代理项序列化为六个字节,因此预算大小的代理项密集字符串能通过下界并进入计量器;匹配列表正是计量器要避免的分配。
Consequences
停止消费回复的子进程现在会在保留 1024 条回复时以 worker-exit 结算运行,无需等待墙钟即可限制宿主内存。完成值计量器对孤立代理项的计数不再产生按代理项计的分,保持其对代理项密集值「计数而不构建」的既有契约。