fix(llm-deepseek): pass reasoning content back on every reasoned turn

An assistant turn that answered without calling a tool serialized no
reasoning_content, so a gateway re-encoding the conversation for another
vendor had no chain of thought to hash and lost that turn's upstream
thinking signature.
This commit is contained in:
Yichen Jiang 2026-08-19 22:02:36 +08:00
parent 75282c92c1
commit 583894f7ae
9 changed files with 104 additions and 25 deletions

View file

@ -0,0 +1,6 @@
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
# side as of the last confirmed-consistent state. Both languages carry equal authority;
# after editing either side, bring the other along and re-record with:
# pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-08-19-deepseek-reasoning-passback-every-turn.md
2026-08-19-deepseek-reasoning-passback-every-turn.md: 0fb9ff3cc0991bb2455bc0ef2ec6916b76e20be0
2026-08-19-deepseek-reasoning-passback-every-turn.zh.md: d01c0c415ca4ce731cc8d9520729fe5b97119a9e

View file

@ -0,0 +1,33 @@
# Agent Note: DeepSeek reasoning passback on every reasoned turn
Status: implemented
English | [中文](2026-08-19-deepseek-reasoning-passback-every-turn.zh.md)
## Problem
`dsh-llm-deepseek` replayed `reasoning_content` in history only on assistant turns that also carried tool calls. DeepSeek's thinking-mode guide requires the field there and ignores it elsewhere, so withholding it on plain turns bought input tokens back with nothing observable lost against `api.deepseek.com`.
That endpoint is not the only one this adapter serves. `Config.baseURL` points it at any OpenAI-compatible endpoint, including a gateway that re-encodes a DeepSeek chat-completions conversation for another vendor. Such a gateway has no wire slot for the upstream thinking signature and recovers it by hashing the replayed chain of thought. A turn the model answered without calling a tool therefore reached the gateway with no reasoning text at all, the signature lookup found nothing, and the reconstructed conversation diverged from the recorded one. Agent runs call tools on most turns, so the loss appeared only at plain-answer turns and looked intermittent.
## Decision
`serializeAssistant` emits `reasoning_content` for every assistant turn whose content carried reasoning, independent of tool calls. An absent reasoning block still emits no field, so a non-thinking turn is unchanged.
The replayed text is byte-exact with what the provider streamed: `translate.ts` accumulates the whole `reasoning_content` channel of one response into a single reasoning block, so the join in `serializeAssistant` concatenates one member and a hash taken over the replay matches a hash taken over the original delivery.
## Alternatives considered
- **A `Config` switch selecting the passback policy.** The two endpoint behaviors are real, but the field is inert where it is unneeded, so the switch only ever buys back one turn's chain of thought in input tokens — against a wrong setting that silently makes a session unreconstructable, with no error at either end to attribute it to. A knob whose wrong position fails silently is worse than the tokens.
- **Deciding from `baseURL`.** Whether an endpoint forwards to another vendor is not readable from its host: an internal endpoint may proxy DeepSeek directly and a public one may forward. The adapter would be guessing at a deployment it cannot see through.
- **Carrying the signature durably instead, as `dsh-llm-pi-ai` does.** That adapter persists `thinkingSignature` per block in its replay state because its providers put the signature on the wire. DeepSeek chat-completions exposes none, so this adapter has nothing to persist and the replayed text is the only channel.
## Consequences
Every reasoned tool-call-free turn now costs its chain of thought in input tokens on later requests. The added text sits at that turn's position and is identical on every subsequent request, so the assembled prefix stays stable and only the first request spanning the change loses cache reuse from that point.
`WireAssistantMessage.reasoning_content` documents both endpoint behaviors, and the package README states the passback rule in the Wire-format notes and the Model Experience token and cache sections.
## Testing
`tests/serialize.spec.ts` pins all three assistant shapes: reasoning beside text with no tool call, reasoning beside a tool call, and a reasoning-only turn whose content stays `""`. Turns carrying no reasoning keep emitting no field, which the content-less and tool-call-only cases cover.

View file

@ -0,0 +1,33 @@
# Agent Note: DeepSeek reasoning passback on every reasoned turn
Status: implemented
[English](2026-08-19-deepseek-reasoning-passback-every-turn.md) | 中文
## Problem
`dsh-llm-deepseek` 只在同时携带工具调用的 assistant 轮次上,才把 `reasoning_content` 回放进历史。DeepSeek 思考模式文档在这类轮次上要求该字段,在其他轮次上会忽略它,因此在普通轮次上不回传能省下输入 token,对 `api.deepseek.com` 而言没有任何可观测的损失。
但该端点不是这个适配器唯一服务的对象。`Config.baseURL` 可以把它指向任何 OpenAI 兼容端点,包括把 DeepSeek chat-completions 对话重新编码转发给其他厂商的网关。这类网关在协议上没有承载上游思考签名的字段,只能对回放的思维链取哈希来恢复它。于是模型未调用工具就作答的轮次到达网关时完全不带推理文本,签名查找落空,重建出的对话与记录中的对话产生分叉。Agent 运行的大多数轮次都会调用工具,所以这个损失只在纯作答轮次上出现,表现为偶发。
## Decision
`serializeAssistant` 对每个内容携带推理的 assistant 轮次都发出 `reasoning_content`,与是否有工具调用无关。没有推理块时仍然不发出该字段,因此非思考轮次的行为不变。
回放文本与提供方流式下发的内容逐字一致:`translate.ts` 会把一次响应的整个 `reasoning_content` 通道累积进单个推理块,因此 `serializeAssistant` 中的拼接只连接一个成员,对回放取的哈希与对原始下发取的哈希相同。
## Alternatives considered
- **用 `Config` 开关选择回传策略。** 两种端点行为都真实存在,但该字段在不需要它的地方是惰性的,所以这个开关最多只换回一个轮次的思维链输入 token —— 代价却是一旦设置错误,会话就会静默地无法重建,两端都不会报错来归因。一个设错就静默失败的旋钮,比那点 token 更糟。
- **根据 `baseURL` 判断。** 一个端点是否会转发给其他厂商,无法从它的主机名读出:内部端点可能直连代理 DeepSeek,公网端点也可能转发。适配器只能对自己看不透的部署方式做猜测。
- **改为持久化签名,如 `dsh-llm-pi-ai` 的做法。** 该适配器在 replay state 中按块持久化 `thinkingSignature`,因为它的提供方会把签名放在协议里。DeepSeek chat-completions 不暴露签名,所以这个适配器没有可持久化的东西,回放文本是唯一通道。
## Consequences
每个含推理且不带工具调用的轮次,如今都会在后续请求中按其思维链计入输入 token。新增文本位于该轮次所在位置,且在此后每次请求中都相同,因此组装出的前缀保持稳定,只有跨越此次变更的第一个请求会从该位置起失去缓存复用。
`WireAssistantMessage.reasoning_content` 记录了两种端点行为,包 README 在协议格式说明以及 Model Experience 的 token 与缓存小节中陈述了该回传规则。
## Testing
`tests/serialize.spec.ts` 固定了三种 assistant 形态:推理与文本并存且无工具调用、推理与工具调用并存、以及内容保持为 `""` 的纯推理轮次。不携带推理的轮次仍不发出该字段,由无内容与仅工具调用两种用例覆盖。

View file

@ -2,5 +2,5 @@
# side as of the last confirmed-consistent state. Both languages carry equal authority;
# after editing either side, bring the other along and re-record with:
# pnpm run verify-translation-pairing --write packages/llm/llm-deepseek/README.md
README.md: 599dc1f530884df58fcc64d8cdf6c62fac80bc1f
README.zh.md: b1a45a1f398fdbb11cddebe34b646edf37f579be
README.md: 1c007d34f329177603305da3858c63960e30a898
README.zh.md: 181b30bd034ca27e983d4e7a54b03676fbaaba29

View file

@ -78,7 +78,7 @@ DeepSeek request identity is separate from app attribution. After credential res
- Streaming only (`stream_options.include_usage` always on). `usage` may arrive attached to the finish chunk or as a trailing usage-only chunk — the translator defers both to `[DONE]`, so `usage` always precedes `finish` and nothing follows `finish`.
- The adapter-owned `off` effort maps to `thinking: {type: 'disabled'}` and never crosses the wire as `reasoning_effort: 'off'`.
- The first thinking-mode chunk carries `reasoning_content: ""` — handled (no spurious reasoning block).
- **Reasoning passback rule**: on assistant turns that carried tool calls, `reasoning_content` is serialized back in history (required by the API in thinking mode); on tool-call-free turns it is dropped (ignored anyway — saves tokens).
- **Reasoning passback rule**: every assistant turn that carried reasoning serializes `reasoning_content` back in history. Thinking mode requires it on tool-call turns; DeepSeek ignores it elsewhere, while a gateway re-encoding the conversation for another vendor recovers that turn's upstream thinking signature by hashing the replayed text.
- Image-capable user messages preserve text/image order. Tool-role content remains a string; consecutive tool-result images are grouped into the following user message with `Attached image(s) from tool result:`.
- Cache accounting: `cacheReadTokens` ← `prompt_cache_hit_tokens` / `prompt_tokens_details.cached_tokens`; DeepSeek reports no cache-write metric.
@ -92,15 +92,15 @@ Non-2xx responses throw `LlmError` with stable codes: `AUTH` (401/403), `QUOTA`
#### What the model sees
The selected DeepSeek model receives the harness system prompt, message history, tool schemas, stop sequences, and call config without adapter-authored prompt prose. The vision model also receives retained user and tool-result images as base64 data URLs; an over-budget older image is represented by the documented placeholder. On a prior assistant turn with tool calls, its reasoning content is passed back as required; reasoning from tool-call-free turns is omitted.
The selected DeepSeek model receives the harness system prompt, message history, tool schemas, stop sequences, and call config without adapter-authored prompt prose. The vision model also receives retained user and tool-result images as base64 data URLs; an over-budget older image is represented by the documented placeholder. Reasoning content from a prior assistant turn is passed back verbatim, whether or not that turn called a tool.
#### Token effect
Provider tokenization governs exact text and image-token input. Conditional reasoning passback increases tool-round-trip context, while dropping other reasoning and over-budget images avoids paying those tokens again; cache-read usage is reported when available.
Provider tokenization governs exact text and image-token input. Reasoning passback carries every reasoned turn's chain of thought into later requests, while dropping over-budget images avoids paying those tokens again; cache-read usage is reported when available.
#### KV Cache effect
An unchanged assembled prefix, including deterministically encoded retained images and placeholders, is eligible for DeepSeek cache reuse, which this adapter reports in usage. A model-route change or any upstream prompt, schema, prefix, history, or image-budget change may prevent reuse from the first changed token; reasoning passback appends during tool round trips.
An unchanged assembled prefix, including deterministically encoded retained images and placeholders, is eligible for DeepSeek cache reuse, which this adapter reports in usage. A model-route change or any upstream prompt, schema, prefix, history, or image-budget change may prevent reuse from the first changed token; reasoning passback appends on every reasoned turn.
### DeepSeek response

View file

@ -78,7 +78,7 @@ DeepSeek 请求身份独立于应用归因。凭据解析成功后,每个提
- 只支持流式输出(`stream_options.include_usage` 始终开启)。`usage` 可能附着在 finish 分片上,也可能作为尾随的纯 usage 分片到达;转换器会将两者都延迟到 `[DONE]`,因此 `usage` 始终位于 `finish` 之前,`finish` 之后不会出现任何内容。
- 适配器持有的 `off` 推理强度映射为 `thinking: {type: 'disabled'}`,绝不会以 `reasoning_effort: 'off'` 通过协议发送。
- 第一个思考模式分片携带 `reasoning_content: ""`,系统会处理它(不会产生多余 reasoning 块)。
- **推理回传规则**:对携带工具调用的 assistant 轮次,会将 `reasoning_content` 序列化回历史(思考模式 API 必需);对不含工具调用的轮次,它会被丢弃(不会使用,可节省 token)。
- **推理回传规则**:每个携带推理内容的 assistant 轮次都会将 `reasoning_content` 序列化回历史。思考模式在工具调用轮次上必需它;DeepSeek 在其他轮次上会忽略它,而将该对话重新编码转发给其他厂商的网关,要靠对回传原文取哈希来恢复该轮次上游的思考签名。
- 支持图片的 user 消息会保留文本/图片顺序。Tool role 内容仍为字符串;连续工具结果中的图片会用 `Attached image(s) from tool result:` 汇总到随后一条 user 消息。
- Cache 计量:`cacheReadTokens` ← `prompt_cache_hit_tokens` / `prompt_tokens_details.cached_tokens`;DeepSeek 不报告 cache-write 指标。
@ -92,15 +92,15 @@ DeepSeek 请求身份独立于应用归因。凭据解析成功后,每个提
#### 模型看到的内容
所选 DeepSeek 模型会收到 harness 系统提示词、消息历史、工具 schema、stop sequence 和调用配置,不含适配器撰写的提示词文本。视觉模型还会通过 base64 data URL 收到保留的 user 与工具结果图片;超出上限的较旧图片由已记录的占位文本表示。当之前的 assistant 轮次包含工具调用时,会按要求回传其推理内容;不含工具调用的轮次会省略推理。
所选 DeepSeek 模型会收到 harness 系统提示词、消息历史、工具 schema、stop sequence 和调用配置,不含适配器撰写的提示词文本。视觉模型还会通过 base64 data URL 收到保留的 user 与工具结果图片;超出上限的较旧图片由已记录的占位文本表示。之前 assistant 轮次的推理内容会原文回传,无论该轮次是否调用了工具。
#### Token 影响
精确文本与图片 token 输入取决于提供方 tokenization。有条件推理回传会增加工具往返上下文,丢弃其他推理和超出上限的图片则避免再次支付这些 token;可用时会报告 cache-read 用量。
精确文本与图片 token 输入取决于提供方 tokenization。推理回传会把每个含推理轮次的思维链带入后续请求,丢弃超出上限的图片则避免再次支付这些 token;可用时会报告 cache-read 用量。
#### KV Cache 影响
未更改的已组装前缀,包括确定性编码的保留图片与占位文本,可使用 DeepSeek cache 复用,适配器会在 usage 中报告它。模型路由变更,或任何上游提示词、schema、前缀、历史或图片上限变更,都可能使从首个发生变化的 token 起的复用失效;推理回传会在工具往返期间追加。
未更改的已组装前缀,包括确定性编码的保留图片与占位文本,可使用 DeepSeek cache 复用,适配器会在 usage 中报告它。模型路由变更,或任何上游提示词、schema、前缀、历史或图片上限变更,都可能使从首个发生变化的 token 起的复用失效;推理回传会在每个含推理的轮次上追加。
### DeepSeek 响应

View file

@ -182,10 +182,12 @@ function serializeAssistant(message: Message): WireMessage {
// the message sits durably in the session log, a null here bricks every
// later turn of that session.
content: text,
// Official passback rule (guides/thinking_mode.mdx): reasoning_content
// must return on tool-call turns; it is ignored on plain turns, so we
// drop it there to save tokens.
...toolCalls.length > 0 && reasoning.length > 0 ? { reasoning_content: reasoning } : {},
// CoT passback on every reasoning-carrying turn. The official rule
// (guides/thinking_mode.mdx) requires it on tool-call turns and ignores it
// elsewhere; a gateway re-encoding the conversation for another vendor
// recovers that turn's upstream thinking signature by hashing this exact
// text, which a tool-call-free turn carries nowhere else.
...reasoning.length > 0 ? { reasoning_content: reasoning } : {},
...toolCalls.length > 0 ? { tool_calls: toolCalls } : {},
}
}

View file

@ -79,9 +79,11 @@ export interface WireAssistantMessage {
role: 'assistant'
content: string | null
/**
* CoT passback. REQUIRED on assistant turns that carried tool calls
* (thinking mode); ignored on tool-call-free turns (we omit it there to
* save tokens). See guides/thinking_mode.mdx § Tool Calls.
* CoT passback, present on every turn whose assistant content carried
* reasoning. REQUIRED on tool-call turns in thinking mode (see
* guides/thinking_mode.mdx § Tool Calls); DeepSeek ignores it elsewhere,
* while a gateway re-encoding for another vendor recovers that turn's
* thinking signature by hashing it.
*/
reasoning_content?: string
tool_calls?: WireToolCall[]

View file

@ -54,7 +54,7 @@ describe('serializeMessages', () => {
expect(wire).toEqual([{ role: 'system', content: 'be brief' }])
})
it('maps plain assistant text without reasoning_content', () => {
it('passes reasoning_content back on tool-call-free turns', () => {
const wire = serializeMessages([
createMessage({
role: 'assistant',
@ -65,8 +65,10 @@ describe('serializeMessages', () => {
source: { kind: 'plugin', plugin: 'test' },
}),
])
// Tool-call-free turn: reasoning is dropped (ignored by the API anyway).
expect(wire).toEqual([{ role: 'assistant', content: 'answer' }])
// A gateway that re-encodes the conversation for another vendor recovers
// the upstream thinking signature by hashing this exact text, and a turn
// that called no tool carries it nowhere else.
expect(wire).toEqual([{ role: 'assistant', content: 'answer', reasoning_content: 'thinking…' }])
})
it('passes reasoning_content back on tool-call turns (official passback rule)', () => {
@ -589,16 +591,17 @@ describe('review fixes: assistant content shapes', () => {
expect(wire).toEqual([{ role: 'assistant', content: '' }])
})
it('serializes a reasoning-ONLY assistant message as "" content with the reasoning dropped', () => {
it('serializes a reasoning-ONLY assistant message as "" content beside its reasoning', () => {
// The model can answer entirely in the reasoning channel (a v4-flash
// greeting did, live). The passback rule keeps reasoning_content off
// plain turns, and content must still be SET — a null here poisoned the
// session log and bricked every later turn of that session.
// greeting did, live). Content must still be SET — a null here poisoned
// the session log and bricked every later turn of that session.
const wire = serializeMessages([createMessage({
role: 'assistant', content: [{ type: 'reasoning', text: '你好!有什么我可以帮你的吗?' }],
source: { kind: 'plugin', plugin: 'test' },
})])
expect(wire).toEqual([{ role: 'assistant', content: '' }])
expect(wire).toEqual([{
role: 'assistant', content: '', reasoning_content: '你好!有什么我可以帮你的吗?',
}])
})
it('serializes tool-call turns with empty string content, not null', () => {