deepseek-harness/snapshots/web/code-mode-round/ui.expected.md

56 lines
2 KiB
Markdown
Raw Normal View History

- banner:
- navigation "Session hierarchy":
- 'button "Using ONE run_code program: run" [disabled]'
- img
- text: Standard mode
- button "Session log":
- text: Session log
- img
- tablist:
- tab "Chat" [selected]
- tab "Trajectory"
- text: "Using ONE run_code program: run bash `echo CODE_ROUND_OK`, then read the file missing.txt catching its error in the program. Return an object with both outcomes. Then reply DONE and stop. {{clock}}"
2026-07-30 15:56:41 +08:00
- button "Copy":
- img
- button "Context injection @deepseek-ai/dsh-system-prompt":
- img
- img
- text: Context injection @deepseek-ai/dsh-system-prompt
- 'button "Think The user wants me to write a single `run_code` program that:"':
- img
- img
- text: "Think The user wants me to write a single `run_code` program that:"
- button "Code Run bash echo and catch missing file read":
- img
- img
- text: Code Run bash echo and catch missing file read
- img
- text: Bash Echo CODE_ROUND_OK Failed
- 'button "Read Error: cannot read \"{{cwd}}/workspace/missing.txt\": not found"':
- img
- text: "Read Error: cannot read \"{{cwd}}/workspace/missing.txt\": not found"
- button "Think The program ran successfully. Let me now reply DONE as instructed.":
- img
- img
- text: Think The program ran successfully. Let me now reply DONE as instructed.
- paragraph: DONE
2026-07-30 15:56:41 +08:00
- button "Copy":
- img
- button "Good response":
- img
- button "Bad response":
- img
2026-07-30 15:56:41 +08:00
- button "Branch into a new conversation":
- img
test(web): stabilize and refresh the aria goldens for the speed readings The TTFT and tok/s readings divide by measured wall time, so they are not reproducible: the same replayed scenario yielded 69 and 70 tok/s on consecutive local runs, and a 3 ms replayed stream reads 26333 tok/s. Baking those into committed goldens made the Web lane flaky by construction, and the goldens for the readings themselves were never refreshed. Three fixes, then a refresh: The footer's decorative dots are `aria-hidden`, so the readings concatenated into one accessible string — `Ran for 13sTTFT 0.2s12 tok/s`. That is a real defect on its own (a reader hears one run-on instead of three facts) and it also denied `{{duration}}` the word boundary it matches on, so even the previously-stable `Ran for` duration started leaking raw. The separators now carry flanking spaces. `normalizeAria` gains `{{throughput}}` beside `{{duration}}`, and its duration alternation accepts the stats line's compact `2m42s` as well as the message-chrome template's `2m 42s` — the compact form had no pattern at all, which is why `LLM 382m39s` survived the first refresh. Refreshed 17 goldens. They also record that the stats line's `LLM` group now renders at all: it folds assistant `timing`, which the live transcript adapter only began attaching in this branch, so the group was previously dead in Chat. Verified by running the lane in replay three times after the refresh: 41/41 files green each time, goldens untouched. Before this change two consecutive runs disagreed on both the values and the failure count.
2026-08-05 17:55:49 +08:00
- text: {{clock}} Ran for {{duration}} TTFT {{duration}} {{throughput}} tok/s
- textbox "Message the agent"
- button "Commands":
- img
- 'button "Access mode, current: Workspace Write"': Workspace Write
- button "Select model, current DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img
test(web): stabilize and refresh the aria goldens for the speed readings The TTFT and tok/s readings divide by measured wall time, so they are not reproducible: the same replayed scenario yielded 69 and 70 tok/s on consecutive local runs, and a 3 ms replayed stream reads 26333 tok/s. Baking those into committed goldens made the Web lane flaky by construction, and the goldens for the readings themselves were never refreshed. Three fixes, then a refresh: The footer's decorative dots are `aria-hidden`, so the readings concatenated into one accessible string — `Ran for 13sTTFT 0.2s12 tok/s`. That is a real defect on its own (a reader hears one run-on instead of three facts) and it also denied `{{duration}}` the word boundary it matches on, so even the previously-stable `Ran for` duration started leaking raw. The separators now carry flanking spaces. `normalizeAria` gains `{{throughput}}` beside `{{duration}}`, and its duration alternation accepts the stats line's compact `2m42s` as well as the message-chrome template's `2m 42s` — the compact form had no pattern at all, which is why `LLM 382m39s` survived the first refresh. Refreshed 17 goldens. They also record that the stats line's `LLM` group now renders at all: it folds assistant `timing`, which the live transcript adapter only began attaching in this branch, so the group was previously dead in Chat. Verified by running the lane in replay three times after the refresh: 41/41 files green each time, goldens untouched. Before this change two consecutive runs disagreed on both the values and the failure count.
2026-08-05 17:55:49 +08:00
- button "7% of context used"
- button "Send message" [disabled]
test(web): stabilize and refresh the aria goldens for the speed readings The TTFT and tok/s readings divide by measured wall time, so they are not reproducible: the same replayed scenario yielded 69 and 70 tok/s on consecutive local runs, and a 3 ms replayed stream reads 26333 tok/s. Baking those into committed goldens made the Web lane flaky by construction, and the goldens for the readings themselves were never refreshed. Three fixes, then a refresh: The footer's decorative dots are `aria-hidden`, so the readings concatenated into one accessible string — `Ran for 13sTTFT 0.2s12 tok/s`. That is a real defect on its own (a reader hears one run-on instead of three facts) and it also denied `{{duration}}` the word boundary it matches on, so even the previously-stable `Ran for` duration started leaking raw. The separators now carry flanking spaces. `normalizeAria` gains `{{throughput}}` beside `{{duration}}`, and its duration alternation accepts the stats line's compact `2m42s` as well as the message-chrome template's `2m 42s` — the compact form had no pattern at all, which is why `LLM 382m39s` survived the first refresh. Refreshed 17 goldens. They also record that the stats line's `LLM` group now renders at all: it folds assistant `timing`, which the live transcript adapter only began attaching in this branch, so the group was previously dead in Chat. Verified by running the lane in replay three times after the refresh: 41/41 files green each time, goldens untouched. Before this change two consecutive runs disagreed on both the values and the failure count.
2026-08-05 17:55:49 +08:00
- text: 1 turns · 2 steps LLM {{duration}} · Tool call {{duration}} TTFT avg {{duration}} · {{throughput}} tok/s Cache hit 52% Input 17.2K tok · Output 252 tok