deepseek-harness/packages/web/tool-web/README.md

130 lines
7.9 KiB
Markdown
Raw Normal View History

# @deepseek-ai/dsh-tool-web
The model-facing web tool suite — `web_search` and `web_fetch` — over the [web capability seam](../web/README.md) (`ctx.web`). It owns model-facing concerns only: tool names, JSON schemas, snake_case argument names, prompt sections, the result-count bound, result formatting, HTML→markdown presentation, and `presentCall`. All web access goes through `ctx.web`; this package never imports a concrete provider. Neither tool exposes a model-facing timeout — each tool's cooperative tool-call budget is declared here via config (`fetchTimeoutMs`/`searchTimeoutMs`, attached as `ToolDefinition.timeoutMs`) and enforced by [`@deepseek-ai/dsh-timeout-policy`](../../timeout/timeout-policy/README.md) (a `tools/execute` wrapper); each tool just forwards `exec.signal` to the seam.
Each tool is registered independently; a product that wants only one disables the other via config (`{ search: false }` / `{ fetch: false }`).
## Tools
| Tool | Args | Behavior |
|---|---|---|
Fix review findings: validate the hooks cap, integer read caps, doc drift, config plumb-through test A Codex review pass on the draft caught four real gaps and two solid suggestions; all addressed except one pushed back on the merits: - hooks-claude/hooks-codex: stderrSummaryMaxChars was the one new knob with NO range validation — a negative/NaN cap would silently misbehave inside slice(). Both bridges now assert a positive integer at the TOP of apply() (before the config-file parse's early return, so a bad value fails the load loudly), with rejection tests. - tool-fs: the read caps count lines/chars/bytes, so positive-FINITE was too loose (a fractional readLimit would flow into windowing arithmetic and the schema description). All four now require a positive integer, matching tool-web's cap. - Doc drift the gates cannot catch: tool-web's README tools table still named WEB_SEARCH_MAX_RESULTS as the mechanism; compact-basic's README/module doc and the compaction-capability-seam RFC still described estimation as fixed char/4 rather than the charsPerToken default. - subagent-acp: the dispose graces were tested only at the startAcpRun level, so a regression that stopped threading plugin config into AcpRunSpec would have survived. A provider-path test now drives the trap-escalation scenario through ctx.subagents.start with small config graces and bounds dispose at 4s. Pushed back on: converting compact-basic's charsPerToken to a schemastery field. The package's whole config is deliberately hand-rolled (resolveConfig, every threshold REQUIRED with no default — a documented design posture); one schemastery field beside it would be incoherent. The knob is cordis.yml-reachable, defaulted, and validated, which is what the convention requires; migrating the package to schemastery wholesale is pre-existing config-surface hygiene out of this change's scope.
2026-07-04 18:06:35 +08:00
| `web_search` | `query` (string) | Discovery. Returns an optional answer plus source URLs. `max_results` is **not** model-facing — the tool sets the bound (the `searchMaxResults` config, default 8) and passes it to the seam. |
| `web_fetch` | `url` (string) | Retrieves a specific URL. HTML bodies are rendered to markdown-ish text; text bodies pass through. A non-2xx status is reported, not an error. The tool-call timeout is deployment policy (`dsh-timeout-policy`), not a model argument. |
2026-07-18 14:59:26 +08:00
Both tools opt into concurrent scheduling because provider reads return content without mutating parent-agent state.
2026-07-21 03:08:35 +08:00
The normalized seam results are also the canonical tool values: `WebSearchResult` and `WebFetchResult`. Native renderers preserve the answer/source and fetched-body text below; provider search/body caps remain acquisition limits rather than presentation-only truncation.
## Config
| Key | Default | Meaning |
|---|---|---|
| `search` | `true` | Register `web_search`. |
| `fetch` | `true` | Register `web_fetch`. |
Expose audited hardcoded tunables as plugin config The audit swept every packages/*/* plugin for the new AGENTS.md convention (no hardcoded tunables in plugins) and exposes each finding as a defaulted, validated Config field. Defaults are the previously hardcoded values throughout, so no deployment or golden changes. - tool-fs (had NO Config): readLimit, readMaxLineLength, readMaxBytes, readStreamMinSize. The caps thread through ReadToolCaps/ReadWindow — read-render already documented that the consumer applies the caps, so they become explicit per-request fields. - tool-web: searchMaxResults (WEB_SEARCH_MAX_RESULTS stays as the schemastery default). Also fixes the stale GREP_LIMIT references in search.ts and the web-capability-seam RFC (no such constant exists). - bash-local: graceMs (SIGTERM->SIGKILL escalation grace). The RunInternals.graceMs test seam is gone: graceMs is now a required SpawnSpec field filled from config, so tests exercise the real config path and the defaults live in exactly one place. - subagent-acp: disposeEofGraceMs / disposeGraceMs. The AcpRunSpec fields become required for the same one-defaulting-layer reason. - session-persistence-sqlite: journalMode ('wal' default; the rollback-journal modes serve filesystems where WAL's shared-memory files do not work, e.g. network mounts). - hooks-claude + hooks-codex: stderrSummaryMaxChars for the persisted hook/result stderr summary. The duplicated summarize() helpers merge into hook-protocol's summarizeStderr(stderr, maxChars), beside the HookResultRecord field it feeds, with the bound parameterized the same way runHook's defaultTimeoutMs already is. - compact-basic: charsPerToken for the token estimator (default 4, the English-text heuristic; CJK-heavy deployments need ~1-2 or compaction fires far too late). Also corrects the BasicCompactService class doc, which claimed defaults the required-field config never had. - fs-local: deletes the dead STREAM_MIN_SIZE constant and the dead FsIoInternals.streamMinSize seam — the read-routing bound lives in the consumer (tool-fs), where it is now config. This is item 1 of the proposed prune-write-only-fs-surface RFC, annotated accordingly. Every new field gets range validation (following the existing assertPositiveFinite pattern), a README row, and tests covering the configured behavior, the schema default, and load-time rejection.
2026-07-04 17:37:23 +08:00
| `searchMaxResults` | `8` | Upper bound on sources returned by one `web_search` call (the seam truncates a longer provider list and flags it). |
| `fetchTimeoutMs` | `30000` | Cooperative tool-call timeout budget (ms) for `web_fetch`. |
| `searchTimeoutMs` | `30000` | Cooperative tool-call timeout budget (ms) for `web_search`. |
`fetchTimeoutMs`/`searchTimeoutMs` declare each tool's cooperative timeout budget (attached as `ToolDefinition.timeoutMs`), enforced by [`@deepseek-ai/dsh-timeout-policy`](../../timeout/timeout-policy/README.md); the model-facing schema exposes no timeout argument.
```yaml
- id: tool-web
name: '@deepseek-ai/dsh-tool-web'
```
## Stable registration
Tool registration follows product **enablement**, not backend availability. A tool stays visible even when its selected provider is missing, misconfigured, ambiguous, or temporarily unavailable; the seam resolves the provider at execution time and execution fails with a structured `WebError` (e.g. `WEB_PROVIDER_UNAVAILABLE`, `WEB_PROVIDER_AMBIGUOUS`), which `ToolRegistry.execute()` turns into an error tool result the model can read and hooks/UI can route on. This keeps the model schema stable without making plugin load order, credential state, or HMR timing part of the model-facing contract. To remove a web tool entirely, disable it here in config.
2026-07-14 04:17:38 +08:00
The tool never calls a provider's `available()` and never enumerates providers — its only execution path is `ctx.web.search()` / `ctx.web.fetch()`, and provider unavailability reaches it as the structured `WebError` codes selection throws at execution time. Provider selection stays entirely inside the seam, with one owner.
Add a gated Known Limitations and Deferred Work section to every package README Every packages/*/* README now carries a canonical '## Known Limitations and Deferred Work' section: condensed, evidence-backed bullets for consumer-visible gaps (unimplemented features, platform caveats, MVP cuts) and consciously postponed work (TODO/FIXME/XXX markers, RFC deferrals still open). The ten pre-existing ad-hoc variants ('What is NOT here (TODO)', 'Deferred', 'Limitations (MVP)', 'Known limitations (tracked TODOs)', ...) are normalized into the canonical heading. A new doc-sync gate, scripts/verify-readme-limitations.ts, enforces the shape: exactly one limitations-like heading per package README, byte-equal to the canonical h2, with at least one bullet; near-miss headings fail so variants cannot creep back. Packages with genuinely nothing to declare (dsh-brand, dsh-timeout, dsh-subagent-mock, dsh-app-boot) are whitelisted in the script and must NOT carry the section; whitelist entries are validated against the scanned package set so a rename fails loud. Wired into the doc-sync chain (package.json) and the run-gates doc-sync leaf set; the standing rule lands in packages/AGENTS.md and the adding-a-package cookbook; decision record in docs/rfc/implemented/process/2026-07-10-readme-known-limitations-gate.md (RFC index regenerated). Also fixes two stale '(deferred)' markers claiming dsh-compact-basic is unimplemented (the dsh-compact seam README's package table and the seam's module doc comment).
2026-07-10 01:51:50 +08:00
2026-07-12 02:55:26 +08:00
## Model Experience
### System prompt
#### What the model sees
Search and fetch contribute the web-search and web-fetch guidance below. A scoped tool restriction does not remove these independently registered sections.
##### Web search guidance
```markdown
Use the web_search tool to discover current information on the web. It returns an optional answer plus a list of source URLs. Follow up with web_fetch when you need the full content of a specific result, and cite the relevant URLs as markdown links.
```
##### Web fetch guidance
```markdown
Use the web_fetch tool to retrieve the content of a specific HTTP(S) URL (for example a result from web_search). It returns the page content decoded to text. Cite the URL as a markdown link when you use its content.
```
#### Token effect
Fixed guidance cost per request for each config-enabled tool, even when a restriction hides its schema.
#### KV Cache effect
Prefix-stable while enabled tools, scope, and guidance text are unchanged. Config enablement or plugin lifecycle may invalidate reuse from the first changed prompt section; scoped schema restrictions do not remove it.
### Tool schemas
#### What the model sees
The model sees the generated [`web_search` and `web_fetch` schemas](../../../docs/tool-catalog.md#deepseek-aidsh-tool-web). Result-count and timeout budgets are deployment settings, not model arguments.
#### Token effect
Fixed schema cost per request; config disablement removes both schema and guidance, while a scoped restriction removes only the schema.
#### KV Cache effect
Prefix-stable while definitions and visibility are unchanged. Config enablement, plugin lifecycle, or scoped restrictions may invalidate reuse from the first changed schema token.
### Search result
#### What the model sees
The optional provider-owned answer is followed by `Sources:` and data-dependent lines shaped exactly `- [<title-or-url>](<url>)`, optionally suffixed ` — <snippet> (<publishedAt>)`. With neither answer nor sources the result says `No results found.` A capped list adds `(Showing the first <count> sources. Refine the query for more.)`; every result ends `Cite the relevant URLs above as markdown links in your answer.`
#### Token effect
Data-dependent results are resent until compaction and sources are capped by `searchMaxResults`.
#### KV Cache effect
Append-only; newly visible content follows the reusable request prefix and does not invalidate existing KV-cache entries.
### Fetch result
#### What the model sees
A successful fetch is exactly `Fetched <finalUrl> (HTTP <statusCode>)`, a blank line, and the provider-owned decoded body. Truncation adds a blank line and `(Content truncated. Fetch a more specific URL or section for the full text.)`; failures become `Error: <message>`. Queries and URLs remain in call history.
#### Token effect
Provider caps bound body size; retained call arguments and results are resent until compaction, and timeout policy can replace a late result with a short error.
#### KV Cache effect
Append-only; newly visible content follows the reusable request prefix and does not invalidate existing KV-cache entries.
### Argument errors
#### What the model sees
Blank inputs become exactly `Error: query must be a non-empty string` or `Error: url must be a non-empty string`.
#### Token effect
Only the failing call adds these retained tokens.
#### KV Cache effect
Append-only; newly visible content follows the reusable request prefix and does not invalidate existing KV-cache entries.
Add a gated Known Limitations and Deferred Work section to every package README Every packages/*/* README now carries a canonical '## Known Limitations and Deferred Work' section: condensed, evidence-backed bullets for consumer-visible gaps (unimplemented features, platform caveats, MVP cuts) and consciously postponed work (TODO/FIXME/XXX markers, RFC deferrals still open). The ten pre-existing ad-hoc variants ('What is NOT here (TODO)', 'Deferred', 'Limitations (MVP)', 'Known limitations (tracked TODOs)', ...) are normalized into the canonical heading. A new doc-sync gate, scripts/verify-readme-limitations.ts, enforces the shape: exactly one limitations-like heading per package README, byte-equal to the canonical h2, with at least one bullet; near-miss headings fail so variants cannot creep back. Packages with genuinely nothing to declare (dsh-brand, dsh-timeout, dsh-subagent-mock, dsh-app-boot) are whitelisted in the script and must NOT carry the section; whitelist entries are validated against the scanned package set so a rename fails loud. Wired into the doc-sync chain (package.json) and the run-gates doc-sync leaf set; the standing rule lands in packages/AGENTS.md and the adding-a-package cookbook; decision record in docs/rfc/implemented/process/2026-07-10-readme-known-limitations-gate.md (RFC index regenerated). Also fixes two stale '(deferred)' markers claiming dsh-compact-basic is unimplemented (the dsh-compact seam README's package table and the seam's module doc comment).
2026-07-10 01:51:50 +08:00
## Known Limitations and Deferred Work
- **`htmlToMarkdown` is a minimal regex converter, not an HTML parser** — it strips script/style/noscript, keeps headings/bullets/links, and decodes about a dozen named entities; tables, images, and nested formatting are lost.
2026-07-19 22:50:49 +08:00
- **The model-facing surface is minimal by design, with promotions deferred** — `max_results` stays a config bound (not a model argument), and `web_fetch` takes only `url` (no `format`/`prompt`/LLM-summarization mode); both are named later steps in [the seam Agent Note](../../../.agents/notes/implemented/architecture/2026-06-24-web-capability-seam.md).
- **No web-specific permission policy** — both tools execute without requesting `ctx.approval`; a deployment that needs confirmation must add a `tools/pre-execute` policy, and the package does not define persistent URL/domain grants.