{"_id":"@creait/dsh-tool-disclosure","name":"@creait/dsh-tool-disclosure","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@creait/dsh-tool-disclosure","description":"Progressive tool disclosure for DeepSeek Harness: rarely-used tool groups cost one catalog line instead of their full schemas until the model loads them with tool_search.","version":"0.1.0","type":"module","main":"lib/index.js","exports":{".":{"default":"./lib/index.js"},"./catalog":{"default":"./lib/catalog.js"},"./client":"./client/client.cjs","./package.json":"./package.json"},"dsh":{"bundle":{"patch":"./cordis.patch.yml"},"client":{"inject":["@deepseek-ai/dsh-client-runtime","@deepseek-ai/dsh-client-locale","@deepseek-ai/dsh-api-remotes"],"platform":"web"}},"dependencies":{"@deepseek-ai/dsh-settings":"^0.1.0-rc.6","@deepseek-ai/schemastery":"^3.18.1"},"peerDependencies":{"@deepseek-ai/cordis":"^4.0.1","@deepseek-ai/dsh-tools":"^0.1.0-rc.7"},"license":"MIT","engines":{"node":">=20"},"keywords":["dsh","deepseek-harness","cordis","plugin","tools","context","progressive-disclosure","mcp"],"repository":{"type":"git","url":"git+https://github.com/CREAIT-nl/dsh-plugins.git","directory":"tool-disclosure"},"homepage":"https://github.com/CREAIT-nl/dsh-plugins/tree/main/tool-disclosure#readme","bugs":{"url":"https://github.com/CREAIT-nl/dsh-plugins/issues"},"author":{"name":"Francesco G","email":"francesco@creait.nl"},"scripts":{"test":"node --test test/*.test.js"},"gitHead":"488ad11d2a6582854dc98b84feacf32e0a2d2337","_id":"@creait/dsh-tool-disclosure@0.1.0","_nodeVersion":"25.8.1","_npmVersion":"11.11.0","dist":{"integrity":"sha512-nd/l3JE+mBESy42xsG9xBy3E6d3vnp9TyH07ZzOe9Fj05QMTxY9/7smAXS0zvXsokelka/DDMeQopi5dg0dw3A==","shasum":"656906226c916535833f3eb13b8f688c9bc574c3","tarball":"https://registry.npmjs.org/@creait/dsh-tool-disclosure/-/dsh-tool-disclosure-0.1.0.tgz","fileCount":8,"unpackedSize":87241,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCICT8WgigrxCfFdcac3pcRXcN6rAccFClT/x0uvjafx1eAiEA3xWJS4ejIYge1JFdt+6BzJZ+Qu6CB4KSvzXXkbFXEyo="}]},"_npmUser":{"name":"creait","email":"francesco@creait.nl"},"directories":{},"maintainers":[{"name":"creait","email":"francesco@creait.nl"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/dsh-tool-disclosure_0.1.0_1787577871198_0.324522828859223"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-24T13:24:31.061Z","0.1.0":"2026-08-24T13:24:31.316Z","modified":"2026-08-24T13:24:31.619Z"},"maintainers":[{"name":"creait","email":"francesco@creait.nl"}],"description":"Progressive tool disclosure for DeepSeek Harness: rarely-used tool groups cost one catalog line instead of their full schemas until the model loads them with tool_search.","homepage":"https://github.com/CREAIT-nl/dsh-plugins/tree/main/tool-disclosure#readme","keywords":["dsh","deepseek-harness","cordis","plugin","tools","context","progressive-disclosure","mcp"],"repository":{"type":"git","url":"git+https://github.com/CREAIT-nl/dsh-plugins.git","directory":"tool-disclosure"},"author":{"name":"Francesco G","email":"francesco@creait.nl"},"bugs":{"url":"https://github.com/CREAIT-nl/dsh-plugins/issues"},"license":"MIT","readme":"# @creait/dsh-tool-disclosure\n\nProgressive tool disclosure for DeepSeek Harness. A rarely-used group of tools\ncosts one line of catalog instead of its full schemas, until the model asks for\nit.\n\n## Why this exists\n\ndsh advertises every registered tool on every request, in every mode, for the\nwhole session. There is no progressive disclosure anywhere in the harness: the\nagent loop reassembles the system prompt at each step and hands\n`assembly.tools` straight to the request builder, so a tool that is mounted is\na tool you pay for — first token to last.\n\nMost of the time that is the right trade. `bash`, `read`, `edit`, `grep`,\n`web_search` are used in the majority of sessions, and a round trip to unlock\nthem would cost more than their schemas ever do.\n\nAn MCP server is where it stops being the right trade. Adding the Playwright\nMCP server to the host plane mounts 24 browser tools whose schemas measure\n**18.5 KB of JSON**, and the input-token count of a live session moved\n**18,322 → 22,765** the moment they appeared: **+4.4k tokens on every request\nof every session**, including the overwhelming majority that never open a\nbrowser.\n\nThe rule this package encodes:\n\n> Defer a group when `schema_cost × sessions_that_never_use_it` beats\n> `one_round_trip × sessions_that_do`.\n\nbash, file, search, web, todo, skill and subagent tools lose that comparison.\nA browser fleet wins it outright.\n\n## What the model sees\n\nInstead of 24 schemas, the system prompt carries one section:\n\n```\nSome of your tools are not listed in your tool schemas yet. Each line below is\na capability you HAVE but cannot call until you load it. Call tool_search with\na group name — or a few words naming the capability you need — and that group's\ntools join your tool list from your next step onward; then call them like any\nother tool. Never tell the user something is beyond you without searching here\nfirst.\n- browser (24 tools): Drive a real Chrome and act in it: open a URL, read the\n  page as an accessibility snapshot, click, type, fill and submit forms, …\n```\n\nThe model calls `tool_search({query: 'browser'})` — or `'I need to click a\nbutton on a page'` — the group's tools join its list from the next step, and it\ncalls them normally. The catalog entry disappears once the group is loaded.\n\nLoads are **per agent session**, held in a `WeakMap` keyed by the agent object,\nso one session opening the browser does not widen anyone else's context, and\nnothing has to be cleaned up when the session ends.\n\n## Measured\n\nOne live Standard-mode session, read out of the session transcript's\n`request/header` records:\n\n| | Tools advertised | Tool schemas | Catalog |\n| --- | --- | --- | --- |\n| Deferred | 34 | 29,816 chars | 956 chars |\n| After `tool_search('browser')` | 58 | 46,870 chars | — |\n\n**17,054 chars (~4.1k tokens) of schema, for 956 chars (~228 tokens) of\ncatalog.** Net saving of roughly 3.9k tokens per request, held for as long as\nthe session does not need a browser — and the first turn of that session\nmeasured 10,699 input tokens rather than the ~14.8k it would otherwise have\nbeen.\n\n## Install\n\n```bash\ndsh plugin --profile web add @creait/dsh-tool-disclosure\n```\n\nThen switch a group on — in **Settings → Tool Disclosure**, or in your profile\npatch, which also lets you give the group a better id and a summary written by\nhand:\n\n```yaml\n- insert:\n    - id: tool-disclosure\n      name: '@creait/dsh-tool-disclosure'\n      config:\n        defer: ['browser']\n        groups:\n          - id: browser\n            match: ['mcp__playwright__*']\n            summary: >-\n              Drive a real Chrome and act in it: open a URL, read the page as an\n              accessibility snapshot, click, type, fill and submit forms, select\n              options, hover, drag, press keys, upload files, handle dialogs,\n              switch tabs, resize, screenshot, read console messages and network\n              requests, and run JavaScript in the page. Reach for it when a page\n              needs JavaScript to render, sits behind a login, or has to be\n              clicked through — web_fetch already handles a plain read.\n```\n\nRestart `dsh` — the boot manifest is assembled at startup.\n\nMount it on the **host plane**, unscoped. dsh dispatches the\n`system-prompt/assemble` waterfall per scope, and an unscoped listener receives\nevery scoped dispatch, so one row covers every agent and every mode.\n\n## Config\n\n| Key | Default | Meaning |\n| --- | --- | --- |\n| `defer` | `[]` | The group ids being held back. The one switch. Written by the settings page. Names, never globs — an entry carrying `*` is dropped. |\n| `keep` | `[]` | Globs never deferred, even when a group's `match` claims them. |\n| `groups` | `[]` | Optional annotations. Each needs `id`, `match` and `summary`. |\n| `groups[].id` | — | What the model passes to `tool_search`, and what goes in `defer`. |\n| `groups[].match` | — | Tool-name globs (`*` only) the group claims. First group claiming a name wins, so config order is the tiebreak. |\n| `groups[].summary` | — | The **only** thing the model knows before loading. Name capabilities, not packages. |\n\nWith `defer` empty the plugin mounts nothing at all — not even its own tool\nschema.\n\n**Every tool the registry holds is already in a group**, whether or not\nanything wrote one down. An MCP server's tools are bucketed under the server's\nname; every other tool stands alone. So `groups` is not what creates a group —\nit *annotates* one that exists either way, with a chosen id, globs of its own\nand a summary in someone's words rather than a derived one. `browser` above is\n`playwright` renamed and described.\n\nThat is why an unannotated group needs no `match`: its id is the MCP server\nname and the globs are derived from it (`<id>` and `mcp__<id>__*`). A server\nthat reconnects carrying three more tools is covered by the same entry, and one\nthat is down at boot still defers when it returns.\n\n`defer` in the patch is a default, not a lock: it lands as the settings **base**\nlayer, so the page opens with that switch on and can still turn it off.\n\nAn id is a name and not a pattern, which is why a `*` in one is dropped rather\nthan escaped: an unannotated group's globs are *derived* from its id, so\n`defer: ['*']` would compile to a matcher claiming every tool the registry\nholds — one line, and the model loses the lot. Globs belong in `match` and\n`keep`, where they are read as globs.\n\nAn annotation whose `match` covers only part of the bucket its id names keeps\nthe id, and the rest stays advertised: the switch defers what the annotation\nclaims and nothing else. The settings page shows those leftovers on the same\nrow and counts them apart from the saving, so a partial `match` reads as what\nit is rather than as a group that costs less than it does.\n\n`keep` is for pulling one tool out of an otherwise deferred group:\n\n```yaml\nkeep: ['mcp__playwright__browser_navigate']\n```\n\n### Writing a summary\n\nThis is the whole interface. The model decides whether to spend a round trip on\nthe strength of one line, so it should read as a list of things it could do —\n\"click, fill forms, read the console\" — not as a description of the package\nthat provides them. Saying when *not* to reach for the group (\"`web_fetch`\nalready handles a plain read\") is worth the words: it stops the group being\nloaded by a session that had a cheaper option.\n\n## Settings\n\nThe settings page owns one thing — whether each group is deferred or advertised\nin full — under **Settings → Tool Disclosure** (the shipped web UI labels the\nfirst of those **设置面板**). A switch applies immediately: the next assembly, in\nevery live session, uses the new value. No restart.\n\nIt is **one list**, costliest first, holding every tool the registry has. There\nis no separate section for the groups someone wrote down, because a group is\nnot created by being written down — an annotation gives one a better id and a\nhand-written summary, and both kinds get the same row and the same switch.\nCostliest first is the order that answers the question the page is for; config\norder would answer \"what did somebody write down\", which nobody is here to\ndecide. Each row names the tools behind it, so no group is only a number.\n\nListing the unannotated groups is the point of listing anything. An MCP server\nthat was mounted and forgotten is invisible until it has a row, and a page that\nshowed only the hand-written half while claiming to show what the harness\ncarries would be the more misleading of the two.\n\n`tool_search` itself never gets a row, and is spared even when a hand-edited\nconfig names it. It is the only call that loads a group back: defer it and the\nmodel loses every group at once, with nothing left that could undo it.\n\nThe page headlines what deferring currently buys, measured from the live\nregistry rather than from the config:\n\n| Figure | What it counts |\n| --- | --- |\n| Saved per request | Characters of schema held back, at ~4.16 chars/token. |\n| Tools deferred | How many tools that is, against the shared registry's total. |\n\nIt shows no \"still advertised\" figure. The measurement reads the **shared**\nregistry, and on the web surface each agent preset mounts its own copies of the\nmode tools per session — a live session advertises more than any global view\ncan see. What is held back is exact; what remains is not knowable from here, so\nthe page does not guess.\n\nOnly the switch list is persisted, as `defer` in the `dsh-tool-disclosure`\nsettings namespace. Group annotations stay in the patch, which is what lets one\nadded there later take effect without being masked by a stale copy of the whole\nlist in the user layer; a group stores nothing but an id, which is what lets an\nMCP server's tools change underneath a switch without the switch going stale. A\nwrite posts the list in full, rebuilt from what is on screen, and compares it\nagainst the **user** layer rather than the merged value — so editing the patch\ncannot silently move a switch you set by hand.\n\nIds are not checked against the registry on the way in. A group exists only\nwhile the thing behind it does, so filtering against what is connected right\nnow would quietly clear the switch of every MCP server that happened to be\ndown, and each would come back advertised. An id matching nothing defers\nnothing and costs nothing.\n\nThe page talks to two plugin-owned loopback routes rather than the settings RPC,\nbecause the harness settings wire only exposes namespaces on its own allowlist,\nwhich a plugin cannot widen:\n\n| Route | Purpose |\n|---|---|\n| `/api/dsh-tool-disclosure/config` | read/write the `defer` list |\n| `/api/dsh-tool-disclosure/groups` | every group the registry holds, measured now |\n\nBoth refuse anything that is not a loopback request, so a page opened over the\nnetwork reads nothing. It says so rather than sitting on a loading line: an\nunreachable route, an absent settings service and a 403 all arrive as the same\nrejected fetch, and a page that rendered them as \"Reading…\" would be claiming a\nrequest is still in flight when none is. The same line stays up beside figures a\nlater read failed to refresh, because a measurement nobody could re-take is a\nnumber from a moment that has passed.\n\nWith every group advertised, `tool_search` unregisters itself: a tool whose\nonly honest answer is \"nothing is deferred\" is a schema charged to every request\nfor no capability at all, which is the exact cost this package exists to remove.\n\n## Design notes\n\n**It filters presentation, not the registry.** dsh has\n`ctx.tools.restrict()`, and it is the wrong lever here — its own docstring says\na single resolver \"feeds presentation, lookup, and dispatch\", so a restricted\ntool is genuinely uncallable, and it validates names against\n`restrictableNames` at install time, which an MCP server registering after its\nhandshake cannot satisfy. This package removes tools from the assembled\nrequest only. A deferred tool stays fully callable: if the model names one from\nmemory, or a hook or a subagent invokes it, it runs.\n\n**Code Mode is left alone.** Under a `code` presentation the schemas render\ninto the generated SDK section of the prompt, which this row cannot filter — so\nit defers nothing there and renders no catalog, rather than claiming a saving\nit did not make.\n\n**Assemblies with no agent are left alone.** They have nowhere to record a\nload, so a catalog they could never act on would be a dead end.\n\n**The catalog counts live registrations.** A group whose MCP server has not\nfinished its handshake renders no line, so the model is never told about a\ncapability that is not there yet.\n\n**An unannotated group's summary is derived at read time, never stored.** It is\nbuilt from what the group is holding back at that assembly — the tools' short\nnames, or a lone tool's own first sentence. A summary written at boot and kept\nwould go on promising a tool the server has since dropped, which is the one\nthing the catalog must not do.\n\n## Query resolution\n\n`tool_search` resolves, in order: the reveal-all words (`*`, `all`,\n`everything`), then an exact group id, then term overlap against each group's\nid and summary with stopwords and sub-3-character terms dropped, best match\nfirst.\n\nLoose matching is deliberate. A query that resolves to nothing costs a whole\nround trip and teaches the model the catalog is unreliable; a query that loads\none group too many costs that group's schemas — which is what it would have\ncost anyway had the model asked for it directly.\n\n## Tests\n\n```bash\nnode --test test/*.test.js\n```\n\nPure logic (globs, partitioning, catalog rendering, query resolution) is tested\nagainst plain data; wiring is driven through a stub Cordis context that\nexercises the real `system-prompt/assemble` listener, the prompt section and\nthe tool — including per-agent isolation, the Code Mode bail-out, and a late\nMCP registration.\n\nThe settings page is driven the same way: the bundle is evaluated as the loader\nevaluates it, `apply` runs against stub slots so the real registration path is\nunder test, and a hook-faithful React stub drives what a browser cannot be made\nto show — a rejected write, a switch that moves before the round trip, the switch\nlist rebuilt from what is rendered, and the numbers coming back from the\nre-measure rather than from the optimistic flip.\n\n## What breaks this\n\n`system-prompt/assemble` is a pre-1.0 internal seam with no compatibility\nguarantee, and this package leans on three properties of it:\n\n- the waterfall's return value is authoritative for `assembly.tools`;\n- an unscoped listener receives every scoped dispatch, which is what makes one\n  host-plane row cover every agent;\n- the agent loop reassembles at **every** step, which is what makes a mid-turn\n  load take effect on the next one.\n\nIf any of those move, the failure is loud in the right direction — tools stop\nbeing deferred, so sessions get wider than they need to be rather than losing a\ncapability. `ctx.tools.modeFor` is probed defensively for the same reason: an\nabsent method is read as a native presentation.\n\nThe settings nav glyph is a deliberate reach past the API. `settings.section`\nhas no icon option — the shell picks the glyph from a hardcoded section-id map\nand falls back to the gear for ids it does not know, ours included. So the\nclient half repaints its own row: it finds the nav cell by label and swaps the\ngear's path geometry for a wrench, mutating the attribute rather than replacing\nthe node so React re-renders over it without restoring the gear. The path is\nhand-drawn — the shipped set has no wrench — at the stroke weight and box\nfootprint of its neighbours. It fails safe: if the shell's markup or the label\nmoves, nothing matches and the row keeps the gear.\n\n`peerDependencies` pins the versions this was built against; a harness upgrade\ncan move them.\n\n## Licence\n\nMIT\n","readmeFilename":"README.md","_rev":"1-5ac2860897353f08f6a9846be9cbf97a"}