{"_id":"@creait/dsh-gen-limit","_rev":"2-7b972bb362824de0a79b21137e2c844b","name":"@creait/dsh-gen-limit","dist-tags":{"latest":"0.2.0"},"versions":{"0.1.0":{"name":"@creait/dsh-gen-limit","version":"0.1.0","keywords":["dsh","deepseek-harness","cordis","plugin","concurrency","rate-limit","subagent"],"author":{"name":"Francesco G","email":"francesco@creait.nl"},"license":"MIT","_id":"@creait/dsh-gen-limit@0.1.0","maintainers":[{"name":"creait","email":"francesco@creait.nl"}],"homepage":"https://github.com/CREAIT-nl/dsh-plugins/tree/main/gen-limit#readme","bugs":{"url":"https://github.com/CREAIT-nl/dsh-plugins/issues"},"dsh":{"bundle":{"patch":"./cordis.patch.yml"},"client":{"inject":["@deepseek-ai/dsh-client-runtime","@deepseek-ai/dsh-client-locale","@deepseek-ai/dsh-api-remotes"],"platform":"web"}},"dist":{"shasum":"5377beac7e616f286244bc9b86bded74169a42c3","tarball":"https://registry.npmjs.org/@creait/dsh-gen-limit/-/dsh-gen-limit-0.1.0.tgz","fileCount":8,"integrity":"sha512-9j4rfXI95I3yAmHQWWWS/dhwCHjEIEOy5k/BahaFQQmy7wnljWI5jath0bkHAfQuZCPXdLOk4HgvI5icdb7Flg==","signatures":[{"sig":"MEYCIQDWXMimVwlPr3e19RIMoH81rKvwNrJGy15fzHEJgEwVcgIhAI/7fjzFgjRwABx/i7jc+pxNrBiJJ3nI5HP01V8TvNlv","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":61970},"main":"lib/index.js","type":"module","engines":{"node":">=20"},"exports":{".":{"default":"./lib/index.js"},"./client":"./client/client.cjs","./package.json":"./package.json"},"gitHead":"488ad11d2a6582854dc98b84feacf32e0a2d2337","scripts":{"test":"node --test test/*.test.js"},"_npmUser":{"name":"creait","email":"francesco@creait.nl"},"repository":{"url":"git+https://github.com/CREAIT-nl/dsh-plugins.git","type":"git","directory":"gen-limit"},"_npmVersion":"11.11.0","description":"Per provider/model concurrency limiter for generating sessions, with a themed settings card.","directories":{},"_nodeVersion":"25.8.1","dependencies":{"@deepseek-ai/schemastery":"^3.18.1","@deepseek-ai/dsh-settings":"^0.1.0-rc.6"},"_hasShrinkwrap":false,"peerDependencies":{"@deepseek-ai/cordis":"^4.0.1","@deepseek-ai/dsh-llm":"^0.1.0-rc.6"},"_npmOperationalInternal":{"tmp":"tmp/dsh-gen-limit_0.1.0_1787577860056_0.705443772958761","host":"s3://npm-registry-packages-npm-production"}},"0.2.0":{"name":"@creait/dsh-gen-limit","description":"Per provider/model concurrency limiter for generating sessions, with a themed settings card.","version":"0.2.0","type":"module","main":"lib/index.js","exports":{".":{"default":"./lib/index.js"},"./client":"./client/client.cjs","./package.json":"./package.json"},"dsh":{"bundle":{"patch":"./cordis.patch.yml"},"client":{"inject":["@deepseek-ai/dsh-client-runtime","@deepseek-ai/dsh-client-locale","@deepseek-ai/dsh-api-remotes"],"platform":"web"}},"dependencies":{"@deepseek-ai/dsh-settings":"^0.1.0-rc.6","@deepseek-ai/schemastery":"^3.18.1","undici":"^8.10.0"},"peerDependencies":{"@deepseek-ai/cordis":"^4.0.1","@deepseek-ai/dsh-llm":"^0.1.0-rc.6"},"license":"MIT","engines":{"node":">=22.19.0"},"keywords":["dsh","deepseek-harness","cordis","plugin","concurrency","rate-limit","subagent"],"repository":{"type":"git","url":"git+https://github.com/CREAIT-nl/dsh-plugins.git","directory":"gen-limit"},"homepage":"https://github.com/CREAIT-nl/dsh-plugins/tree/main/gen-limit#readme","bugs":{"url":"https://github.com/CREAIT-nl/dsh-plugins/issues"},"author":{"name":"Francesco G","email":"francesco@creait.nl"},"scripts":{"test":"node --test test/*.test.js"},"gitHead":"c6086194b046a24cdd782ddba46961151a1e5388","_id":"@creait/dsh-gen-limit@0.2.0","_nodeVersion":"24.19.0","_npmVersion":"12.0.2","dist":{"integrity":"sha512-xjCTJkWFQV5Uc99HOd6lCsuSKy8Caqh/pyABOySdrAE4zSRs//rZTxOGmoh/lTXdLvmI2oerqhMvMIW6edclkQ==","shasum":"cab2f9106d6b55aeb25ff436ae79cf14b89251c7","tarball":"https://registry.npmjs.org/@creait/dsh-gen-limit/-/dsh-gen-limit-0.2.0.tgz","fileCount":10,"unpackedSize":80414,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@creait%2fdsh-gen-limit@0.2.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIG6NX7B2C0HTiCijvCpVCqpiRWtpKs1uRgZY+fJ0rAO1AiBozplK+Uitb6ROOrCUf95Y5EhL8auUFttRu1mxo4CXeQ=="}]},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:22817eb7-6614-496b-a784-5a65ad6efef9"}},"directories":{},"maintainers":[{"name":"creait","email":"francesco@creait.nl"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/dsh-gen-limit_0.2.0_1787733820809_0.14375665616304656"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-24T13:24:19.844Z","modified":"2026-08-26T08:43:41.220Z","0.1.0":"2026-08-24T13:24:20.219Z","0.2.0":"2026-08-26T08:43:40.940Z"},"bugs":{"url":"https://github.com/CREAIT-nl/dsh-plugins/issues"},"author":{"name":"Francesco G","email":"francesco@creait.nl"},"license":"MIT","homepage":"https://github.com/CREAIT-nl/dsh-plugins/tree/main/gen-limit#readme","keywords":["dsh","deepseek-harness","cordis","plugin","concurrency","rate-limit","subagent"],"repository":{"type":"git","url":"git+https://github.com/CREAIT-nl/dsh-plugins.git","directory":"gen-limit"},"description":"Per provider/model concurrency limiter for generating sessions, with a themed settings card.","maintainers":[{"name":"creait","email":"francesco@creait.nl"}],"readme":"# @creait/dsh-gen-limit\n\nPer provider/model concurrency limits for DeepSeek Harness, with a settings card.\n\nSome backends fall over — or bill hard — when several sessions generate against\nthem at once. A self-hosted GPU serving one model has a real ceiling; a metered\nAPI has a financial one. dsh has no per-model concurrency control, so a single\nagent that fans out subagents can saturate either.\n\nThis caps how many sessions may generate concurrently on a given\nprovider/model. Work past the cap **waits in a FIFO queue** rather than failing:\na fan-out of eight researchers against a limit of three is a pacing problem, and\nbouncing five of them does not make the brief smaller, it just spends the retry\nbudget re-asking.\n\n## How it enforces\n\n**The count — `llm/stream` waterfall.** Every streaming model call is capped by\nthe number of **distinct sessions actually generating**. A session reentering\nthe loop is not counted twice, and a session parked waiting on a subagent holds\nno stream — the child is what counts. This is where the limit is enforced, and\nit is enforced the same way no matter how a session started: a subagent from a\ntool call, a child the workflow engine started directly, or a plain\nconversation.\n\n**The pacing — `tools/pre-execute`.** When an agent calls a subagent spawn tool\n(`subagent`, `subagent_fork`) and the target provider/model is already full, the\nspawn joins the same queue and is admitted as soon as there is room, so a\nfan-out is paced instead of piling more sessions onto a saturated backend.\n\nThis gate deliberately holds **no slot of its own**. Reserving one per admitted\nspawn is the obvious design and it double-counts: the child then takes a second\nslot the instant it generates, so every live child costs two. There is no clean\nseam to hand a reservation over either — `subagent/start` carries no\nback-reference to the spawn that caused it, and it also fires for children the\nworkflow engine starts without any tool call. So a child is counted exactly\nonce, where it can be counted consistently.\n\n**What a wait costs.** `queueTimeoutMs` bounds how long a request waits (`0`\nwaits indefinitely) and `maxQueued` bounds how many may wait at once — an\nunbounded queue in front of a slow backend is a memory leak that presents as a\nhang. Only a request that exhausts its wait fails, and it fails with the code\n`GEN_CAPACITY_EXCEEDED`; for a spawn that means the tool call is denied.\nReaching that point means the backend has been saturated for a sustained\nperiod, not that a request was unlucky with timing.\n\n**The consequence — transport timeouts.** Waiting for a slot means a stream may\nlegitimately go quiet for a long time, so the limiter also makes sure the socket\nagrees. `llm-pi-ai` lets a provider declare `streamIdleTimeoutMs`, but the SSE\nstream rides Node's built-in `fetch`, whose `bodyTimeout` defaults to five\nminutes and which nothing in the harness configures — so any value above\n300000ms is unreachable. Raising it just moves the kill from the harness\nwatchdog (`TIMEOUT`) to undici (`TypeError: terminated`, classified `TRANSPORT`,\nequally retryable), and each retry restarts the step from scratch and discards\neverything it had generated.\n\nSo `transport.js` reads the timeout **the provider already declares** and\ninstalls a dispatcher that applies it to that provider's origin, plus a\n30-second margin so the harness watchdog stays the one that reports a dead\nstream. There is nothing new to configure, and no other origin is affected — MCP\nservers, web fetches and the update check keep Node's defaults. A provider that\ndeclares no `streamIdleTimeoutMs`, or one under five minutes, is left alone.\n\n## Install\n\n```sh\ndsh plugin --profile web add @creait/dsh-gen-limit\n```\n\nThe package ships its own `cordis.patch.yml`, so it inserts its roster row on\nits own — no manual profile edit. Add it to `dsh.profile.bundles` to activate\nthe browser half.\n\n## Configure\n\nLimits live in the `dsh-gen-limit` settings namespace, one row per\nprovider/model. **`max: -1` means unlimited, and any pair without a row defaults\nto unlimited** — the plugin is inert until you give it a limit.\n\nSeed them from the row in your profile patch — `provider` and `model` are\nwhatever ids your own routes publish:\n\n```yaml\n- id: gen-limit\n  config:\n    limits:\n      - { provider: local-gpu, model: deepseek-v4-flash, max: 2 }\n      - { provider: anthropic, model: claude-opus-4, max: 1 }\n    queueTimeoutMs: 120000   # how long a request waits for a slot; 0 = forever\n    maxQueued: 64            # how many may wait at once\n```\n\nOr edit it in the GUI: **Settings → Plugins → Plugin config → Generation\nConcurrency** (the shipped web UI labels those **设置面板 → 插件 → 插件配置**).\nThe card lists the live providers and models from the same `llm` service the\nconversation uses, so the rows are pickable rather than typed from memory.\n\n## Routes\n\nThe card talks to three plugin-owned loopback routes rather than the settings\nRPC — the harness settings wire only exposes namespaces on its own allowlist,\nwhich a plugin cannot widen:\n\n| Route | Purpose |\n|---|---|\n| `/api/dsh-gen-limit/config` | read/write the limit rows |\n| `/api/dsh-gen-limit/catalog` | live provider/model list |\n| `/api/dsh-gen-limit/stats` | what is generating right now |\n\n## What breaks this\n\n`llm/stream` and `tools/pre-execute` are pre-1.0 internal seams with no\ncompatibility guarantee. `peerDependencies` pins the versions this was\nbuilt against; a harness upgrade can move them.\n\nThe transport half rests on the seam Node leaves for proxies: built-in `fetch`\ntakes no per-call timeout options and reads its dispatcher from a global that\n`undici`'s `setGlobalDispatcher` writes. Verified on Node 25.8.1 with undici\n8.10.0 — a 3s `bodyTimeout` installed this way killed a built-in `fetch` body at\n3.5s, where the default had taken 301s. It is a convention, not a contract; a\nruntime that stopped honouring it would put the five-minute ceiling back, which\nis where things stood before this existed.\n\nThe settings nav glyph is a deliberate reach past the API. `settings.section`\nhas no icon option — the shell picks the glyph from a hardcoded section-id map\n(ui-settings-general `navIcon`) and falls back to the gear for ids it does not\nknow, ours included. So the client half repaints its own row: it finds the nav\ncell by label and swaps the gear's path geometry for the official\n`IconBranchOutline16` path, mutating the attribute rather than replacing the\nnode so React re-renders over it without restoring the gear. It fails safe — if\nthe shell's markup moves, nothing matches and the row keeps the gear.\n","readmeFilename":"README.md"}