{"_id":"@argszero/cordis-plugin-overflow-classifier-guard","_rev":"2-f9040375497dd7e2a30d95bf7544d32e","name":"@argszero/cordis-plugin-overflow-classifier-guard","dist-tags":{"latest":"0.1.1"},"versions":{"0.1.0":{"name":"@argszero/cordis-plugin-overflow-classifier-guard","version":"0.1.0","keywords":["cordis","deepseek-harness","dsh","plugin","guard","context-overflow","compaction","max_prompt_tokens","openai-compatible"],"license":"MIT","_id":"@argszero/cordis-plugin-overflow-classifier-guard@0.1.0","maintainers":[{"name":"argszero","email":"argszero.reg@gmail.com"}],"dsh":{"bundle":{"patch":"./cordis.patch.yml"}},"dist":{"shasum":"bb95dded23986f093b95eb45ff43c637955dccbe","tarball":"https://registry.npmjs.org/@argszero/cordis-plugin-overflow-classifier-guard/-/cordis-plugin-overflow-classifier-guard-0.1.0.tgz","fileCount":6,"integrity":"sha512-iNsywIZNmgEM0lDGdEQ1tf50Oc+zjBmgNQubM5jIyMIk72MpQ1yzfyNrpR/DPQXQzbYfrUBaUc4gM7NwrVNTKw==","signatures":[{"sig":"MEQCIADbSjRZcZkvv0GUfYfXZLKKhfi58gcuLzYvfBtn6jz1AiAjCv80mnyO38a1i0m/D+W+oMLptN29a9bQD1gE7P+nPw==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":39347},"main":"lib/index.js","type":"module","types":"lib/types/index.d.ts","exports":{".":{"types":"./lib/types/index.d.ts","default":"./lib/index.js"},"./src/*":"./src/*","./package.json":"./package.json"},"scripts":{"test":"tsc && node --test \"test/*.test.js\"","build":"tsc","prepublishOnly":"tsc"},"_npmUser":{"name":"argszero","email":"argszero.reg@gmail.com"},"_npmVersion":"11.17.0","description":"Context-overflow classifier guard for dsh: observes the llm/stream waterfall and reclassifies a provider request-size rejection that names an *input* budget (max_prompt_tokens / max_input_tokens) rather than the model's context window. Unrecognized, such ","directories":{},"_nodeVersion":"26.5.0","dependencies":{"@deepseek-ai/schemastery":"^3.18.1"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.5.0","@deepseek-ai/cordis":"^4.0.2","@deepseek-ai/dsh-llm":"0.1.5-rc.2"},"peerDependencies":{"@deepseek-ai/cordis":"^4.0.2","@deepseek-ai/dsh-llm":">=0.1.2-rc.1 <0.2.0 || >=0.1.5-alpha.1 <0.2.0"},"peerDependenciesMeta":{"@deepseek-ai/dsh-llm":{"optional":false}},"_npmOperationalInternal":{"tmp":"tmp/cordis-plugin-overflow-classifier-guard_0.1.0_1789148009211_0.9004061561196328","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"_id":"@argszero/cordis-plugin-overflow-classifier-guard@0.1.1","dsh":{"bundle":{"patch":"./cordis.patch.yml"}},"dist":{"shasum":"e17cea49a08e355c59546f8815a2f62aea4c8fc5","tarball":"https://registry.npmjs.org/@argszero/cordis-plugin-overflow-classifier-guard/-/cordis-plugin-overflow-classifier-guard-0.1.1.tgz","fileCount":6,"integrity":"sha512-0Q/tmefCTm3iI+hEtQBupp8OhbJERXqQ2y8OyKDNXyMxvjk/tsd079MoKB4p8LRnAvficxAolH8fH4lK3Nb71Q==","signatures":[{"sig":"MEUCIE3/gCXbpslv35rLLa86WeiWDOdOrTNVlELTzRdYIK00AiEAzKox3LLN+Lx9xWh0G2vZpBGY6pprD6hmskUEq9y29+8=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"},{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIHUi/JWT0C57cNfiksO+xr9khdT6f5ce55VpZfdTgKYkAiAit4WckR7adHPEtdIOFKE6IiUtzIqzJIYEz3Jnu6Jbhg=="}],"unpackedSize":40418},"main":"lib/index.js","name":"@argszero/cordis-plugin-overflow-classifier-guard","type":"module","types":"lib/types/index.d.ts","exports":{".":{"types":"./lib/types/index.d.ts","default":"./lib/index.js"},"./src/*":"./src/*","./package.json":"./package.json"},"gitHead":"dc909bb82fa6a9a856a35751c51fadaeefa1a62b","license":"MIT","scripts":{"test":"tsc && node --test \"test/*.test.js\"","build":"tsc","prepublishOnly":"tsc"},"version":"0.1.1","_npmUser":{"name":"argszero","email":"argszero.reg@gmail.com"},"keywords":["cordis","deepseek-harness","dsh","plugin","guard","context-overflow","compaction","max_prompt_tokens","openai-compatible"],"_npmVersion":"11.17.0","description":"Context-overflow classifier guard for dsh: observes the llm/stream waterfall and reclassifies a provider request-size rejection that names an *input* budget (max_prompt_tokens / max_input_tokens) rather than the model's context window. Unrecognized, such ","directories":{},"maintainers":[{"name":"argszero","email":"argszero.reg@gmail.com"}],"_nodeVersion":"26.5.0","dependencies":{"@deepseek-ai/schemastery":"^3.18.1"},"_hasShrinkwrap":false,"devDependencies":{"semver":"^7.8.5","typescript":"^5.5.0","@deepseek-ai/cordis":"^4.0.2","@deepseek-ai/dsh-llm":"0.1.5-rc.2"},"peerDependencies":{"@deepseek-ai/cordis":"^4.0.2","@deepseek-ai/dsh-llm":">=0.1.3-alpha.2 <0.1.4 || >=0.1.5-alpha.1 <0.2.0 || >=0.1.6-alpha.1 <0.2.0"},"peerDependenciesMeta":{"@deepseek-ai/dsh-llm":{"optional":false}},"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/cordis-plugin-overflow-classifier-guard_0.1.1_1789693719183_0.15256544913547243"}}},"time":{"created":"2026-09-11T17:33:29.052Z","modified":"2026-09-18T01:08:39.412Z","0.1.0":"2026-09-11T17:33:29.357Z","0.1.1":"2026-09-18T01:08:39.259Z"},"license":"MIT","keywords":["cordis","deepseek-harness","dsh","plugin","guard","context-overflow","compaction","max_prompt_tokens","openai-compatible"],"description":"Context-overflow classifier guard for dsh: observes the llm/stream waterfall and reclassifies a provider request-size rejection that names an *input* budget (max_prompt_tokens / max_input_tokens) rather than the model's context window. Unrecognized, such ","maintainers":[{"name":"argszero","email":"argszero.reg@gmail.com"}],"readme":"# @argszero/cordis-plugin-overflow-classifier-guard\n\n**A request-size rejection the harness cannot read, reported in the vocabulary it does read — so the compaction that exists for it can run.**\n\nA provider can reject a request as too large in two vocabularies. The harness\nunderstands one of them.\n\n## The defect (discussion #6361)\n\n**Vocabulary 1 — the model's context window.** Recognized: the classifier\nrequires the literal word *context*.\n\n```js\n// packages/llm/llm/src/error.ts:80-90 — every pattern names the model context\nSTRUCTURED_CONTEXT_OVERFLOW  // \"context length exceeded\" / \"context window …\"\nMAXIMUM_CONTEXT_LENGTH       // \"maximum context length\"\nTOO_LARGE_FOR_CONTEXT        // \"request too large for this model's context\"\nEXCEEDS_MODEL_CONTEXT        // \"input exceeds the model's context\"\n```\n\nA provider using this wording maps to `CONTEXT_WINDOW_EXCEEDED`\n(`llm-deepseek/src/adapter.ts:337-351`).\n\n**Vocabulary 2 — an *input*-token budget, named as a parameter.** Not\nrecognized: the word \"context\" never appears.\n\n```json\n{\"code\":20015,\"message\":\"number of input tokens (150019) has exceeded max_prompt_tokens (98304) limit.\"}\n```\n\nHTTP 400 falls through to `INVALID_REQUEST`. On the pi-ai route the catch-all\n`PI_AI_ERROR` (`llm-pi-ai/src/stream.ts:41-67`) does the same — and its\nusage-based `isContextOverflow` cannot help either, because it compares against\nthe *catalog's* context window, which a relay station does not publish.\n\n**The verdict is not cosmetic — it severs the recovery path:**\n\n1. **Compaction never runs.** The overflow hook opens by refusing anything that\n   is not the overflow code — `if (failure.code !== CONTEXT_WINDOW_EXCEEDED_CODE\n   || signal.aborted) return next()` (`compaction-basic/src/index.ts:184`).\n2. **Retry never runs.** `INVALID_REQUEST` is not in `DEFAULT_RETRYABLE_CODES`\n   (`llm/src/retry-policy.ts:17-23`) — and resending an identical over-large\n   body could not succeed anyway.\n3. **The turn hard-aborts** (`agent-loop/src/agent.ts:447-468`) instead of\n   compacting.\n\nAnd the overflow branch is *exactly* the one built for this: it **bypasses the\nnormal pressure threshold and retained-tail policy** so it can force one useful\nbalanced reduction. A percentage-of-window trigger cannot fire when the\nprovider's real ceiling is a separate, smaller budget — so restoring the\nclassification is what makes compaction reachable in the first place.\n\n## Install and mount\n\n```sh\nnpm install @argszero/cordis-plugin-overflow-classifier-guard\n```\n\nThe bundle patch mounts it; no configuration is required:\n\n```yaml\n- insert:\n    - id: overflow-classifier-guard\n      name: '@argszero/cordis-plugin-overflow-classifier-guard'\n```\n\n## What it does\n\nObserves the `llm/stream` waterfall — the around-dispatch seam, which the\nin-tree `guard/` family cannot reach (`guard/timeout-policy` is a per-tool\ndeadline; `guard/repeat-tool-reminder` only arms once a tool call exists) — and\nrewrites the terminal `error` finish when its message names an input budget and\nthe harness classified it as a terminal request-shape failure.\n\n- **Streaming is preserved.** The verdict needs only the terminal chunk, and the\n  terminal chunk is the last a provider sends — so nothing is buffered. Every\n  chunk is forwarded the moment it arrives; live token streaming and the durable\n  log are untouched. A test fails if the guard ever starts buffering.\n- **The provider's text is preserved verbatim**, with the reason appended. Only\n  the classification changes — which matters, because a gateway's wording is the\n  whole problem here.\n- **`CONTEXT_WINDOW_EXCEEDED` is the harness' own constant**, imported rather\n  than restated, so a synthesized verdict is indistinguishable from an adapter's\n  own and can never drift from the key the recovery path tests.\n- **Bounded recovery.** The loop routes the verdict to `agent/request-error`,\n  where `dsh-compaction-basic` — mounted in `bundle/base` and `bundle/web-app` —\n  compacts and returns `{ kind: 'retry' }`, bounded by its `maxOverflowRetries`\n  (default 1, `config.ts:93`). No loop is possible.\n\n## Deliberate narrowness\n\nReclassifying too eagerly would compact a session for a failure compaction\ncannot fix, so **four** conditions must all hold:\n\n| Condition | Why |\n| --- | --- |\n| The finish is an `error`, never `aborted` | The caller cancelled; there is nothing to recover |\n| No visible content has streamed | Re-running after real output would duplicate it |\n| The code is a **terminal request-shape** code | Transient codes belong to the retry policy; `AUTH`/`QUOTA`/`EMPTY_RESPONSE` have different remedies |\n| The message names an **input** budget | Not the model context (the harness already matched), not quota, not output `max_tokens` |\n\nThe last row is where the judgement lives. `max_prompt_tokens` /\n`max_input_tokens` are caps stated **separately from** the model's context\nwindow, so they are the signal; an output-side `max_tokens` complaint is\nexcluded, because shortening the request does not help it.\n\nThe wording check is a **bounded window** — either side of the budget name,\nwithin 160 characters — because providers state the two halves in either order:\n\n```\nnumber of input tokens (150019) has exceeded max_prompt_tokens (98304)   ← verb first\nmax_input_tokens exceeded (98304)                                        ← name first\n```\n\nA name-first-only rule would miss the one live evidence we have; an unbounded\nsearch would fire on a message that names a budget and then reports an unrelated\nsurplus further down.\n\n## Configuration\n\n| key | default | meaning |\n| --- | --- | --- |\n| `mode` | `'error'` | `'error'` = report as `CONTEXT_WINDOW_EXCEEDED`; `'warn'` = detect and log only; `'off'` = pure pass-through |\n| `eligibleCodes` | `[]` | extra terminal request-shape codes your gateway emits (built-ins: `INVALID_REQUEST`, `PI_AI_ERROR`, any `HTTP_4xx`) |\n| `extraPatterns` | `[]` | your gateway's wording, as regex sources; each is a standalone assertion, so the English vocabulary is not also required |\n| `requireExceedWording` | `true` | demand an \"exceeded\"-style word near the budget name; set `false` for a gateway that names the budget without asserting the exceedance |\n\nSet a code that is refused by design (`AUTH`, `QUOTA`, `ABORTED`,\n`EMPTY_RESPONSE`) and the refusal still wins — those are terminal for a\ndifferent remedy.\n\n```yaml\n- set:\n    - id: overflow-classifier-guard\n      config:\n        mode: warn\n        extraPatterns: ['prompt budget overrun']\n```\n\nStart with `mode: warn` if you want to see the detections without changing\nbehaviour; the log line names the code it would have replaced.\n\n## Relationship to other guards\n\n- `cordis-plugin-empty-response-guard` corrects a **successful turn that\n  produced nothing**. This one corrects a **failed turn that was recoverable**.\n  Both sit on the same `llm/stream` seam and share its streaming discipline.\n- `cordis-plugin-credential-rotate` re-tries a *transient* failure that the\n  retry policy refuses. This one re-classifies a *terminal* failure that the\n  recovery path refuses. Both hold the same law: **never leak a discarded\n  attempt's terminal frame**.\n- The in-tree `guard/` family cannot see either shape.\n\n## Scope, honestly\n\nThis is a **seam** fix, not the core fix. The real repair is to widen the\nharness' own classifier for named input budgets — the adapter already has the\ndetail string and the place to do it. Until upstream decides, this keeps the\nfailure classified where the machinery already looks.\n\nIt is also **not** a fix for token estimation. `CHARS_PER_TOKEN = 4`\n(`token-meter/src/estimate.ts:13`) misprices CJK, but that is a separate\ncalibration question; what this plugin restores is the path that does not depend\non the estimate at all.\n\n## Compatibility\n\nPeer-compatible with `@deepseek-ai/dsh-llm` on the **0.1.3**, **0.1.5** and\n**0.1.6** lines — each one has had the full suite run against it.\n\n**0.1.1 moved the floor up from `0.1.2-rc.1` to `0.1.3-alpha.2`,** for a reason\nthe range could not see: this plugin imports `chunkHasVisibleText` from the\nharness, and that function lives in `assistant-stream`, which first exists in\n`0.1.3-alpha.2`. On `0.1.2-rc.1` the package installs happily and then dies at\nmount:\n\n```\nSyntaxError: The requested module '@deepseek-ai/dsh-llm'\ndoes not provide an export named 'chunkHasVisibleText'\n```\n\n**0.1.1 also admits the 0.1.6 line**, which the previous range silently excluded:\na comparator only admits a prerelease when some comparator shares that\nprerelease's `major.minor.patch`, so `>=0.1.5-alpha.1 <0.2.0` is not \"0.1.5 and\nlater\" — it drops every `0.1.6-*` release, one `npm install` after\n`0.1.6-alpha.2` became the `alpha` dist-tag.\n\n`test/peer-range.test.js` computes admission over the published version list\nwith `semver` and asserts the admitted set exactly, so neither an untested extra\nline nor a missing one can pass a regex-shaped test again.\n\n## Tests\n\n```sh\nnpm test\n```\n\n50 tests, including the reported SiliconFlow body verbatim, the verb-first and\nname-first word orders, an incrementality test that fails if the guard ever\nbuffers, and a packaging test that fails if the codes are ever hardcoded instead\nof imported from the harness.\n\n## License\n\nMIT\n","readmeFilename":"README.md"}