{"_id":"@aprilwizard/dsh-multi-cot","name":"@aprilwizard/dsh-multi-cot","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@aprilwizard/dsh-multi-cot","description":"Multi-sampled test-time compute plugin for DeepSeek Harness: parallel planning, internal voting, and a plan/execute/review workflow","version":"0.1.0","publishConfig":{"access":"public"},"repository":{"type":"git","url":"git+https://github.com/aprilwizard/dsh-multi-cot.git","directory":"."},"keywords":["dsh-plugin","deepseek-harness","multi-cot","test-time-compute"],"type":"module","main":"lib/types/index.js","types":"lib/types/index.d.ts","exports":{".":{"types":"./lib/types/index.d.ts","default":"./lib/types/index.js"},"./invariant":{"types":"./lib/types/invariant.d.ts","default":"./lib/types/invariant.js"},"./package.json":"./package.json"},"license":"MIT","engines":{"node":">=22"},"dependencies":{"@deepseek-ai/schemastery":">=0.0.1-rc.1"},"peerDependencies":{"@deepseek-ai/cordis":"^4.0.1","@deepseek-ai/dsh-agent":">=0.0.1-rc.1","@deepseek-ai/dsh-invariants":">=0.0.1-rc.1","@deepseek-ai/dsh-llm":">=0.0.1-rc.1","@deepseek-ai/dsh-session":">=0.0.1-rc.1","@deepseek-ai/dsh-system-prompt":">=0.0.1-rc.1"},"devDependencies":{"@deepseek-ai/cordis":"4.0.1","@deepseek-ai/dsh-agent":"0.0.1-rc.1","@deepseek-ai/dsh-agent-loop":"0.0.1-rc.1","@deepseek-ai/dsh-agent-loop-testkit":"0.0.1-rc.1","@deepseek-ai/dsh-attachment":"0.0.1-rc.1","@deepseek-ai/dsh-brand":"0.0.1-rc.1","@deepseek-ai/dsh-code-runtime":"0.0.1-rc.1","@deepseek-ai/dsh-invariants":"0.0.1-rc.1","@deepseek-ai/dsh-llm":"0.0.1-rc.1","@deepseek-ai/dsh-llm-pi-ai":"0.0.1-rc.1","@deepseek-ai/dsh-scope":"0.0.1-rc.1","@deepseek-ai/dsh-session":"0.0.1-rc.1","@deepseek-ai/dsh-session-persistence":"0.0.1-rc.1","@deepseek-ai/dsh-system-prompt":"0.0.1-rc.1","@deepseek-ai/dsh-timeout":"0.0.1-rc.1","@deepseek-ai/dsh-tools":"0.0.1-rc.1","@deepseek-ai/dsh-user-approval":"0.0.1-rc.1","@deepseek-ai/schemastery":"^3.18.1","@types/node":"^22.0.0","tsx":"^4.22.4","typescript":"^6.0.3","vitest":"^4.1.8"},"scripts":{"build":"tsc","typecheck":"tsc --noEmit -p tsconfig.test.json","test":"vitest run","prepublishOnly":"npm run build && npm run typecheck && npm test","cache-check":"node scripts/multi-cot-cache-check.mjs","live-run":"tsx scripts/multi-cot-live-run.ts"},"gitHead":"1a7ca2407db07cd8297829a0f539a0394e04a6fe","_id":"@aprilwizard/dsh-multi-cot@0.1.0","bugs":{"url":"https://github.com/aprilwizard/dsh-multi-cot/issues"},"homepage":"https://github.com/aprilwizard/dsh-multi-cot#readme","_nodeVersion":"22.23.1","_npmVersion":"11.18.0","dist":{"integrity":"sha512-5QUEeSvZuMcviNPxCpd/j2UB20uWogrVmHpu4e1VF4irvpcv5g0ZoyamUIRm3NOUUgPLDr75tRmlngABoeTZlQ==","shasum":"d8150a34532c66ffe3d99f47945ac6bea6b089e3","tarball":"https://registry.npmjs.org/@aprilwizard/dsh-multi-cot/-/dsh-multi-cot-0.1.0.tgz","fileCount":24,"unpackedSize":95732,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCICanXHd7ERoAeVzWGh9cO5gI3u88WVfzHtcp/WgW+StEAiBvsIvin6SN+TNlxc+DSRgtKq/aKhFhxNfAgzBP1KqvLg=="}]},"_npmUser":{"name":"aprilwizard","email":"1638509496@qq.com"},"directories":{},"maintainers":[{"name":"aprilwizard","email":"1638509496@qq.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/dsh-multi-cot_0.1.0_1786641020831_0.5164361996725852"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-13T17:10:20.630Z","0.1.0":"2026-08-13T17:10:21.015Z","modified":"2026-08-13T17:10:21.221Z"},"maintainers":[{"name":"aprilwizard","email":"1638509496@qq.com"}],"description":"Multi-sampled test-time compute plugin for DeepSeek Harness: parallel planning, internal voting, and a plan/execute/review workflow","homepage":"https://github.com/aprilwizard/dsh-multi-cot#readme","keywords":["dsh-plugin","deepseek-harness","multi-cot","test-time-compute"],"repository":{"type":"git","url":"git+https://github.com/aprilwizard/dsh-multi-cot.git","directory":"."},"bugs":{"url":"https://github.com/aprilwizard/dsh-multi-cot/issues"},"license":"MIT","readme":"# @aprilwizard/dsh-multi-cot\n\n[English](README.md) | 中文\n\nMulti-CoT 通过重复生成并采样思维链，为 dsh 提供近似“并行测试时计算”的能力：同一输入\n并行跑 `samples` 条推理链，由内部投票选出最佳结果，再用它驱动“先计划后执行”或完整的\n“计划 → 执行 → 复检”三阶段工作流。插件只包含纯函数选择核心和稳定的协议提示词，接在\n`agent/pre-step`、`llm/stream`、`agent/turn-stopping` 三个扩展点上，不需要修改 dsh\n源码。\n\n近期通过修改 codex 源码，用 `deepseek-v4-flash` 在隔离环境下用 Terminal-Bench 2.1\n做了开/关对比（baseline = 普通模式，workflow = 多采样三阶段）：\n\n| 题目 | baseline | workflow |\n|---|---|---|\n| write-compressor (hard) | ✅ 685.6s / 2475B / 1.83M in | ✅ 643.5s / 2231B / 1.17M in |\n| cancel-async-tasks (hard) | ❌ 219.8s（5/6 测试过） | ✅ 647.3s |\n| polyglot-rust-c (hard) | ✅ 522.0s | ✅ 1,097.9s |\n| regex-chess (hard) | ✅ 1,075.0s / 6.93M in | ✅ 产物合格，进程 1800s 超时被杀 |\n| sqlite-db-truncate (medium) | ✅ 129.2s | ✅ 598.5s |\n\n单题 GSM8K 对比：正确率不变（26/26），耗时约 29×（4.4s → 126.8s），输入 token\n11.1k → ~219k——多采样主要用成本换稳定性。\n\n> 注：以上为近期 codex 实验的参考数据；本 dsh 插件尚未在同条件下跑过这套对比。\n\n## 设计思路\n\n### 采样与投票\n\n普通回合只问模型一次，返回什么就是什么。这个插件用相同输入问 `samples` 次，再让\n`samples` 个内部投票者给候选打分。输入相同让成本可控：provider 会缓存共享前缀，并行\n采样不会真的付出 `samples` 倍的输入成本。结果全部相同就跳过投票。\n\n### 三阶段工作流\n\n`workflow` 模式把任务分成三个阶段：信息收集、具体实现、编写报告。每阶段三步：计划\n（离线采样 + 投票）、执行（普通工具循环）、复检（再一轮离线采样 + 投票）。复检失败就\n重新计划，最多 `reviewMaxRetries` 次，再失败则强制推进并标记“存疑”。阶段状态由插件\n管理，不靠模型打印标记。\n\n### 离线工作放在哪里\n\n计划必须在模型请求构建前选好。`llm/stream` 要求同步返回流，计划步也不能声明空批\n（空批会让 loop 关回合）。所以：\n\n- 首个计划在 `agent/pre-step` 选择；\n- 后续计划在 `agent/turn-stopping` 选择；\n- 选中的计划存进进程内表，`llm/stream` 只做同步短接，返回存好的计划流。\n\n因此后续阶段的计划之前会多一行 user 角色“为阶段 N 制定计划”。\n\n### 缓存\n\n协议文本是系统提示词里稳定的 section（order 50）。采样请求字节完全相同；投票共享\n一个前缀（所有候选的压缩链），只追加自己的完整链和打分指令。opencode-go 与官方\nDeepSeek 都实测过：\n\n| 实验 | opencode-go | 官方 DeepSeek |\n|---|---|---|\n| 相同请求重复 | 0 命中 | 0 命中 |\n| 共享前缀 voter 1+ | ~90% | ~90% |\n| 真机工作流缓存命中 | 83.1%（40 次） | 62.7%（32 次） |\n\n相同请求重复不命中（推理模型把 reasoning tokens 算进输出端缓存单元），所以设计\n依赖的是共享前缀复用。\n\n### 推理链\n\n投票用的是模型真实推理的压缩版。chat completions wire 上 `reasoning_content`\n经 pi-ai 变成 thinking 事件，dsh adapter 再映射成 `reasoning` 块。有个坑：\nopencode-go 内置 provider 默认走 DeepSeek thinking 方言，未指定 effort 时会发\n`thinking: {type: \"disabled\"}`；设 `compat.thinkingFormat: openai` 即可保持推理开启。\n\n### 日志与降级\n\n每次离线请求、usage、决策和阶段切换都作为会话事件落盘，选中的计划由 loop 写成普通\nassistant 消息，会话日志是唯一事实来源。失败一律降级而不是卡死：选择为空 → 普通模型\n调用；结果相同 → 跳过投票；票无效 → 多数优先/第一条；复检解析失败 → PASS；流异常 →\n普通循环。\n\n## 配置\n\n| 键 | 默认 | 含义 |\n|---|---|---|\n| `mode` | `off` | `off` 关闭；`first-plan` 先离线计划再执行；`workflow` 对三个阶段分别运行计划 → 执行 → 复检。 |\n| `samples` | `3` | 每次离线选择的并行采样数（也是内部投票者数）；启用模式下为 2–16。 |\n| `votePoints` | `100` | 每张内部票分配给各采样的整数总分。 |\n| `compressedChainMaxTokens` | `300` | 压缩思考链的绝对 token 上限。 |\n| `compressedChainRatio` | `0.2` | 压缩链相对完整链长度的比例上限。 |\n| `nearTopDistance` | `0.05` | 低于最高分该比例内的候选仍参与加权随机。 |\n| `reviewMaxRetries` | `2` | 阶段被强制推进前允许的复检失败次数。 |\n\n所有值在插件加载时校验；非法范围响亮失败，不静默回退默认值。\n\n## 事件\n\n本包声明四个 log-only `SessionEventMap` 成员：\n\n| 事件 | 用途 |\n|---|---|\n| `multi-cot/phase` | 按轮次保存的阶段状态（最后一条生效）；resume/fork 通过 fold 恢复。 |\n| `multi-cot/request` | 一次离线采样/投票/复检请求的完整 system 与 messages，保证可重建。 |\n| `multi-cot/usage` | 一次离线请求的 provider usage。 |\n| `multi-cot/decision` | 一次离线选择的选中下标、归一化分数与复检结论。 |\n\n## Model Experience\n\n### 协议 section\n\n#### 模型看到什么\n\n启用模式时，每个请求在 prompt order 50 处携带稳定的 `multi-cot:protocol` section。\n\n##### 工作流模式\n\n```markdown\nThe agent completes a task in three phases: information gathering, implementation, and report writing. Each phase runs a plan step, an execution step, and a review step. During execution you may gather missing information directly, but you must not change the plan; when the plan must change, return through review and re-plan. A phase review may fail at most twice before the phase advances anyway; when that happens, mark anything uncertain as doubtful in the final report.\n```\n\n##### 先计划模式\n\n```markdown\nBefore executing a task, produce one concrete plan, then follow it during execution.\n```\n\n#### Token 影响\n\n插件组合且模式启用时为固定按请求成本；`off` 不贡献任何文本。\n\n#### KV Cache 影响\n\nsection 文本与顺序不变时前缀稳定；启用、停用或切换模式会使 order 50 起的复用失效。\n\n### 计划、执行与复检步骤\n\n#### 模型看到什么\n\n计划步呈现为离线选出的 assistant 计划；后续阶段的计划步与执行步携带稳定的 user 角色请求/指令行。\n\n##### 后续阶段计划请求与执行指令\n\n```markdown\nProduce a plan for phase 2 (implementation). Do not execute it yet; it will be selected and reviewed before execution.\n\nExecute the plan above for phase 2 (implementation). Gather missing information directly during execution, but do not change the plan unless a review requires it.\n```\n\n#### Token 影响\n\n每个后续阶段计划前多一条 user 行、每阶段一条执行指令；离线采样与投票的 token 只计入\n`multi-cot/usage` 事件，不计入循环的 assistant usage。\n\n#### KV Cache 影响\n\n同一轮内的离线采样共享字节相同的前缀；执行指令追加在可复用的计划前缀之后。\n\n## 已知限制与后续工作\n\n- **离线请求不参与上下文压缩** —— 长会话中采样输入可能膨胀；离线输入压缩边界待定。\n- **阶段状态不做中断恢复** —— `multi-cot/phase` 可折叠出持久状态，但进行中的计划/复检\n  步骤本身无检查点；中断的轮次从下一个新轮次重新开始。\n- **投票注入无完整链上限** —— 压缩链有界，但投票者自己的完整链按原样注入；该上限待定。\n- **后续阶段计划请求多一条 user 行** —— 计划请求消息会留在转录中，位于离线选出的计划\n  之前；无历史痕迹的内部唤醒方案待定。\n\n## 安装\n\n先安装包，再在 `cordis.yml` 中组合：\n\n```sh\nnpm install @aprilwizard/dsh-multi-cot\n```\n\n```yaml\n- id: multi-cot\n  name: '@aprilwizard/dsh-multi-cot'\n  config:\n    mode: workflow   # off | first-plan | workflow\n    samples: 3\n```\n\npeer 依赖为 `@deepseek-ai/cordis`、`@deepseek-ai/dsh-agent`、\n`@deepseek-ai/dsh-invariants`、`@deepseek-ai/dsh-llm`、\n`@deepseek-ai/dsh-session` 与 `@deepseek-ai/dsh-system-prompt`。\n\n安装注意事项：已发布的 dsh rc.1 包 peer 依赖未发布的纯类型包\n`@deepseek-ai/dsh-type-meta`，npm 自动安装 peer 会失败。在上游发布它之前，请用\npnpm 安装（`pnpm-workspace.yaml` 里设 `autoInstallPeers: false`），或等待新版 dsh\n包发布。\n\n## 开发\n\n插件已在 npm 发布的 `@deepseek-ai/dsh-*`（`0.0.1-rc.1` +\n`@deepseek-ai/cordis` 4.0.1）上验证通过，不需要 dsh checkout 或 link：\n\n```sh\npnpm install\npnpm build         # 生成 lib/types 供发布\npnpm test          # 单测 + 集成 + baseline/workflow 对比\npnpm typecheck\npnpm cache-check   # opencode-go 端点的真实缓存实验\npnpm live-run      # opencode-go 端点的真实端到端工作流\n```\n\n一个安装注意事项：已发布的 rc.1 dsh 包 peer 依赖纯类型包\n`@deepseek-ai/dsh-type-meta`，而它从未发布。`pnpm-workspace.yaml` 因此设置了\n`autoInstallPeers: false`，所有运行时 peer 都在 devDependencies 中显式列出。\n等更新的 dsh 包发布后，升级这些 devDependency 版本即可。\n\n`cache-check` 与 `live-run` 读取 `OPENCODE_GO_API_KEY`（回退\n`OPENCODE_API_KEY`），两者都不会打印凭证。\n\n## Provider 说明（推理捕获）\n\nDeepSeek 系模型在 chat completions wire 上返回 `reasoning_content`。pi-ai 内置的\n`opencode-go` 与 `deepseek` catalog provider 默认走 DeepSeek thinking 方言，未指定\nreasoning effort 时会发送 `thinking: {type: \"disabled\"}`，导致 reasoning 内容到不了\nharness。将 route 配置为 `compat.thinkingFormat: openai` 可保持 provider 默认思考\n开启：\n\n```yaml\n- id: llm\n  name: '@deepseek-ai/dsh-llm-pi-ai'\n  config:\n    providers:\n      opencode-go:\n        apiKeyEnv: OPENCODE_GO_API_KEY\n        api: openai-completions\n        baseURL: https://opencode.ai/zen/go/v1\n        compat:\n          thinkingFormat: openai\n        models:\n          - id: deepseek-v4-flash\n            contextWindow: 1000000\n          - id: deepseek-v4-pro\n            contextWindow: 1000000\n      deepseek:\n        apiKeyEnv: DEEPSEEK_API_KEY\n        api: openai-completions\n        baseURL: https://api.deepseek.com\n        compat:\n          thinkingFormat: openai\n        models:\n          - id: deepseek-v4-flash\n            contextWindow: 1000000\n          - id: deepseek-v4-pro\n            contextWindow: 1000000\n```\n\n两个 route 都用该配置验证过（官方 DeepSeek 还提供 `deepseek-chat`，它不返回\nreasoning）。推理捕获不需要任何 dsh 源码改动：`thinkingFormat: openai` 下 pi-ai\n会发出 thinking 事件，dsh 自带的 adapter 会映射成 `reasoning` 块。自带 adapter 的\nusage 映射不转发 reasoning token 计数，因此 `multi-cot/usage` 事件会省略\n`reasoningTokens`；链捕获与压缩不受影响。\n","readmeFilename":"README.zh.md","_rev":"1-25697e17f44c57fc702bc5b76bbc28f6"}