{"_id":"@aispin/plugin-verifier","_rev":"3-005f458a27f29b7a63e067bb4576a113","name":"@aispin/plugin-verifier","dist-tags":{"latest":"0.3.3"},"versions":{"0.3.1":{"name":"@aispin/plugin-verifier","version":"0.3.1","keywords":["dsh-plugin","deepseek-harness","llm-verifier","llm-as-a-verifier","best-of-n","best-of-5","test-time-scaling","test-time-compute","verifier","selection","agent-evaluation","sampling","deepseek","deepseek-v4-flash"],"license":"MIT","_id":"@aispin/plugin-verifier@0.3.1","maintainers":[{"name":"aispin_dev","email":"dev@aispin.dev"}],"homepage":"https://github.com/aispin-dev/llm-as-a-Verifier-dsh#readme","bugs":{"url":"https://github.com/aispin-dev/llm-as-a-Verifier-dsh/issues"},"dsh":{"bundle":{"patch":"./cordis.patch.yml"},"client":{"inject":["@deepseek-ai/dsh-client-runtime","@deepseek-ai/dsh-client-ui-settings","@deepseek-ai/dsh-client-web-react"],"platform":"web"}},"dist":{"shasum":"38030dc702fab743799e276d94fae18e7dabbb3a","tarball":"https://registry.npmjs.org/@aispin/plugin-verifier/-/plugin-verifier-0.3.1.tgz","fileCount":16,"integrity":"sha512-gPzsRYK4t7AD1XDH9J3y0HDTOYpkwLWOArc7MzstfmQk/n+rJrNeM9op4hRna4gweo7x0JiyGVEMhXxey38xLQ==","signatures":[{"sig":"MEYCIQD23r4/IpCSssgn8MxijyGKlbeS6nFdvPjj97Ee/F+7gQIhAOuHYWql5KeeydAQhkdv2usEu6OYjZCKTCdSJryo/I66","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":145989},"main":"lib/index.js","type":"module","types":"lib/types/index.d.ts","aispin":{"plugin":{"trust":"first-party","provides":["verifier"],"seamDeps":["tools"]}},"exports":{".":{"types":"./lib/types/index.d.ts","default":"./lib/index.js"},"./src/*":"./src/*","./client":{"types":"./lib/types/client.d.ts","default":"./lib/client.js"},"./package.json":"./package.json"},"gitHead":"62939a442fd33e87077caadd829823cc418495d0","scripts":{"test":"vitest run","build":"tsc -p tsconfig.json && tsdown && cp src/client.js lib/client.js","typecheck":"tsc -p tsconfig.json --noEmit && tsc -p tsconfig.test.json --noEmit"},"_npmUser":{"name":"aispin_dev","email":"dev@aispin.dev"},"repository":{"url":"git+https://github.com/aispin-dev/llm-as-a-Verifier-dsh.git","type":"git"},"_npmVersion":"11.6.2","description":"LLM-as-a-Verifier for dsh with the Best-of-N conversation mode built in: rank N candidates with a fine-grained verifier (expected grade over the logprob distribution), on demand via the verify tool or automatically on every turn of a Best-of-N session. In","directories":{},"_nodeVersion":"24.13.0","dependencies":{},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"tsdown":"^0.22.14","vitest":"^4.1.8","typescript":"^6.0.3","@types/node":"^22.20.0","@deepseek-ai/cordis":"4.0.1","@deepseek-ai/dsh-llm":"^0.1.0-rc.6","@deepseek-ai/dsh-tools":"^0.1.0-rc.6","@deepseek-ai/schemastery":"^3.18.1","@deepseek-ai/dsh-settings":"^0.1.0-rc.7","@deepseek-ai/dsh-credentials":"^0.1.0-rc.7","@deepseek-ai/dsh-client-runtime":"^0.1.0-rc.7","@deepseek-ai/dsh-client-web-react":"^0.1.0-rc.7","@deepseek-ai/dsh-client-ui-settings":"^0.1.0-rc.7"},"peerDependencies":{"react":"^18.2.0","@deepseek-ai/cordis":"4.0.1","@deepseek-ai/dsh-llm":"^0.1.0-rc.6","@deepseek-ai/dsh-tools":"^0.1.0-rc.6","@deepseek-ai/schemastery":"^3.18.1","@deepseek-ai/dsh-settings":"^0.1.0-rc.7","@deepseek-ai/dsh-credentials":"^0.1.0-rc.7","@deepseek-ai/dsh-client-runtime":"^0.1.0-rc.7","@deepseek-ai/dsh-client-web-react":"^0.1.0-rc.7","@deepseek-ai/dsh-client-ui-settings":"^0.1.0-rc.7"},"_npmOperationalInternal":{"tmp":"tmp/plugin-verifier_0.3.1_1787313962374_0.28199127339388674","host":"s3://npm-registry-packages-npm-production"}},"0.3.2":{"name":"@aispin/plugin-verifier","version":"0.3.2","keywords":["dsh-plugin","deepseek-harness","llm-verifier","llm-as-a-verifier","best-of-n","best-of-5","test-time-scaling","test-time-compute","verifier","selection","agent-evaluation","sampling","deepseek","deepseek-v4-flash"],"license":"MIT","_id":"@aispin/plugin-verifier@0.3.2","maintainers":[{"name":"aispin_dev","email":"dev@aispin.dev"}],"homepage":"https://github.com/aispin-dev/llm-as-a-Verifier-dsh#readme","bugs":{"url":"https://github.com/aispin-dev/llm-as-a-Verifier-dsh/issues"},"dsh":{"bundle":{"patch":"./cordis.patch.yml"},"client":{"inject":["@deepseek-ai/dsh-client-runtime","@deepseek-ai/dsh-client-ui-settings","@deepseek-ai/dsh-client-web-react"],"platform":"web"}},"dist":{"shasum":"2d3c145b9a1959b8bdcccfcaf37f19672e16c9ef","tarball":"https://registry.npmjs.org/@aispin/plugin-verifier/-/plugin-verifier-0.3.2.tgz","fileCount":16,"integrity":"sha512-WYmbPOmTD/BAxbw0Z1zvesEXbVhpPwmfuOXOlIG4tLL5XpPgPP3eqZPYfh1/b5cxM1nt1aFfrZuj6UwYZPOpeA==","signatures":[{"sig":"MEQCIGVvcpjG6dRJ63m8jd8pJFckxXdDhppnavBbRC9K1tA0AiB63lURcLVV0Op3OTdbOj1Lxu/KNqJINIyJjDrKyEpOZw==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":147012},"main":"lib/index.js","type":"module","types":"lib/types/index.d.ts","aispin":{"plugin":{"trust":"first-party","provides":["verifier"],"seamDeps":["tools"]}},"exports":{".":{"types":"./lib/types/index.d.ts","default":"./lib/index.js"},"./src/*":"./src/*","./client":{"types":"./lib/types/client.d.ts","default":"./lib/client.js"},"./package.json":"./package.json"},"gitHead":"2c9bf1de713fe5ec5519d15f012cd6da16ae6f03","scripts":{"test":"vitest run","build":"tsc -p tsconfig.json && tsdown && cp src/client.js lib/client.js","typecheck":"tsc -p tsconfig.json --noEmit && tsc -p tsconfig.test.json --noEmit"},"_npmUser":{"name":"aispin_dev","email":"dev@aispin.dev"},"repository":{"url":"git+https://github.com/aispin-dev/llm-as-a-Verifier-dsh.git","type":"git"},"_npmVersion":"11.6.2","description":"LLM-as-a-Verifier for dsh with the Best-of-N conversation mode built in: rank N candidates with a fine-grained verifier (expected grade over the logprob distribution), on demand via the verify tool or automatically on every turn of a Best-of-N session. In","directories":{},"_nodeVersion":"24.13.0","dependencies":{},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"tsdown":"^0.22.14","vitest":"^4.1.8","typescript":"^6.0.3","@types/node":"^22.20.0","@deepseek-ai/cordis":"4.0.1","@deepseek-ai/dsh-llm":"^0.1.0-rc.6","@deepseek-ai/dsh-tools":"^0.1.0-rc.6","@deepseek-ai/schemastery":"^3.18.1","@deepseek-ai/dsh-settings":"^0.1.0-rc.7","@deepseek-ai/dsh-credentials":"^0.1.0-rc.7","@deepseek-ai/dsh-client-runtime":"^0.1.0-rc.7","@deepseek-ai/dsh-client-web-react":"^0.1.0-rc.7","@deepseek-ai/dsh-client-ui-settings":"^0.1.0-rc.7"},"peerDependencies":{"react":"^18.2.0","@deepseek-ai/cordis":"4.0.1","@deepseek-ai/dsh-llm":"^0.1.0-rc.6","@deepseek-ai/dsh-tools":"^0.1.0-rc.6","@deepseek-ai/schemastery":"^3.18.1","@deepseek-ai/dsh-settings":"^0.1.0-rc.7","@deepseek-ai/dsh-credentials":"^0.1.0-rc.7","@deepseek-ai/dsh-client-runtime":"^0.1.0-rc.7","@deepseek-ai/dsh-client-web-react":"^0.1.0-rc.7","@deepseek-ai/dsh-client-ui-settings":"^0.1.0-rc.7"},"_npmOperationalInternal":{"tmp":"tmp/plugin-verifier_0.3.2_1787626026334_0.17721863897414614","host":"s3://npm-registry-packages-npm-production"}},"0.3.3":{"name":"@aispin/plugin-verifier","description":"LLM-as-a-Verifier for dsh with the Best-of-N conversation mode built in: rank N candidates with a fine-grained verifier (expected grade over the logprob distribution), on demand via the verify tool or automatically on every turn of a Best-of-N session. In","version":"0.3.3","publishConfig":{"access":"public"},"type":"module","main":"lib/index.js","types":"lib/types/index.d.ts","exports":{".":{"types":"./lib/types/index.d.ts","default":"./lib/index.js"},"./client":{"types":"./lib/types/client.d.ts","default":"./lib/client.js"},"./src/*":"./src/*","./package.json":"./package.json"},"aispin":{"plugin":{"trust":"first-party","seamDeps":["tools"],"provides":["verifier"]}},"dsh":{"client":{"inject":["@deepseek-ai/dsh-client-runtime","@deepseek-ai/dsh-client-ui-settings","@deepseek-ai/dsh-client-web-react"],"platform":"web"},"bundle":{"patch":"./cordis.patch.yml"}},"repository":{"type":"git","url":"git+https://github.com/aispin-dev/llm-as-a-Verifier-dsh.git"},"homepage":"https://github.com/aispin-dev/llm-as-a-Verifier-dsh#readme","bugs":{"url":"https://github.com/aispin-dev/llm-as-a-Verifier-dsh/issues"},"keywords":["dsh-plugin","deepseek-harness","llm-verifier","llm-as-a-verifier","best-of-n","best-of-5","test-time-scaling","test-time-compute","verifier","selection","agent-evaluation","sampling","deepseek","deepseek-v4-flash"],"license":"MIT","scripts":{"build":"tsc -p tsconfig.json && tsdown && cp src/client.js lib/client.js","typecheck":"tsc -p tsconfig.json --noEmit && tsc -p tsconfig.test.json --noEmit","test":"vitest run"},"dependencies":{},"peerDependencies":{"@deepseek-ai/cordis":"4.0.1","@deepseek-ai/dsh-client-runtime":"^0.1.0-rc.7","@deepseek-ai/dsh-client-ui-settings":"^0.1.0-rc.7","@deepseek-ai/dsh-client-web-react":"^0.1.0-rc.7","@deepseek-ai/dsh-credentials":"^0.1.0-rc.7","@deepseek-ai/dsh-llm":"^0.1.0-rc.6","@deepseek-ai/dsh-settings":"^0.1.0-rc.7","@deepseek-ai/dsh-tools":"^0.1.0-rc.6","@deepseek-ai/schemastery":"^3.18.1","react":"^18.2.0"},"devDependencies":{"@deepseek-ai/cordis":"4.0.1","@deepseek-ai/dsh-client-runtime":"^0.1.0-rc.7","@deepseek-ai/dsh-client-ui-settings":"^0.1.0-rc.7","@deepseek-ai/dsh-client-web-react":"^0.1.0-rc.7","@deepseek-ai/dsh-credentials":"^0.1.0-rc.7","@deepseek-ai/dsh-llm":"^0.1.0-rc.6","@deepseek-ai/dsh-settings":"^0.1.0-rc.7","@deepseek-ai/dsh-tools":"^0.1.0-rc.6","@deepseek-ai/schemastery":"^3.18.1","@types/node":"^22.20.0","tsdown":"^0.22.14","typescript":"^6.0.3","vitest":"^4.1.8"},"gitHead":"f3710b870347bb9c478c5ab0e6e936bcb5274abe","_id":"@aispin/plugin-verifier@0.3.3","_nodeVersion":"24.13.0","_npmVersion":"11.6.2","dist":{"integrity":"sha512-imvQ9BgjxQVxcMlNOD0P2cSY8Jko+VcZCtgvB96cV4keL/9MVadunZfduFqiLkMHqJ6bqbkEwudV5s8d7PBTSQ==","shasum":"c8e606dd07b3ffe321272421289c968ffc0bdedb","tarball":"https://registry.npmjs.org/@aispin/plugin-verifier/-/plugin-verifier-0.3.3.tgz","fileCount":16,"unpackedSize":153490,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIGzzC5GHtKMhkmPLdcQ4QB8wo3ooBXzOYkl35xumHx8oAiEAzdpzo7f6cUuktEghaye2mvEMMKDuczCQi+rXCPAkxWQ="}]},"_npmUser":{"name":"aispin_dev","email":"dev@aispin.dev"},"directories":{},"maintainers":[{"name":"aispin_dev","email":"dev@aispin.dev"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/plugin-verifier_0.3.3_1787632428724_0.9800572263069651"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-21T12:06:02.080Z","modified":"2026-08-25T04:33:49.057Z","0.3.1":"2026-08-21T12:06:02.520Z","0.3.2":"2026-08-25T02:47:06.461Z","0.3.3":"2026-08-25T04:33:48.869Z"},"bugs":{"url":"https://github.com/aispin-dev/llm-as-a-Verifier-dsh/issues"},"license":"MIT","homepage":"https://github.com/aispin-dev/llm-as-a-Verifier-dsh#readme","keywords":["dsh-plugin","deepseek-harness","llm-verifier","llm-as-a-verifier","best-of-n","best-of-5","test-time-scaling","test-time-compute","verifier","selection","agent-evaluation","sampling","deepseek","deepseek-v4-flash"],"repository":{"type":"git","url":"git+https://github.com/aispin-dev/llm-as-a-Verifier-dsh.git"},"description":"LLM-as-a-Verifier for dsh with the Best-of-N conversation mode built in: rank N candidates with a fine-grained verifier (expected grade over the logprob distribution), on demand via the verify tool or automatically on every turn of a Best-of-N session. In","maintainers":[{"name":"aispin_dev","email":"dev@aispin.dev"}],"readme":"# LLM-as-a-Verifier for dsh — Best-of-N (Bo5) conversation mode\n\n**English** | [中文文档](#中文文档)\n\n> **Give DeepSeek V4 Flash test-time scaling: V4 Flash + Bo5 self-verification reaches Fable-5-level scores** — 88% on Terminal-Bench 2.1, frontier-model accuracy at a fraction of the cost (≈11× cheaper).\n\nAn independent dsh-native implementation of the test-time selection method from [LLM-as-a-Verifier](https://arxiv.org/abs/2607.05391) (arXiv:2607.05391, MIT). Method by the paper's authors; this implementation by [Aispin](https://github.com/aispin-dev).\n\n## The paper's idea, in one minute\n\nCheap models can *generate* great answers — they just can't *recognize* which one is great. LLM-as-a-Verifier closes that gap:\n\n1. **Sample N candidates** from a cheap model (DeepSeek V4 Flash): slightly different attempts at the same task.\n2. **Grade with a fine-grained verifier** — the same cheap model, asked to grade *pairs* of candidates on an A–T letter scale. The score is not the sampled letter: it is the **expectation over the grade token's logprob distribution**, Σ p(token)·φ(letter) — the model's full belief, not one dice throw.\n3. **Both orderings per pair** cancel the verifier's position bias; repeated evaluations alternate slots.\n4. **Select the best** — the paper's core result: *V4 Flash sampling 5 candidates + self-verification matches Fable-5-level frontier scores on Terminal-Bench 2.1 (88.0%) at ~1/11 the cost.*\n\nThis plugin packages that pipeline as a dsh plugin with a conversational twist: **every assistant turn becomes Best-of-N automatically** — you see one answer, the model produced five.\n\n## One plugin, three faces\n\n| Face | Entry | Use |\n|---|---|---|\n| **Tool** | `verify` tool | On demand — \"use the verify tool to compare A/B/C\", the agent calls it |\n| **Service** | `ctx.verifier.verify({ task, candidates })` | For code — orchestration lines, other plugins |\n| **Mode** | Best-of-N conversation mode | Invisible — Bo-N sessions sample every turn N ways, verify, replay only the winner |\n\n## Install\n\nFrom npm (the recommended path — resolves every dependency through your profile):\n\n```bash\ndsh plugin --profile <your-profile> add @aispin/plugin-verifier\n```\n\nOr plain npm:\n\n```bash\nnpm install @aispin/plugin-verifier\n```\n\n**Zero-config**: the verifier inherits dsh's configured provider state (credentials + settings seams) — if you've configured DeepSeek on the Models page, it just works. Try it locally:\n\n```bash\ngit clone https://github.com/aispin-dev/llm-as-a-Verifier-dsh.git\ndsh plugin --profile <your-profile> add /path/to/llm-as-a-Verifier-dsh\n```\n\n## Best-of-N: three-state switch (hot)\n\n```\n① settings global (Web UI panel) → ② session preset (\"Bo-N mode\") → ③ profile config default → off\n```\n\n**Two independent tiers in the Web settings panel**: the global tier (what the global switch turns on for every session) and the \"Bo-N 模式\" preset tier (what sessions that selected the Bo-N preset use — defaults to Bo5, set independently). Plus a user-adjustable **verify timeout** (default 90s — the ranking's own budget, never borrowed by sampling).\n\n**Sampling degrade chain**: each rollout carries the sampling budget as its own wall-clock cap — a Bo5 whose 2 rollouts overrun degrades to Bo3 and still ranks the survivors (the footer says: `采样 5 路 2 路未完成 · 3 选 1`); below 2 survivors the turn fails open to a normal answer.\n\nThe Web settings panel offers the tiers with **transparent cost cards**:\n\n| Tier | Model calls | Tokens | Latency |\n|---|---|---|---|\n| Off | 1 | 1× | 1× |\n| Fast · Bo-3 | ~9 | 2–3× | ~7–15s |\n| Precise · Bo-5 | ~16 | 3–5× | ~12–30s |\n| Custom | 2–8 ways | linear | linear |\n\nEvery turn's footer meters the real spend: `⚡ Best-of-N · 5 选 1 → 候选 #2 · 20.0/20 · 24.3s · 10.8K tok`\n\n## What's inside (implementation parity with the paper)\n\n- **Fine-grained reward**: expected grade over the top-20 logprob distribution, A=20…T=1 grouped band scale, grading at temperature 1.0 (the natural belief distribution — never collapsed)\n- **PPT pivot tournament** (the paper's O(N·k) selection): random Hamiltonian ring (each candidate lands exactly once per slot — bias cancels inside the ring) → top-k pivots → only non-pivot×pivot pairs graded. Live-verified: Bo-5 grading calls 20 → 11 (−45%)\n- **Prefix-cache prompt layout** (paper v0.2.0, −3.4× uncached tokens): criteria at the prompt tail; role + scale + task + candidates form the shared prefix\n- **Capability-adaptive grading**: logprobs endpoints (DeepSeek official) get expected-grade scoring; logprob-less endpoints (MiniMax, various gateway providers) auto-degrade to letter-sampling grading with double evaluation — any OpenAI-compatible endpoint works (`autoDegrade: false` for strict mode)\n- **Fail-open discipline**: any breakdown degrades to a normal answer with an explanatory footer — never a dead turn\n\n## License\n\nMIT © 2026 Aispin. The method is from [LLM-as-a-Verifier](https://arxiv.org/abs/2607.05391) (arXiv:2607.05391, MIT). Not affiliated with the paper's authors or DeepSeek.\n\n---\n\n# 中文文档\n\n> **给 DeepSeek V4 Flash 测试时扩展能力：V4 Flash + Bo5 自验证达到 Fable 5 级评分**——Terminal-Bench 2.1 上 88%，以前沿模型级别的准确率、约 1/11 的成本完成任务。\n\n[LLM-as-a-Verifier](https://arxiv.org/abs/2607.05391)（arXiv:2607.05391, MIT）测试时选择方法的 dsh 原生独立实现。方法归论文作者，实现归 [Aispin](https://github.com/aispin-dev)。\n\n## 论文的思想，一分钟讲清\n\n便宜模型能*生成*好答案——只是认不出*哪个*是好答案。LLM-as-a-Verifier 补上这一环：\n\n1. **采样 N 个候选**（DeepSeek V4 Flash）：同一任务的多个略有差异的尝试。\n2. **细粒度验证器评分**——同一个便宜模型，对候选**两两成对**按 A–T 字母量表打分。分数不是采样出的那个字母，而是 **grade token 对数概率分布上的期望值** Σ p(token)·φ(letter)——模型的完整信念，不是掷一次骰子。\n3. **每对双向各评一次**抵消验证器的位置偏置；重复评估交替 A/B 槽位。\n4. **选出最佳**——论文核心结论：*V4 Flash 采样 5 条候选 + 自验证择优，在 Terminal-Bench 2.1 上达到 Fable 5 级前沿评分（88.0%），成本约 1/11*。\n\n本插件把这套管线做成 dsh 插件，并加上对话形态：**每个回答自动变成 Best-of-N**——你看到一条答案，模型实际做了五条。\n\n## 一个插件，三张面孔\n\n| 面孔 | 入口 | 用法 |\n|---|---|---|\n| **工具面** | `verify` 工具 | 有感——对话里说\"用 verify 工具比较 A/B/C\"，模型主动调用 |\n| **服务面** | `ctx.verifier.verify({ task, candidates })` | 代码消费（编排线、其他插件） |\n| **模式面** | Best-of-N 对话模式 | 无感——选中模式的会话，每轮后台 N 路采样 + 择优，只把胜者呈现给用户 |\n\n## 安装\n\nnpm 安装（推荐——依赖经你的 profile 完整解析）：\n\n```bash\ndsh plugin --profile <your-profile> add @aispin/plugin-verifier\n```\n\n或直接 npm：\n\n```bash\nnpm install @aispin/plugin-verifier\n```\n\n**零配置**：验证器继承 dsh 已配置的 provider 状态（credentials + settings seam）——在 Models 页面配过 DeepSeek 即可直接用。本地试用：\n\n```bash\ngit clone https://github.com/aispin-dev/llm-as-a-Verifier-dsh.git\ndsh plugin --profile <your-profile> add /path/to/llm-as-a-Verifier-dsh\n```\n\n## Best-of-N：三态开关（热生效）\n\n```\n① settings 全局（Web 设置面板）→ ② session preset（\"Bo-N 模式\"）→ ③ profile config 默认 → 关\n```\n\n**Web 设置面板的两层独立档位**：全局档位（全局开关开启时所有会话用）+ \"Bo-N 模式\"档位（选中该 preset 的会话用，默认 Bo5，单独设置互不影响）。另有用户可调的**评分超时**（默认 90 秒——评审阶段的独立预算，不被采样挤占）。\n\n**采样降级链**：每路采样以采样预算为自身时限——Bo5 有 2 路超时则降级为 Bo3 继续对存活者择优（footer 明示：`采样 5 路 2 路未完成 · 3 选 1`）；存活不足 2 路才 fail-open 为普通回答。\n\nWeb 设置面板的档位卡直接标注**消耗透明**：\n\n| 档位 | 模型调用 | token | 延迟 |\n|---|---|---|---|\n| 关闭 | 1 次 | 1× | 1× |\n| 快速经济 · Bo-3 | ~9 次 | 2–3× | ~7–15s |\n| 精准 · Bo-5 | ~16 次 | 3–5× | ~12–30s |\n| 自定义 | 2–8 路 | 线性 | 线性 |\n\n每轮回答尾部 footer 显示实际开销：`⚡ Best-of-N · 5 选 1 → 候选 #2 · 20.0/20 · 24.3s · 10.8K tok`\n\n## 实现要点（与论文对齐）\n\n- **细粒度奖励**：top-20 logprob 分布上的期望分（A=20…T=1 分组带量表），评分温度 1.0（读自然信念分布，绝不坍缩）\n- **PPT 概率枢轴锦标赛**（论文 O(N·k) 选择算法）：随机哈密顿环（每候选恰好在 A/B 槽各一次——环内天然消位置偏置）→ top-k 枢轴 → 只补非枢轴×枢轴对。实测 Bo-5 评分调用 20 → 11（−45%）\n- **前缀缓存布局**（论文 v0.2.0，未缓存 token −3.4×）：criteria 置于 prompt 尾部，角色+量表+任务+候选构成跨调用共享前缀\n- **能力自适应评分**：有 logprobs 的端点（DeepSeek 官方）用期望分；没有的（MiniMax、部分网关）自动降级采样评分（每对双评补偿方差）——任何 OpenAI 兼容端点都能当评审（`autoDegrade: false` 切严格模式）\n- **Fail-open 纪律**：任何断裂降级为普通回答并在 footer 说明原因，绝不杀死对话轮\n\n## 许可\n\nMIT © 2026 Aispin。方法来自 [LLM-as-a-Verifier](https://arxiv.org/abs/2607.05391)（arXiv:2607.05391, MIT）。与论文作者及 DeepSeek 无隶属关系。\n","readmeFilename":"README.md"}