{"_id":"@dsh-trading/verdict","name":"@dsh-trading/verdict","dist-tags":{"latest":"0.2.0"},"versions":{"0.2.0":{"name":"@dsh-trading/verdict","version":"0.2.0","description":"Evaluation harness for backtest artifacts: lookahead lint, fill validation against real candles, seeded random baseline, sample-size power — deterministic checks that say NOT PROVEN when the evidence can't","type":"module","license":"MIT","main":"lib/index.js","types":"lib/index.d.ts","exports":{".":{"types":"./lib/index.d.ts","default":"./lib/index.js"},"./package.json":"./package.json"},"peerDependencies":{"@deepseek-ai/cordis":"^4.0.1","@deepseek-ai/dsh-tools":"0.1.0-rc.6","@deepseek-ai/schemastery":"^3.18.1","@dsh-trading/market-data":"^0.2.0"},"devDependencies":{"@deepseek-ai/cordis":"^4.0.1","@deepseek-ai/dsh-tools":"0.1.0-rc.6","@deepseek-ai/schemastery":"^3.18.1","@types/node":"^22.10.0","typescript":"^5.7.2","@dsh-trading/market-data":"^0.2.0"},"repository":{"type":"git","url":"git+https://github.com/maddogfinance/dsh-trading.git","directory":"packages/verdict"},"homepage":"https://github.com/maddogfinance/dsh-trading#readme","bugs":{"url":"https://github.com/maddogfinance/dsh-trading/issues"},"publishConfig":{"access":"public"},"keywords":["dsh","dsh-plugin","deepseek-harness","trading","backtest","evaluation","lookahead-bias","research"],"scripts":{"build":"tsc -p tsconfig.json","typecheck":"tsc -p tsconfig.json --noEmit"},"_id":"@dsh-trading/verdict@0.2.0","_integrity":"sha512-iXSBAY2hExWmSu0rm/b1PDt3WN54ugb8WGTMpOgRtRIusxLzyABzmhFGPANx/DE01DkjtDZR1r3mX5I4ZSUjBg==","_resolved":"/private/var/folders/h6/m6c5_kx15bdgwstddtmw39h80000gn/T/1c870fbc447b8de9593063d2645a1f1c/dsh-trading-verdict-0.2.0.tgz","_from":"file:dsh-trading-verdict-0.2.0.tgz","_nodeVersion":"24.12.0","_npmVersion":"11.6.2","dist":{"integrity":"sha512-iXSBAY2hExWmSu0rm/b1PDt3WN54ugb8WGTMpOgRtRIusxLzyABzmhFGPANx/DE01DkjtDZR1r3mX5I4ZSUjBg==","shasum":"49809f94eac7f8f58da87f868e266de702604a8a","tarball":"https://registry.npmjs.org/@dsh-trading/verdict/-/verdict-0.2.0.tgz","fileCount":19,"unpackedSize":69944,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQCRK/5X0U9urO9wPbZ/fcjlpV3nXbjARTfCIxyXY44RDQIhANpWSi7zWYJ9wKQBOP5r6av/0p7SQa/E98jUABku9GNB"}]},"_npmUser":{"name":"celineycn","email":"yucenning@gmail.com"},"directories":{},"maintainers":[{"name":"celineycn","email":"yucenning@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/verdict_0.2.0_1787147664220_0.2765655302494401"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-19T13:54:23.987Z","0.2.0":"2026-08-19T13:54:24.411Z","modified":"2026-08-19T13:54:24.640Z"},"maintainers":[{"name":"celineycn","email":"yucenning@gmail.com"}],"description":"Evaluation harness for backtest artifacts: lookahead lint, fill validation against real candles, seeded random baseline, sample-size power — deterministic checks that say NOT PROVEN when the evidence can't","homepage":"https://github.com/maddogfinance/dsh-trading#readme","keywords":["dsh","dsh-plugin","deepseek-harness","trading","backtest","evaluation","lookahead-bias","research"],"repository":{"type":"git","url":"git+https://github.com/maddogfinance/dsh-trading.git","directory":"packages/verdict"},"bugs":{"url":"https://github.com/maddogfinance/dsh-trading/issues"},"license":"MIT","readme":"# @dsh-trading/verdict\n\nThe evaluation harness: deterministic checks over backtest artifacts. The\nmodel proposes results; this plugin disposes of the ones that could not be\ntrue — and refuses to certify the ones the sample cannot support.\n\nLLM coding agents reproduce the leaky backtest tutorials they were trained on\n(\"LLM lookahead bias\"): anyone can now produce a professional-looking,\nself-deceiving backtest in an afternoon. This package is the detection layer.\n\n## Tools\n\n- **`audit_backtest`** — takes a backtest artifact (see CONTRACTS.md §6),\n  fetches the same candles from the mounted market-data provider (window\n  padded by whole bars so the first trade's entry bar is never cropped),\n  and runs:\n  - *Fill validation*: every entry/exit price must sit inside its bar's true\n    low..high — a price the bar never printed is a fill that never happened;\n    times beyond the data window are flagged, never validated against the\n    last bar. Same philosophy as `annotate_chart`'s trust gate. An optional\n    `priceTolerancePct` absorbs cross-vendor OHLC disagreement.\n  - *Trade independence*: duplicated or overlapping same-side trades\n    collapse to effective samples — one lucky trade copied 800 times is one\n    observation, not a track record.\n  - *Random baseline*: a seeded Monte-Carlo of random-entry twins (same\n    bars, holding periods, side mix), ranked against the strategy's\n    close-to-close SHADOW returns — optimistic intrabar fill assumptions\n    cannot buy a percentile. Results inside the luck distribution get\n    NOT PROVEN, whatever the headline return says.\n  - *Fill model*: self-reported vs shadow mean return; a material gap means\n    the edge lives in fill placement this timeframe cannot verify.\n  - *Sample-size power*: can `n` independent trades support the claim on\n    BOTH the win-rate and mean-return dimensions? Underpowered means\n    NOT PROVEN, not a prettier metric.\n  - *Plausibility*: win-rate CIs implausibly close to 100% get named.\n  - *Costs*: artifacts that don't declare fees/slippage included are flagged.\n- **`lint_strategy_code`** — pattern rules for the classic leaks:\n  `shift(-n)`, `rolling(center=True)`, `np.roll(x, -n)`, backfill,\n  full-data scaler fits, `lead()`. Hits are suspect constructs to inspect;\n  silence is not proof.\n\n## Verdict semantics\n\n`DEFECTS_FOUND` | `NOT_PROVEN` | `NO_DEFECTS_FOUND` — none of these means\n\"this strategy works\". The report says so in its own headline, caps itself\nat NOT_PROVEN whenever a substantive check could not run, and ends every\nreport by naming its own blind spots (in-sample selection, omitted trades,\nintrabar fills).\n\nEvery check is plain seeded code; identical inputs produce identical\nreports, so a session-log replay reproduces the audit exactly. No LLM grades\nits own homework here.\n\n## Artifact contract (v1)\n\n```jsonc\n{\n  \"version\": 1,\n  \"symbol\": \"DEMO-EQ\",          // exactly as the provider reports it\n  \"timeframe\": \"1d\",\n  \"trades\": [{\n    \"entryTime\": \"2026-01-02T10:00:00Z\",  // any instant inside the entry bar\n    \"exitTime\":  \"2026-01-05T14:00:00Z\",\n    \"side\": \"long\",                        // or \"short\"\n    \"entryPrice\": 100.0,\n    \"exitPrice\": 104.2\n  }],\n  \"costs\": { \"included\": false }           // optional; omit = gross returns\n}\n```\n\nCanonical copy lives in CONTRACTS.md §6. Additive-only within version 1,\nlike every dsh-trading contract.\n\n## Boundary\n\nResearch only. Verdicts judge evidence quality; nothing here forecasts,\nrecommends, or trades. Nothing is investment advice.\n\nBoth tools read the files they are pointed at (artifact JSON, strategy\nsources) from disk as given, relative to the working directory or absolute —\npoint them only at files you mean to show the model.\n","readmeFilename":"README.md","_rev":"1-fed8463ab8e4541645f5fca039d20b43"}