{"_id":"@aglegg/forge-harness","_rev":"5-b0ca747228d9f85c5493edc47f9b455f","name":"@aglegg/forge-harness","dist-tags":{"latest":"0.2.2"},"versions":{"0.1.1":{"name":"@aglegg/forge-harness","version":"0.1.1","keywords":["coding-agent","local-llm","cli","developer-tools"],"author":{"name":"Andre Glegg"},"license":"Apache-2.0","_id":"@aglegg/forge-harness@0.1.1","maintainers":[{"name":"aglegg","email":"andreglegg@me.com"}],"homepage":"https://github.com/andreglegg/forge#readme","bugs":{"url":"https://github.com/andreglegg/forge/issues"},"bin":{"forge":"bin/forge"},"dist":{"shasum":"53626d19af0d1d7074b69a18dbeeb4aaa7ea5288","tarball":"https://registry.npmjs.org/@aglegg/forge-harness/-/forge-harness-0.1.1.tgz","fileCount":95,"integrity":"sha512-b7z+/48PgIvP0MYiJ5zBAJvA2spvS9VZewNBFof4m0PQii31WL+4qeam9fhvwMx9e1HTagIoD467ZraXBK+rgA==","signatures":[{"sig":"MEYCIQDKRRCaRVNCGc3U79xD6YM4rOykPwkWkltsdLcH3Euq0QIhALZ4DYPro1N8ovI/ITWF4ICuP3jLiR+7TX0hUwxRpDBu","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":1218622},"type":"module","engines":{"node":">=22.12.0"},"gitHead":"fcbfa36718939c107ff3391d26d1e383478d6965","scripts":{"lint":"biome check .","test":"vitest run","build":"tsc -p tsconfig.build.json","check":"npm run lint && npm run typecheck && npm run test","format":"biome check --write .","prepack":"npm run check && npm run build","typecheck":"tsc -p tsconfig.json"},"_npmUser":{"name":"aglegg","email":"andreglegg@me.com"},"repository":{"url":"git+https://github.com/andreglegg/forge.git","type":"git"},"_npmVersion":"10.9.3","description":"Local-first coding-agent harness optimized for local and small language models.","directories":{},"_nodeVersion":"22.19.0","dependencies":{"zod":"^4.4.3","typescript-compiler":"npm:typescript@5.9.3"},"publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"devDependencies":{"vitest":"^4.1.10","typescript":"^7.0.2","@types/node":"^22.20.1","@biomejs/biome":"^2.5.6"},"_npmOperationalInternal":{"tmp":"tmp/forge-harness_0.1.1_1787867792561_0.38808491196251405","host":"s3://npm-registry-packages-npm-production"}},"0.1.2":{"name":"@aglegg/forge-harness","version":"0.1.2","keywords":["coding-agent","local-llm","cli","developer-tools"],"author":{"name":"Andre Glegg"},"license":"Apache-2.0","_id":"@aglegg/forge-harness@0.1.2","maintainers":[{"name":"aglegg","email":"andreglegg@me.com"}],"homepage":"https://github.com/andreglegg/forge#readme","bugs":{"url":"https://github.com/andreglegg/forge/issues"},"bin":{"forge":"bin/forge"},"dist":{"shasum":"6ccccf4ddabc65fd3217aa2c57f1074182754359","tarball":"https://registry.npmjs.org/@aglegg/forge-harness/-/forge-harness-0.1.2.tgz","fileCount":94,"integrity":"sha512-MpU9fJv2EqIMCcFmdlVQfrmnZOoTWcPLzbLTCBY9vV8uP5CKQLgMHbA9eGkUWjHydeL3Qm+Bz+wAwmdz10QpsA==","signatures":[{"sig":"MEYCIQDmBZqouriUS8xiVSPAXoR3ClSefS/YAo/SssUzfQt4agIhAMJwAeioY088RKbd7vAzzCj8xk6qwlbYHZdk2qYjmXlw","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@aglegg%2fforge-harness@0.1.2","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":1218927},"type":"module","engines":{"node":">=22.12.0"},"gitHead":"b8fc3fa1be1f174679a09182aa68722ab7bdb7d7","scripts":{"lint":"biome check .","test":"vitest run","build":"tsc -p tsconfig.build.json","check":"npm run lint && npm run typecheck && npm run test","format":"biome check --write .","prepack":"npm run check && npm run build","typecheck":"tsc -p tsconfig.json"},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:da90f081-82c7-4e86-a3c8-bfadd33acbab"}},"repository":{"url":"git+https://github.com/andreglegg/forge.git","type":"git"},"_npmVersion":"11.17.0","description":"Local-first coding-agent harness optimized for local and small language models.","directories":{},"_nodeVersion":"24.19.0","dependencies":{"zod":"^4.4.3","typescript-compiler":"npm:typescript@5.9.3"},"publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"devDependencies":{"vitest":"^4.1.10","typescript":"^7.0.2","@types/node":"^22.20.1","@biomejs/biome":"^2.5.6"},"_npmOperationalInternal":{"tmp":"tmp/forge-harness_0.1.2_1787869967733_0.11758585697583546","host":"s3://npm-registry-packages-npm-production"}},"0.2.0":{"name":"@aglegg/forge-harness","version":"0.2.0","keywords":["coding-agent","local-llm","cli","developer-tools"],"author":{"name":"Andre Glegg"},"license":"Apache-2.0","_id":"@aglegg/forge-harness@0.2.0","maintainers":[{"name":"aglegg","email":"andreglegg@me.com"}],"homepage":"https://github.com/andreglegg/forge#readme","bugs":{"url":"https://github.com/andreglegg/forge/issues"},"bin":{"forge":"bin/forge"},"dist":{"shasum":"2bc82b12de67b7e202e78d69bee0ca2d1fdab705","tarball":"https://registry.npmjs.org/@aglegg/forge-harness/-/forge-harness-0.2.0.tgz","fileCount":96,"integrity":"sha512-5U68PlogQpRBpn3cuJU0mfC0512QxhWV88kVTqSVs3Gj9sZZbjtyZDerDkK9E9Yp4KJOo0wQEJllOm+rO2Rl8Q==","signatures":[{"sig":"MEUCICithh/Z6nMyBWY99/1eMoEKRdTkKuE6eggVOS45dOttAiEAmbMF0S5fjaaSx34iCfN+XGol/KH3bGV4EaokYDItDaY=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@aglegg%2fforge-harness@0.2.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":1253426},"type":"module","engines":{"node":">=22.12.0"},"gitHead":"1d733a85b4cc0ffafdcd9286f33cdaf1f9c40769","scripts":{"lint":"biome check .","test":"vitest run","build":"tsc -p tsconfig.build.json","check":"npm run lint && npm run typecheck && npm run test","format":"biome check --write .","prepack":"npm run check && npm run build","typecheck":"tsc -p tsconfig.json"},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:da90f081-82c7-4e86-a3c8-bfadd33acbab"}},"repository":{"url":"git+https://github.com/andreglegg/forge.git","type":"git"},"_npmVersion":"11.17.0","description":"Local-first coding-agent harness optimized for local and small language models.","directories":{},"_nodeVersion":"24.19.0","dependencies":{"zod":"^4.4.3","typescript-compiler":"npm:typescript@5.9.3"},"publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"devDependencies":{"vitest":"^4.1.11","typescript":"^7.0.2","@types/node":"^22.20.1","@biomejs/biome":"^2.5.11"},"_npmOperationalInternal":{"tmp":"tmp/forge-harness_0.2.0_1787872452732_0.2800340315683063","host":"s3://npm-registry-packages-npm-production"}},"0.2.1":{"name":"@aglegg/forge-harness","version":"0.2.1","keywords":["coding-agent","local-llm","cli","developer-tools"],"author":{"name":"Andre Glegg"},"license":"Apache-2.0","_id":"@aglegg/forge-harness@0.2.1","maintainers":[{"name":"aglegg","email":"andreglegg@me.com"}],"homepage":"https://github.com/andreglegg/forge#readme","bugs":{"url":"https://github.com/andreglegg/forge/issues"},"bin":{"forge":"bin/forge"},"dist":{"shasum":"3fee92d207a49c2664f89ac12206792db1287199","tarball":"https://registry.npmjs.org/@aglegg/forge-harness/-/forge-harness-0.2.1.tgz","fileCount":96,"integrity":"sha512-A2OtIHDuV4A20GMTrn7nNCOD8agE1uyeq6aQcmJHJ/CpvAo5GXumMcLPTGIkpzs3JOqceEyDKC/uEpQJAIj83A==","signatures":[{"sig":"MEQCIBlD4iLgTDztrMpg0gF8CoR1rPNDr1/uQowZ9weaNcXLAiA6dMoW3wxqPYRRFImUVNfgtrLO6T3efXs1QhKAsiN68g==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@aglegg%2fforge-harness@0.2.1","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":1255538},"type":"module","engines":{"node":">=22.12.0"},"gitHead":"03d51c51020796c200c04205895bcaab76813871","scripts":{"lint":"biome check .","test":"vitest run","build":"tsc -p tsconfig.build.json","check":"npm run lint && npm run typecheck && npm run test","format":"biome check --write .","prepack":"npm run check && npm run build","typecheck":"tsc -p tsconfig.json"},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:da90f081-82c7-4e86-a3c8-bfadd33acbab"}},"repository":{"url":"git+https://github.com/andreglegg/forge.git","type":"git"},"_npmVersion":"11.17.0","description":"Local-first coding-agent harness optimized for local and small language models.","directories":{},"_nodeVersion":"24.19.0","dependencies":{"zod":"^4.4.3","typescript-compiler":"npm:typescript@5.9.3"},"publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"devDependencies":{"vitest":"^4.1.11","typescript":"^7.0.2","@types/node":"^22.20.1","@biomejs/biome":"^2.5.11"},"_npmOperationalInternal":{"tmp":"tmp/forge-harness_0.2.1_1787875428563_0.20589784382504184","host":"s3://npm-registry-packages-npm-production"}},"0.2.2":{"name":"@aglegg/forge-harness","version":"0.2.2","description":"Local-first coding-agent harness optimized for local and small language models.","author":{"name":"Andre Glegg"},"repository":{"type":"git","url":"git+https://github.com/andreglegg/forge.git"},"homepage":"https://github.com/andreglegg/forge#readme","bugs":{"url":"https://github.com/andreglegg/forge/issues"},"type":"module","license":"Apache-2.0","bin":{"forge":"bin/forge"},"keywords":["coding-agent","local-llm","cli","developer-tools"],"engines":{"node":">=22.12.0"},"publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"scripts":{"build":"tsc -p tsconfig.build.json","typecheck":"tsc -p tsconfig.json","lint":"biome check .","format":"biome check --write .","test":"vitest run","check":"npm run lint && npm run typecheck && npm run test","prepack":"npm run check && npm run build"},"dependencies":{"typescript-compiler":"npm:typescript@5.9.3","zod":"^4.4.3"},"devDependencies":{"@biomejs/biome":"^2.5.11","@types/node":"^22.20.1","typescript":"^7.0.2","vitest":"^4.1.11"},"gitHead":"9b5171462d9673c805c8657b3012953486464688","_id":"@aglegg/forge-harness@0.2.2","_nodeVersion":"24.19.0","_npmVersion":"11.17.0","dist":{"integrity":"sha512-d9focxEktwmcYdz2BbRvB4tJ+2f0Xo2zsooHHhN6OFAgMjfefKAWE02UoDlx1/oX04ar1inuD6wWjIj3s5Qx6g==","shasum":"34c006ac395b54a4aab8697817001f2c69689016","tarball":"https://registry.npmjs.org/@aglegg/forge-harness/-/forge-harness-0.2.2.tgz","fileCount":96,"unpackedSize":1255785,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@aglegg%2fforge-harness@0.2.2","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQD6HLIqSDLYGCAlEm2VGz7sK8lOM8xDvXUP94UcYDnuwgIhAL8TvrJCQcLReoLe1tVqEUMWZ8flwRoBw4HelmHgKInI"}]},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:da90f081-82c7-4e86-a3c8-bfadd33acbab"}},"directories":{},"maintainers":[{"name":"aglegg","email":"andreglegg@me.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/forge-harness_0.2.2_1787876287547_0.774184930481602"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-27T21:56:32.240Z","modified":"2026-08-28T00:18:08.047Z","0.1.1":"2026-08-27T21:56:32.694Z","0.1.2":"2026-08-27T22:32:47.890Z","0.2.0":"2026-08-27T23:14:12.905Z","0.2.1":"2026-08-28T00:03:48.753Z","0.2.2":"2026-08-28T00:18:07.735Z"},"bugs":{"url":"https://github.com/andreglegg/forge/issues"},"author":{"name":"Andre Glegg"},"license":"Apache-2.0","homepage":"https://github.com/andreglegg/forge#readme","keywords":["coding-agent","local-llm","cli","developer-tools"],"repository":{"type":"git","url":"git+https://github.com/andreglegg/forge.git"},"description":"Local-first coding-agent harness optimized for local and small language models.","maintainers":[{"name":"aglegg","email":"andreglegg@me.com"}],"readme":"# forge — a coding-agent harness for local and small models\n\n<p align=\"center\">\n  <img src=\"docs/assets/forge-wordmark.png\" alt=\"Forge\" width=\"420\">\n</p>\n\n[![CI](https://github.com/andreglegg/forge/actions/workflows/ci.yml/badge.svg)](https://github.com/andreglegg/forge/actions/workflows/ci.yml)\n[![npm](https://img.shields.io/npm/v/%40aglegg%2Fforge-harness)](https://www.npmjs.com/package/@aglegg/forge-harness)\n[![license](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)\n\nA local-first, chat-first coding-agent harness written in TypeScript and optimized\nfor local and small language models.\n\n**New here? Read [`USING_FORGE.md`](USING_FORGE.md).**\n\n## Install\n\nForge requires Node.js 22.12 or newer. From npm:\n\n```sh\nnpm install --global @aglegg/forge-harness\ncd /path/to/your/project\nforge doctor\nforge init\nforge\n```\n\nFrom this repository:\n\n```sh\nnpm ci\nnpm run build\nnode bin/forge --version\n```\n\nForge is a **public alpha**. See [`docs/STATUS.md`](docs/STATUS.md) for the exact shipped/planned boundary, [`docs/PRODUCT_PLAN.md`](docs/PRODUCT_PLAN.md) for the ordered product milestones, and [`docs/SECURITY.md`](docs/SECURITY.md) before using autonomous approval on valuable code.\n\nGlobal npm installs check the npm registry for a newer stable Forge release at\nmost once every 12 hours. When a compatible update exists, Forge installs that\nexact version before loading the CLI, so the current invocation continues on the\nnew build. Source checkouts, project-local installs, `npx`, and CI are never\nself-modified; an offline or failed update check never blocks Forge.\n\nForce a check with `forge update`. Disable automatic checks with\n`FORGE_AUTO_UPDATE=0`. If a self-update cannot be applied, the recovery command\nis still:\n\n```sh\nnpm install --global @aglegg/forge-harness@latest\n```\n\n## Status\n\n`forge` on PATH runs this. It indexes bounded Git-aware project trees, can\nnavigate deep monorepos with scoped list/glob/search, ranged reads, static\nTypeScript/JavaScript module relationships, and revision-bound symbol declarations, supports read-only and plan modes,\nand applies anchored edits plus transactional file and directory operations. It also runs approved commands, verifies completion, and\nasks before workspace effects unless explicitly automated. Driven end to end\nagainst local Qwen coder models.\n\n**Measured, as of 2026-08-04:**\n\n| what | result |\n|------|--------|\n| Aider Polyglot, 225 cases | **60.00%** (135/225) |\n| Pinned 42-case model screen | **28/42** and **27/42** across two Forge runs |\n| Local multi-file suite, 14 tasks | **14/14**, 0 false successes, 0 damaged |\n| Greenfield project, unsupervised | **builds it, then over-reports** — see `bench/PROJECT_TRIAL.md` |\n| Model floor | 30B-MoE 64-67%, 14B 12-14%, 7B 4.8% — a cliff, not a gradient |\n\n**To use it on a real project, read [`USING_FORGE.md`](USING_FORGE.md).** It\ncovers setup, what the numbers above do and do not license, and the failure\nmodes to watch for.\n\nFull evidence, including two interventions that were measured and rejected, is\nin [`bench/`](bench/).\n\nThe honest caveat: the pty test that drives the approval prompt through a real\nterminal **skips** where no controlling terminal is available — including the\nshell it was developed in. On such a machine that one path is genuinely\nunverified, and the test says so rather than passing emptily.\n\n## Development approach\n\nForge is designed, maintained, and evaluated by Andre Glegg using AI-assisted\ndevelopment under a test-first, benchmark-gated workflow. The maintainer owns\nthe architecture, safety constraints, evaluation methodology, and release\ndecisions. Contributions are judged by observable behavior, tests, security\ninvariants, and reproducible evidence—not by who or what typed the first draft.\n\n## The instrument\n\nEvery run records what the model actually said, to `.forge/traces/`. `forge\nreplay` re-decodes that fixed text through the current decoder and reports how\nmuch of it became a usable action:\n\n```\n$ forge replay\n9 of 9 replies converted (100.00%), 1 repaired\nrepairs applied\n  orphan_search_block      1\n  stray_marker             1\n```\n\nNo model is contacted, no repository is read, nothing is timed. So any movement\nin that number is this code's doing and nothing else — which makes it something\nyou can run per commit and bisect on. A coding benchmark cannot do that: it\nmoves the model and the harness together, and its own variance swamps the\ndifference you were looking for.\n\nRecording is unconditional rather than behind a flag, because a corpus you have\nto remember to switch on is empty exactly when you need it: after the run that\nwent wrong.\n\nWhat it measures is *conversion* — did the reply become an action. Not whether\nthe resulting edit was any good. A confident, wrong patch counts as converted.\nConversion is a floor on quality, not a substitute for it.\n\n## Running it\n\n```\nforge                     interactive chat in the current directory\nforge run \"<task>\"        one shot, exit 0 on success\nforge run \"<task>\" --yes  approve workspace effects (CI, batch)\nforge run \"<task>\" --yes --isolate\n                          run in a detached clean Git worktree and retain a patch\nforge run \"<task>\" --yes --isolate --promote\n                          apply the verified patch after conflict/risk checks\nforge plan \"<task>\"       inspect and return a plan without effects\nforge run \"<question>\" --read-only\nforge doctor              provider/model/verifier diagnostics\nforge init                create an idempotent forge.json\nforge config --json       resolved configuration\nforge profiles            list named local-model profiles\nforge update              force an npm update check for global installs\nforge continue [id]       reopen interactive chat with retained history\nforge replay              score the decoder on everything recorded here\nforge sessions            what has been run here\nforge show <id>           replay a recorded session\nforge undo [id]           put back what a session changed\nforge --version            print the installed version\nforge polyglot <dataset> --name <run>\n                          run/resume Aider's 225-case Polyglot benchmark\nforge compare <a> <b>     paired comparison of two Polyglot report files\n\n  --profile <name>        select a named profile from forge.json\n  --native                use the provider's tool-calling instead of the text\n                          protocol (OpenAI and Anthropic wires, auto-detected)\n  --task-packet           include bounded student-facing exercise docs\n  --batch-actions         invite bounded independent reads/known-file edits\n  --isolate               require a clean Git root and protect the original checkout\n  --promote               apply a verified isolated patch to the original\n  --allow-risk            override reviewed critical patch-risk findings\n  --hooks                 explicitly enable project lifecycle hooks (headless)\n  --mode <mode>           workspace, read-only, or plan\n  --stream-json           durable Run events as JSONL plus a final result\n```\n\nIt defaults to `http://127.0.0.1:8790/v1` and asks the endpoint what it serves\nvia `/v1/models`. Coding commands also perform a minimal completion preflight,\nso a gateway that advertises an inactive model profile fails before the agent\nstarts. Override with `--url` / `FORGE_URL` and `--model` / `FORGE_MODEL`.\n\nNamed profiles keep local-model-specific settings together:\n\n```json\n{\n  \"profile\": \"local-30b\",\n  \"profiles\": {\n    \"local-30b\": {\n      \"url\": \"http://127.0.0.1:44100/v1\",\n      \"model\": \"qwen3-coder-30b-a3b-instruct-q4_k_m\",\n      \"contextWindow\": 65536,\n      \"maxTokens\": 4096,\n      \"temperature\": 0.1,\n      \"native\": false,\n      \"maxTurns\": 12\n    }\n  },\n  \"verify\": [[\"npm\", \"test\"]]\n}\n```\n\nUse `forge profiles`, `forge config --json`, or override for one run with\n`--profile <name>`. Explicit CLI flags and `FORGE_URL` / `FORGE_MODEL` take\nprecedence over a profile.\n\nVerification detection understands npm, pnpm, Yarn, and Bun root projects. It\nprefers a root `check` script when present and otherwise uses `test`. Monorepos\ncan configure package-specific checks without a shell:\n\n```json\n{\n  \"verify\": [\n    [\"cd\", \"packages/api\", \"&&\", \"pnpm\", \"test\"],\n    [\"cd\", \"apps/web\", \"&&\", \"pnpm\", \"check\"]\n  ]\n}\n```\n\nThe `cd <repository-directory> && <one command>` spelling is parsed into a\nvalidated working directory and a token-array command; it is never passed to a\nshell.\n\n`--isolate` is the safest headless repository mode currently shipped. It requires\nthe selected path to be the Git root and completely clean, including untracked\nfiles. The model and verifier run in a detached temporary worktree. Forge writes\nthe resulting binary patch to `.forge/isolated/`; without `--promote`, the\noriginal remains unchanged. Promotion requires verification and rechecks HEAD,\ncleanliness, and `git apply --check`. Added patch lines are also scanned for\nlikely credentials/private keys, package install lifecycle scripts, dangerous\nworkflows, and dependency metadata changes. Critical findings retain the patch\nand block promotion unless you inspect it and explicitly add `--allow-risk`.\nThis heuristic scan is not proof that a patch is safe. Worktree isolation\nprotects repository mutations but is not an OS, network, or process sandbox.\n\nHeadless project hooks never run merely because a repository defines them.\nEnable them explicitly with `--hooks`:\n\n```json\n{\n  \"hooks\": {\n    \"sessionStart\": [[\"node\", \"scripts/forge-start.mjs\"]],\n    \"beforeVerify\": [[\"npm\", \"run\", \"format:check\"]],\n    \"afterVerify\": [[\"node\", \"scripts/forge-audit.mjs\"]],\n    \"sessionEnd\": [[\"node\", \"scripts/forge-notify.mjs\"]]\n  }\n}\n```\n\nHooks are sequential shell-free token arrays with a scrubbed environment,\n60-second timeout, bounded output, and fail-closed exit semantics. They receive\n`FORGE_HOOK_EVENT`, `FORGE_SESSION_ID`, and post-verification hooks receive\n`FORGE_VERIFIED`. Hooks remain arbitrary repository-controlled programs;\ninspect them before opting in.\n\n`FORGE_TRACE=<path>` writes every turn's exact bytes plus the decoded proposals\nand repairs, one JSON object per line. Reach for it first when a run misbehaves:\nevery early guess about a live failure in this package has been wrong, and the\ntrace has settled each one in a single run.\n\n## Benchmarking and promotion criteria\n\nThe target is not “one run scored higher.” It is a paired, reproducible gain on\nthe same cases, model weights, endpoint, budgets, verifier and dataset commit.\nEvery Polyglot run writes those inputs plus a fingerprint of the built Forge\nexecutable to `.forge/benchmarks/polyglot/<name>/run.json`; a changed setting\ncannot resume into an old run.\n\nInstall the language toolchains required by the selected cases. Python\nverification uses `uv` to provide an isolated cached pytest. Java needs a JDK\nand a valid `JAVA_HOME`; Forge contains `GRADLE_USER_HOME` inside the run so a\nbroken or shared user cache cannot corrupt the verdict. CMake, Go, npm and Cargo\nmust likewise be on `PATH` for their language cases.\n\nStart with a deterministic 42-case screen: seven evenly spaced exercises from\neach of the six official languages. Run the conservative control first, then\nthe candidate profile:\n\n```\nnpm run build\n\nforge polyglot /path/to/polyglot-benchmark \\\n  --name model-screen-control --per-language 7 --model-digest <weights-digest>\n\nforge polyglot /path/to/polyglot-benchmark \\\n  --name model-screen-candidate --per-language 7 --model-digest <weights-digest> \\\n  --task-packet --batch-actions\n\nforge compare \\\n  .forge/benchmarks/polyglot/model-screen-control/report.json \\\n  .forge/benchmarks/polyglot/model-screen-candidate/report.json\n```\n\nFor idea generation, do not pay for a paired 42-case screen after every edit.\n`--discover` defaults to two evenly spaced cases per language, one attempt and\neight turns. Narrow it further with `--case` when a change targets a retained\nfailure; explicit selection and budget flags override the preset:\n\n```\nforge polyglot /path/to/polyglot-benchmark \\\n  --name idea-context-v1 --discover --task-packet --jobs 2\n```\n\nThis mode is evidence for choosing the next experiment, not for making a score\nclaim. Use the paired 42-case screen only when a candidate survives its focused\nfailure replays and discovery sample. `--jobs N` runs isolated cases concurrently;\nstart with two and watch model-server memory and throughput. A single-GPU server\nthat serializes requests will not benefit, so the default remains one worker.\n\n`--task-packet` admits only a small allowlist of student-facing specifications,\nincluding `.docs/instructions.md`; `.meta` reference solutions are never copied\ninto the worktree. `--batch-actions` changes prompting, not mutation safety:\nreads can be proposed together, while edits still go through preview, approval,\nrevision revalidation and serialized commit.\n\nPromote a candidate only when it has more paired gains than regressions, does\nnot increase false-success attempts, and does not trade a small score gain for\na large runtime regression. Then run all 225 official cases by omitting\n`--per-language`. Long runs are atomic and resumable. `--batch-size 6` processes\nsix pending cases per invocation if the endpoint needs short supervised batches.\nThe default protocol uses a 12-turn first attempt, independent verification,\nthen a fresh eight-turn retry using the test failure. A promotion claim should\nrequire a positive paired delta over the pinned control, an exact McNemar\np-value below 0.05, no increase in false-success attempts, and no material\nruntime regression on the same model and hardware.\n\nFailures retain the worktree, agent logs, verification output, timeout state,\nfailure class, turns, actions and elapsed time. Inspect classes and paired\nregressions before changing prompts. Syntax/protocol failures call for codec\nwork; no-progress cases call for retrieval or steering; test failures call for\nlanguage-level guidance. That containment is important: broad prompt growth is\nusually a context and latency regression for a small model.\n\nProject-specific guidance can live in root `AGENTS.md` or `CLAUDE.md`. Small\nreusable instruction packs go in `.forge/instructions/*.md`:\n\n```\n---\nkeywords: [rust, borrow, lifetime]\n---\nPrefer the smallest ownership change. Re-run the narrow failing test first.\n```\n\nAt most two keyword-matched packs are included. Add one only after retained\nfailures show a repeated, general error class, then rerun the same paired screen.\n\n## Project-scale repository navigation\n\nForge builds a deterministic repository index before each task. In a Git\nrepository it uses tracked plus non-ignored untracked paths; outside Git it uses\na bounded filesystem walk. The initial prompt receives a compact project map\nwith top-level areas and important manifests instead of a shallow file dump.\n\nThe text and native protocols expose the same inspection operations:\n\n```text\nLIST packages/api\nGLOB **/*.test.ts\nGREP RequestHandler\nSEARCH exact literal text\nRELATED packages/api/src/server.ts\nSYMBOL RequestHandler\nREFERENCES RequestHandler\nCALLERS RequestHandler\nREAD packages/api/src/server.ts:120-220\n```\n\n`LIST` shows one directory level. `GLOB`, `GREP`, and `SEARCH` operate across the\ncomplete bounded index, including paths deeper than three levels and beyond 200\nfiles. `RELATED` reports the nearest package root, direct relative module\ndependencies, inbound dependents, and related tests for TypeScript/JavaScript\nfiles. Tasks that explicitly name code-shaped symbols automatically rank and inline matching declaration, caller, reference, and one-hop dependency files, even when filenames do not share the task wording. This automatic pre-turn analysis is capped at 200 supported source files; larger repositories keep the lightweight lexical path and can use the explicit semantic tools. `SYMBOL` reports exact declarations, `REFERENCES` reports syntax occurrences, and `CALLERS` uses the TypeScript checker to resolve direct calls and constructor calls across relative-import aliases and lexical scopes. Every result carries an exact range and source revision. `READ path:start-end` gives an exact line range, while an unrestricted\nlarge read is clipped with a continuation instruction. Search skips binary files\nand files above 2 MiB; reads are bounded to 16,000 characters per action. A\nranged, clipped, or failed read never authorizes a wholesale replacement of an\nexisting file; large files must be changed with anchored edits. The index is\ncapped at 50,000 entries.\n\nRelationship and symbol scans are bounded to 10,000 supported source files and 512 KiB per\nfile. Module relationships resolve relative imports, export-from declarations, `require`, string-literal dynamic imports, common source extensions, TypeScript `.js` specifiers, and directory indexes. `CALLERS` resolves direct checker-visible calls and constructor calls, but not dynamic dispatch, reflection, package/path aliases, inferred runtime targets, or untyped calls.\n\nGenerated dependency/build/cache directories, common credential files,\n`.docs`, and `.meta` are excluded from ordinary context. Exercise specifications\nunder `.docs` remain available only through the explicit `--task-packet` mode.\nAll resolved paths remain constrained to the selected repository.\n\n## Filesystem operations\n\nThe text and native protocols expose the same first-class workspace operations:\n\n```text\nDELETE path/to/file-or-directory\nMKDIR path/to/new/directory\nMOVE source/path -> destination/path\nCOPY source/path -> destination/path\nRENAME old/path -> new/path\n```\n\n`DELETE` can remove a file, symlink, empty directory, or bounded non-empty\ndirectory tree. `MOVE`, `COPY`, and `RENAME` accept files, directories, binary\ncontent, and symlinks without following them. A final symlink may be targeted\nas an entry; structured mutations never traverse a symlinked parent directory.\nDestinations must not already exist; replacement requires a separately visible\n`DELETE` first.\n\nEvery operation is previewed before approval, revalidates exact source and\ndestination snapshots before commit, and records each affected entry for\nbinary-safe guarded `forge undo`. The repository root and protected metadata\nroots `.git`, `.forge`, and `.codex-bridge` cannot be targeted. First-class tree\noperations are bounded to 10,000 entries and 128 MiB of retained content. Larger\nor specialized operations require an explicitly approved `RUN` command and do\nnot gain the same structured preview or per-entry undo guarantees.\n\n## The shape\n\n```\nCommands ──▶ Run actor (single writer) ──▶ Effects\n   ▲               │                          │\n   │               ▼                          ▼\napprove/cancel  durable event journal    model / tools / fs\n   │               │                          │\n   └───────────────┴───── results ◀───────────┘\n                   │\n        projections (pure, many)\n        ├─ interactive renderer\n        └─ headless renderer\n```\n\nAn **actor**, not an async generator. A generator that yields events and takes\napprovals through `.next(value)` makes approval awkward, cancels the run when a\nconsumer stops iterating, and forces consumers to *compete* for values — but the\nrenderer, the journal and a session recorder all need to watch independently.\n\nDurable domain events are journalled and replay to identical `RunState`. Token\ndeltas are ephemeral presentation signals: coalesced, droppable under\nbackpressure, never journalled. A replay reconstructs decisions, not typing.\n\n**`submit()` returns its results.** Callers must never rebuild them by watching\nthe event stream — subscribers are asynchronous, so an array they populate has\nnot necessarily been filled when the `await` returns. Events are for observing;\nthe return value is causal.\n\n## One semantic protocol, several codecs\n\n```\nprovider bytes → deltas → text codec | native codec → ActionProposal → …\n```\n\nThe boundary is `ActionProposal`. A SEARCH/REPLACE block and a native tool call\nare two spellings of one intent, and a test asserts they decode identically.\nDrift is prevented structurally, not by discipline: both codec surfaces *and*\nthe prompt that teaches them are generated from one `TOOLS` registry, so a tool\ncannot exist in one and not the other.\n\nBoth codecs are incremental and neither emits until a terminator arrives. Half an\nedit must never become an action.\n\n## The seam everything else composes onto\n\n```\npropose → preview (in memory, nothing touched) → approve → revalidate → commit\n```\n\nThe preview computes the diff without writing, so the user approves something\nthat exists before it happens. The commit re-checks the file's sha256 and refuses\na stale proposal: approval is consent for a *specific version*, and a file can\nmove between proposing and committing.\n\n`always` is scoped to an action class for the session. Approving an edit must\nnever silently approve command execution.\n\n## What live small models taught this code\n\nEach of these is a defect that unit tests could not see and a frontier model\nwould probably have hidden.\n\n- **The model will declare victory over a failed edit.** A `final` in a turn\n  whose own actions failed and committed nothing is rejected. A false success is\n  worse than a failure, because nothing downstream checks again.\n- **It writes anchors before it looks.** A turn that both reads and mutates has\n  the mutation deferred — the anchor was composed before the file was seen.\n- **It cannot tell whether an edit landed.** A commit returns the resulting file,\n  not just \"applied\". Without that it edits again; it once appended the same\n  function four times and called the run a success.\n- **Loop safety must be precise, not absolute.** Reads are keyed by path *plus\n  content revision*, so re-reading an unchanged file is refused with advice while\n  re-reading after an edit is allowed. A mutation resets the repeat counts too:\n  the world changed.\n- **Never feed the raw reply back.** The transcript carries a canonical\n  re-rendering of the decoded turn. Otherwise a malformed marker teaches the\n  model the malformation is fine, and an invented SEARCH body stays in the\n  conversation and anchors every later attempt to the invention.\n- **Marker direction is noise.** `>>>>>>> SEARCH` is accepted as readily as\n  `<<<<<<< SEARCH` and counted as a repair. There is nothing else that line\n  could mean.\n\n## Commands\n\n```\nnpm run check    lint + typecheck + test\nnpm test\nnpm run build\n```\n","readmeFilename":"README.md"}