{"_id":"@metaharness/weight-eft","_rev":"3-dc605a193c2e851efbed68d1ee13a256","name":"@metaharness/weight-eft","dist-tags":{"latest":"0.1.2"},"versions":{"0.1.0":{"name":"@metaharness/weight-eft","version":"0.1.0","keywords":["llm","lora","fine-tuning","evolutionary-fine-tuning","weight-eft","sft","dpo","distillation","cost-optimization","swe-bench","metaharness","darwin-mode","contamination-guard"],"author":{"name":"rUv","email":"ruv@ruv.net"},"license":"MIT","_id":"@metaharness/weight-eft@0.1.0","maintainers":[{"name":"ruvnet","email":"ruv@ruv.net"}],"homepage":"https://github.com/ruvnet/agent-harness-generator","bugs":{"url":"https://github.com/ruvnet/agent-harness-generator/issues"},"bin":{"weight-eft":"dist/cli.js"},"dist":{"shasum":"0211bb9993fff72a1edfc92951b48d88dc99e00e","tarball":"https://registry.npmjs.org/@metaharness/weight-eft/-/weight-eft-0.1.0.tgz","fileCount":35,"integrity":"sha512-5Gta1vUWUDO+9KhjUYkwMG1HGBLLD9GagDL87bHDOI6ynZ8ACscWtQaH8ksUVeIH3ex/oLFG01XIgu1ktr0pIQ==","signatures":[{"sig":"MEYCIQCEH+2ALCoRhUgn5oBgBOXkRdqzZXZLzn9bTqqQLgXAQAIhAKeh+NFeK9Y8hLTZXLjbZhcALY6fyB4Fnw19IutdU1Ue","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":106204},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","engines":{"node":">=20.0.0"},"exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"},"./cli":{"types":"./dist/cli.d.ts","import":"./dist/cli.js"}},"gitHead":"d473ccab8593ca8ebc4ef585cd9e541836f7b308","scripts":{"lint":"tsc --noEmit","test":"vitest run","build":"tsc"},"_npmUser":{"name":"ruvnet","email":"ruv@ruv.net"},"repository":{"url":"git+https://github.com/ruvnet/agent-harness-generator.git","type":"git","directory":"packages/weight-eft"},"_npmVersion":"10.9.7","description":"Evolutionary fine-tuning — distill the harness's archival success into the open cheap tier (GLM/Qwen) via LoRA so the cost-cascade escalates to a frontier model less often. SFT-distills ALL gold-resolved trajectories; on-policy DPO on GLM-vs-GLM pairs onl","directories":{},"_nodeVersion":"22.22.2","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"vitest":"^2.0.0","typescript":"^5.4.0"},"_npmOperationalInternal":{"tmp":"tmp/weight-eft_0.1.0_1782577619706_0.6636351404574989","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"name":"@metaharness/weight-eft","version":"0.1.1","keywords":["llm","lora","fine-tuning","peft","sft","dpo","rlhf-alternative","model-distillation","knowledge-distillation","ai-agents","coding-agent","agentic","llm-agent","swe-bench","llm-routing","model-cascade","cost-optimization","llm-cost","openrouter","qwen","deepseek","training-data","jsonl","trl","axolotl","unsloth","self-improving","weight-eft","metaharness"],"author":{"name":"rUv","email":"ruv@ruv.net"},"license":"MIT","_id":"@metaharness/weight-eft@0.1.1","maintainers":[{"name":"ruvnet","email":"ruv@ruv.net"}],"homepage":"https://github.com/ruvnet/agent-harness-generator/tree/main/packages/weight-eft#readme","bugs":{"url":"https://github.com/ruvnet/agent-harness-generator/issues"},"bin":{"weight-eft":"dist/cli.js"},"dist":{"shasum":"3cff5978274d0d91814a0708402aa7a0fa2e1272","tarball":"https://registry.npmjs.org/@metaharness/weight-eft/-/weight-eft-0.1.1.tgz","fileCount":35,"integrity":"sha512-GSg0APPAbRK93OzrzlE+R8hfEK+I5+Zhmh0Z28RC9Mk5/MjhPo3shqINO7ye8VPGYHIO4rars9FwCWbe/V4cEQ==","signatures":[{"sig":"MEUCIGUAjmS52f3ok5v6iybdpJuPAdF1SfVw71DSmNJ2jZwRAiEA+xrOe+AwXe+amXs+epazs8pHZ9LNbu9WfirZl6h2MjE=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":107774},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","engines":{"node":">=20.0.0"},"exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"},"./cli":{"types":"./dist/cli.d.ts","import":"./dist/cli.js"}},"gitHead":"e0ad8a4d1cc7e0949c2ac08f4b63d7f6a1bee2f5","scripts":{"lint":"tsc --noEmit","test":"vitest run","build":"tsc"},"_npmUser":{"name":"ruvnet","email":"ruv@ruv.net"},"repository":{"url":"git+https://github.com/ruvnet/agent-harness-generator.git","type":"git","directory":"packages/weight-eft"},"_npmVersion":"10.9.7","description":"Fine-tune cheap open-source LLMs (GLM, Qwen, DeepSeek) on your AI coding agent's successful runs with LoRA (SFT + DPO) so your model cascade escalates to expensive frontier models (GPT, Claude) less often — cutting cost-per-resolved. Turns run history int","directories":{},"_nodeVersion":"22.22.2","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"vitest":"^2.0.0","typescript":"^5.4.0"},"_npmOperationalInternal":{"tmp":"tmp/weight-eft_0.1.1_1782579446544_0.16030619335668983","host":"s3://npm-registry-packages-npm-production"}},"0.1.2":{"_id":"@metaharness/weight-eft@0.1.2","bin":{"weight-eft":"dist/cli.js"},"bugs":{"url":"https://github.com/ruvnet/agent-harness-generator/issues"},"dist":{"shasum":"b679cd98f284132db754189d80d71d9a853e0e86","tarball":"https://registry.npmjs.org/@metaharness/weight-eft/-/weight-eft-0.1.2.tgz","fileCount":35,"integrity":"sha512-2C6Ivgz1yJgRrklG0NqsNAjm+THyFsfX4DR2VyeqkLHJeR+LnVEOmzdUPQPp0yBJcXvoO9+du8R6HzZ4vDItrQ==","signatures":[{"sig":"MEYCIQDGMIvX52Hbnpbu3D67Op3VrlzrpiEPLREyaA9Y/Rd58wIhALzSoDbWP4Lc9wKSPK4mtnDICy8NNUZbN9rcFCpx1wSe","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"},{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQCmXSVP9LbQjsS3hgCEyBbxZyfWtiedtKpjU1nmLR6c4wIgWCnJ7CZd4ttz/SvFtwIn3kkJsTsKVfV1I1PXfjT3S0o="}],"unpackedSize":115755},"main":"./dist/index.js","name":"@metaharness/weight-eft","type":"module","types":"./dist/index.d.ts","author":{"name":"rUv","email":"ruv@ruv.net"},"engines":{"node":">=20.0.0"},"exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"},"./cli":{"types":"./dist/cli.d.ts","import":"./dist/cli.js"}},"license":"MIT","scripts":{"lint":"tsc --noEmit","test":"vitest run","build":"tsc"},"version":"0.1.2","_npmUser":{"name":"ruvnet","email":"ruv@ruv.net"},"homepage":"https://github.com/ruvnet/agent-harness-generator/tree/main/packages/weight-eft#readme","keywords":["llm","lora","fine-tuning","peft","sft","dpo","rlhf-alternative","model-distillation","knowledge-distillation","ai-agents","coding-agent","agentic","llm-agent","swe-bench","llm-routing","model-cascade","cost-optimization","llm-cost","openrouter","qwen","deepseek","training-data","jsonl","trl","axolotl","unsloth","self-improving","weight-eft","metaharness"],"repository":{"url":"git+https://github.com/ruvnet/agent-harness-generator.git","type":"git","directory":"packages/weight-eft"},"_npmVersion":"10.9.8","description":"Fine-tune cheap open-source LLMs (GLM, Qwen, DeepSeek) on your AI coding agent's successful runs with LoRA (SFT + DPO) so your model cascade escalates to expensive frontier models (GPT, Claude) less often — cutting cost-per-resolved. Turns run history int","directories":{},"maintainers":[{"name":"ruvnet","email":"ruv@ruv.net"}],"_nodeVersion":"22.23.2","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"vitest":"^2.0.0","typescript":"^5.4.0"},"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/weight-eft_0.1.2_1790437871438_0.6305596131896707"}}},"time":{"created":"2026-06-27T16:26:59.525Z","modified":"2026-09-26T15:51:11.789Z","0.1.0":"2026-06-27T16:27:00.004Z","0.1.1":"2026-06-27T16:57:26.706Z","0.1.2":"2026-09-26T15:51:11.529Z"},"bugs":{"url":"https://github.com/ruvnet/agent-harness-generator/issues"},"author":{"name":"rUv","email":"ruv@ruv.net"},"license":"MIT","homepage":"https://github.com/ruvnet/agent-harness-generator/tree/main/packages/weight-eft#readme","keywords":["llm","lora","fine-tuning","peft","sft","dpo","rlhf-alternative","model-distillation","knowledge-distillation","ai-agents","coding-agent","agentic","llm-agent","swe-bench","llm-routing","model-cascade","cost-optimization","llm-cost","openrouter","qwen","deepseek","training-data","jsonl","trl","axolotl","unsloth","self-improving","weight-eft","metaharness"],"repository":{"url":"git+https://github.com/ruvnet/agent-harness-generator.git","type":"git","directory":"packages/weight-eft"},"description":"Fine-tune cheap open-source LLMs (GLM, Qwen, DeepSeek) on your AI coding agent's successful runs with LoRA (SFT + DPO) so your model cascade escalates to expensive frontier models (GPT, Claude) less often — cutting cost-per-resolved. Turns run history int","maintainers":[{"name":"ruvnet","email":"ruv@ruv.net"}],"readme":"# @metaharness/weight-eft\n\n> **Make cheap open-source LLMs solve more coding tasks on their own.** Fine-tune them (LoRA) on your AI agent's *past successful runs*, so your pipeline calls expensive frontier models (GPT, Claude) **less often** — and your cost-per-fix drops.\n\n[![npm version](https://img.shields.io/npm/v/@metaharness/weight-eft.svg)](https://www.npmjs.com/package/@metaharness/weight-eft)\n[![license: MIT](https://img.shields.io/npm/l/@metaharness/weight-eft.svg)](./LICENSE)\n[![node](https://img.shields.io/node/v/@metaharness/weight-eft.svg)](https://nodejs.org)\n\n```bash\nnpm i @metaharness/weight-eft\n```\n\n## What is this? (plain language)\n\nIf you run an **AI coding agent**, you probably use a **model cascade**: a cheap\nmodel (GLM / Qwen / DeepSeek) tries first, and only the hard problems\n**escalate** to an expensive frontier model (GPT / Claude). Every escalation\ncosts real money.\n\n**`weight-eft` makes the cheap model smarter** by fine-tuning it with **LoRA** on\nthe trajectories your agent *already solved* — turning your run history into\ntraining data. The cheap model then resolves more issues by itself, so you\n**escalate less and pay less per solved task.**\n\nIt's a self-improving loop: **your agent's wins become the next model's training set.**\n\n- **Input:** your agent's run archive (successful + failed trajectories).\n- **Output:** portable LoRA training data — **SFT + DPO** in standard formats\n  (OpenAI chat JSONL / TRL / axolotl / unsloth) **+ a GPU training plan**.\n- **Goal:** lower **cost-per-resolved**, not a leaderboard score.\n\n## Why it exists (the honest, bounded thesis)\n\nWe attack the **cost axis, not the capability ceiling.** A small (7-14B) local\nfine-tune **will not** out-reason a frontier model on the hardest problems —\nthat's a model-capability ceiling (measured: clean-eval ~37.3%, ADR-198 / §53).\nThe win is **fewer escalations** (lower cost), and the tooling keeps the\ntelemetry honest about exactly that: the eval metric is\n**escalation-rate-reduction + cost/resolved**, *never* \"we beat the frontier.\"\n\nUnder the hood this is the gradient/weight counterpart to Darwin's gradient-free\npolicy evolution (*freeze the model, evolve the harness*) — here we **also**\nevolve the cheap model's *weights*, on the open tier, from the harness's own\narchive.\n\n## The data recipe (on/off-policy)\n\n| Set | Contents | Why |\n|-----|----------|-----|\n| **SFT** | **ALL** gold-resolved trajectories — cheap-OWN *and* frontier-escalation | SFT (max-likelihood) is off-policy-stable, so a frontier success on an issue the cheap model couldn't solve is **off-policy-safe DISTILLATION**. |\n| **DPO** | **ON-POLICY cheap-vs-cheap pairs ONLY** — `chosen` = a resolved sample, `rejected` = an empty/failed sample by the **same cheap model on the same instance** (BoN-derived) | A frontier-chosen-vs-cheap-rejected pair is **off-policy and unstable** (the reference policy never produced the chosen completion). That signal goes to SFT instead. |\n\n### Output formats (canonical / portable)\n\nExported files use **standard** schemas (portable to TRL / axolotl / unsloth /\nruvllm-MicroLoRA), never a custom format. A thin runner-adapter at the training\nboundary maps standard → whatever the runner ingests.\n\n- **SFT** — OpenAI chat JSONL:\n  `{\"messages\":[{role:system},{role:user},{role:assistant,tool_calls:[…]},{role:tool,…},…,{role:assistant}]}`.\n  **`tool_calls` are preserved** — the ReAct loop is **not** flattened to plain\n  text; the model learns real tool-use trajectories.\n- **DPO** — TRL/HF conversational preference:\n  `{\"prompt\":[system+issue], \"chosen\":[resolved trajectory], \"rejected\":[failed trajectory]}`.\n  ReAct diverges from the first action, so `prompt` is the shared system+issue\n  and chosen/rejected are full trajectories from there.\n\n## The guards\n\n1. **Contamination guard (the headline correctness property).** Strict\n   **train/eval instance-ID disjointness.** The exporter excludes any\n   trajectory whose `instance_id` is in the caller's `evalHoldout`, and\n   `assertTrainEvalDisjoint` **throws** on any overlap. *Training on eval\n   instances is fake lift — the exact contamination we debunk elsewhere.*\n2. **Reward-hacking filter** (Ornith-1.0 borrow). A **deterministic monitor**\n   drops any \"success\" that read a withheld gold/test path, modified the\n   verification harness, or escaped the sandbox. An archived reward-hack would\n   teach the model to reward-hack — this is the **training-data analog of the\n   conformance firewall**, separate from and *in addition to* the disjointness\n   guard.\n3. **Long-context filter.** SWE/ReAct trajectories can blow past a 7-14B\n   context window (~32k). Over-budget trajectories are **dropped (or truncated\n   with `--truncate`) and REPORTED** — never silently lost.\n\nEvery drop is surfaced in the export report (`droppedRewardHacked`,\n`excludedByHoldout`, `droppedOverLength`, `truncatedOverLength`).\n\n## The `weightAdapter` genome gene (prune-the-overfitter safety net)\n\nA LoRA tune can overfit. Rather than trust it blindly, the adapter is a **gene**\nin the Darwin genome (`packages/darwin-mode/bench/swebench/evolve-config.mjs`):\n\n- `weightAdapter: null` = **BASE** (no adapter) — the default and the control.\n  A genome that never opts in is **byte-identical (by key) to a pre-gene\n  genome.**\n- `weightAdapter: 'sft'` = SFT-distilled adapter.\n- `weightAdapter: 'sft-dpo'` = SFT then on-policy DPO.\n\nBase competes against the tuned variants under the **same conformant fitness**,\nso **evolution prunes an adapter that doesn't actually lift held-out resolve.**\nThe gene is inert until an adapter is trained (a GPU job) — it only *names* an\nadapter; it does not create one.\n\n## The training runner (GPU-gated)\n\n`weight-eft train` is **$0 by default** — it emits a training **plan** (config +\nthe exact `ruvllm microlora …` command). A **real** run requires **BOTH** an\nexplicit `--train` flag **AND** a detected GPU / endpoint; otherwise it dry-runs\nor refuses. Target is **7-14B** (Qwen2.5-Coder-7B / GLM-4-9B class) — *not* 32B\n(§59: 32B q4 spills a 16GB GPU). Stages: SFT first, then optional on-policy DPO\nfrom the SFT checkpoint.\n\n## CLI\n\n```bash\n# Status / recipe summary\nweight-eft status\nmetaharness weight-eft status         # via the umbrella CLI\n\n# Export training sets ($0). evalHoldout enforces the contamination guard.\nweight-eft export --archive archive.json --eval-holdout holdout.json --out-dir ./out\n\n# Emit the training plan ($0 dry-run). Add --train on a GPU host to run.\nweight-eft train --base Qwen/Qwen2.5-Coder-7B-Instruct --params-b 7 \\\n  --sft ./out/sft.jsonl --dpo ./out/dpo.jsonl --adapter glm5.2\n\n# Measure the cost-Pareto delta (base vs adapter cascade runs).\nweight-eft eval --base-outcomes base.json --adapter-outcomes adapter.json\n```\n\n### The exact (later, GPU) command to train + eval\n\n```bash\n# 1) Export ($0) — disjoint train/eval, reward-hack-filtered, long-context-filtered\nweight-eft export --archive darwin-archive.json --eval-holdout clean-eval-ids.json --out-dir ./eft\n\n# 2) Train (GPU host) — SFT then on-policy DPO. ruvllm/MicroLoRA executes plan.command.\nweight-eft train --base Qwen/Qwen2.5-Coder-7B-Instruct --params-b 7 \\\n  --sft ./eft/sft.jsonl --dpo ./eft/dpo.jsonl --adapter glm5.2 --train\n# (refuses unless WEIGHT_EFT_BASE_URL / CUDA_VISIBLE_DEVICES is set)\n\n# 3) Run the conformant cascade twice (base vs glm5.2-sft-dpo adapter) on the\n#    HELD-OUT clean set via the existing darwin eval path, collect per-instance\n#    CascadeOutcome[] for each, then:\nweight-eft eval --base-outcomes base-outcomes.json --adapter-outcomes adapter-outcomes.json\n```\n\n## Input contract\n\nThe exporter codes against `DarwinTrajectory[]` (see `src/types.ts`) —\nreconstructable from Firestore `darwin_runs` + the local prediction/trajectory\nartifacts (`predictions-*.jsonl` rows carry `instance_id` + `model_patch`; the\nagentic loop carries the `messages` array with `tool_calls`, see\n`darwin-mode/bench/swebench/solve-agentic.mjs`). A tiny mock fixture archive\nlives in `__tests__/fixtures/`.\n\n## Status (honest)\n\n- **Runnable, $0:** exporter (with all three guards), training-plan emission,\n  cost-Pareto eval folding, the `weightAdapter` gene (wired into darwin's\n  evolve-config genome + the umbrella `metaharness weight-eft` CLI).\n- **Scaffolded, GPU-gated:** the actual LoRA training (`spawn(plan.command)` on\n  a GPU host implementing the ruvllm/MicroLoRA seam). No training run, no GPU\n  job, no paid model call has been executed.\n\nSee **ADR-198** for the full rationale, the SFT-distill / on-policy-DPO recipe,\nthe disjointness invariant, and the self-scaffolding RL roadmap (Ornith-1.0).\n\n## License\n\nMIT\n","readmeFilename":"README.md"}