{"_id":"@avee1234/flywheel","name":"@avee1234/flywheel","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@avee1234/flywheel","version":"0.1.0","description":"A zero-dependency CLI for turning agent traces into measured, guarded improvements.","type":"module","bin":{"flywheel":"bin/flywheel.js"},"engines":{"node":">=18"},"scripts":{"test":"node --test"},"repository":{"type":"git","url":"git+https://github.com/abhid1234/flywheel.git"},"keywords":["ai","agent","cli","self-improvement","traces","evaluation","zero-dependency"],"license":"MIT","_id":"@avee1234/flywheel@0.1.0","gitHead":"0ea7d7d786b45b2c1603c8a6b8c4b738159e16fc","bugs":{"url":"https://github.com/abhid1234/flywheel/issues"},"homepage":"https://github.com/abhid1234/flywheel#readme","_nodeVersion":"22.22.3","_npmVersion":"10.9.8","dist":{"integrity":"sha512-O4RqqObJzwbpIeePwOTUhrTYReJ+qReG3oG6xCIX5kB57jf7Cj3nRvPU2vLktSv85pWBOFTfrDZQbj5I5n9+Pg==","shasum":"4e3b2efaa1961e3878ba6e9756dcf69b05a59e33","tarball":"https://registry.npmjs.org/@avee1234/flywheel/-/flywheel-0.1.0.tgz","fileCount":44,"unpackedSize":206223,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQDqrpA7Bh041J6w/4edAmTYNLUwReVyVIHmnJkbEAqJSAIgZDam1H2sqW/P1FHs4SAQDO/fAcahxALMjxVurmVoseI="}]},"_npmUser":{"name":"avee1234","email":"das.abhijit34@gmail.com"},"directories":{},"maintainers":[{"name":"avee1234","email":"das.abhijit34@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/flywheel_0.1.0_1784860590810_0.8615936185076942"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-24T02:36:30.582Z","0.1.0":"2026-07-24T02:36:30.993Z","modified":"2026-07-24T02:36:31.208Z"},"maintainers":[{"name":"avee1234","email":"das.abhijit34@gmail.com"}],"description":"A zero-dependency CLI for turning agent traces into measured, guarded improvements.","homepage":"https://github.com/abhid1234/flywheel#readme","keywords":["ai","agent","cli","self-improvement","traces","evaluation","zero-dependency"],"repository":{"type":"git","url":"git+https://github.com/abhid1234/flywheel.git"},"bugs":{"url":"https://github.com/abhid1234/flywheel/issues"},"license":"MIT","readme":"# flywheel\n\n**Agents that improve from their own production traces.** Harvest what an agent actually did, label which turns failed, cluster the recurring failures, propose a fix, gate it, and — the part everyone skips — *measure whether the fix actually helped* by replaying the exact command that failed in production.\n\nZero dependencies. Pure core, thin CLI seam. Built to be honest about what it can and can't prove.\n\n```\nharvest → label → cluster → propose → gate → measure → attest → report\n```\n\n> **▶ [Interactive playground](playground/index.html)** — step the real pipeline through a live failure, and paste your own errors into the Signature Lab (the actual `classifyError` / `normalizeErrorText` source, running in your browser). Open the file locally, or serve `playground/` via GitHub Pages.\n\n---\n\n## The one rule everything is built around\n\n> **The LLM writes the fix. It never writes the success criterion.**\n\nIf a model can author its own passing test, a self-improvement loop optimizes for the model's imagination instead of reality. So the eval contract is derived *deterministically from the observed production failure* — a **witness**: the exact command, in the exact directory, that failed. A fix is only \"verified\" when replaying that witness goes **red → green**. Three guards enforce this in code, not prompt:\n\n1. The success criterion comes from a recorded failure — the model can't move the goalposts.\n2. A proposed edit's anchor must appear **verbatim, exactly once** in the target file, or it's rejected before gating.\n3. The improvement claim is falsifiable by replaying the command that failed.\n\n## What it composes\n\nflywheel doesn't reinvent the safety machinery — it wires together three existing zero-dep packages:\n\n| stage | package | role |\n|---|---|---|\n| gate | [`selfpatch`](https://github.com/abhid1234/selfpatch) | every self-edit is gated (blocks secrets/`settings.json`), verified before/after, and revertible |\n| oracle | [`truecall`](https://github.com/abhid1234/truecall) | deterministic post-conditions (τ²-bench: ~0 false positives) — licenses trusting a red→green at n=1 |\n| ledger | [`provenant`](https://github.com/abhid1234/provenant) | each applied fix is attested, recording what it derived from |\n\nIntegrating them took a four-fix adapter (artifact shape, timestamp seam, parents-are-attestation-ids, checks-are-objects) — concrete evidence these packages were built in isolation and had never actually been composed. A self-verifying contract test imports the real `provenant` validator so the integration can't silently rot.\n\n## Install & run\n\n```bash\ngit clone https://github.com/abhid1234/flywheel && cd flywheel\nnode --test                                   # 216 tests, zero deps\n\n# harvest your own Claude Code transcripts into episodes, then work the pipeline\nnode bin/flywheel.js harvest  ~/.claude/projects --out ~/.flywheel\nnode bin/flywheel.js label    --in ~/.flywheel/episodes --out ~/.flywheel/episodes\nnode bin/flywheel.js clusters --in ~/.flywheel/episodes --top 15\nnode bin/flywheel.js longtail --in ~/.flywheel/episodes          # the rare/singleton failures\nnode bin/flywheel.js report   --in ~/.flywheel/episodes --out atlas.html   # the failure atlas\nnode bin/flywheel.js trend                                        # corpus over time (compounding)\n\n# propose fixes — auto-eligible (witness-verifiable) vs human-review (behavioural)\nnode bin/flywheel.js loop --mode review --in ~/.flywheel/episodes --llm echo --dry-run\nnode bin/flywheel.js trial-run --cluster <sig> --in ~/.flywheel/episodes --agent fake   # statistical arm (free/deterministic)\n\n# mint gold labels from GitHub merge-status; run --help for the full surface\nnode bin/flywheel.js gold --in ~/.flywheel/episodes --repos owner/repo\nnode bin/flywheel.js --help\n```\n\n## The pieces\n\n- **harvest** — segments transcripts into episodes by `promptId` (98.4% coverage on real data), joins tool calls to results, parses exit codes and Claude Code tool-level errors, and applies the *recovery rule*: a failure the agent later fixes in the same episode doesn't count. Filters benign non-zero exits (a `grep` with no match is not a failure) and non-agent-faults (a user declining a tool is not a defect the agent could fix).\n- **label** — assigns trust tiers: `gold` (adjudicated / structured outcome), `strong` (deterministic post-condition or unrecovered terminal error), `weak` (proxy signals), `unknown`. `tier` is a required, non-defaulting field — a label without it fails validation, so weak signals can never silently become training signal.\n- **cluster** — groups failures by exact signature, then merges near-identical ones (token-set Jaccard, same tool + error class). No embeddings — machine-generated error text is already near-canonical. `isProposable` gates: size ≥ 3, ≥ 3 gold/strong, ≥ 1 replayable witness.\n- **propose** — a deterministic brief → the LLM writes a fix → strict parse → a witness-derived eval contract, targeted at the real `CLAUDE.md`/config where the failures happened (creating it human-gated if absent). `weights` layer throws (unbuildable, honestly unconstructible rather than fake-pluggable).\n- **measure** — the deterministic arm (witness replay, ~0 variance) ships first. `trial-run` executes the statistical arm against an injectable agent (`fake` = free/deterministic; `codex`/`claude` = real, spawned under hard cost bounds) with a sealed held-out split and a hard n ≥ 60 floor — it reports `inconclusive`/`powered:no` below that rather than fabricating significance.\n- **loop** — `--mode auto` applies *only* gated S1 fixes that pass witness replay; `--mode review` proposes, gates, and **queues** behavioural fixes (the ones that actually cluster but can't be witness-verified) for a human, never auto-applying. Append-only hash-linked `ledger.jsonl`.\n- **gold / trend / longtail** — `gold` mints ground-truth labels from GitHub merge-status; `trend` reads the per-tick `history.jsonl` into a compounding time series that can't be built retroactively; `longtail` groups the singleton failures too rare to cluster into a read-only human-triage view.\n\n## Honest status\n\nThis is a research build, and its findings include what *doesn't* work yet:\n\n- ✅ **M0 — harvest** works on a real 440 MB / 610-file corpus in ~2 s → 837 episodes.\n- ✅ **M1 — the labeler is validated.** Run against the author's real agent-factory transcripts and cross-checked with GitHub ground truth, it recovered **3 of 4** documented failures. The one it missed was a review-stage *semantic* defect (a test that passes without testing the thing) — invisible from exit codes, and it correctly did not fabricate a signal it couldn't see.\n- ✅ **M2 — the closed loop is proven,** both ways: a proposed context note that *couldn't* fix a missing dependency was reported `helped:false` (the gate refused to certify it), and a recorded production witness, replayed unchanged after the real cause was fixed, went red → green.\n- ⏳ **M3 — the statistical arm was blocked on gold-label volume** the original corpus didn't have (the pre-registered [KC-6 finding](#): *the mechanism works; that it moves the outcome is unproven*). The [`integrations/daytona`](integrations/daytona) extension attacks exactly this: it generates gold-labeled trials *by construction* by running controlled tasks in isolated cloud sandboxes, then runs the statistical arm as **RL on agent trajectories** — a live agent improving from its own graded failures. Result ([writeup](integrations/daytona/learn/FINDINGS.md)): the climb **reproduces** across independent live runs (n=3, every run +27–48pp on a sealed held-out set), while individual per-lesson attribution stays underpowered at that scale — reported honestly, not rounded up.\n\n## Extensions\n\n- **[`integrations/daytona`](integrations/daytona)** — the statistical arm made runnable: a controlled-task benchmark (6 business scenarios, A/A-gated on real sandboxes) and a full **RL-on-trajectories loop** (rollout → verifiable reward → distilled lesson → repeat), proven live. Isolated dependency; the zero-dep core is untouched.\n\n## Design invariants\n\n- **Zero dependencies.** Nothing under `src/` imports a `node:` module except `src/hash.js`; `bin/flywheel.js` is the only I/O seam. The entire core is pure and testable offline.\n- **No clock, no randomness in `src/`.** Content-hash ids make every stage idempotent.\n- **Deterministic bootstrap** (seeded), so statistical results are reproducible.\n\nBuilt primarily by Codex (`gpt-5.6-sol`), specified and independently verified chunk by chunk. MIT.\n","readmeFilename":"README.md","_rev":"1-59085901442a45244a9680363e289217"}