{"_id":"@baublet/piqwy","_rev":"6-6aea11fd7aeda8fb32463bd89681ece8","name":"@baublet/piqwy","dist-tags":{"latest":"0.6.0"},"versions":{"0.1.0":{"name":"@baublet/piqwy","version":"0.1.0","keywords":["qwythos","llm","coding-agent","runpod","llama.cpp","pi.dev","openai-proxy","on-demand-gpu"],"author":{"name":"baublet"},"license":"MIT","_id":"@baublet/piqwy@0.1.0","maintainers":[{"name":"baublet","email":"baublet@gmail.com"}],"homepage":"https://github.com/baublet/ryanmpoe#readme","bugs":{"url":"https://github.com/baublet/ryanmpoe/issues"},"bin":{"piqwy":"piqwy.mjs"},"dist":{"shasum":"19a57034ade8ef6674ca5cfc8fb0b01626c04835","tarball":"https://registry.npmjs.org/@baublet/piqwy/-/piqwy-0.1.0.tgz","fileCount":8,"integrity":"sha512-mPPV9vjLDFNE8phuCWK0K/0x482nMFBP6oKy4n6NImYcRZU30nnsbN62ChK98x1J1WZ8RtKZRMbIx9eoFSvInA==","signatures":[{"sig":"MEUCIQCg6cdNqPyLxRoHBQK6tUx/ANR/DUn/gnhyolc0HqFhvQIgN2nqFPp5Ttfzp0PIKo9EfPhWv49gKAAmTX0zHIIzVjQ=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":40111},"type":"module","engines":{"node":">=18"},"exports":"./piqwy.mjs","gitHead":"bbe144e3ea4eeaa609e84c9d031d080b6a0aca7f","scripts":{"build":"node build.mjs","start":"node piqwy.mjs","prepublishOnly":"node build.mjs"},"_npmUser":{"name":"baublet","email":"baublet@gmail.com"},"repository":{"url":"git+https://github.com/baublet/ryanmpoe.git","type":"git","directory":"piqwy"},"_npmVersion":"10.9.4","description":"At-will Qwythos coding agent: spins up Qwythos-9B on a cheap on-demand RunPod GPU behind a local OpenAI-compatible proxy, wires the pi.dev harness to it, and tears the pod down when idle.","directories":{},"_nodeVersion":"24.11.1","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/piqwy_0.1.0_1782912486059_0.10651285299502589","host":"s3://npm-registry-packages-npm-production"}},"0.2.0":{"name":"@baublet/piqwy","version":"0.2.0","keywords":["qwythos","llm","coding-agent","runpod","llama.cpp","pi.dev","openai-proxy","on-demand-gpu"],"author":{"name":"baublet"},"license":"MIT","_id":"@baublet/piqwy@0.2.0","maintainers":[{"name":"baublet","email":"baublet@gmail.com"}],"homepage":"https://github.com/baublet/ryanmpoe#readme","bugs":{"url":"https://github.com/baublet/ryanmpoe/issues"},"bin":{"piqwy":"piqwy.mjs"},"dist":{"shasum":"4a520f8e333e45c87392f8acbf73a263d1d86dd6","tarball":"https://registry.npmjs.org/@baublet/piqwy/-/piqwy-0.2.0.tgz","fileCount":8,"integrity":"sha512-tLOGKT3ien9L9W1/UZd3ks7WgNTF1YcPvctr+I86DNat9N3Vl3FQBHsE6z40lYGfq7XWO669KJplrnA4wcmcig==","signatures":[{"sig":"MEUCIAUI1dO+rSwPVygGTBfbxNTIy64a2FoFQolbjbrAk6yPAiEAhkAEN2LU9a1PjgWDcNsU0sl0IA94sdhp5AATEf1iicM=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":50988},"type":"module","engines":{"node":">=18"},"exports":"./piqwy.mjs","gitHead":"acf3d717875a943e58d27b7c687438bb2c47204b","scripts":{"build":"node build.mjs","start":"node piqwy.mjs","prepublishOnly":"node build.mjs"},"_npmUser":{"name":"baublet","email":"baublet@gmail.com"},"repository":{"url":"git+https://github.com/baublet/ryanmpoe.git","type":"git","directory":"piqwy"},"_npmVersion":"10.9.4","description":"At-will Qwythos coding agent: spins up Qwythos-9B on a cheap on-demand RunPod GPU behind a local OpenAI-compatible proxy, wires the pi.dev harness to it, and tears the pod down when idle.","directories":{},"_nodeVersion":"24.11.1","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/piqwy_0.2.0_1782936797476_0.3478072790689737","host":"s3://npm-registry-packages-npm-production"}},"0.3.0":{"name":"@baublet/piqwy","version":"0.3.0","keywords":["qwythos","llm","coding-agent","runpod","llama.cpp","pi.dev","openai-proxy","on-demand-gpu"],"author":{"name":"baublet"},"license":"MIT","_id":"@baublet/piqwy@0.3.0","maintainers":[{"name":"baublet","email":"baublet@gmail.com"}],"homepage":"https://github.com/baublet/ryanmpoe#readme","bugs":{"url":"https://github.com/baublet/ryanmpoe/issues"},"bin":{"piqwy":"piqwy.mjs"},"dist":{"shasum":"42bf48f4149e5fb629858c17e87c7530e44cfe16","tarball":"https://registry.npmjs.org/@baublet/piqwy/-/piqwy-0.3.0.tgz","fileCount":8,"integrity":"sha512-dK+kOjOt8QbxVOorDxA8CI/zxMVYGNmQSvLAWWp2M0O5LnttIx2oc7bFuwsqOnGwWFcfyByY53vngNHx+wtlKQ==","signatures":[{"sig":"MEQCIC6pwL5LfRF7B0EWGuIpEPu+tTYlp6Ab/nfcQBW4XhaUAiBNLO2+OwHUzbPyIqpycQgFDIlsEDpnGXQcAr9qMWUSjg==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":61780},"type":"module","engines":{"node":">=18"},"exports":"./piqwy.mjs","gitHead":"a60bc6e3e1c9c65e84f0291da0213f18674d2cc3","scripts":{"build":"node build.mjs","start":"node piqwy.mjs","prepublishOnly":"node build.mjs"},"_npmUser":{"name":"baublet","email":"baublet@gmail.com"},"repository":{"url":"git+https://github.com/baublet/ryanmpoe.git","type":"git","directory":"piqwy"},"_npmVersion":"10.9.4","description":"At-will Qwythos coding agent: spins up Qwythos-9B on a cheap on-demand RunPod GPU behind a local OpenAI-compatible proxy, wires the pi.dev harness to it, and tears the pod down when idle.","directories":{},"_nodeVersion":"24.11.1","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/piqwy_0.3.0_1782952507920_0.14869521960498444","host":"s3://npm-registry-packages-npm-production"}},"0.4.0":{"name":"@baublet/piqwy","version":"0.4.0","keywords":["qwythos","llm","coding-agent","runpod","llama.cpp","pi.dev","openai-proxy","on-demand-gpu"],"author":{"name":"baublet"},"license":"MIT","_id":"@baublet/piqwy@0.4.0","maintainers":[{"name":"baublet","email":"baublet@gmail.com"}],"homepage":"https://github.com/baublet/ryanmpoe#readme","bugs":{"url":"https://github.com/baublet/ryanmpoe/issues"},"bin":{"piqwy":"piqwy.mjs"},"dist":{"shasum":"bb0930c3a973e02d4c136b01b8782af219f60b09","tarball":"https://registry.npmjs.org/@baublet/piqwy/-/piqwy-0.4.0.tgz","fileCount":8,"integrity":"sha512-lBdbheJuESus7QVTLPzOAMZA3n0W5t6vzcjOh+og/CPavx9Y6FFnbHYUNdsA5wgJUMHwwvIeLybpHIYJVEnFvA==","signatures":[{"sig":"MEUCIAu9Fi9NdNcauz4pNYV0oAfORNBKvM22CHqDsfHt7RSCAiEApHos4XzkYhe0RcFuKa9Mum6CNHWcVW7RIG2h190kAp0=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":69326},"type":"module","engines":{"node":">=18"},"exports":"./piqwy.mjs","gitHead":"219578a27100091f1de010bcc5eecd16709370f6","scripts":{"build":"node build.mjs","check":"biome check","start":"node piqwy.mjs","format":"biome check --write","prepublishOnly":"node build.mjs"},"_npmUser":{"name":"baublet","email":"baublet@gmail.com"},"repository":{"url":"git+https://github.com/baublet/ryanmpoe.git","type":"git","directory":"piqwy"},"_npmVersion":"10.9.4","description":"At-will Qwythos coding agent: spins up Qwythos-9B on a cheap on-demand RunPod GPU behind a local OpenAI-compatible proxy, wires the pi.dev harness to it, and tears the pod down when idle.","directories":{},"_nodeVersion":"24.11.1","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"@biomejs/biome":"^2.5.2"},"_npmOperationalInternal":{"tmp":"tmp/piqwy_0.4.0_1783011233333_0.08988635147438284","host":"s3://npm-registry-packages-npm-production"}},"0.5.0":{"name":"@baublet/piqwy","version":"0.5.0","keywords":["qwythos","llm","coding-agent","runpod","llama.cpp","pi.dev","openai-proxy","on-demand-gpu"],"author":{"name":"baublet"},"license":"MIT","_id":"@baublet/piqwy@0.5.0","maintainers":[{"name":"baublet","email":"baublet@gmail.com"}],"homepage":"https://github.com/baublet/ryanmpoe#readme","bugs":{"url":"https://github.com/baublet/ryanmpoe/issues"},"bin":{"piqwy":"piqwy.mjs"},"dist":{"shasum":"e056b8c98b0969b40711380230f7eb78b5a5026a","tarball":"https://registry.npmjs.org/@baublet/piqwy/-/piqwy-0.5.0.tgz","fileCount":9,"integrity":"sha512-rpG38CUwXYFfm98V5Uag7T5XuT37+6TU1lCKW7FjVjPRSOcu75usKerxu+G2rKVJqKuXig0gaXKgawEYu7h2tw==","signatures":[{"sig":"MEQCIDFHt2ryqH6WlRJpTQwF/fzEG52UVeQGcMOTflp6JIxFAiB2gUKoaKo5I1/dprqmHksBjDcLL2/pI+mEpcGlmwgi+g==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":122491},"type":"module","engines":{"node":">=18"},"exports":"./piqwy.mjs","gitHead":"063d6e13b5d0d02252a0e4dd7c8bb0f899b698fc","scripts":{"build":"node build.mjs","check":"biome check","smoke":"node scripts/smoke.mjs","start":"node piqwy.mjs","format":"biome check --write","prepublishOnly":"node build.mjs"},"_npmUser":{"name":"baublet","email":"baublet@gmail.com"},"repository":{"url":"git+https://github.com/baublet/ryanmpoe.git","type":"git","directory":"piqwy"},"_npmVersion":"10.9.4","description":"At-will open-weights coding agent: spins up a selectable model (Qwythos default) on a cheap on-demand RunPod GPU behind a local OpenAI-compatible proxy, wires the pi.dev / Claude Code harness to it, and tears the pod down when idle.","directories":{},"_nodeVersion":"24.11.1","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"@biomejs/biome":"^2.5.2"},"_npmOperationalInternal":{"tmp":"tmp/piqwy_0.5.0_1783133329054_0.8113305814022196","host":"s3://npm-registry-packages-npm-production"}},"0.6.0":{"name":"@baublet/piqwy","version":"0.6.0","description":"At-will open-weights coding agent: spins up a selectable model (Qwythos default) on a cheap on-demand RunPod GPU behind a local OpenAI-compatible proxy, wires the pi.dev / Claude Code harness to it, and tears the pod down when idle.","type":"module","bin":{"piqwy":"piqwy.mjs"},"exports":"./piqwy.mjs","engines":{"node":">=18"},"keywords":["qwythos","llm","coding-agent","runpod","llama.cpp","pi.dev","openai-proxy","on-demand-gpu"],"license":"MIT","author":{"name":"baublet"},"publishConfig":{"access":"public"},"repository":{"type":"git","url":"git+https://github.com/baublet/ryanmpoe.git","directory":"piqwy"},"scripts":{"start":"node piqwy.mjs","build":"node build.mjs","smoke":"node scripts/smoke.mjs","check":"biome check","format":"biome check --write","prepublishOnly":"node build.mjs"},"devDependencies":{"@biomejs/biome":"^2.5.2"},"_id":"@baublet/piqwy@0.6.0","gitHead":"e4257917bc814014ce3293e51d67e71430e300fd","bugs":{"url":"https://github.com/baublet/ryanmpoe/issues"},"homepage":"https://github.com/baublet/ryanmpoe#readme","_nodeVersion":"24.11.1","_npmVersion":"10.9.4","dist":{"integrity":"sha512-iuABO/Q87gML7WxnJCls+oLYgQJHcu0SCrvKG0F+K+0Yhpb0sz+O6zW57ISxHg/k/MKOhuVylXYuIkDc/9K/6A==","shasum":"29b8b7d134fe9fa4fdd7c8a4e99edf66975360f7","tarball":"https://registry.npmjs.org/@baublet/piqwy/-/piqwy-0.6.0.tgz","fileCount":9,"unpackedSize":123308,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQDMEswKQzsAMAE/nVVoEc1el6j2QfyyT03FELOPnTK2qwIgBW8gZaG375pWZZYBcdjGxQUy18LAUMKncDnfWhBKS/A="}]},"_npmUser":{"name":"baublet","email":"baublet@gmail.com"},"directories":{},"maintainers":[{"name":"baublet","email":"baublet@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/piqwy_0.6.0_1783176282237_0.06095933694387723"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-01T13:28:05.945Z","modified":"2026-07-04T14:44:42.456Z","0.1.0":"2026-07-01T13:28:06.206Z","0.2.0":"2026-07-01T20:13:17.629Z","0.3.0":"2026-07-02T00:35:08.064Z","0.4.0":"2026-07-02T16:53:53.486Z","0.5.0":"2026-07-04T02:48:49.181Z","0.6.0":"2026-07-04T14:44:42.359Z"},"bugs":{"url":"https://github.com/baublet/ryanmpoe/issues"},"author":{"name":"baublet"},"license":"MIT","homepage":"https://github.com/baublet/ryanmpoe#readme","keywords":["qwythos","llm","coding-agent","runpod","llama.cpp","pi.dev","openai-proxy","on-demand-gpu"],"repository":{"type":"git","url":"git+https://github.com/baublet/ryanmpoe.git","directory":"piqwy"},"description":"At-will open-weights coding agent: spins up a selectable model (Qwythos default) on a cheap on-demand RunPod GPU behind a local OpenAI-compatible proxy, wires the pi.dev / Claude Code harness to it, and tears the pod down when idle.","maintainers":[{"name":"baublet","email":"baublet@gmail.com"}],"readme":"# piqwy\n\nAt-will **Qwythos** coding agent. Run `piqwy` in any directory and you get the [pi.dev](https://pi.dev)\nagent wired to **Qwythos-9B** running on a cheap on-demand **RunPod** GPU that spins up when you need it\nand **86's itself when idle**. One file, zero dependencies, Node ≥18.\n\n```bash\npiqwy                 # ensure the model is serving, then launch the pi.dev agent in this dir\npiqwy -p \"add a test for utils.ts\"   # headless one-shot\npiqwy claude          # launch Claude Code on the model instead of pi.dev\npiqwy models          # list the model catalog (active one marked)\npiqwy model devstral  # switch models (persists); optional tier: piqwy model devstral smart\npiqwy status          # daemon + live pods + stats (reuses / creates / est cost)\npiqwy down            # terminate pods this daemon created\npiqwy down --all      # terminate every piqwy pod on the account\npiqwy stop            # stop the local daemon\n```\n\nQwythos is the default; `piqwy models` lists the rest and `piqwy model <name>` (or `-m <name>` on any\nrun) switches. See [Models](#models).\n\n### Flags\n\n- `--json` — machine-readable JSON instead of the formatted output (`status`, `down`). Everything else\n  prints for humans; pass this when you're scripting.\n- `--all` — with `down`, also stop piqwy pods this daemon didn't create. Default `down` only touches\n  pods this daemon started; if it sees other piqwy pods it leaves them and tells you.\n- `-- <args...>` — forward everything after `--` verbatim to the underlying harness (pi.dev or Claude\n  Code). Use it for harness flags piqwy doesn't know about:\n\n  ```bash\n  piqwy -- --add-dir /some/repo                     # pass --add-dir to pi.dev\n  piqwy claude -- --dangerously-skip-permissions    # pass a flag to Claude Code\n  ```\n\n  For `claude`, flags typed right after `claude` also pass through, so the `--` is optional there.\n\n## Two harnesses\n\nThe daemon exposes **both** APIs on the same local port, so you can drive Qwythos with either agent:\n\n- **pi.dev** (`piqwy`) — OpenAI-native; the daemon proxies `/v1/chat/completions` straight through.\n- **Claude Code** (`piqwy claude`) — Claude Code speaks the Anthropic Messages API, so the daemon\n  translates `/v1/messages` ⇄ OpenAI on the fly (deterministic mapping, tool-use IDs preserved, real\n  SSE streaming). `piqwy claude` sets `ANTHROPIC_BASE_URL` at the daemon and launches `claude` here.\n  Heads up: Qwythos-9B is a small model — Claude Code leans hard on strong tool-use, so expect a\n  rougher ride than a frontier Claude model.\n\n## Models\n\nOne model serves at a time. Each catalog entry is a self-contained recipe — GGUF repo, quant ladder,\nnative context, model-card sampling — pinned to the GPU order it runs on best, so switching a model\nre-points the whole pipeline (and cycles the pod, since the pod identity includes the repo/quant/ctx).\n\n| name | what it is | best GPU | ctx |\n|---|---|---|---|\n| `qwythos` *(default)* | 9B hybrid-attention reasoner, cheapest | RTX 4090 (24GB) | 256k |\n| `ornith-9b` | Ornith-1.0-9B, agentic-coding specialist (SWE-bench 69%) | RTX 4090 (24GB) | 128k |\n| `gemma-4-12b` | Gemma 4 12B dense (Google), QAT + MTP 3× decode | RTX 4090 (24GB) | 256k |\n| `qwen3-coder` | Qwen3-Coder-30B-A3B MoE, fast agentic coder | L40S (48GB) | 32k |\n| `devstral` | Devstral-Small-2507, 24B dense SWE/tool-use specialist | L40S (48GB) | 128k |\n| `gemma-4-26b` | Gemma 4 26B-A4B MoE (Google), fast agentic decode | RTX 5090 (32GB) | 256k |\n| `qwen36-27b` | Qwen3.6-27B dense, flagship coding + MTP 2× decode | L40S (48GB) | 128k |\n| `qwen36-35b` | Qwen3.6-35B-A3B MoE, fast agentic coder (~140 t/s) | L40S / 5090 | 64k |\n| `ornith-35b` | Ornith-1.0-35B MoE, **top open agentic coder** (SWE-bench 76%) | L40S (48GB) | 128k |\n\nEach runs on a single GPU sized to the model (24GB for the 9-12B, 48GB for the 24-35B), and every one\nis measured green — see **MODELS.md** for the red/green snake test and serving numbers\n(decode/prefill/cached) on real RunPod hardware. Contexts use a `q8_0` KV cache (half the memory of\n`f16`) so most models reach 128k-256k on a single card; Gemma's sliding-window attention fits **256k on\na 32GB 5090**. The exceptions are `qwen3-coder` and `devstral`, whose architectures crash on quantized\nKV and so need `f16` (capping qwen3-coder at 32k).\n\n**Best for agentic coding:** the two `ornith-*` models are purpose-built for it (tool-calling +\n`<think>` reasoning). `ornith-35b` is the strongest open agentic coder here (SWE-bench Verified ~76%,\n~160 t/s on an L40S); `ornith-9b` is the cheapest good coder (fits a 24GB 4090). Several models use MTP\n(multi-token-prediction) speculative decoding for a 2-3× decode speedup with no quality loss:\n`gemma-4-12b` (QAT+MTP) and `qwen36-27b`.\n\n```bash\npiqwy models                 # list them, active one marked, with GPU order + context\npiqwy model qwen3-coder      # select (persists to ~/.piqwy/config.json)\npiqwy model qwen3-coder smart # ...at the smart tier (Q8) instead of the default fast (Q4)\npiqwy -m devstral            # use devstral for this run (also persists the choice)\n```\n\nTiers trade quality for cost/speed: `fast` (~Q4), `mid` (~Q6), `smart` (~Q8). `ctx` is set to fit the\nfirst GPU in each model's preference list; bump it per-model in `~/.piqwy/config.json` for a bigger card.\nFull-attention models (devstral, qwen25-coder, qwen3-coder) serve with an `f16` KV cache — quantized KV +\nflash-attention crashes them on the first request. If the preferred GPU is dry or too small, piqwy falls\nthrough the rest of the list (community → secure) and tears down anything that places but can't serve. Add\nyour own model by dropping an entry into `models` in the config — same shape as the built-ins.\n\n## How it works\n\n```\npiqwy (CLI) ──http──> piqwyd (daemon, local OpenAI PROXY) ──REST──> RunPod ──> model pod (llama.cpp)\n   starts daemon if down        routes every request to a pod                  self-destructs when idle\n   installs + wires pi.dev      (reuse live / create + warm if down)           proxy URL = OpenAI endpoint\n   launches pi in cwd           tracks pods, records stats, prunes idle-killed\n```\n\nThe daemon is the **data plane**: everything goes through `http://127.0.0.1:8787/v1`, and it routes each\nrequest to a Qwythos pod (strategy `live` today — the one running pod; pluggable later). When a pod\nself-destructs on idle, the next request transparently spins up a fresh one — clients never re-wire.\nMissing key / no capacity come back as OpenAI-shaped errors with a fix hint.\n\n`piqwy status` shows an **estimated** cost, accrued by the daemon while it's polling, at each pod's\nactual RunPod rate (`adjustedCostPerHr`, after Savings Plans). It only counts pods this daemon tracked,\nand because it accrues on a poll interval it can be off by a few billing minutes — treat it as a ballpark,\nnot an invoice.\n\n## Setup\n\n```bash\nexport RUNPOD_API_KEY=...        # or put RUNPOD_API_KEY=... in ~/.piqwy/.env\n```\n\nGet a key at https://www.runpod.io/console/user/settings. No GPU quota forms. First run downloads the\nmodel (~1-3 min); later runs in the idle window reuse the warm pod instantly.\n\n### Run it\n\n```bash\nnpx @baublet/piqwy               # once published (scoped, public)\n# or, from a checkout:\nnode piqwy.mjs                   # single file, run anywhere\n```\n\n## Config\n\nBuilt-in defaults cover everything; override any of them in `~/.piqwy/config.json` (deep-merged):\n\n- `model_key` — the active model (default `qwythos`); `models` — the catalog (see [Models](#models)).\n  Each entry carries its own repo, quant ladder, `ctx`, `native_ctx`, `gpu_prefs`, and `sampling`.\n- `runpod` — image, Community/Secure, disk, pod label, and the fallback `gpu_prefs`/`gpu_types` map.\n- `serve` — `ctx`, `np` (continuous-batch slots), `kv` cache quant, `idle_seconds` (per-model `ctx`\n  and `sampling` override these).\n- `proxy` — host/port, routing `strategy`.\n- `harness` — pi.dev provider name + install URL.\n\nSee **MODELS.md** for each model's red/green smoke test and measured serving performance on RunPod\n(regenerate it with `npm run smoke -- --md`).\n\n## Pod scripts & build\n\nThe pod-side shell is the source of truth in `scripts/` (`onstart.sh` serves llama.cpp; `idle_watchdog.sh`\nself-destructs the pod on idle). They're base64-embedded into `piqwy.mjs` so the tool stays a single\nself-contained file. After editing either, regenerate the embed:\n\n```bash\nnpm run build      # node build.mjs — re-injects scripts/*.sh into piqwy.mjs (also runs on publish)\n```\n\n## Code style\n\n[Biome](https://biomejs.dev) with stock defaults — no custom rules, nothing to argue about.\n\n```bash\nnpm run format     # biome check --write — format + safe lint fixes\nnpm run check      # biome check — verify only (for CI)\n```\n\n`build.mjs` writes the base64 embeds in Biome's canonical shape, so the formatter and the build step\ndon't fight over those lines.\n\n## Publishing\n\nNot published yet. To publish (review first):\n\n```bash\nnpm pack            # inspect the tarball contents\nnpm publish         # ship it (bump version in package.json first)\n```\n","readmeFilename":"README.md"}