{"_id":"@3thfloor/engine","_rev":"3-8079a1c6a6425ae4f996f2cc8fe72489","name":"@3thfloor/engine","dist-tags":{"latest":"0.1.2"},"versions":{"0.1.0":{"name":"@3thfloor/engine","version":"0.1.0","keywords":["llm","ai","inference","local","gguf","llama","embedding"],"author":{"name":"Justin Bench","email":"justin.bench@gmail.com"},"license":"PolyForm-Noncommercial-1.0.0","_id":"@3thfloor/engine@0.1.0","maintainers":[{"name":"ziptslug","email":"justin@3thfloor.com"}],"homepage":"https://3thfloor.com","bugs":{"url":"https://github.com/3thfloor/engine/issues"},"dist":{"shasum":"c0c32248b573580a6116f0954c58ecb5c09bee0c","tarball":"https://registry.npmjs.org/@3thfloor/engine/-/engine-0.1.0.tgz","fileCount":29,"integrity":"sha512-yDnY7u5Vt55+0MCaOKVSU+jz1KRkGa2q14E9q3pCcpe8E5iA8YuIs0Bfjt3tcDu2b0kzjUqsyK8xpnoLyJddXA==","signatures":[{"sig":"MEYCIQD9wBwobSDy+0TmI6Gt7qargeHcVQY7cZ98CFzQjEMqYQIhAMICr/NvZTj8SFtqLy3dN3GaFC7Uto7+vDgPNVhIttzA","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":91538},"main":"./dist/cjs/index.js","type":"module","types":"./dist/index.d.ts","module":"./dist/index.js","engines":{"node":">=18.0.0"},"exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js","require":"./dist/cjs/index.js"}},"scripts":{"dev":"tsc --watch","build":"tsc && tsc -p tsconfig.cjs.json && node -e \"require('fs').writeFileSync('dist/cjs/package.json',JSON.stringify({type:'commonjs'}))\"","prepublishOnly":"npm run build"},"_npmUser":{"name":"ziptslug","email":"justin@3thfloor.com"},"repository":{"url":"git+https://github.com/3thfloor/engine.git","type":"git"},"_npmVersion":"11.9.0","description":"Local AI inference SDK. npm install, load, chat.","directories":{},"_nodeVersion":"22.14.0","dependencies":{"node-llama-cpp":"^3.0.0"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.4.0","@types/node":"^20.0.0"},"_npmOperationalInternal":{"tmp":"tmp/engine_0.1.0_1786747789487_0.025475101430465497","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"name":"@3thfloor/engine","version":"0.1.1","keywords":["llm","ai","inference","local","gguf","llama","embedding"],"author":{"name":"Justin Bench","email":"justin.bench@gmail.com"},"license":"PolyForm-Noncommercial-1.0.0","_id":"@3thfloor/engine@0.1.1","maintainers":[{"name":"ziptslug","email":"justin@3thfloor.com"}],"homepage":"https://3thfloor.com","bugs":{"url":"https://github.com/3thfloor/laconic/issues"},"dist":{"shasum":"3cd2c500eb6bc13e2ec05cec00744c7ae1b0354a","tarball":"https://registry.npmjs.org/@3thfloor/engine/-/engine-0.1.1.tgz","fileCount":29,"integrity":"sha512-fbBBSgW+VNVnigvm3v77l3XixlEj5wgY7+o+ju93BE7BfLvs6RK0PXc9yMaXtDMuZMAPuoEWcoOz4hzAIeIwwA==","signatures":[{"sig":"MEUCIGxXux2OWDEXil5BocRI6idMKHVuw1YbrPloqrBUjcXGAiEA1/6aKxoOxIqJ590srSj8kR2/xrfRmwHxzwpgR0THbJs=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":91539},"main":"./dist/cjs/index.js","type":"module","types":"./dist/index.d.ts","module":"./dist/index.js","engines":{"node":">=18.0.0"},"exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js","require":"./dist/cjs/index.js"}},"gitHead":"7bf8f7a0fb4518bb2125ead34c008be0298f6193","scripts":{"dev":"tsc --watch","build":"tsc && tsc -p tsconfig.cjs.json && node -e \"require('fs').writeFileSync('dist/cjs/package.json',JSON.stringify({type:'commonjs'}))\"","prepublishOnly":"npm run build"},"_npmUser":{"name":"ziptslug","email":"justin@3thfloor.com"},"repository":{"url":"git+https://github.com/3thfloor/laconic.git","type":"git"},"_npmVersion":"11.9.0","description":"Local AI inference SDK. npm install, load, chat.","directories":{},"_nodeVersion":"22.14.0","dependencies":{"node-llama-cpp":"^3.0.0"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.4.0","@types/node":"^20.0.0"},"_npmOperationalInternal":{"tmp":"tmp/engine_0.1.1_1787077595698_0.3852225921534753","host":"s3://npm-registry-packages-npm-production"}},"0.1.2":{"name":"@3thfloor/engine","version":"0.1.2","description":"Local AI inference SDK. npm install, load, chat.","author":{"name":"Justin Bench","email":"justin.bench@gmail.com"},"license":"PolyForm-Noncommercial-1.0.0","homepage":"https://3thfloor.com","type":"module","main":"./dist/cjs/index.js","module":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"import":"./dist/index.js","require":"./dist/cjs/index.js","types":"./dist/index.d.ts"}},"scripts":{"build":"tsc && tsc -p tsconfig.cjs.json && node -e \"require('fs').writeFileSync('dist/cjs/package.json',JSON.stringify({type:'commonjs'}))\"","dev":"tsc --watch","prepublishOnly":"npm run build"},"dependencies":{"node-llama-cpp":"^3.0.0"},"devDependencies":{"typescript":"^5.4.0","@types/node":"^20.0.0"},"engines":{"node":">=18.0.0"},"keywords":["llm","ai","inference","local","gguf","llama","embedding"],"repository":{"type":"git","url":"git+https://github.com/3thfloor/laconic.git"},"_id":"@3thfloor/engine@0.1.2","gitHead":"ee32fed49de6c9afbe42de7ef87b8a1b0817669e","bugs":{"url":"https://github.com/3thfloor/laconic/issues"},"_nodeVersion":"22.14.0","_npmVersion":"10.9.2","dist":{"integrity":"sha512-OBcpmo/0sLrxDfVryG/YFqHFs3jixjfGYL/zkjEnhfMaRGtmiBz0KgOQe3lXCjX8kxMePGDM/CZZrQ10nu67gA==","shasum":"d9330e464b2254a77c277423d458f7c72a554721","tarball":"https://registry.npmjs.org/@3thfloor/engine/-/engine-0.1.2.tgz","fileCount":29,"unpackedSize":95690,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIARhs8sNAVSQJEvlipD99l3cMQ/GeFBYRhxVtf+qubUUAiEA8RfbyY4hsI79FYLbZIuZTFoUDNapeY0usKJHsVQ33U0="}]},"_npmUser":{"name":"ziptslug","email":"justin@3thfloor.com"},"directories":{},"maintainers":[{"name":"ziptslug","email":"justin@3thfloor.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/engine_0.1.2_1787436745871_0.9847833453903498"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-14T22:49:49.282Z","modified":"2026-08-22T22:12:26.160Z","0.1.0":"2026-08-14T22:49:49.630Z","0.1.1":"2026-08-18T18:26:35.851Z","0.1.2":"2026-08-22T22:12:26.006Z"},"bugs":{"url":"https://github.com/3thfloor/laconic/issues"},"author":{"name":"Justin Bench","email":"justin.bench@gmail.com"},"license":"PolyForm-Noncommercial-1.0.0","homepage":"https://3thfloor.com","keywords":["llm","ai","inference","local","gguf","llama","embedding"],"repository":{"type":"git","url":"git+https://github.com/3thfloor/laconic.git"},"description":"Local AI inference SDK. npm install, load, chat.","maintainers":[{"name":"ziptslug","email":"justin@3thfloor.com"}],"readme":"# @3thfloor/engine\n\nLocal AI inference for Node.js and Electron. One install. No daemon. No API keys.\n\nAsk a question, get a string back. That is the whole API surface you need to learn.\n\n```js\nconst answer = await engine.chat(\"assistant\", \"What is a smoke test?\");\nconsole.log(answer); // a plain string\n```\n\nNo `.choices[0].message.content`. No response envelopes. No server to babysit. The model is an object in memory that lives and dies with your process.\n\n## Install\n\n```bash\nnpm install @3thfloor/engine\n```\n\nThe install downloads the right prebuilt native binary for your platform automatically (macOS arm64/x64, Linux x64, Windows x64). No compiler toolchain required. GPU acceleration (Metal on macOS, CUDA and Vulkan on Linux/Windows) is detected and used when available.\n\nRequires Node 18 or later.\n\n## Quick Start\n\nTypeScript:\n\n```ts\nimport { Engine } from \"@3thfloor/engine\";\n\nconst engine = new Engine();\nawait engine.load(\"assistant\", \"./models/qwen3-4b-q4.gguf\");\n\nconst answer: string = await engine.chat(\"assistant\", \"Write one unit test case name for a login form.\");\nconsole.log(answer);\n```\n\nJavaScript (ESM):\n\n```js\nimport { Engine } from \"@3thfloor/engine\";\n\nconst engine = new Engine();\nawait engine.load(\"assistant\", \"./models/qwen3-4b-q4.gguf\");\n\nconst answer = await engine.chat(\"assistant\", \"Write one unit test case name for a login form.\");\nconsole.log(answer);\n```\n\n`load()` gives the model an alias. Every other call refers to the model by that alias. `chat()` returns a string. Done.\n\n## Sessions (Conversation History)\n\nA `Session` keeps conversation history for you. Each `send()` includes everything said so far.\n\n```ts\nimport { Engine } from \"@3thfloor/engine\";\n\nconst engine = new Engine();\nawait engine.load(\"assistant\", \"./models/qwen3-4b-q4.gguf\");\n\nconst session = engine.session(\"assistant\", {\n  system: \"You are a QA lead. Answer in two sentences or fewer.\",\n});\n\nconst first = await session.send(\"What is boundary value analysis?\");\nconsole.log(first);\n\nconst second = await session.send(\"Give me one example using an age field.\");\nconsole.log(second); // the model remembers the first exchange\n```\n\nCall `session.clear()` to clear history, or `session.history` to inspect it as an array of `{ role, content }` objects.\n\n## Streaming\n\n`stream()` returns an async iterator of token strings. Print them as they arrive:\n\n```ts\nfor await (const token of engine.stream(\"assistant\", \"Explain regression testing in one paragraph.\")) {\n  process.stdout.write(token);\n}\nprocess.stdout.write(\"\\n\");\n```\n\nSessions stream too:\n\n```ts\nfor await (const token of session.stream(\"Now explain it to a new hire.\")) {\n  process.stdout.write(token);\n}\n```\n\n## Multiple Models\n\nLoad as many models as your memory allows. Route between them by alias:\n\n```ts\nconst engine = new Engine();\nawait engine.load(\"fast\", \"./models/qwen3-4b-q4.gguf\");\nawait engine.load(\"smart\", \"./models/qwen3-32b-q4.gguf\");\n\n// Small model triages, big model handles the hard stuff\nconst triage = await engine.chat(\"fast\", `Is this question simple or complex? Answer one word.\\n\\n${question}`);\nconst alias = triage.toLowerCase().includes(\"complex\") ? \"smart\" : \"fast\";\n\nconst answer = await engine.chat(alias, question);\n```\n\nUnload a model when you are done with it:\n\n```ts\nawait engine.unload(\"fast\");\n```\n\n## Agents\n\nDefine tools as plain functions, then let the model call them:\n\n```ts\nimport { Engine, tool, runAgent } from \"@3thfloor/engine\";\n\nconst engine = new Engine();\nawait engine.load(\"assistant\", \"./models/qwen3-32b-q4.gguf\");\n\nconst getTestCount = tool(\n  \"get_test_count\",\n  \"Returns the number of tests in the suite\",\n  async () => {\n    return { total: 412, passing: 398, failing: 14 };\n  }\n);\n\nconst result = await runAgent(engine, \"assistant\", \"How many tests are failing right now?\", {\n  tools: [getTestCount],\n});\n\nconsole.log(result); // a string, same as chat()\n```\n\nThe agent loop handles tool selection, execution, and feeding results back to the model. You get the final answer as a string.\n\n## Model Management\n\nThe engine ships with a model registry so you are not passing file paths around your codebase:\n\n```ts\n// Register a local file under an alias\nengine.models.add(\"qwen-4b\", \"./models/qwen3-4b-q4.gguf\");\n\n// See what is registered\nconsole.log(engine.models.list());\n// [ { alias: \"qwen-4b\", path: \"/abs/path/to/qwen3-4b-q4.gguf\", sizeMb: 2548.3, ctx: 4096, ... } ]\n\n// Download straight from HuggingFace (auto-picks Q4_K_M by default)\nawait engine.models.download(\"Qwen/Qwen3-4B-Instruct-GGUF\");\n\n// Or request a specific quantization\nawait engine.models.download(\"Qwen/Qwen3-4B-Instruct-GGUF\", { quantization: \"Q4_K_M\" });\n\n// Or a specific filename\nawait engine.models.download(\"Qwen/Qwen3-4B-Instruct-GGUF\", { filename: \"qwen3-4b-instruct-q4_k_m.gguf\" });\n\n// Load by registered alias (alias = filename without .gguf)\nconst entry = engine.models.info(\"qwen3-4b-instruct-q4_k_m\");\nawait engine.load(\"assistant\", entry.path);\n```\n\nDownloads go to `~/.3thfloor/models/` and are skipped if the file already exists.\n\n## Use in Electron\n\nThe engine runs inside the Electron main process. No daemon, no localhost port, no HTTP. Expose it to your renderer with a single IPC handler:\n\n```ts\n// main.ts\nimport { app, ipcMain } from \"electron\";\nimport { Engine } from \"@3thfloor/engine\";\n\nconst engine = new Engine();\n\napp.whenReady().then(async () => {\n  await engine.load(\"assistant\", \"./models/qwen3-4b-q4.gguf\");\n\n  ipcMain.handle(\"ai:chat\", async (_event, message: string) => {\n    return engine.chat(\"assistant\", message);\n  });\n});\n```\n\n```ts\n// renderer (through your preload bridge)\nconst answer = await window.api.invoke(\"ai:chat\", \"Summarize this bug report.\");\n```\n\nBecause inference happens in the main process, your app ships as one binary with no background services to install, no ports to fight over, and nothing left running when the user quits.\n\n## Use in Any Node App\n\nThe engine works anywhere Node 18+ runs: Express and Fastify APIs, CLI tools, test runners, cron jobs, queue workers. Create the engine once at startup, load your models, and the engine stays alive for the lifetime of your process. Every request after the first load is answered from the model already sitting in memory, so there is no cold start and no connection handling to write.\n\n```ts\n// express example\napp.post(\"/ask\", async (req, res) => {\n  const answer = await engine.chat(\"assistant\", req.body.question);\n  res.json({ answer });\n});\n```\n\n---\n\n## License\n\nFree for personal projects, research, experiments, and noncommercial use under the [PolyForm Noncommercial License 1.0.0](./LICENSE).\n\nIf you are shipping a product with this engine embedded, a commercial license is required. Contact [justin@3thfloor.com](mailto:justin@3thfloor.com).\n\n---\n\nBuilt by Justin Bench, [3th Floor AI](https://3thfloor.com).\n","readmeFilename":"README.md"}