{"_id":"@archimonde12/llm-proxy","name":"@archimonde12/llm-proxy","dist-tags":{"latest":"1.0.0"},"versions":{"1.0.0":{"name":"@archimonde12/llm-proxy","version":"1.0.0","description":"A lightweight gateway that routes OpenAI-compatible requests to local and remote LLM backends.","main":"dist/index.js","exports":{".":"./dist/index.js","./package.json":"./package.json","./cli":"./dist/cli/bin.js"},"bin":{"llm-proxy":"dist/cli/bin.js"},"engines":{"node":">=20"},"scripts":{"dev":"tsx watch src/index.ts","build":"vite build --config ui/vite.config.ts && tsc -p tsconfig.build.json","build:ui":"vite build --config ui/vite.config.ts","prepublishOnly":"npm run build","start":"node dist/index.js","cli":"node dist/cli/bin.js","test":"vitest run","test:watch":"vitest"},"keywords":["llm","proxy","gateway","openai","ollama","cli","ai"],"author":{"name":"llmog"},"license":"MIT","repository":{"type":"git","url":"git+https://github.com/llmog/llm-proxy.git"},"bugs":{"url":"https://github.com/llmog/llm-proxy/issues"},"homepage":"https://github.com/llmog/llm-proxy#readme","dependencies":{"@fastify/static":"^9.1.0","commander":"^14.0.1","dotenv":"^17.2.2","fastify":"^5.8.4","prom-client":"^15.1.3","prompts":"^2.4.2","zod":"^4.1.5"},"packageManager":"pnpm@10.33.0+sha512.10568bb4a6afb58c9eb3630da90cc9516417abebd3fabbe6739f0ae795728da1491e9db5a544c76ad8eb7570f5c4bb3d6c637b2cb41bfdcdb47fa823c8649319","devDependencies":{"@types/node":"^25.5.2","@types/prompts":"^2.4.9","@types/react":"^19.2.14","@types/react-dom":"^19.2.3","@vitejs/plugin-react":"^6.0.1","react":"^19.2.4","react-dom":"^19.2.4","tsx":"^4.21.0","typescript":"^6.0.2","vite":"^8.0.7","vitest":"^3.2.4"},"_id":"@archimonde12/llm-proxy@1.0.0","_nodeVersion":"23.7.0","_npmVersion":"9.5.0","dist":{"integrity":"sha512-zRES0Ew5mhNPo+CNtiAq+A99m5gr4Xa1i6gr8u55VuctCOFH8y4PwAXV52IdIr6D24lLQ76kEWm4ffrgTHi8FQ==","shasum":"7ff1bf6d5a6573d81586e988942b53ca3e9469a7","tarball":"https://registry.npmjs.org/@archimonde12/llm-proxy/-/llm-proxy-1.0.0.tgz","fileCount":42,"unpackedSize":360153,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIFP18qOleaaXnYice99EIAVniLDFPhOBgo07kw+gsXiFAiEAyYxXA5MwDgrq2CawExm817FOslBVk2dwjP60sp/I0x8="}]},"_npmUser":{"name":"archimonde12","email":"hoanbkhn@gmail.com"},"directories":{},"maintainers":[{"name":"archimonde12","email":"hoanbkhn@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/llm-proxy_1.0.0_1775812734037_0.05289030897440883"},"_hasShrinkwrap":false}},"time":{"created":"2026-04-10T09:18:53.947Z","1.0.0":"2026-04-10T09:18:54.183Z","modified":"2026-04-10T09:18:54.412Z"},"maintainers":[{"name":"archimonde12","email":"hoanbkhn@gmail.com"}],"description":"A lightweight gateway that routes OpenAI-compatible requests to local and remote LLM backends.","homepage":"https://github.com/llmog/llm-proxy#readme","keywords":["llm","proxy","gateway","openai","ollama","cli","ai"],"repository":{"type":"git","url":"git+https://github.com/llmog/llm-proxy.git"},"author":{"name":"llmog"},"bugs":{"url":"https://github.com/llmog/llm-proxy/issues"},"license":"MIT","readme":"# llm-proxy\n\n**llm-proxy** is a lightweight, high-performance gateway designed to unify multiple LLM backends (Ollama, vLLM, OpenAI-compatible servers, etc.) into a single, standardized OpenAI-compatible API endpoint.\n\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)\n[![Node.js Version](https://img.shields.io/badge/node-%3E%3D20-brightgreen)](https://nodejs.org/)\n[![Package Manager](https://img.shields.io/badge/package-pnpm-ff6302.svg)](https://pnpm.io/)\n[![TypeScript](https://img.shields.io/badge/TypeScript-%233178c6.svg?style=flat&logo=typescript&logoColor=white)](https://www.typescriptlang.org/)\n[![Fastify](https://img.shields.io/badge/Fastify-black?logo=fastify)](https://www.fastify.io/)\n[![Build Status](https://img.shields.io/badge/build%20passing-%234caf50.svg)](#)\n\n---\n\n## Key features\n\n*   **Unified interface**: Access different backends (Ollama, vLLM, etc.) using the standard OpenAI SDK and payload format.\n*   **Multi-backend support**: Seamlessly route requests to various providers via a simple `models.json` configuration.\n*   **Dynamic configuration**: Update your model list and backend URLs on-the-fly via the **Admin API** without restarting the server.\n*   **Built-in observability**:\n    *   **Prometheus metrics**: Monitor request counts, latency, and token usage.\n    *   **Request history**: Track recent requests via an in-memory buffer.\n*   **Web UI**: A built-in dashboard to manage configuration, environment variables, and monitoring (when `ui/dist` is present).\n*   **Streaming**: Full support for Server-Sent Events (SSE) for real-time chat completions.\n*   **Lightweight and fast**: Built with **Fastify** and **TypeScript** for minimal overhead and maximum throughput.\n\n---\n\n## Quick start\n\n### Prerequisites\n\n*   [Node.js](https://nodejs.org/) (**v20+ required**)\n*   [pnpm](https://pnpm.io/)\n\n### Installation\n\n**From npm (global CLI)** — no clone required; the published package includes `dist/` and `ui/dist` (dashboard at `/ui` when you run the server).\n\n```bash\nnpm install -g @archimonde12/llm-proxy\n```\n\nRequires **Node.js 20+** (see `engines` in `package.json`). After installing, run `llm-proxy --help` or `llm-proxy start`.\n\n**Without global install (npx):**\n\n```bash\nnpx @archimonde12/llm-proxy --help\n```\n\nConfiguration file resolution is described under [Configuration](#configuration) (important when using a global install from arbitrary working directories).\n\n**From source** (development):\n\n```bash\n# Clone the repository (replace OWNER with your fork or upstream)\ngit clone https://github.com/OWNER/llm-proxy.git\ncd llm-proxy\n\n# Install dependencies\npnpm install\n```\n\n### Running the server\n\n**Development** (TypeScript; no production UI bundle required for API-only work):\n\n```bash\npnpm dev\n```\n\nIn **development** (`pnpm dev`), the default bind is **0.0.0.0:8787** (override with `HOST` and `PORT`). In **`llm-proxy start`** / **`pnpm start`**, the default bind is **127.0.0.1:8787** unless you set `HOST`, `PORT`, or pass `--host` / `--port`.\n\n**Production** (compiled server and built web UI):\n\n```bash\nnpm run build   # or: pnpm build — Vite UI bundle + TypeScript; produces ui/dist and dist/\npnpm start      # or: npm start — runs node dist/index.js\n```\n\nA full build always runs the UI step first, so `ui/dist` is present afterward. Maintainers can use **npm** or **pnpm** for `build` / `start`; the build script does not invoke `pnpm` internally.\n\n### CLI (`llm-proxy`)\n\nWith **`npm install -g @archimonde12/llm-proxy`**, use the `llm-proxy` command on your `PATH`.\n\nFrom a **source** tree after `npm run build` / `pnpm build`:\n\n```bash\npnpm cli -- <subcommand> [options]\n# equivalent:\nnode dist/cli/bin.js <subcommand> [options]\n```\n\n| Command | Purpose |\n| :--- | :--- |\n| `init` | Create a starter `models.json` (wizard when interactive; use `-y` for defaults). Options: `--file <path>`, `-y` / `--yes`. |\n| `start` | Start the HTTP server. Options: `--models <path>`, `--port <port>` (default **8787**), `--host <host>` (default **127.0.0.1**). |\n| `status` | Call public **`GET /healthz`**. Options: `--url <baseUrl>` (default `http://127.0.0.1:8787`). |\n| `doctor` | Validate config, check listen port, optionally ping upstreams. Options: `--models <path>`, `--host`, `--port`, `--deep`. |\n| `config validate` | Validate `models.json` against the schema. Options: `--file <path>` (default `./models.json`). |\n\n### Environment variables (server / CLI)\n\n| Variable | Used by | Description |\n| :--- | :--- | :--- |\n| `MODELS_PATH` | `pnpm dev`, `start`, `doctor` | Path to `models.json`. For `start` and `doctor`, only used when `--models` is omitted; `pnpm dev` always resolves the file from env / defaults (no `--models` flag). |\n| `PORT` | `pnpm dev`, `start` | Listen port if unset (default **8787**); `start` also accepts `--port`. |\n| `HOST` | `pnpm dev`, `start` | Bind address if unset — **`pnpm dev`** defaults to **0.0.0.0**, **`start`** defaults to **127.0.0.1**; `start` also accepts `--host`. |\n| `LOG_LEVEL` | Server | Optional: `debug`, `info`, `warn`, or `error`; omit to log at all levels. |\n\n---\n\n## Configuration\n\nThe proxy uses a `models.json` file to map your custom model IDs to specific backends.\n\n**Resolution order** (see [`src/config/load.ts`](src/config/load.ts)):\n\n1. **`llm-proxy start --models <path>`** — explicit file path.\n2. **`MODELS_PATH`** — if set, that path is used (a starter file is created if missing).\n3. **Otherwise:** **`./models.json`** relative to the **current working directory** — if it exists, it is used.\n4. **Otherwise:** **`~/.config/llm-proxy/models.json`** — canonical user config; if it exists, it is used (good default for a **global** install when you are not in a project directory).\n5. **Otherwise:** **`~/.config/llm-open-gateway/models.json`** — legacy path for backward compatibility with older installs; used only if the canonical path above does not exist.\n6. **Otherwise:** a starter `models.json` is created at **`./models.json`** in the current working directory.\n\nThis means a global install does not require a checkout: use a file in the cwd, set `MODELS_PATH`, pass `--models`, or keep your config under `~/.config/llm-proxy/models.json`.\n\n### Example `models.json`\n\n```json\n{\n  \"models\": [\n    {\n      \"id\": \"ollama-llama3\",\n      \"adapter\": \"ollama\",\n      \"baseUrl\": \"http://localhost:11434\",\n      \"model\": \"llama3\"\n    },\n    {\n      \"id\": \"vllm-mixtral\",\n      \"adapter\": \"openai_compatible\",\n      \"baseUrl\": \"http://localhost:8000\",\n      \"model\": \"mixtral-8x7b\",\n      \"apiKey\": \"your-secret-api-key\"\n    }\n  ]\n}\n```\n\n### Supported adapters\n\n| Adapter | Description |\n| :--- | :--- |\n| `ollama` | Optimized for the [Ollama](https://ollama.com/) HTTP API. |\n| `openai_compatible` | Any server implementing the OpenAI `/v1/chat/completions` contract. |\n| `deepseek` | DeepSeek-compatible HTTP API (see [`src/adapters/deepseek.ts`](src/adapters/deepseek.ts)). |\n\n### Optional fields (per model)\n\n| Field | Description |\n| :--- | :--- |\n| `apiKey` | Secret sent to the upstream (see **Secrets** below). |\n| `apiKeyHeader` | Header name for the API key (requires `apiKey`). |\n| `headers` | Extra static headers as string key/value pairs. |\n| `timeoutMs` | Upstream request timeout in milliseconds. |\n\n### Secrets: `apiKey` and `${ENV_VAR}`\n\nTo avoid committing raw keys, you can set `apiKey` to a **single** environment placeholder: the string must be exactly `${` + the variable name + `}`, where the name uses only letters, digits, and underscores. On load, the server replaces that string with the value from `process.env`. If the variable is unset or empty, `apiKey` is dropped so the rest of the entry can still pass validation.\n\nAfter changing environment variables (for example via **`PUT /admin/env`**), call **`POST /admin/reload`** so `${...}` placeholders are resolved again from the updated process environment.\n\n---\n\n## Web UI (dashboard)\n\nThe UI is a static SPA served from **`ui/dist`**. The server registers it **only if that directory exists**; otherwise there is no **`GET /`** redirect to the app (see [`src/server.ts`](src/server.ts)).\n\n**Prerequisite:** run **`pnpm build:ui`** or a full **`pnpm build`** before starting the server if you want the dashboard.\n\n**URLs**\n\n* **`GET /`** → **302** to **`/ui/`** when the UI bundle is present.\n* Static assets are served under **`/ui/`**.\n\n**In-app routes** (hash-based):\n\n* **`#/configuration`** — Edit `models.json` and environment variables through the Admin API.\n* **`#/monitoring`** — Metrics overview with time ranges **15m**, **1h**, and **24h** (aligned with `/admin/metrics/overview`).\n* **`#/models`** — Model list, request logs, and debug message capture.\n\n**Security:** When the process binds to a non-loopback address, **`/admin/*`** and **`/ui/*`** are restricted to localhost clients. To use the dashboard remotely, use local access, SSH port forwarding, or a tunnel you trust.\n\n**UI development:** There is no separate `package.json` under `ui/`; for a local Vite dev server you can run e.g. `pnpm exec vite --config ui/vite.config.ts` from the repo root. Production assets are built with **`pnpm build:ui`**.\n\n---\n\n## Admin API, health, and metrics\n\n*Note: Admin routes are restricted to `localhost` when the server is bound to a non-loopback interface.*\n\n### Public health (no admin guard)\n\n| Method | Path | Description |\n| :--- | :--- | :--- |\n| `GET` | `/healthz` | Liveness: returns `{ ok: true }` while the process is running. Used by **`llm-proxy status`**. |\n| `GET` | `/readyz` | Readiness. With **`?deep=1`**, probes each model's upstream; may return **503** if any probe fails. |\n\n### Admin API (localhost when exposed on a public bind)\n\n**Config and environment**\n\n| Method | Path | Description |\n| :--- | :--- | :--- |\n| `GET` | `/admin/health` | Process/version hint and active config path (not the same as public `/healthz`). |\n| `GET` | `/admin/config` | Current configuration, write target, and metadata. |\n| `PUT` | `/admin/config` | Replace configuration (validated); writes atomically. |\n| `POST` | `/admin/reload` | Reload `models.json` from disk. |\n| `GET` | `/admin/env` | List relevant env keys and `.env` file metadata. |\n| `PUT` | `/admin/env` | Apply env updates to `.env` and `process.env`. |\n\n**Connectivity and metrics**\n\n| Method | Path | Description |\n| :--- | :--- | :--- |\n| `POST` | `/admin/test-connection` | Probe a configured **`modelId`** or an inline adapter/baseUrl/model payload. |\n| `GET` | `/admin/metrics/summary` | JSON snapshot from in-process metrics. |\n| `GET` | `/admin/metrics/overview` | Overview for **`range=15m`**, **`1h`**, or **`24h`** (query param). |\n\n**Logs and requests**\n\n| Method | Path | Description |\n| :--- | :--- | :--- |\n| `GET` | `/admin/requests` | Recent proxy request history (`limit` query param, capped). |\n| `GET` | `/admin/requests/:requestId` | Single request record by id. |\n| `GET` | `/admin/logs/models` | Model-scoped logs with **`range`**, optional **`modelId`**, **`status`**, **`limit`**. |\n| `GET` | `/admin/models/:modelId/debug/messages` | Recent captured **system** / **user** messages for debugging (`limit`, **`roles`**). |\n\n### Prometheus\n\n| Method | Path | Description |\n| :--- | :--- | :--- |\n| `GET` | `/metrics` | Prometheus text exposition format. |\n\n---\n\n## Security and exposure\n\nIf you intend to expose `llm-proxy` to the internet (e.g., via **ngrok** or a reverse proxy), please follow these best practices:\n\n1.  **Use HTTPS**: Always use an SSL/TLS tunnel or terminating proxy.\n2.  **Admin and UI**: The server restricts `/admin/*` and `/ui/*` to localhost when it detects a non-loopback bind.\n3.  **Upstream credentials**: Use `apiKey` / `${ENV_VAR}` in `models.json` for provider authentication.\n\n**Example with ngrok:**\n\n```bash\nngrok http 8787\n```\n\n---\n\n## Contributing\n\nPull requests are welcome; please keep changes focused and consistent with existing patterns.\n\n---\n\n## License\n\nDistributed under the **MIT License**. See [`LICENSE`](LICENSE) for more information.\n","readmeFilename":"README.md","_rev":"1-0bf976a4757e483d9b5fc99483a067e5"}