{"_id":"omp-makora-provider","_rev":"3-fb9fe8c3bd6bce4fe5354d178c42698c","name":"omp-makora-provider","dist-tags":{"latest":"1.0.4"},"versions":{"1.0.0":{"name":"omp-makora-provider","version":"1.0.0","keywords":["omp","plugin","provider","makora","ai","llm","deepseek","glm","kimi","llama","qwen"],"author":{"name":"Ryan Brosas","email":"ryanbrosas32834@outlook.com"},"license":"MIT","_id":"omp-makora-provider@1.0.0","maintainers":[{"name":"ryanjoserbrosas","email":"ryanjoserbrosas@gmail.com"}],"omp":{"extensions":["./index.ts"]},"dist":{"shasum":"5955708902c4140220cbcf9174d52a143117ca94","tarball":"https://registry.npmjs.org/omp-makora-provider/-/omp-makora-provider-1.0.0.tgz","fileCount":7,"integrity":"sha512-J0QGZDrR+mxFYRdJFh3DDLtZYztbJ2wUpCg7TaYJCEe23IiWZ1r0kzzo2JVPZ0OIUHfCNJonaeZU1+2GLjZ83A==","signatures":[{"sig":"MEYCIQDQ5vtQUFZL/UP7R5pTfYPWIlvc6JNmXOqXQ/WSOCIHRwIhANjZK1VLL/+Q6m23XjEfnq3ZkzF9kKBUs2EyvOVyRvWj","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":48742},"main":"index.ts","type":"module","gitHead":"18039dee7800075e0aa878b94ab5a7a189499aa7","scripts":{"build":"echo 'nothing to build'","check":"echo 'nothing to check'","clean":"echo 'nothing to clean'","update-models":"node scripts/update-models.js"},"_npmUser":{"name":"ryanjoserbrosas","email":"ryanjoserbrosas@gmail.com"},"_npmVersion":"10.9.8","description":"Makora provider plugin for OMP — Access DeepSeek V4, GLM 5.1, Kimi K2.6, Llama 3.3, Qwen 3.6, and more through the Makora inference API","directories":{},"_nodeVersion":"22.22.3","_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/omp-makora-provider_1.0.0_1781777473970_0.4141183028639046","host":"s3://npm-registry-packages-npm-production"}},"1.0.1":{"name":"omp-makora-provider","version":"1.0.1","keywords":["omp","plugin","provider","makora","ai","llm","deepseek","glm","kimi","llama","qwen"],"author":{"name":"Ryan Brosas","email":"ryanbrosas32834@outlook.com"},"license":"MIT","_id":"omp-makora-provider@1.0.1","maintainers":[{"name":"ryanjoserbrosas","email":"ryanjoserbrosas@gmail.com"}],"omp":{"extensions":["./index.ts"]},"dist":{"shasum":"50d80ab010b6ca34ee8cff15b5e9a87b828f28aa","tarball":"https://registry.npmjs.org/omp-makora-provider/-/omp-makora-provider-1.0.1.tgz","fileCount":7,"integrity":"sha512-HlgewO7f7jJ161KD7TTYi/Zg4BGDfMXDxSEJGDG4h9AnXrnUxLcE9E60waDxy9epbzJ2SjRoRf7MvKMefuwWzA==","signatures":[{"sig":"MEQCIEvO9a/PA0dX+AZ7xpITdb13ByXwzoe0mImYj8Vec2yJAiBvTvlCwY56SQn+g8KDP60tL2WGjMRQoTy6qV8l86Uz+Q==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":49013},"main":"index.ts","type":"module","gitHead":"18039dee7800075e0aa878b94ab5a7a189499aa7","scripts":{"build":"echo 'nothing to build'","check":"echo 'nothing to check'","clean":"echo 'nothing to clean'","update-models":"node scripts/update-models.js"},"_npmUser":{"name":"ryanjoserbrosas","email":"ryanjoserbrosas@gmail.com"},"_npmVersion":"10.9.8","description":"Makora provider plugin for OMP — Access DeepSeek V4, GLM 5.1, Kimi K2.6, Llama 3.3, Qwen 3.6, and more through the Makora inference API","directories":{},"_nodeVersion":"22.22.3","_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/omp-makora-provider_1.0.1_1781777627804_0.7127000587945793","host":"s3://npm-registry-packages-npm-production"}},"1.0.4":{"name":"omp-makora-provider","version":"1.0.4","description":"Makora provider plugin for OMP — Access DeepSeek V4, GLM 5.1, Kimi K2.6, Llama 3.3, Qwen 3.6, and more through the Makora inference API","type":"module","main":"index.ts","repository":{"type":"git","url":"git+https://github.com/ryan-brosas/omp-makora-provider.git"},"scripts":{"clean":"echo 'nothing to clean'","build":"echo 'nothing to build'","check":"echo 'nothing to check'","test":"npx vitest run","update-models":"node scripts/update-models.js","test:watch":"npx vitest"},"keywords":["omp","plugin","provider","makora","ai","llm","deepseek","glm","kimi","llama","qwen"],"author":{"name":"Ryan Brosas","email":"ryanbrosas32834@outlook.com"},"license":"MIT","omp":{"extensions":["./index.ts"]},"devDependencies":{"vitest":"^4.1.9"},"_id":"omp-makora-provider@1.0.4","gitHead":"9126d4c883eb5d51dcb3f08e21fb5285d97f2022","bugs":{"url":"https://github.com/ryan-brosas/omp-makora-provider/issues"},"homepage":"https://github.com/ryan-brosas/omp-makora-provider#readme","_nodeVersion":"22.22.3","_npmVersion":"10.9.8","dist":{"integrity":"sha512-2Jo0hQM9FI01Xf/s6wM9se/CYIFNcyZ9V42ckg+p3OGuaThAixb15ppiD/hfRHoQMgJXEdge9EOvHmg8FYe1eQ==","shasum":"1a1a64151aa3c9d74c50ebc00aeb369fee59a3eb","tarball":"https://registry.npmjs.org/omp-makora-provider/-/omp-makora-provider-1.0.4.tgz","fileCount":7,"unpackedSize":52326,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/omp-makora-provider@1.0.4","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQCkoB1yNNwX1drAxuAuQ5qVSGBNRxsG1WEF5Us0Pt8SjwIgRdkQXcpczFHWDv9EV1UUKZBqn5ftVLIjfi1f9ETgywo="}]},"_npmUser":{"name":"ryanjoserbrosas","email":"ryanjoserbrosas@gmail.com"},"directories":{},"maintainers":[{"name":"ryanjoserbrosas","email":"ryanjoserbrosas@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/omp-makora-provider_1.0.4_1782162210752_0.04398586520108427"},"_hasShrinkwrap":false}},"time":{"created":"2026-06-18T10:11:13.905Z","modified":"2026-06-22T21:03:31.216Z","1.0.0":"2026-06-18T10:11:14.116Z","1.0.1":"2026-06-18T10:13:47.997Z","1.0.4":"2026-06-22T21:03:30.917Z"},"author":{"name":"Ryan Brosas","email":"ryanbrosas32834@outlook.com"},"license":"MIT","keywords":["omp","plugin","provider","makora","ai","llm","deepseek","glm","kimi","llama","qwen"],"description":"Makora provider plugin for OMP — Access DeepSeek V4, GLM 5.1, Kimi K2.6, Llama 3.3, Qwen 3.6, and more through the Makora inference API","maintainers":[{"name":"ryanjoserbrosas","email":"ryanjoserbrosas@gmail.com"}],"readme":"<div align=\"center\">\n\n# 🔁 omp-makora-provider\n\n**Open-weight models through [Makora](https://inference.makora.com)**\n\n_DeepSeek V4, Kimi K2.6, GLM 5.1 / 5.2, Qwen 3.6 — with client-side tool call repair for [OMP](https://github.com/oh-my-pi/oh-my-pi) / [pi](https://github.com/earendil-works/pi-coding-agent)._\n\n[![OMP plugin](https://img.shields.io/badge/OMP-plugin-blueviolet)](https://github.com/oh-my-pi/oh-my-pi)\n[![license](https://img.shields.io/badge/license-MIT-blue)](./LICENSE)\n\n</div>\n\n---\n\n## Models\n\n<!-- MODELS_TABLE_START -->\n| Model | ID | Reasoning | Notes |\n|-------|----|-----------|-------|\n| DeepSeek V4 Flash | `deepseek-ai/DeepSeek-V4-Flash` | Yes | maxTokens 32768; `include_reasoning` + `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field |\n| DeepSeek V4 Pro | `deepseek-ai/DeepSeek-V4-Pro` | Yes | maxTokens 32768; `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning_content` field |\n| GLM 5.1 FP8 | `zai-org/GLM-5.1-FP8` | Yes | maxTokens 16384; `enable_thinking` via `qwen-chat-template`; returns `reasoning_content` field; client-side tool call parsing (vLLM streaming parser bypass) |\n| GLM 5.2 FP8 | `zai-org/GLM-5.2-FP8` | Yes | maxTokens 16384; `enable_thinking` via `qwen-chat-template`; returns `reasoning` field; native tool calls work in both stream and non-stream (no client-side repair needed) |\n| GPT-OSS 120B | `openai/gpt-oss-120b` | Yes | maxTokens 16384; reasoning always on |\n| Kimi K2.6 NVFP4 | `nvidia/Kimi-K2.6-NVFP4` | Yes | maxTokens 16384; vision maxImagesPerRequest 5; reasoning on by default; client-side tool call parsing (vLLM streaming parser bypass) |\n| Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | Yes | maxTokens 16384; vision maxImagesPerRequest 5; reasoning on by default; client-side tool call parsing (vLLM streaming parser bypass) |\n| Llama 3.3 70B FP8 | `amd/Llama-3.3-70B-Instruct-FP8-KV` | No | maxTokens 16384; custom per-slug endpoint |\n| Llama 3.3 70B Instruct | `meta-llama/Llama-3.3-70B-Instruct` | No | maxTokens 8192; non-reasoning text-only model |\n| MiniMax M3 MXFP8 | `MiniMaxAI/MiniMax-M3-MXFP8` | Yes | maxTokens 16384; vision maxImagesPerRequest 5; reasoning via `chat_template_kwargs.enable_thinking`; returns `reasoning_content` field |\n| Qwen 3.6 27B NVFP4 | `unsloth/Qwen3.6-27B-NVFP4` | Yes | maxTokens 16384; `enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass) |\n| Qwen 3.6 35B A3B NVFP4 | `unsloth/Qwen3.6-35B-A3B-NVFP4` | Yes | maxTokens 16384; `enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass) |\n<!-- MODELS_TABLE_END -->\n\n## Quickstart\n\n```bash\n# 1. Install\nomp plugin install omp-makora-provider\n\n# 2. Add your API key\nomp\n/login makora\n\n# 3. Pick a model and go\n/model makora\n```\n\nThat's it. Makora models now appear in `/model`. No `-e` flag, no manual clone, no config files.\n\n### API key\n\n`/login makora` prompts for your Makora API key, validates it, and stores it. Or set it explicitly:\n\n```bash\nexport MAKORA_OPTIMIZE_TOKEN=your-api-key\n```\n\nGet a key at [inference.makora.com](https://inference.makora.com).\n\n### Other install paths\n\n```bash\n# From GitHub\nomp plugin install https://github.com/ryan-brosas/omp-makora-provider\n\n# Local development\ngit clone https://github.com/ryan-brosas/omp-makora-provider.git\nomp plugin link ./omp-makora-provider\n```\n\n\n## Model Resolution\n\nModels are discovered from the Makora `/v1/models` API and stored in `models.json`. Custom definitions and overrides are layered via `patch.json` and `custom-models.json`.\n\n| File | Purpose |\n|---|---|\n| `models.json` | Auto-generated from Makora API (model discovery). Regenerated by `node scripts/update-models.js` — do not edit manually |\n| `patch.json` | Manual overrides (reasoning, compat, notes, limits, etc.) applied on top of `models.json` |\n| `custom-models.json` | Models not available via the API (e.g. per-slug endpoint models) |\n\nModels are loaded by merging `models.json` → apply `patch.json` → merge `custom-models.json`.\n\n### Patch metadata fields\n\n`patch.json` supports the same model metadata fields consumed by the provider, including `reasoning`, `input`, `contextWindow`, `maxTokens`, `vision`, `notes`, `thinkingLevelMap`, and `compat`. Use `maxTokens` for safe output caps because Makora model discovery does not report max output tokens. Use `vision.maxImagesPerRequest` for multimodal request limits when a model declares `input: [\"text\", \"image\"]`.\n\n## Adding Custom Models\n\nDo **not** edit `models.json` directly — it is auto-generated from the API. To customize:\n\n- **Override an existing model**: Add entries to `patch.json` (reasoning, compat, notes, maxTokens, etc.)\n- **Add new models not in the API**: Add entries to `custom-models.json`:\n\n```json\n[\n  {\n    \"id\": \"my-org/my-model\",\n    \"name\": \"My Custom Model\",\n    \"reasoning\": false,\n    \"input\": [\"text\"],\n    \"cost\": { \"input\": 0, \"output\": 0, \"cacheRead\": 0, \"cacheWrite\": 0 },\n    \"contextWindow\": 131072,\n    \"maxTokens\": 16384,\n    \"baseUrl\": \"https://inference.makora.com/my-model-slug/v1\"\n  }\n]\n```\n\n## API Notes\n\n- Each model is accessible at `https://inference.makora.com/v1/chat/completions` (unified endpoint)\n- Models with a `baseUrl` override use their per-slug endpoint instead\n- The API is OpenAI-compatible (chat completions format)\n- All models are hosted on vLLM\n- The `developer` role is not supported (prompts are silently dropped); `supportsDeveloperRole` is set to `false` for all models\n\n## vLLM Caveats\n\nThese issues are common to all vLLM-hosted providers and affect Makora models:\n\n- **GLM 5.1 tool calling**: vLLM's streaming tool call handling is broken for GLM — the model outputs Zhipu's native `<tool_call>` XML format as raw text. The `message_end` hook parses this into `toolCall` blocks so OMP / pi can execute the tools. A `context` hook then strips `tool_calls` from assistant messages before follow-up requests, converting them back to `<tool_call>` text to avoid a ZAI/vLLM server crash (500: `'str object' has no attribute 'items'`) that occurs when any assistant message contains a `tool_calls` field. If upstream fixes both the streaming parser and the 500 crash, the `message_end` hook gracefully skips (existing valid `toolCall` blocks are preserved), and the `context` hook's text-stripping is harmless (GLM natively understands `<tool_call>` text).\n\n  - **Kimi K2.6 + Qwen 3.6 tool calling**: vLLM's streaming tool call handling is broken or missing for these models. The `before_provider_request` hook sets `tool_choice: \"none\"` and `skip_special_tokens: false` so the model's tool call tokens pass through as plain text. The `message_end` hook then re-parses into `toolCall` blocks:\n\n    - **Kimi K2.6**: Uses `<|tool_call_begin|>...<|tool_call_end|>` tokens. Makora's vLLM is missing both `--enable-auto-tool-choice` and `--tool-call-parser` for this model.\n    - **Qwen 3.6**: Uses hermes-style `<function=...>` XML, sometimes with `█` delimiters. Same vLLM flag limitation as Kimi.\n\n- **GLM 5.1 CoT leak**: On some vLLM builds, disabling reasoning may still leak chain-of-thought into `content` terminated by a ``` marker. See [vllm-project/vllm#31319](https://github.com/vllm-project/vllm/issues/31319).\n\n- **DeepSeek V4 reasoning**: The official DeepSeek API uses `thinking: { type: \"enabled\" }` which Makora's vLLM silently ignores. The `before_provider_request` hook rewrites the payload to use vLLM-native params instead:\n  - **DS V4 Pro**: `chat_template_kwargs: { thinking: true }`. Returns `reasoning_content`.\n  - **DS V4 Flash**: `include_reasoning: true` + `chat_template_kwargs: { thinking: true }`. `include_reasoning` alone returns `reasoning: null` on this vLLM build — both params are required. Returns `reasoning`.\n- **GLM 5.1 reasoning**: Returns `reasoning_content` (not `reasoning`). OMP / pi's OpenAI completions handler checks `reasoning_content` first, so this is handled correctly.\n- **MiniMax M3 reasoning**: Uses `chat_template_kwargs.enable_thinking` to toggle thinking (not `chat_template_kwargs.thinking` like DeepSeek). The `before_provider_request` hook rewrites the DeepSeek API-style `thinking` param into vLLM-native `chat_template_kwargs: { enable_thinking: true }`. Returns `reasoning_content` field.\n","readmeFilename":"README.md","homepage":"https://github.com/ryan-brosas/omp-makora-provider#readme","repository":{"type":"git","url":"git+https://github.com/ryan-brosas/omp-makora-provider.git"},"bugs":{"url":"https://github.com/ryan-brosas/omp-makora-provider/issues"}}