{"_id":"@bre1470/wllama","name":"@bre1470/wllama","dist-tags":{"latest":"2.3.7-b8860"},"versions":{"2.3.7-b8860":{"private":false,"version":"2.3.7-b8860","name":"@bre1470/wllama","license":"MIT","description":"Custom WebAssembly binding for llama.cpp - Enabling on-browser LLM inference","author":{"name":"Xuan Son NGUYEN","email":"contact@ngxson.com"},"repository":{"type":"git","url":"git+https://github.com/ngxson/wllama.git"},"homepage":"https://github.com/ngxson/wllama#readme","bugs":{"url":"https://github.com/ngxson/wllama/issues"},"main":"index.js","type":"module","types":"esm/index.d.ts","contributors":[{"name":"Sophie Bremer","email":"sophie.bremer@highsoft.com"}],"devDependencies":{"@vitest/browser":"^4.1.5","@vitest/browser-playwright":"^4.1.5","@vitest/browser-webdriverio":"^4.1.5","express":"^4.18.3","mime-types":"^2.1.35","playwright":"^1.49.0","prettier":"^3.3.3","terser":"^5.39.0","tsup":"^8.4.0","typedoc":"^0.27.2","typescript":"^5.4.2","webdriverio":"^9.4.1"},"directories":{"example":"examples"},"keywords":["wasm","webassembly","llama","llm","ai","rag","embeddings","generation"],"prettier":{"trailingComma":"es5","tabWidth":2,"semi":true,"singleQuote":true,"bracketSameLine":false},"scripts":{"serve":"node ./scripts/http_server.js","serve:mt":"MULTITHREAD=1 node ./scripts/http_server.js","clean":"rm -rf ./esm && rm -rf ./docs && rm -rf ./wasm","build:worker":"./scripts/build_worker.sh","build:glue":"node ./cpp/generate_glue_prototype.js","build:wasm":"./scripts/build_wasm.sh && npm run build:glue","build:tsup":"tsup src/index.ts --format cjs,esm --clean","build:minified":"terser esm/index.js -o esm/index.min.js --compress --mangle --source-map","build:typedef":"tsc --emitDeclarationOnly --declaration -p tsconfig.build.json","build":"npm run clean && npm run build:worker && npm run build:tsup && npm run build:minified && npm run build:typedef","postbuild":"./scripts/post_build.sh && npm run docs","docs":"typedoc --tsconfig tsconfig.build.json src/index.ts","upload":"npm run format && npm run build && npm publish --access public","format":"prettier --write .","test":"vitest","test:firefox":"BROWSER=firefox vitest","test:safari":"BROWSER=safari vitest"},"_id":"@bre1470/wllama@2.3.7-b8860","gitHead":"f3255fbaee5993ac6b5904a3903ee133683f7c44","_nodeVersion":"24.13.1","_npmVersion":"10.8.1","dist":{"integrity":"sha512-vfhtST/5R3b+tHo7rGLRpZRd1NkKJq3Vdmr0lpssQ1+pfSd+Fn//9Z9ME4x8Dbxz/NeuwXf0W/0k1kI+UCF3FA==","shasum":"ad9cfb5e08067ba771221a8233a81e739bca12ea","tarball":"https://registry.npmjs.org/@bre1470/wllama/-/wllama-2.3.7-b8860.tgz","fileCount":21,"unpackedSize":6447249,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIBVAa8pxsqNCH7aXqf5sibttuRFmOzJixv/TuVmgM4PUAiEA036UCpN8L/E+eqjPUWovPpOwUZv3quA+jjT3QD0K9Fo="}]},"_npmUser":{"name":"bre1470","email":"sophie.bremer@highsoft.com"},"maintainers":[{"name":"bre1470","email":"sophie.bremer@highsoft.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/wllama_2.3.7-b8860_1776785210920_0.23428339080625182"},"_hasShrinkwrap":false}},"time":{"created":"2026-04-21T15:26:50.630Z","2.3.7-b8860":"2026-04-21T15:26:51.332Z","modified":"2026-04-21T15:26:51.594Z"},"maintainers":[{"name":"bre1470","email":"sophie.bremer@highsoft.com"}],"description":"Custom WebAssembly binding for llama.cpp - Enabling on-browser LLM inference","homepage":"https://github.com/ngxson/wllama#readme","keywords":["wasm","webassembly","llama","llm","ai","rag","embeddings","generation"],"repository":{"type":"git","url":"git+https://github.com/ngxson/wllama.git"},"contributors":[{"name":"Sophie Bremer","email":"sophie.bremer@highsoft.com"}],"author":{"name":"Xuan Son NGUYEN","email":"contact@ngxson.com"},"bugs":{"url":"https://github.com/ngxson/wllama/issues"},"license":"MIT","readme":"# Custom wllama - WASM binding for llama.cpp\n\n\n## Features\n\n* Updated llama.cpp integration\n* Pre-built npm package [@bre1470/wllama](https://www.npmjs.com/package/@bre1470/wllama)\n\n\n## Usage\n\nInstall it:\n\n```bash\nnpm i @bre1470/wllama\n```\n\nThen, import the module:\n\n```ts\nimport { Wllama } from '@bre1470/wllama';\nlet wllamaInstance = new Wllama(WLLAMA_CONFIG_PATHS, ...);\n// (the rest is the same with earlier example)\n```\n\n\n## Development\n\n* `git config --global http.postBuffer 524288000`\n* Switch to desired llama.cpp in `.gitmodules`\n* `git submodule update --init --recursive`\n* `npm i`\n* `npm audit fix`\n* `git remote add original https://github.com/ngxson/wllama.git`\n* `git pull original master`\n* `git merge master`\n* `git submodule update --recursive`\n* Make sure to have **Docker Desktop** running!\n* `npm run build:wasm`\n* `npm run build`\n* `npm pack` or `npm publish --access public`\n\n---\n\n# wllama - Wasm binding for llama.cpp\n\n![](./README_banner.png)\n\nWebAssembly binding for [llama.cpp](https://github.com/ggerganov/llama.cpp)\n\n👉 [Try the demo app](https://huggingface.co/spaces/ngxson/wllama)\n\n📄 [Documentation](https://github.ngxson.com/wllama/docs/)\n\nFor changelog, please visit [releases page](https://github.com/ngxson/wllama/releases)\n\n> [!IMPORTANT]  \n> Version 2.0 is released 👉 [read more](./guides/intro-v2.md)\n\n![](./assets/screenshot_0.png)\n\n## Features\n\n- Typescript support\n- Can run inference directly on browser (using [WebAssembly SIMD](https://emscripten.org/docs/porting/simd.html)), no backend or GPU is needed!\n- No runtime dependency (see [package.json](./package.json))\n- High-level API: completions, embeddings\n- Low-level API: (de)tokenize, KV cache control, sampling control,...\n- Ability to split the model into smaller files and load them in parallel (same as `split` and `cat`)\n- Auto switch between single-thread and multi-thread build based on browser support\n- Inference is done inside a worker, does not block UI render\n- Pre-built npm package [@wllama/wllama](https://www.npmjs.com/package/@wllama/wllama)\n\nLimitations:\n- To enable multi-thread, you must add `Cross-Origin-Embedder-Policy` and `Cross-Origin-Opener-Policy` headers. See [this discussion](https://github.com/ffmpegwasm/ffmpeg.wasm/issues/106#issuecomment-913450724) for more details.\n- No WebGPU support, but maybe possible in the future\n- Max file size is 2GB, due to [size restriction of ArrayBuffer](https://stackoverflow.com/questions/17823225/do-arraybuffers-have-a-maximum-length). If your model is bigger than 2GB, please follow the **Split model** section below.\n\n## Code demo and documentation\n\n📄 [Documentation](https://github.ngxson.com/wllama/docs/)\n\nDemo:\n- Basic usages with completions and embeddings: https://github.ngxson.com/wllama/examples/basic/\n- Embedding and cosine distance: https://github.ngxson.com/wllama/examples/embeddings/\n- For more advanced example using low-level API, have a look at test file: [wllama.test.ts](./src/wllama.test.ts)\n\n## How to use\n\n### Use Wllama inside React Typescript project\n\nInstall it:\n\n```bash\nnpm i @wllama/wllama\n```\n\nThen, import the module:\n\n```ts\nimport { Wllama } from '@wllama/wllama';\nlet wllamaInstance = new Wllama(WLLAMA_CONFIG_PATHS, ...);\n// (the rest is the same with earlier example)\n```\n\nFor complete code example, see [examples/main/src/utils/wllama.context.tsx](./examples/main/src/utils/wllama.context.tsx)\n\nNOTE: this example only covers completions usage. For embeddings, please see [examples/embeddings/index.html](./examples/embeddings/index.html)\n\n### Prepare your model\n\n- It is recommended to split the model into **chunks of maximum 512MB**. This will result in slightly faster download speed (because multiple splits can be downloaded in parallel), and also prevent some out-of-memory issues.  \n  See the \"Split model\" section below for more details.\n- It is recommended to use quantized Q4, Q5 or Q6 for balance among performance, file size and quality. Using IQ (with imatrix) is **not** recommended, may result in slow inference and low quality.\n\n### Simple usage with ES6 module\n\nFor complete code, see [examples/basic/index.html](./examples/basic/index.html)\n\n```javascript\nimport { Wllama } from './esm/index.js';\n\n(async () => {\n  const CONFIG_PATHS = {\n    'single-thread/wllama.wasm': './esm/single-thread/wllama.wasm',\n    'multi-thread/wllama.wasm' : './esm/multi-thread/wllama.wasm',\n  };\n  // Automatically switch between single-thread and multi-thread version based on browser support\n  // If you want to enforce single-thread, add { \"n_threads\": 1 } to LoadModelConfig\n  const wllama = new Wllama(CONFIG_PATHS);\n  // Define a function for tracking the model download progress\n  const progressCallback =  ({ loaded, total }) => {\n    // Calculate the progress as a percentage\n    const progressPercentage = Math.round((loaded / total) * 100);\n    // Log the progress in a user-friendly format\n    console.log(`Downloading... ${progressPercentage}%`);\n  };\n  // Load GGUF from Hugging Face hub\n  // (alternatively, you can use loadModelFromUrl if the model is not from HF hub)\n  await wllama.loadModelFromHF(\n    'ggml-org/models',\n    'tinyllamas/stories260K.gguf',\n    {\n      progressCallback,\n    }\n  );\n  const outputText = await wllama.createCompletion(elemInput.value, {\n    nPredict: 50,\n    sampling: {\n      temp: 0.5,\n      top_k: 40,\n      top_p: 0.9,\n    },\n  });\n  console.log(outputText);\n})();\n```\n\nAlternatively, you can use the `*.wasm` files from CDN:\n\n```js\nimport WasmFromCDN from '@wllama/wllama/esm/wasm-from-cdn.js';\nconst wllama = new Wllama(WasmFromCDN);\n// NOTE: this is not recommended, only use when you can't embed wasm files in your project\n```\n\n### Split model\n\nCases where we want to split the model:\n- Due to [size restriction of ArrayBuffer](https://stackoverflow.com/questions/17823225/do-arraybuffers-have-a-maximum-length), the size limitation of a file is 2GB. If your model is bigger than 2GB, you can split the model into small files.\n- Even with a small model, splitting into chunks allows the browser to download multiple chunks in parallel, thus making the download process a bit faster.\n\nWe use `llama-gguf-split` to split a big gguf file into smaller files. You can download the pre-built binary via [llama.cpp release page](https://github.com/ggerganov/llama.cpp/releases):\n\n```bash\n# Split the model into chunks of 512 Megabytes\n./llama-gguf-split --split-max-size 512M ./my_model.gguf ./my_model\n```\n\nThis will output files ending with `-00001-of-00003.gguf`, `-00002-of-00003.gguf`, and so on.\n\nYou can then pass to `loadModelFromUrl` or `loadModelFromHF` the URL of the first file and it will automatically load all the chunks:\n\n```js\nconst wllama = new Wllama(CONFIG_PATHS, {\n  parallelDownloads: 5, // optional: maximum files to download in parallel (default: 3)\n});\nawait wllama.loadModelFromHF(\n  'ngxson/tinyllama_split_test',\n  'stories15M-q8_0-00001-of-00003.gguf'\n);\n```\n\n### Custom logger (suppress debug messages)\n\nWhen initializing Wllama, you can pass a custom logger to Wllama.\n\nExample 1: Suppress debug message\n\n```js\nimport { Wllama, LoggerWithoutDebug } from '@wllama/wllama';\n\nconst wllama = new Wllama(pathConfig, {\n  // LoggerWithoutDebug is predefined inside wllama\n  logger: LoggerWithoutDebug,\n});\n```\n\nExample 2: Add emoji prefix to log messages\n\n```js\nconst wllama = new Wllama(pathConfig, {\n  logger: {\n    debug: (...args) => console.debug('🔧', ...args),\n    log: (...args) => console.log('ℹ️', ...args),\n    warn: (...args) => console.warn('⚠️', ...args),\n    error: (...args) => console.error('☠️', ...args),\n  },\n});\n```\n\n## How to compile the binary yourself\n\nThis repository already come with pre-built binary from llama.cpp source code. However, in some cases you may want to compile it yourself:\n- You don't trust the pre-built one.\n- You want to try out latest - bleeding-edge changes from upstream llama.cpp source code.\n\nYou can use the commands below to compile it yourself:\n\n```shell\n# /!\\ IMPORTANT: Require having docker compose installed\n\n# Clone the repository with submodule\ngit clone --recurse-submodules https://github.com/ngxson/wllama.git\ncd wllama\n\n# Optionally, you can run this command to update llama.cpp to latest upstream version (bleeding-edge, use with your own risk!)\n# git submodule update --remote --merge\n\n# Install the required modules\nnpm i\n\n# Firstly, build llama.cpp into wasm\nnpm run build:wasm\n# Then, build ES module\nnpm run build\n```\n\n## TODO\n\n- Add support for LoRA adapter\n- Support GPU inference via WebGL\n- Support multi-sequences: knowing the resource limitation when using WASM, I don't think having multi-sequences is a good idea\n- Multi-modal: Waiting for refactoring LLaVA implementation from llama.cpp\n","readmeFilename":"README.md","_rev":"1-6214bd73e192d937a126e590402fbbea"}