{"_id":"@ashvardanian/uform","name":"@ashvardanian/uform","dist-tags":{"latest":"2.0.2"},"versions":{"2.0.2":{"name":"@ashvardanian/uform","type":"module","version":"2.0.2","description":"Pocket-Sized Multimodal AI for Content Understanding and Generation","dependencies":{"@huggingface/hub":"^0.14.8","@xenova/transformers":"^2.17.0","node-fetch":"^3.3.2","onnxruntime-node":"^1.17.0","onnxruntime-web":"^1.17.3"},"devDependencies":{"nodemon":"^2.0.15"},"scripts":{"start":"node javascript/encoders.mjs","test":"node javascript/encoders_test.js"},"main":"javascript/index.mjs","directories":{"doc":"docs"},"keywords":["AI","multimodal","content generation","huggingface"],"author":{"name":"Ash Vardanian, Unum Cloud"},"license":"Apache-2.0","_id":"@ashvardanian/uform@2.0.2","gitHead":"99acc2dcc6185925aba7a220a2ab3110ac8d9730","_nodeVersion":"20.12.2","_npmVersion":"10.5.0","dist":{"integrity":"sha512-Ou7XXkK1I9fodY4gYCbyJvn+P41oBIUEJGHOl8Aw7B1WaZFJPG/an6p0C4An44/Dl9q52YDFKqhfzz3iMKV7yg==","shasum":"668a0f961ebc2d07ab5950e64ae6fa050b892dfe","tarball":"https://registry.npmjs.org/@ashvardanian/uform/-/uform-2.0.2.tgz","fileCount":5,"unpackedSize":36659,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEUCIGlSTlAXZg9oc1KRIlqOK1rwJy8w1Q6nq6jApHEhhzElAiEA/A8C1PXgdA3DIfOUUGOFrgwsKao27RDN3u2TEuwLrmw="}]},"_npmUser":{"name":"ashvardanian","email":"ashvardanian@gmail.com"},"maintainers":[{"name":"ashvardanian","email":"ashvardanian@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/uform_2.0.2_1714016725995_0.6808738169746964"},"_hasShrinkwrap":false}},"time":{"created":"2024-04-25T03:45:25.884Z","2.0.2":"2024-04-25T03:45:26.203Z","modified":"2024-04-25T03:45:26.489Z"},"maintainers":[{"name":"ashvardanian","email":"ashvardanian@gmail.com"}],"description":"Pocket-Sized Multimodal AI for Content Understanding and Generation","keywords":["AI","multimodal","content generation","huggingface"],"author":{"name":"Ash Vardanian, Unum Cloud"},"license":"Apache-2.0","readme":"<h1 align=\"center\">UForm</h1>\n<h3 align=\"center\">\nPocket-Sized Multimodal AI<br/>\nFor Content Understanding and Generation<br/>\n</h3>\n<br/>\n\n<p align=\"center\">\n<a href=\"https://discord.gg/jsMURnSFM2\"><img height=\"25\" src=\"https://github.com/unum-cloud/.github/raw/main/assets/discord.svg\" alt=\"Discord\"></a>\n&nbsp; &nbsp; &nbsp;\n<a href=\"https://www.linkedin.com/company/unum-cloud/\"><img height=\"25\" src=\"https://github.com/unum-cloud/.github/raw/main/assets/linkedin.svg\" alt=\"LinkedIn\"></a>\n&nbsp; &nbsp; &nbsp;\n<a href=\"https://twitter.com/unum_cloud\"><img height=\"25\" src=\"https://github.com/unum-cloud/.github/raw/main/assets/twitter.svg\" alt=\"Twitter\"></a>\n&nbsp; &nbsp; &nbsp;\n<a href=\"https://unum.cloud/post\"><img height=\"25\" src=\"https://github.com/unum-cloud/.github/raw/main/assets/blog.svg\" alt=\"Blog\"></a>\n&nbsp; &nbsp; &nbsp;\n<a href=\"https://github.com/unum-cloud/uform\"><img height=\"25\" src=\"https://github.com/unum-cloud/.github/raw/main/assets/github.svg\" alt=\"GitHub\"></a>\n</p>\n\n<p align=\"center\">\nMultimodal Embeddings from 64 to 768 Dimensions • 1B Parameter Chat\n<br/>\nShort Texts • Images • 🔜 Video Clips • 🔜 Long Documents\n<br/>\nONNX • CoreML • PyTorch\n<br/>\n<a href=\"https://github.com/unum-cloud/uform/blob/main/python/README.md\">Python</a>\n • \n<a href=\"https://github.com/unum-cloud/uform/blob/main/javascript/README.md\">JavaScript</a>\n • \n<a href=\"https://github.com/unum-cloud/uform/blob/main/swift/README.md\">Swift</a>\n</p>\n\n---\n\n![UForm Chat Preview](https://github.com/ashvardanian/usearch-images/blob/main/assets/uform-gen-preview.jpg?raw=true)\n\nWelcome to UForm, a __multimodal__ AI library that's as versatile as it is efficient.\nUForm [tiny embedding models](#encoder) will help you understand and search visual and textual content across various languages.\nUForm [small generative models](#decoder), on the other hand, don't only support conversational and chat use-cases, but are great for fast image captioning and Visual Question Answering (VQA).\nWith compact __custom pre-trained transformer models__, this can run anywhere from your server farm down to your smartphone.\n\n## Features\n\n- __Tiny Embeddings__: 64-dimensional [Matryoshaka][matryoshka]-style embeddings for extremely fast [search][usearch].\n- __Throughput__: Thanks to the small size, the inference speed is [2-4x faster](#speed) than competitors.\n- __Portable__: Models come with native ONNX support, making them easy to deploy on any platform.\n- __Quantization Aware__: Down-cast embeddings from `f32` to `i8` without losing much recall.\n- __Multilingual__: Trained on a balanced dataset, the recall is great across over 20 languages.\n\n[usearch]: https://github.com/unum-cloud/usearch\n[matryoshka]: https://arxiv.org/abs/2205.13147\n\n## Models\n\nFor accuracy and speed benchmarks refer to the [evaluation page](https://github.com/unum-cloud/uform/blob/main/BENCHMARKS.md).\n\n### Embedding Models\n\n<table style=\"width:100%; border-collapse:collapse;\">\n    <thead>\n        <tr>\n            <th>Model</th>\n            <th style=\"text-align:right;\">Parameters</th>\n            <th style=\"text-align:right;\">Languages</th>\n            <th style=\"text-align:right;\">Architecture</th>\n        </tr>\n    </thead>\n    <tbody>\n        <tr>\n            <td><code><a href=\"https://huggingface.co/unum-cloud/uform-vl-english-large/\">uform3-image-text-english-large</a></code>  🆕</td>\n            <td style=\"text-align:right;\">365 M</td>\n            <td style=\"text-align:right;\">1</td>\n            <td style=\"text-align:right;\">12 layer BERT, ViT-L/14</td>\n        </tr>\n        <tr>\n            <td><code><a href=\"https://huggingface.co/unum-cloud/uform-vl-english/\">uform3-image-text-english-base</a></code></td>\n            <td style=\"text-align:right;\">143 M</td>\n            <td style=\"text-align:right;\">1</td>\n            <td style=\"text-align:right;\">4 layer BERT, ViT-B/16</td>\n        </tr>\n        <tr>\n            <td><code><a href=\"https://huggingface.co/unum-cloud/uform-vl-english-small/\">uform3-image-text-english-small</a></code>  🆕</td>\n            <td style=\"text-align:right;\">79 M</td>\n            <td style=\"text-align:right;\">1</td>\n            <td style=\"text-align:right;\">4 layer BERT, ViT-S/16</td>\n        </tr>\n        <tr>\n            <td><code><a href=\"https://huggingface.co/unum-cloud/uform-vl-multilingual-v2/\">uform3-image-text-multilingual-base</a></code></td>\n            <td style=\"text-align:right;\">206M</td>\n            <td style=\"text-align:right;\">21</td>\n            <td style=\"text-align:right;\">12 layer BERT, ViT-B/16</td>\n        </tr>\n    </tbody>\n</table>\n\n### Generative Models\n\n<table style=\"width:100%; border-collapse:collapse;\">\n    <thead>\n        <tr>\n            <th>Model</th>\n            <th style=\"text-align:right;\">Parameters</th>\n            <th style=\"text-align:right;\">Purpose</th>\n            <th style=\"text-align:right;\">Architecture</th>\n        </tr>\n    </thead>\n    <tbody>\n        <tr>\n            <td><code><a href=\"https://huggingface.co/unum-cloud/uform-gen2-dpo/\">uform-gen2-dpo</a></code>  🆕</td>\n            <td style=\"text-align:right;\">1.2 B</td>\n            <td style=\"text-align:right;\">Chat, Image Captioning, VQA</td>\n            <td style=\"text-align:right;\">qwen1.5-0.5B, ViT-H/14</td>\n        </tr>\n        <tr>\n            <td><code><a href=\"https://huggingface.co/unum-cloud/uform-gen2-qwen-500m/\">uform-gen2-qwen-500m</a></code></td>\n            <td style=\"text-align:right;\">1.2 B</td>\n            <td style=\"text-align:right;\">Chat, Image Captioning, VQA</td>\n            <td style=\"text-align:right;\">qwen1.5-0.5B, ViT-H/14</td>\n        </tr>\n        <tr>\n            <td><code><a href=\"https://huggingface.co/unum-cloud/uform-gen/\">uform-gen</a></code> ⚠️</td>\n            <td style=\"text-align:right;\">1.5 B</td>\n            <td style=\"text-align:right;\">Image Captioning, VQA</td>\n            <td style=\"text-align:right;\">llama-1.3B, ViT-B/16</td>\n        </tr>\n    </tbody>\n</table>\n\n## Quick Start Examples\n\n### Embedding Models\n\nFirst, `pip install uform`.\nThen, load the model:\n\n```py\nfrom uform import get_model, Modality\n\nprocessors, models = get_model('unum-cloud/uform3-image-text-english-small')\n\nmodel_text = models[Modality.TEXT_ENCODER]\nmodel_image = models[Modality.IMAGE_ENCODER]\nprocessor_text = processors[Modality.TEXT_ENCODER]\nprocessor_image = processors[Modality.IMAGE_ENCODER]\n```\n\nEmbed images:\n\n```py\nimport requests\nfrom io import BytesIO\nfrom PIL import Image\n\nimage_url = 'https://media-cdn.tripadvisor.com/media/photo-s/1b/28/6b/53/lovely-armenia.jpg'\nimage_url = Image.open(BytesIO(requests.get(image_url).content))\nimage_data = processor_image(image)\nimage_features, image_embedding = model_image.encode(image_data, return_features=True)\n```\n\nEmbed queries:\n\n```py\ntext = 'a cityscape bathed in the warm glow of the sun, with varied architecture and a towering, snow-capped mountain rising majestically in the background'\ntext_data = processor_text(text)\ntext_features, text_embedding = model_text.encode(text_data, return_features=True)\n```\n\nFor more details check out:\n\n- Python docs on embedding models in [python/README.md](https://github.com/unum-cloud/uform/blob/main/python/README.md#embedding-models)\n- JavaScript docs on embedding models in [javascript/README.md](https://github.com/unum-cloud/uform/blob/main/javascript/README.md#embedding-models)\n- Swift docs on embedding models in [swift/README.md](https://github.com/unum-cloud/uform/blob/main/swift/README.md#embedding-models)\n\n### Generative Models\n\nThe generative models are natively compatible with \n\n```python\nfrom transformers import AutoModel, AutoProcessor\n\nmodel = AutoModel.from_pretrained('unum-cloud/uform-gen2-dpo', trust_remote_code=True)\nprocessor = AutoProcessor.from_pretrained('unum-cloud/uform-gen2-dpo', trust_remote_code=True)\n\nprompt = 'Question or Instruction'\nimage = Image.open('image.jpg')\n\ninputs = processor(text=[prompt], images=[image], return_tensors='pt')\n\nwith torch.inference_mode():\n     output = model.generate(\n        **inputs,\n        do_sample=False,\n        use_cache=True,\n        max_new_tokens=256,\n        eos_token_id=151645,\n        pad_token_id=processor.tokenizer.pad_token_id\n    )\nprompt_len = inputs['input_ids'].shape[1]\ndecoded_text = processor.batch_decode(output[:, prompt_len:])[0]\n```\n\nFor more details check out:\n\n- Python docs on generative models in [python/README.md](https://github.com/unum-cloud/uform/blob/main/python/README.md#generative-models)\n- JavaScript docs on generative models 🔜\n- Swift docs on generative models 🔜\n\n## Technical Details\n\n### Down-casting, Quantization, Matryoshka, and Slicing\n\nDepending on the application, the embeddings can be down-casted to smaller numeric representations without losing much recall.\nSwitching from `f32` to `f16` is recommended in almost all cases, unless you are running on very old hardware without half-precision support.\nSwitching to `i8` with linear scaling is also possible, but will be noticeable in the recall on larger collections with millions of searchable entries.\nSimilarly, for higher-dimensional embeddings (512 or 768), a common strategy is to quantize them into single-bit representations for faster search.\n\n```python\nimport numpy as np\n\nf32_embedding: np.ndarray = model.encode_text(text_data, return_features=False)\nf16_embedding: np.ndarray = f32_embedding.astype(np.float16)\ni8_embedding: np.ndarray = (f32_embedding * 127).astype(np.int8)\nb1_embedding: np.ndarray = np.packbits((f32_embedding > 0).astype(np.uint8))\n```\n\nAlternative approach to quantization is to use the Matryoshka embeddings, where the embeddings are sliced into smaller parts, and the search is performed in a hierarchical manner.\n\n```python\nimport numpy as np\n\nlarge_embedding: np.ndarray = model.encode_text(text_data, return_features=False)\nsmall_embedding: np.ndarray = large_embedding[:, :256]\ntiny_embedding: np.ndarray = large_embedding[:, :64]\n```\n\nBoth approaches are natively supported by the [USearch][github-usearch] vector-search engine and the [SimSIMD][github-simsimd] numerics libraries.\nWhen dealing with small collections (up to millions of entries) and looking for low-latency cosine distance calculations, you can [achieve 5x-2500x performance improvement][report-simsimd] over Torch, NumPy, SciPy, and vanilla Python using SimSIMD.\n\n```python\nfrom simsimd import cosine, hamming\n\ndistance: float = cosine(f32_embedding, f32_embedding) # 32x SciPy performance on Apple M2 CPU\ndistance: float = cosine(f16_embedding, f16_embedding) # 79x SciPy performance on Apple M2 CPU\ndistance: float = cosine(i8_embedding, i8_embedding) # 133x SciPy performance on Apple M2 CPU\ndistance: float = hamming(b1_embedding, b1_embedding) # 17x SciPy performance on Apple M2 CPU\n```\n\nSimilarly, when dealing with large collections (up to billions of entries per server) and looking for high-throughput search, you can [achieve 100x performance improvement][report-usearch] over FAISS and other vector-search solutions using USearch.\nHere are a couple of examples:\n\n```python\nfrom usearch.index import Index\n\nf32_index = Index(ndim=64, metric='cos', dtype='f32') # for Matryoshka embeddings\nf16_index = Index(ndim=64, metric='cos', dtype='f16') # for Matryoshka embeddings\ni8_index = Index(ndim=256, metric='cos', dtype='i8') # for quantized embeddings\nb1_index = Index(ndim=768, metric='hamming', dtype='b1') # for binary embeddings\n```\n\n[github-usearch]: https://github.com/unum-cloud/usearch\n[github-simsimd]: https://github.com/ashvardanian/simsimd\n[report-usearch]: https://www.unum.cloud/blog/2023-11-07-scaling-vector-search-with-intel\n[report-simsimd]: https://ashvardanian.com/posts/python-c-assembly-comparison/\n\n### Compact Packaging\n\nPyTorch is a heavy dependency to carry, especially if you run on Edge or IoT devices.\nUsing vanilla ONNX runtime, one can significantly reduce memory consumption and deployment latency.\n\n```sh\n$ conda create -n uform_torch python=3.10 -y\n$ conda create -n uform_onnx python=3.10 -y\n$ conda activate uform_torch && pip install -e \".[torch]\" && conda deactivate\n$ conda activate uform_onnx && pip install -e \".[onnx]\" && conda deactivate\n$ du -sh $(conda info --envs | grep 'uform_torch' | awk '{print $2}')\n> 5.2G    ~/conda/envs/uform_torch\n$ du -sh $(conda info --envs | grep 'uform_onnx' | awk '{print $2}')\n> 461M    ~/conda/envs/uform_onnx\n```\n\nMost of that weight can be further reduced down to 100 MB for both the model and the runtime.\nYou can pick one of many supported [ONNX execution providers][onnx-providers], which includes XNNPACK, CUDA and TensorRT for Nvidia GPUs, OpenVINO on Intel, DirectML on Windows, ROCm on AMD, CoreML on Apple devices, and more to come.\n\n[onnx-providers]: https://onnxruntime.ai/docs/execution-providers/\n\n### Multimodal Chat in CLI\n\nThe generative models can be used for chat-like experiences in the command line.\nFor that, you can use the `uform-chat` CLI tool, which is available in the UForm package.\n\n```bash\n$ pip install uform\n$ uform-chat --model unum-cloud/uform-gen2-dpo --image=zebra.jpg\n$ uform-chat --model unum-cloud/uform-gen2-dpo \\\n>     --image=\"https://bit.ly/3tIVg9M\" \\\n>     --device=\"cuda:0\" \\\n>     --fp16\n```\n","readmeFilename":"README.md"}