{"_id":"@arclabs561/rank-refine","name":"@arclabs561/rank-refine","dist-tags":{"latest":"0.7.36"},"versions":{"0.7.36":{"name":"@arclabs561/rank-refine","collaborators":["Arc <attobop@gmail.com>"],"description":"SIMD-accelerated MaxSim (ColBERT/ColPali), cosine similarity, diversity (MMR/DPP), token alignment/highlighting for vector search and RAG. Supports text and multimodal late interaction.","version":"0.7.36","license":"MIT OR Apache-2.0","repository":{"type":"git","url":"git+https://github.com/arclabs561/rank-refine.git"},"main":"rank_refine.js","types":"rank_refine.d.ts","keywords":["vector-search","similarity","colbert","colpali","rag","simd","multimodal","late-interaction","maxsim","token-alignment"],"gitHead":"79a7c4361882a85624ec468ab4303bdc9857893a","_id":"@arclabs561/rank-refine@0.7.36","bugs":{"url":"https://github.com/arclabs561/rank-refine/issues"},"homepage":"https://github.com/arclabs561/rank-refine#readme","_nodeVersion":"20.19.5","_npmVersion":"11.6.4","dist":{"integrity":"sha512-jeeiIMfQW8EyM+ZX4VLvEx5z8EkZqhjCUFWJ57kmwZwJHjcxCZZQZsLmPZDECA5ndEdFtQywuH3tIra11KIWkw==","shasum":"210c10b6a03a5ee1aff8ae682099fba6aa983d38","tarball":"https://registry.npmjs.org/@arclabs561/rank-refine/-/rank-refine-0.7.36.tgz","fileCount":5,"unpackedSize":16497,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIG0KCLz6rSExEHqzQGXlTdj9VzJu+kZgjGIDjUlK/5EdAiBZhzVj82NB8e5+3eQjBhRHNdo/p+jt3222GnfO4HkexQ=="}]},"_npmUser":{"name":"arclabs561","email":"femtobop@gmail.com"},"directories":{},"maintainers":[{"name":"arclabs561","email":"femtobop@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/rank-refine_0.7.36_1764715303172_0.11310224410817904"},"_hasShrinkwrap":false}},"time":{"created":"2025-12-02T22:41:43.036Z","0.7.36":"2025-12-02T22:41:43.367Z","modified":"2025-12-02T22:41:43.773Z"},"maintainers":[{"name":"arclabs561","email":"femtobop@gmail.com"}],"description":"SIMD-accelerated MaxSim (ColBERT/ColPali), cosine similarity, diversity (MMR/DPP), token alignment/highlighting for vector search and RAG. Supports text and multimodal late interaction.","homepage":"https://github.com/arclabs561/rank-refine#readme","keywords":["vector-search","similarity","colbert","colpali","rag","simd","multimodal","late-interaction","maxsim","token-alignment"],"repository":{"type":"git","url":"git+https://github.com/arclabs561/rank-refine.git"},"bugs":{"url":"https://github.com/arclabs561/rank-refine/issues"},"license":"MIT OR Apache-2.0","readme":"# rank-refine\n\nSIMD-accelerated similarity scoring for vector search and RAG. Provides MaxSim (ColBERT/ColPali), cosine similarity, diversity selection (MMR, DPP), token pooling, token-level alignment/highlighting, and Matryoshka refinement. Supports both text (ColBERT) and multimodal (ColPali) late interaction.\n\n[![CI](https://github.com/arclabs561/rank-refine/actions/workflows/ci.yml/badge.svg)](https://github.com/arclabs561/rank-refine/actions)\n[![Crates.io](https://img.shields.io/crates/v/rank-refine.svg)](https://crates.io/crates/rank-refine)\n[![Docs](https://docs.rs/rank-refine/badge.svg)](https://docs.rs/rank-refine)\n\n```\ncargo add rank-refine\n```\n\n## Why Late Interaction?\n\nDense retrieval encodes each document as a single vector. This works for broad matching but loses token-level alignment.\n\n**Problem**: A query like \"capital of France\" might match a document about \"France's economic capital\" with high similarity, even if it never mentions \"capital\" in the geographic sense.\n\n**Solution**: Late interaction (ColBERT-style) keeps one vector per token instead of pooling. At query time, each query token finds its best-matching document token, then we sum those matches.\n\n```\nDense:           \"the quick brown fox\" → [0.1, 0.2, ...]  (1 vector)\nLate Interaction: \"the quick brown fox\" → [[...], [...], [...], [...]]  (4 vectors)\n```\n\nThis preserves token-level semantics that single-vector embeddings lose. Useful for reranking in RAG pipelines.\n\n## What This Is\n\nScoring primitives for retrieval systems:\n\n| You need | This crate provides |\n|----------|---------------------|\n| Score pre-computed embeddings | `cosine`, `dot`, `maxsim` |\n| ColBERT/late interaction | `maxsim_vecs`, `maxsim_batch` |\n| Token-level alignment/highlighting | `maxsim_alignments`, `highlight_matches` |\n| Batch alignment/highlighting | `maxsim_alignments_batch`, `highlight_matches_batch` |\n| Alignment utilities | `top_k_alignments`, `filter_alignments`, `alignment_stats` |\n| Diversity selection | `mmr_cosine`, `dpp` |\n| Compress token embeddings | `pool_tokens`, `pool_tokens_adaptive` |\n| Two-stage refinement | `matryoshka::refine` |\n\n**What this is NOT**: embedding generation, model weights, or storage systems. This crate scores embeddings you provide; it does not run model inference. See [fastembed-rs](https://github.com/Anush008/fastembed-rs) for inference. Trait-based interfaces available for custom models (see `crossencoder` module).\n\n## Usage\n\n```rust\nuse rank_refine::simd::{cosine, maxsim_vecs};\n\n// Dense similarity\nlet score = cosine(&query_embedding, &doc_embedding);\n\n// Late interaction (ColBERT)\nlet score = maxsim_vecs(&query_tokens, &doc_tokens);\n```\n\n### Realistic Example\n\n```rust\nuse rank_refine::colbert;\n\n// Query: \"capital of France\" (32 tokens, 128-dim embeddings)\nlet query = vec![\n    vec![0.12, -0.45, 0.89, ...],  // \"capital\" token\n    vec![0.34, 0.67, -0.23, ...],  // \"of\" token\n    vec![0.78, -0.12, 0.45, ...],  // \"France\" token\n    // ... 29 more tokens\n];\n\n// Document: \"Paris is the capital of France\" (100 tokens)\nlet doc = vec![\n    vec![0.11, -0.44, 0.90, ...],  // \"Paris\" token\n    vec![0.35, 0.66, -0.24, ...],  // \"is\" token\n    // ... 98 more tokens\n];\n\n// MaxSim finds best matches for each query token\nlet score = colbert::maxsim_vecs(&query, &doc);\n// \"capital\" matches \"capital\" (0.95), \"France\" matches \"France\" (0.92)\n// Score = 0.95 + 0.92 + ... (sum of best matches per query token)\n```\n\n## API\n\n### Similarity (SIMD-accelerated)\n\n| Function | Input | Notes |\n|----------|-------|-------|\n| `cosine(a, b)` | `&[f32]` | Normalized, -1 to 1 |\n| `dot(a, b)` | `&[f32]` | Unnormalized |\n| `maxsim(q, d)` | `&[&[f32]]` | Sum of max similarities |\n| `maxsim_cosine(q, d)` | `&[&[f32]]` | Cosine variant |\n| `maxsim_weighted(q, d, w)` | `&[&[f32]], &[f32]` | Per-token weights |\n| `maxsim_batch(q, docs)` | `&[Vec<f32>], &[Vec<Vec<f32>>]` | Batch scoring |\n| `maxsim_alignments(q, d)` | `&[&[f32]]` | Token-level alignments |\n| `maxsim_alignments_batch(q, docs)` | `&[Vec<f32>], &[Vec<Vec<f32>>]` | Batch alignments |\n| `highlight_matches(q, d, t)` | `&[&[f32]], f32` | Highlighted token indices |\n| `highlight_matches_batch(q, docs, t)` | `&[Vec<f32>], &[Vec<Vec<f32>>], f32` | Batch highlights |\n| `top_k_alignments(aligns, k)` | `&[(usize, usize, f32)], usize` | Top-k alignments |\n| `filter_alignments(aligns, min)` | `&[(usize, usize, f32)], f32` | Filter by score |\n| `alignment_stats(aligns)` | `&[(usize, usize, f32)]` | Min/max/mean/sum stats |\n\n### Token Pooling\n\n| Function | Tokens Kept | Notes |\n|----------|-------------|-------|\n| `pool_tokens(t, 2)` | 50% | Safe default |\n| `pool_tokens(t, 4)` | 25% | Use `hierarchical` feature |\n| `pool_tokens_adaptive(t, f)` | varies | Auto-selects greedy vs ward |\n| `pool_tokens_with_protected(t, f, n)` | varies | Keeps first n tokens unpooled |\n\n### Diversity\n\n| Function | Algorithm |\n|----------|-----------|\n| `mmr_cosine(candidates, embeddings, config)` | Maximal Marginal Relevance |\n| `dpp(candidates, embeddings, config)` | Determinantal Point Process |\n\n### Token Alignment & Highlighting\n\n| Function | Purpose |\n|----------|---------|\n| `maxsim_alignments(query, doc)` | Get (query_idx, doc_idx, score) alignment pairs |\n| `highlight_matches(query, doc, threshold)` | Extract highlighted doc token indices |\n| `colbert::alignments(query, doc)` | Convenience wrapper for alignments |\n| `colbert::highlight(query, doc, threshold)` | Convenience wrapper for highlighting |\n\n### Utilities\n\n| Function | Purpose |\n|----------|---------|\n| `normalize_maxsim(score, qlen)` | Scale to [0,1] |\n| `softmax_scores(scores)` | Probability distribution |\n| `top_k_indices(scores, k)` | Top-k by score |\n| `blend(a, b, α)` | Linear interpolation |\n| `idf_weights(doc_freqs, total_docs)` | IDF weighting for query tokens |\n| `bm25_weights(doc_freqs, query_freqs, total_docs, k1)` | BM25-style weighting |\n| `extract_snippet_indices(alignments, context, max)` | Extract text snippet indices |\n| `patches_to_regions(patch_indices, w, h, patches)` | Convert patches to image regions |\n\n## How It Works\n\n### MaxSim (Late Interaction)\n\nMaxSim scores token-level alignment. For each query token, find its best-matching document token, then sum:\n\n$$\\text{score}(Q, D) = \\sum_{i=1}^{|Q|} \\max_{j=1}^{|D|} (q_i \\cdot d_j)$$\n\nwhere $|Q|$ is the number of query tokens and $|D|$ is the number of document tokens.\n\n**Visual example**:\n\n```\nQuery tokens:     [q1]  [q2]  [q3]\n                    \\    |    /\n                     \\   |   /     Each query token searches\n                      v  v  v      for its best match\nDocument tokens:  [d1] [d2] [d3] [d4]\n                   ↑         ↑\n                  0.9       0.8    (best matches)\n\nMaxSim = 0.9 + 0.8 + ... (sum of best matches)\n```\n\n**Example**: Query \"capital of France\" (2 tokens) vs document \"Paris is the capital of France\" (6 tokens):\n- Query token \"capital\" finds best match: `dot(\"capital\", \"capital\") = 0.95`\n- Query token \"France\" finds best match: `dot(\"France\", \"France\") = 0.92`\n- MaxSim = 0.95 + 0.92 = 1.87\n\nThis captures token-level alignment: \"capital\" and \"France\" both have strong matches, even if they appear in different parts of the document. Single-vector embeddings average these signals and lose precision.\n\n**Token-level alignment and highlighting**: Unlike single-vector embeddings, ColBERT can show exactly which document tokens match each query token. Use `maxsim_alignments()` to get alignment pairs, or `highlight_matches()` to extract highlighted token indices for snippet extraction.\n\n**Multimodal support (ColPali)**: The same alignment functions work for vision-language retrieval. In ColPali-style systems, image patches are treated as \"tokens\"—query text tokens align with image patch embeddings. This enables visual snippet extraction: identify which image regions (patches) are relevant to a query, then extract those regions as visual snippets for display. Use `patches_to_regions()` to convert patch indices to pixel coordinates.\n\n**Query augmentation with [MASK] tokens**: ColBERT uses [MASK] tokens for soft query expansion. These tokens are added during encoding and should be weighted lower (typically 0.2-0.4) than original query tokens when using `maxsim_weighted()`. See `examples/mask_token_weighting.rs` for a complete example.\n\n**IDF and BM25 weighting**: Use `idf_weights()` or `bm25_weights()` to compute importance weights for query tokens based on document frequency. These can be passed to `maxsim_weighted()` to boost rare terms and improve retrieval quality by ~2-5%.\n\n**When to use**:\n- Second-stage reranking (after dense retrieval)\n- Precision-critical applications (legal, medical)\n- Queries with multiple important terms\n\n**When not to use**:\n- First-stage retrieval (too slow for millions of docs)\n- Storage-constrained (typically 10-50x larger than dense, depends on document length)\n\n### MMR (Diversity)\n\n**Problem**: Top-k by relevance returns near-duplicates. A search for \"async programming\" might return 10 Python asyncio tutorials instead of examples in Python, Rust, JavaScript, and Go.\n\n**Solution**: Balance relevance with diversity by penalizing similarity to already-selected items:\n\n$$\\text{MMR}(d) = \\lambda \\cdot \\text{rel}(d) - (1-\\lambda) \\cdot \\max_{s \\in S} \\text{sim}(d, s)$$\n\nwhere:\n- $\\lambda$ is the relevance-diversity tradeoff parameter (range [0, 1])\n- $\\text{rel}(d)$ is the relevance score of document $d$\n- $S$ is the set of already-selected documents\n- $\\text{sim}(d, s)$ is the similarity between document $d$ and selected document $s$\n\n**Lambda parameter**:\n- `λ = 1.0`: Pure relevance (equivalent to top-k)\n- `λ = 0.5`: Balanced (common default for RAG)\n- `λ = 0.3`: Strong diversity (exploration mode)\n- `λ = 0.0`: Maximum diversity (ignore relevance)\n\n**When to use**:\n- RAG pipelines (diverse context helps LLMs)\n- Recommendation systems (avoid redundancy)\n- Search result diversification\n\n**When not to use**:\n- Known-item search (single correct answer)\n- When relevance is paramount\n\n### Token Pooling\n\n**Storage**: ColBERT stores one vector per token. For 10M documents with 100 tokens each:\n- Storage = 10M × 100 × 128 × 4 bytes = 512 GB (vs ~5 GB for dense: 10M × 128 × 4 bytes)\n\n**Token pooling**: Cluster similar document tokens and store only cluster centroids:\n\n```\nBefore:  [tok1] [tok2] [tok3] [tok4] [tok5] [tok6]  (6 vectors)\n         similar--^       similar------^\nAfter:   [mean(1,2)]  [tok3] [mean(4,5)] [tok6]    (4 vectors = 33% reduction)\n```\n\n**Why it works**: Many tokens are redundant. Function words cluster together, as do related content words. Merging similar tokens loses little discriminative information.\n\n| Factor | Tokens Kept | MRR@10 Loss | When to Use |\n|--------|-------------|-------------|-------------|\n| 2 | 50% | 0.1–0.3% | Default choice |\n| 3 | 33% | 0.5–1.0% | Good tradeoff |\n| 4 | 25% | 1.5–3.0% | Storage-constrained (use `hierarchical` feature) |\n| 8+ | 12% | 5–10% | Extreme storage constraints only |\n\nNumbers from MS MARCO dev (Clavie et al., 2024). Pool at index time. Re-score with unpooled embeddings at query time if needed.\n\n**When pooling hurts more**:\n- Long documents with many distinct concepts\n- Queries that need to distinguish between similar tokens\n- Short passages (less redundancy to exploit)\n\n## Benchmarks\n\nMeasured on Apple M3 Max with `cargo bench`:\n\n| Operation | Dim | Time |\n|-----------|-----|------|\n| `dot` | 128 | 13ns |\n| `dot` | 768 | 126ns |\n| `cosine` | 128 | 40ns |\n| `cosine` | 768 | 380ns |\n| `maxsim` | 32q×128d×128dim | 49μs |\n| `maxsim` (pooled 2x) | 32q×64d×128dim | 25μs |\n| `maxsim` (pooled 4x) | 32q×32d×128dim | 12μs |\n\nThese timings enable real-time reranking of 100-1000 candidates.\n\n## Vendoring\n\nIf you prefer not to add a dependency:\n\n- `src/simd.rs` is self-contained (~600 lines)\n- AVX2+FMA / NEON with portable fallback\n- No dependencies\n\n## Quick Decision Guide\n\nWhat problem are you solving?\n\n1. **Scoring embeddings** → Use `cosine` or `dot`\n   - Dense embeddings: `cosine(&query, &doc)`\n   - Pre-normalized: `dot(&query, &doc)`\n\n2. **Reranking with token-level precision** → Use `maxsim_vecs`\n   - After dense retrieval (100-1000 candidates)\n   - ColBERT-style token embeddings\n   - Precision-critical applications\n\n3. **Diversity selection** → Use `mmr_cosine` or `dpp`\n   - RAG pipelines (diverse context)\n   - Recommendation systems\n   - Search result diversification\n\n4. **Compressing token embeddings** → Use `pool_tokens`\n   - Storage-constrained deployments\n   - Factor 2-3 recommended (50-66% reduction)\n\n5. **Two-stage refinement with Matryoshka** → Use `matryoshka::refine`\n   - After coarse search using head dimensions\n   - Refine candidates using tail dimensions\n   - Matryoshka embeddings (hierarchical dimension packing)\n\n**When not to use**:\n- First-stage retrieval over millions of docs (use dense + ANN)\n- Very short documents (<10 tokens, little benefit over dense)\n- Latency-critical first-stage (dense embeddings are faster)\n- Storage-constrained without pooling (typically 10-50x larger than dense, depends on document length)\n\nSee [REFERENCE.md](REFERENCE.md) for algorithm details and edge cases.\n\n## Multimodal Support (ColPali)\n\nThis crate supports both text-only (ColBERT) and multimodal (ColPali) late interaction:\n\n- **Text-to-text**: Query text tokens align with document text tokens (standard ColBERT)\n- **Text-to-image**: Query text tokens align with image patch embeddings (ColPali-style)\n\nFor ColPali systems, document images are split into patches (e.g., 32×32 grid = 1024 patches per page).\nEach patch becomes a \"token\" embedding. The same `maxsim_alignments()` and `highlight_matches()`\nfunctions work for both text and multimodal retrieval, enabling visual snippet extraction.\n\n**Visual snippet extraction**: In ColPali, identifying which image patches match query tokens enables\nextracting visual regions (snippets) from document images. These snippets can be displayed to users\nshowing exactly which parts of a document image are relevant to their query.\n\nSee `examples/multimodal_alignment.rs` for a demonstration.\n\n## Features\n\n| Feature | Dependency | Purpose |\n|---------|------------|---------|\n| `hierarchical` | kodama | Ward's clustering (better at 4x+) |\n\n## See Also\n\n- [rank-fusion](https://crates.io/crates/rank-fusion): merge ranked lists (no embeddings)\n- [fastembed-rs](https://github.com/Anush008/fastembed-rs): generate embeddings\n- [DESIGN.md](DESIGN.md): architecture decisions\n- [REFERENCE.md](REFERENCE.md): algorithm reference\n\n## License\n\nMIT OR Apache-2.0\n","readmeFilename":"README.md","_rev":"1-ae4741fda61de65afc89b9cf4cd5903e"}