{"_id":"@beaconbay/ck-search","_rev":"3-e3b4aadb73586a07e2853b8c92c35678","name":"@beaconbay/ck-search","dist-tags":{"latest":"0.7.11"},"versions":{"0.7.0":{"name":"@beaconbay/ck-search","version":"0.7.0","keywords":["code-search","semantic-search","search","grep","ripgrep","code","ast","tree-sitter","cli","rust","embeddings","ai","ml","vector-search","developer-tools","devtools","code-analysis","refactoring","mcp","claude"],"author":{"url":"https://github.com/BeaconBay","name":"Mike Renwick"},"license":"MIT OR Apache-2.0","_id":"@beaconbay/ck-search@0.7.0","maintainers":[{"name":"runonthespot","email":"mike.renwick@gmail.com"}],"homepage":"https://github.com/BeaconBay/ck#readme","bugs":{"url":"https://github.com/BeaconBay/ck/issues"},"bin":{"ck":"cli/ck.js"},"dist":{"shasum":"9a6f69e443c01e8b995d3da489e77524574a5376","tarball":"https://registry.npmjs.org/@beaconbay/ck-search/-/ck-search-0.7.0.tgz","fileCount":7,"integrity":"sha512-xU7Ltv0QEcgI+XE91XlZgRCtkisrrlFv9UNSTiRzt6+a1ZpBWROq6jwadlz+Ic9sipA2Gn8DbLDGfkz6vBl/vg==","signatures":[{"sig":"MEQCICIIBPEIwajtSl8UVtenEI+lVLxjJWVAFX/9XLa1beriAiBjFouUs1oyGA2d9I/H8LoKdGUVO4YmwP20cSBjs28TSg==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":41048},"main":"index.js","engines":{"node":">=16"},"funding":{"url":"https://github.com/sponsors/BeaconBay","type":"github"},"gitHead":"40e3d01c3e0f130f0c37821a00f93ffaac850af1","scripts":{"test":"node scripts/test.js","install":"node scripts/install.js"},"_npmUser":{"name":"runonthespot","email":"mike.renwick@gmail.com"},"repository":{"url":"git+https://github.com/BeaconBay/ck.git","type":"git"},"_npmVersion":"10.9.2","description":"Semantic code search - grep that understands what you're looking for. Find code by meaning, not just keywords. Powered by AI embeddings with tree-sitter AST parsing.","directories":{},"_nodeVersion":"22.14.0","dependencies":{"tar":"^6.2.1"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/ck-search_0.7.0_1760563112739_0.44681918330131354","host":"s3://npm-registry-packages-npm-production"}},"0.7.9":{"name":"@beaconbay/ck-search","version":"0.7.9","keywords":["code-search","semantic-search","search","grep","ripgrep","code","ast","tree-sitter","cli","rust","embeddings","ai","ml","vector-search","developer-tools","devtools","code-analysis","refactoring","mcp","claude"],"author":{"url":"https://github.com/BeaconBay","name":"Mike Renwick"},"license":"MIT OR Apache-2.0","_id":"@beaconbay/ck-search@0.7.9","maintainers":[{"name":"runonthespot","email":"mike.renwick@gmail.com"}],"homepage":"https://github.com/BeaconBay/ck#readme","bugs":{"url":"https://github.com/BeaconBay/ck/issues"},"bin":{"ck":"cli/ck.js"},"dist":{"shasum":"a0b4e15e0f036a53a48788114805592dae5f29b3","tarball":"https://registry.npmjs.org/@beaconbay/ck-search/-/ck-search-0.7.9.tgz","fileCount":12,"integrity":"sha512-4M7jBrWdByuglhXKBAUveY1hoE6Cp7ktdyLKuAVK+YeixKYozG91rNXw9l0bu1M8QSylYm6qxO0DRklhLnk62g==","signatures":[{"sig":"MEUCIGGFAdXmmuY8LnkgTnfqhvq/4nS3Pk4+lsgMK54MdHIsAiEAtxGh5iP6SyDQhPYMxfMIQq44X2Ipv0qSfEEGcA5SbHA=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":67364347},"engines":{"node":">=18"},"funding":{"url":"https://github.com/sponsors/BeaconBay","type":"github"},"gitHead":"00364435a121bffc5c3c140538f562945c5c5736","scripts":{"test":"node scripts/test.js","install":"node scripts/install.js"},"_npmUser":{"name":"runonthespot","email":"mike.renwick@gmail.com"},"repository":{"url":"git+https://github.com/BeaconBay/ck.git","type":"git"},"_npmVersion":"11.14.1","description":"Semantic code search - grep that understands what you're looking for. Find code by meaning, not just keywords. Powered by AI embeddings with tree-sitter AST parsing.","directories":{},"_nodeVersion":"22.14.0","dependencies":{"tar":"^7.5.15"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/ck-search_0.7.9_1779634949631_0.10747586617852822","host":"s3://npm-registry-packages-npm-production"}},"0.7.11":{"name":"@beaconbay/ck-search","version":"0.7.11","description":"Semantic code search - grep that understands what you're looking for. Find code by meaning, not just keywords. Powered by AI embeddings with tree-sitter AST parsing.","license":"MIT OR Apache-2.0","author":{"name":"Mike Renwick","url":"https://github.com/BeaconBay"},"repository":{"type":"git","url":"git+https://github.com/BeaconBay/ck.git"},"bugs":{"url":"https://github.com/BeaconBay/ck/issues"},"homepage":"https://github.com/BeaconBay/ck#readme","funding":{"type":"github","url":"https://github.com/sponsors/BeaconBay"},"bin":{"ck":"cli/ck.js"},"scripts":{"install":"node scripts/install.js","test":"node scripts/test.js"},"engines":{"node":">=18"},"keywords":["code-search","semantic-search","search","grep","ripgrep","code","ast","tree-sitter","cli","rust","embeddings","ai","ml","vector-search","developer-tools","devtools","code-analysis","refactoring","mcp","claude"],"dependencies":{"tar":"^7.5.15"},"publishConfig":{"access":"public"},"gitHead":"23215108e4a3a36c8731a85429c4562ff4cbac2e","_id":"@beaconbay/ck-search@0.7.11","_nodeVersion":"22.14.0","_npmVersion":"11.14.1","dist":{"integrity":"sha512-wXBEtFXo8DPtmatrT+eW5yV2i/ecDo1AvyyUClk64y2d1EeQjCjfXH5A68ZPB6B4Bs0QLpRrZyYxPQJnvCvbKw==","shasum":"66f60862e628079e9859d050290975fa54bdeca6","tarball":"https://registry.npmjs.org/@beaconbay/ck-search/-/ck-search-0.7.11.tgz","fileCount":12,"unpackedSize":67364590,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIGm6ApIYrIJeTe1ENlgao8Obzh78nvlUwPx/JwZilMWoAiEA7RRzp1OjybO+QvvMM5kfTSknf3FAzXvUXBaNKviHICc="}]},"_npmUser":{"name":"runonthespot","email":"mike.renwick@gmail.com"},"directories":{},"maintainers":[{"name":"runonthespot","email":"mike.renwick@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/ck-search_0.7.11_1779645677313_0.09241237461374086"},"_hasShrinkwrap":false}},"time":{"created":"2025-10-15T21:18:32.653Z","modified":"2026-05-24T18:01:18.124Z","0.7.0":"2025-10-15T21:18:32.937Z","0.7.9":"2026-05-24T15:02:30.153Z","0.7.11":"2026-05-24T18:01:18.000Z"},"bugs":{"url":"https://github.com/BeaconBay/ck/issues"},"author":{"name":"Mike Renwick","url":"https://github.com/BeaconBay"},"license":"MIT OR Apache-2.0","homepage":"https://github.com/BeaconBay/ck#readme","keywords":["code-search","semantic-search","search","grep","ripgrep","code","ast","tree-sitter","cli","rust","embeddings","ai","ml","vector-search","developer-tools","devtools","code-analysis","refactoring","mcp","claude"],"repository":{"type":"git","url":"git+https://github.com/BeaconBay/ck.git"},"description":"Semantic code search - grep that understands what you're looking for. Find code by meaning, not just keywords. Powered by AI embeddings with tree-sitter AST parsing.","maintainers":[{"name":"runonthespot","email":"mike.renwick@gmail.com"}],"readme":"# ck - Semantic Code Search\n\n[![CI](https://github.com/BeaconBay/ck/actions/workflows/ci.yaml/badge.svg)](https://github.com/BeaconBay/ck/actions/workflows/ci.yaml)\n[![Crates.io](https://img.shields.io/crates/v/ck-search.svg)](https://crates.io/crates/ck-search)\n[![Downloads](https://img.shields.io/crates/d/ck-search.svg)](https://crates.io/crates/ck-search)\n[![License](https://img.shields.io/badge/license-MIT%2FApache--2.0-blue.svg)](LICENSE-MIT)\n[![MSRV](https://img.shields.io/badge/rust-1.88%2B-blue.svg)](https://www.rust-lang.org)\n[![Documentation](https://img.shields.io/badge/docs-beaconbay.github.io%2Fck-blue)](https://beaconbay.github.io/ck/)\n\n**ck (seek)** finds code by meaning, not just keywords. It's grep that understands what you're looking for — search for \"error handling\" and find try/catch blocks, error returns, and exception handling code even when those exact words aren't present.\n\n## 🚀 Quick Start\n\n```bash\n# Install from crates.io\ncargo install ck-search\n\n# Just search — ck builds and updates indexes automatically\nck --sem \"error handling\" src/\nck --sem \"authentication logic\" src/\nck --sem \"database connection pooling\" src/\n\n# Traditional grep-compatible search still works\nck -n \"TODO\" *.rs\nck -R \"TODO|FIXME\" .\n\n# Combine both: semantic relevance + keyword filtering\nck --hybrid \"connection timeout\" src/\n```\n\n> **📚 [Full Documentation](https://beaconbay.github.io/ck/)** — Installation guides, tutorials, feature deep-dives, and API reference\n\n## ✨ Headline Features\n\n### 🤖 **AI Agent Integration (MCP Server)**\nConnect ck directly to Claude Desktop, Cursor, or any MCP-compatible AI client for seamless code search integration:\n\n```bash\n# Start MCP server for AI agent integration\nck --serve\n```\n\n**Claude Desktop Setup:**\n\n```bash\n# Install via Claude Code CLI (recommended)\nclaude mcp add ck-search -s user -- ck --serve\n\n# Note: You may need to restart Claude Code after installation\n# Verify installation with:\nclaude mcp list  # or use /mcp in Claude Code\n```\n\n**Manual Configuration (alternative):**\n```json\n{\n  \"mcpServers\": {\n    \"ck\": {\n      \"command\": \"ck\",\n      \"args\": [\"--serve\"],\n      \"cwd\": \"/path/to/your/codebase\"\n    }\n  }\n}\n```\n\n**Tool Permissions:** When prompted by Claude Code, approve permissions for ck-search tools (semantic_search, regex_search, hybrid_search, etc.)\n\n**Available MCP Tools:**\n- `semantic_search` - Find code by meaning using embeddings\n- `regex_search` - Traditional grep-style pattern matching\n- `hybrid_search` - Combined semantic and keyword search\n- `index_status` - Check indexing status and metadata\n- `reindex` - Force rebuild of search index\n- `health_check` - Server status and diagnostics\n\n**Built-in Pagination:** Handles large result sets gracefully with page_size controls, cursors, and snippet length management.\n\n### 🎨 **Interactive TUI (Terminal User Interface)**\nLaunch an interactive search interface with real-time results and multiple preview modes:\n\n```bash\n# Start TUI for current directory\nck --tui\n\n# Start with initial query\nck --tui \"error handling\"\n```\n\n**Features:**\n- **Multiple Search Modes**: Toggle between Semantic, Regex, and Hybrid search with `Tab`\n- **Preview Modes**: Switch between Heatmap, Syntax highlighting, and Chunk view with `Ctrl+V`\n- **View Options**: Toggle between snippet and full-file view with `Ctrl+F`\n- **Multi-select**: Select multiple files with `Ctrl+Space`, open all in editor with `Enter`\n- **Search History**: Navigate with `Ctrl+Up/Down`\n- **Editor Integration**: Opens files in `$EDITOR` with line numbers (Vim, VS Code, Cursor, etc.)\n- **Progress Tracking**: Live indexing progress with file and chunk counts\n- **Config Persistence**: Preferences saved to `~/.config/ck/tui.json`\n\nSee [TUI.md](TUI.md) for keyboard shortcuts and detailed usage.\n\n### 🔍 **Semantic Search**\nFind code by concept, not keywords. Understands synonyms, related terms, and conceptual similarity:\n\n```bash\n# These find related code even without exact keywords:\nck --sem \"retry logic\"           # finds backoff, circuit breakers\nck --sem \"user authentication\"   # finds login, auth, credentials\nck --sem \"data validation\"       # finds sanitization, type checking\n\n# Get complete functions/classes containing matches\nck --sem --full-section \"error handling\"  # returns entire functions\n```\n\n### ⚡ **Drop-in grep Compatibility**\nAll your muscle memory works. Same flags, same behavior, same output format:\n\n```bash\nck -i \"warning\" *.log              # Case-insensitive\nck -n -A 3 -B 1 \"error\" src/       # Line numbers + context\nck -l \"error\" src/                  # List files with matches only\nck -L \"TODO\" src/                   # List files without matches\nck -R --exclude \"*.test.js\" \"bug\"  # Recursive with exclusions\n```\n\n### 🎯 **Hybrid Search**\nCombine keyword precision with semantic understanding using Reciprocal Rank Fusion:\n\n```bash\nck --hybrid \"async timeout\" src/    # Best of both worlds\nck --hybrid --scores \"cache\" src/   # Show relevance scores with color highlighting\nck --hybrid --threshold 0.02 query  # Filter by minimum relevance\n```\n\n### ⚙️ **Automatic Delta Indexing with Chunk-Level Caching**\nSemantic and hybrid searches transparently create and refresh their indexes before running. The first search builds what it needs; subsequent searches intelligently reuse cached embeddings:\n\n- **Chunk-level incremental indexing**: Only changed chunks are re-embedded (80-90% cache hit rate for typical code changes)\n- **Content-aware invalidation**: Doc comments and whitespace changes properly invalidate cache\n- **Model consistency**: Prevents silent embedding corruption when switching models\n- **Smart caching**: Hash-based invalidation using blake3(text + trivia) for reliable change detection\n\n### 📁 **Smart File Filtering**\nAutomatically excludes cache directories, build artifacts, and respects `.gitignore` and `.ckignore` files:\n\n```bash\n# ck respects multiple exclusion layers (all are additive):\nck \"pattern\" .                           # Uses .gitignore + .ckignore + defaults\nck --no-ignore \"pattern\" .               # Skip .gitignore (still uses .ckignore)\nck --no-ckignore \"pattern\" .             # Skip .ckignore (still uses .gitignore)\nck --exclude \"dist\" --exclude \"logs\" .   # Add custom exclusions\n\n# .ckignore file (created automatically on first index):\n# - Excludes images, videos, audio, binaries, archives by default\n# - Excludes JSON/YAML config files (issue #27)\n# - Uses same syntax as .gitignore (glob patterns, ! for negation)\n# - Persists across searches (issue #67)\n# - Located at repository root, editable for custom patterns\n\n# Exclusion patterns use .gitignore syntax:\nck --exclude \"node_modules\" .            # Exclude directory and all contents\nck --exclude \"*.test.js\" .                # Exclude files matching pattern\nck --exclude \"build/\" --exclude \"*.log\" . # Multiple exclusions\n# Note: Patterns are relative to the search root\n```\n\n**Why .ckignore?** While `.gitignore` handles version control exclusions, many files that *should* be in your repo aren't ideal for semantic search. Config files (`package.json`, `tsconfig.json`), images, videos, and data files add noise to search results and slow down indexing. `.ckignore` lets you focus semantic search on actual code while keeping everything else in git. Think of it as \"what should I search\" vs \"what should I commit\".\n\n## 🛠 Advanced Usage\n\n### AI Agent Integration\n\n#### MCP Server (Recommended)\n```python\n# Example usage in AI agents\nresponse = await client.call_tool(\"semantic_search\", {\n    \"query\": \"authentication logic\",\n    \"path\": \"/path/to/code\",\n    \"page_size\": 25,\n    \"top_k\": 50,           # Limit total results (default: 100 for MCP)\n    \"snippet_length\": 200\n})\n\n# Handle pagination\nif response[\"pagination\"][\"next_cursor\"]:\n    next_response = await client.call_tool(\"semantic_search\", {\n        \"query\": \"authentication logic\",\n        \"path\": \"/path/to/code\",\n        \"cursor\": response[\"pagination\"][\"next_cursor\"]\n    })\n```\n\n#### JSONL Output (Custom Workflows)\nPerfect structured output for LLMs, scripts, and automation:\n\n```bash\n# JSONL format - one JSON object per line (recommended for agents)\nck --jsonl --sem \"error handling\" src/\nck --jsonl --no-snippet \"function\" .        # Metadata only\nck --jsonl --topk 5 --threshold 0.7 \"auth\"  # High-confidence results\n\n# Traditional JSON (single array)\nck --json --sem \"error handling\" src/ | jq '.file'\n```\n\n**Why JSONL for AI agents?**\n- ✅ **Streaming friendly**: Process results as they arrive\n- ✅ **Memory efficient**: Parse one result at a time\n- ✅ **Error resilient**: One malformed line doesn't break entire response\n- ✅ **Standard format**: Used by OpenAI API, Anthropic API, and modern ML pipelines\n\n### Search & Filter Options\n\n```bash\n# Threshold filtering\nck --sem --threshold 0.7 \"query\"           # Only high-confidence matches\nck --hybrid --threshold 0.01 \"concept\"     # Low-confidence (exploration)\n\n# Limit results\nck --sem --topk 5 \"authentication patterns\"\n\n# Complete code sections\nck --sem --full-section \"database queries\"  # Complete functions\nck --full-section \"class.*Error\" src/       # Complete classes (works with regex too)\n\n# Relevance scoring\nck --sem --scores \"machine learning\" docs/\n# [0.847] ./ai_guide.txt: Machine learning introduction...\n# [0.732] ./statistics.txt: Statistical learning methods...\n```\n\n\n### Language Coverage\n\n| Language | Indexing | Chunking | AST-aware | Notes |\n|----------|----------|----------|-----------|-------|\n| Markdown | ✅ | ✅ | ✅ | Headings, sections, code blocks |\n| Zig | ✅ | ✅ | ✅ | contributed by [@Nevon](https://github.com/Nevon) (PR #72) |\n\n### Model Selection\n\nChoose the right embedding model for your needs:\n\n```bash\n# Default: BGE-Small (fast, precise chunking)\nck --index .\n\n# Mixedbread xsmall: Optimized for local semantic search (4K context, 384 dims)\nck --index --model mxbai-xsmall .\n\n# Enhanced: Nomic V1.5 (8K context, optimal for large functions)\nck --index --model nomic-v1.5 .\n\n# Code-specialized: Jina Code (optimized for programming languages)\nck --index --model jina-code .\n```\n\n**Model Comparison:**\n- **`bge-small`** (default): 400-token chunks, fast indexing, good for most code\n- **`mxbai-xsmall`**: 4K context window, 384 dimensions, optimized for local inference (Mixedbread)\n- **`nomic-v1.5`**: 1024-token chunks with 8K model capacity, better for large functions\n- **`jina-code`**: 1024-token chunks with 8K model capacity, specialized for code understanding\n\n### Index Management\n\n```bash\n# Check index status\nck --status .\n\n# Clean up and rebuild / switch models\nck --clean .\nck --switch-model mxbai-xsmall .\nck --switch-model nomic-v1.5 .\nck --switch-model nomic-v1.5 --force .     # Force rebuild\n\n# Add single file to index\nck --add new_file.rs\n\n# File inspection (analyze chunking and token usage)\nck --inspect src/main.rs\nck --inspect --model bge-small src/main.rs  # Test different models\n```\n\n**Interrupting Operations:** Indexing can be safely interrupted with Ctrl+C. The partial index is saved, and the next operation will resume from where it stopped, only processing new or changed files.\n\n## 📚 Language Support\n\n| Language | Indexing | Tree-sitter Parsing | Semantic Chunking |\n|----------|----------|-------------------|------------------|\n| Python | ✅ | ✅ | ✅ Functions, classes |\n| JavaScript/TypeScript | ✅ | ✅ | ✅ Functions, classes, methods |\n| Rust | ✅ | ✅ | ✅ Functions, structs, traits |\n| Go | ✅ | ✅ | ✅ Functions, types, methods |\n| C | ✅ | ✅ | ✅ Functions, structs, enums, unions |\n| C++ | ✅ | ✅ | ✅ Classes, structs, namespaces, templates |\n| Markdown | ✅ | ✅ | ✅ Headings, sections, code blocks |\n| Ruby | ✅ | ✅ | ✅ Classes, methods, modules |\n| Haskell | ✅ | ✅ | ✅ Functions, types, instances |\n| C# | ✅ | ✅ | ✅ Classes, interfaces, methods |\n| Dart | ✅ | ✅ | ✅ Classes, mixins, methods |\n\n**Text Formats:** JSON, YAML, TOML, XML, HTML, CSS, shell scripts, SQL, log files, config files, and any other text format.\n\n**Smart Binary Detection:** Uses ripgrep-style content analysis, automatically indexing any text file while correctly excluding binary files.\n\n**Unsupported File Types:** Text files with unrecognized extensions (like `.org`, `.adoc`, etc.) are automatically indexed as plain text. ck detects text vs binary based on file contents, not extensions.\n\n## 🏗 Installation\n\n### From crates.io\n```bash\ncargo install ck-search\n```\n\n### From Source\n```bash\ngit clone https://github.com/BeaconBay/ck\ncd ck\ncargo install --path ck-cli\n```\n\n### Package Managers\n```bash\n# Currently available:\ncargo install ck-search    # ✅ Available now via crates.io\n\n# Coming soon:\nbrew install ck-search     # 🚧 In development (use cargo for now)\napt install ck-search      # 🚧 In development\n```\n\n## 💡 Examples\n\n### Finding Code Patterns\n```bash\n# Find authentication/authorization code\nck --sem \"user permissions\" src/\nck --sem \"access control\" src/\nck --sem \"login validation\" src/\n\n# Find error handling strategies\nck --sem \"exception handling\" src/\nck --sem \"error recovery\" src/\nck --sem \"fallback mechanisms\" src/\n\n# Find performance-related code\nck --sem \"caching strategies\" src/\nck --sem \"database optimization\" src/\nck --sem \"memory management\" src/\n```\n\n### Team Workflows\n```bash\n# Find related test files\nck --sem \"unit tests for authentication\" tests/\nck -l --sem \"test\" tests/           # List test files by semantic content\n\n# Identify refactoring candidates\nck --sem \"duplicate logic\" src/\nck --sem \"code complexity\" src/\nck -L \"test\" src/                   # Find source files without tests\n\n# Security audit\nck --hybrid \"password|credential|secret\" src/\nck --sem \"input validation\" src/\n```\n\n### Integration Examples\n```bash\n# Git hooks\ngit diff --name-only | xargs ck --sem \"TODO\"\n\n# CI/CD pipeline\nck --json --sem \"security vulnerability\" . | security_scanner.py\n\n# Code review prep\nck --hybrid --scores \"performance\" src/ > review_notes.txt\n\n# Documentation generation\nck --json --sem \"public API\" src/ | generate_docs.py\n```\n\n## ⚡ Performance\n\n**Field-tested on real codebases:**\n\n- **Indexing:** ~1M LOC in under 2 minutes\n- **Incremental indexing:** 80-90% cache hit rate for typical code changes (only changed chunks re-embedded)\n- **Search:** Sub-500ms queries on typical codebases\n- **Index size:** ~2x source code size with compression\n- **Memory:** Efficient streaming for large repositories\n- **Token precision:** HuggingFace tokenizers for exact model-specific token counting\n\n## 🔧 Architecture\n\nck uses a modular Rust workspace:\n\n- **`ck-cli`** - Command-line interface and MCP server\n- **`ck-tui`** - Interactive terminal user interface (ratatui-based)\n- **`ck-core`** - Shared types, configuration, and utilities\n- **`ck-engine`** - Search engine implementations (regex, semantic, hybrid)\n- **`ck-index`** - File indexing, hashing, and sidecar management\n- **`ck-embed`** - Text embedding providers (FastEmbed, API backends)\n- **`ck-ann`** - Approximate nearest neighbor search indices\n- **`ck-chunk`** - Text segmentation and language-aware parsing ([query-based chunking](docs/explanation/query-based-chunking.md))\n- **`ck-models`** - Model registry and configuration management\n\n### Index Storage\n\nIndexes are stored in `.ck/` directories alongside your code:\n\n```\nproject/\n├── src/\n├── docs/\n└── .ck/           # Semantic index (can be safely deleted)\n    ├── embeddings.json\n    ├── ann_index.bin\n    └── tantivy_index/\n```\n\nThe `.ck/` directory is a cache — safe to delete and rebuild anytime.\n\n## 🧪 Testing\n\n```bash\n# Run the full test suite\ncargo test --workspace\n\n# Test with each feature combination\ncargo hack test --each-feature --workspace\n```\n\n## 🤝 Contributing\n\nck is actively developed and welcomes contributions:\n\n1. **Issues:** Report bugs, request features\n2. **Code:** Submit PRs for bug fixes, new features\n3. **Documentation:** Improve examples, guides, tutorials\n4. **Testing:** Help test on different codebases and languages\n\n### Development Setup\n```bash\ngit clone https://github.com/BeaconBay/ck\ncd ck\ncargo build --workspace\ncargo test --workspace\n./target/debug/ck --index test_files/\n./target/debug/ck --sem \"test query\" test_files/\n```\n\n### CI Requirements\nBefore submitting a PR, ensure your code passes all CI checks:\n\n```bash\n# Format code (required)\ncargo fmt --all\n\n# Run clippy linter (required - must have no warnings)\ncargo clippy --workspace --all-features --all-targets -- -D warnings\n\n# Run tests (required)\ncargo test --workspace\n\n# Check minimum supported Rust version (MSRV)\ncargo hack check --each-feature --locked --rust-version --workspace\n```\n\nThe CI pipeline runs on Ubuntu, Windows, and macOS to ensure cross-platform compatibility.\n\n## 🗺 Roadmap\n\n### Current (v0.7+)\n- ✅ MCP (Model Context Protocol) server for AI agent integration\n- ✅ Chunk-level incremental indexing with smart embedding reuse\n- ✅ grep-compatible CLI with semantic search and file listing flags\n- ✅ FastEmbed integration with BGE models and enhanced model selection\n- ✅ File exclusion patterns and glob support\n- ✅ Threshold filtering and relevance scoring with visual highlighting\n- ✅ Tree-sitter parsing and intelligent chunking for 7+ languages\n- ✅ Complete code section extraction (`--full-section`)\n- ✅ Clean stdout/stderr separation for reliable scripting\n- ✅ Token-aware chunking with HuggingFace tokenizers\n- ✅ Published to crates.io (`cargo install ck-search`)\n\n### Next (v0.6+)\n- 🚧 Configuration file support\n- 🚧 Package manager distributions (brew, apt)\n- 🚧 Enhanced MCP tools (file writing, refactoring assistance)\n- 🚧 VS Code extension\n- 🚧 JetBrains plugin\n- 🚧 Additional language chunkers (Java, PHP, Swift)\n\n## ❓ FAQ\n\n**Q: How is this different from grep/ripgrep/silver-searcher?**\nA: ck includes all the features of traditional search tools, but adds semantic understanding. Search for \"error handling\" and find relevant code even when those exact words aren't used.\n\n**Q: Does it work offline?**\nA: Yes, completely offline. The embedding model runs locally with no network calls.\n\n**Q: How big are the indexes?**\nA: Typically 1-3x the size of your source code. The `.ck/` directory can be safely deleted to reclaim space.\n\n**Q: Is it fast enough for large codebases?**\nA: Yes. The first semantic search builds the index automatically; after that only changed files are reprocessed, keeping searches sub-second even on large projects.\n\n**Q: Can I use it in scripts/automation?**\nA: Absolutely. The `--json` and `--jsonl` flags provide structured output perfect for automated processing and AI agent integration.\n\n**Q: What about privacy/security?**\nA: Everything runs locally. No code or queries are sent to external services. The embedding model is downloaded once and cached locally.\n\n**Q: Where are the embedding models cached?**\nA: Models are cached in platform-specific directories:\n- Linux/macOS: `~/.cache/ck/models/`\n- Windows: `%LOCALAPPDATA%\\ck\\cache\\models\\`\n- Fallback: `.ck_models/models/` in current directory\n\n## 📄 License\n\nLicensed under either of:\n- Apache License, Version 2.0 ([LICENSE-APACHE](LICENSE-APACHE))\n- MIT License ([LICENSE-MIT](LICENSE-MIT))\n\nat your option.\n\n## 🙏 Credits\n\nBuilt with:\n- [Rust](https://rust-lang.org) - Systems programming language\n- [FastEmbed](https://github.com/Anush008/fastembed-rs) - Fast text embeddings\n- [Tantivy](https://github.com/quickwit-oss/tantivy) - Full-text search engine\n- [clap](https://github.com/clap-rs/clap) - Command line argument parsing\n\nInspired by the need for better code search tools in the age of AI-assisted development.\n\n---\n\n**Start finding code by what it does, not what it says.**\n\n```bash\ncargo install ck-search\nck --sem \"the code you're looking for\"\n```\n","readmeFilename":"README.md"}