{"_id":"@adhd/sox-ingest","_rev":"2-63a4b859117205fab642e32630f5529e","name":"@adhd/sox-ingest","dist-tags":{"latest":"0.1.1"},"versions":{"0.1.0":{"name":"@adhd/sox-ingest","version":"0.1.0","license":"MIT","_id":"@adhd/sox-ingest@0.1.0","maintainers":[{"name":"pseudosky","email":"skywinston.sk@gmail.com"}],"sox":{"area":"data","group":"ingest","concerns":["content-hash (SHA-256 of normalized content — used for graph-store dedup)","extractive summary (lead-N sentences — first summaryMaxSentences sentences, no scoring; zero LLM; content under 100 chars is returned unchanged via content.trim())","deterministic tag extraction (noun phrases, high-frequency terms, tagMaxCount)","chunking (maxChars / overlapChars sliding window)","per-chunk contentHash for dedup at the chunk level"],"invariants":["zero-LLM, zero-I/O, synchronous — ingest() is a pure function, always safe to call in the write path without latency budget concerns (it lives in core.ts and reaches no chunker; the AstChunker WASM grammar load IS disk I/O but is not on ingest()'s path)","deterministic + byte-reproducible — same input always produces the same hash, summary, and tags (no random or time-based components)","[BL-231] dist/core.js is the CJS-SAFE entry — its module graph must never contain a top-level await. dist/index.js (the root) is ESM-ONLY and cannot be require()d: it re-exports AstChunker, whose module-scope `await Parser.init()` is required to keep chunk()/estimate() synchronous. CommonJS consumers import @adhd/sox-ingest/core. Guarded by tools/test-bl231-cjs-boundary.mjs"],"entrypoints":["dist/index.js","dist/core.js"]},"dist":{"shasum":"fe206c59ec30797c1e5cc00ded65aef406c4976e","tarball":"https://registry.npmjs.org/@adhd/sox-ingest/-/sox-ingest-0.1.0.tgz","fileCount":31,"integrity":"sha512-CSqltEFoOe+kCyO17q2jKhzEjaboEop5TD8PNdnPdf0RxZlNcQ8S6is44IDt9rPy/BUARdKFF6tbfwbsXNG9fQ==","signatures":[{"sig":"MEYCIQDXdGz20oxwJHiNTcFT1gciO5xjuEtmCjIf2ha8xuaLBwIhAIEsatr8E0ABEAUzirpy+aBhingbRpPKHil5yqaDkHHR","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":102416},"main":"./dist/index.js","type":"module","_from":"file:adhd-sox-ingest-0.1.0.tgz","types":"./dist/index.d.ts","engines":{"node":">=20"},"exports":{".":{"types":"./dist/index.d.ts","default":"./dist/index.js"},"./core":{"types":"./dist/core.d.ts","default":"./dist/core.js"}},"private":false,"_npmUser":{"name":"pseudosky","email":"skywinston.sk@gmail.com"},"_resolved":"/private/tmp/soxpub2/adhd-sox-ingest-0.1.0.tgz","_integrity":"sha512-CSqltEFoOe+kCyO17q2jKhzEjaboEop5TD8PNdnPdf0RxZlNcQ8S6is44IDt9rPy/BUARdKFF6tbfwbsXNG9fQ==","_npmVersion":"11.6.2","description":"Write-path single-item transforms for the memory domain — content-hash (SHA-256), extractive summary (lead-N sentences, zero LLM), deterministic tag extraction, and chunking/normalization. Pure stateless functions; no storage deps.","directories":{},"_nodeVersion":"24.11.1","dependencies":{"web-tree-sitter":"0.25.10","tree-sitter-wasms":"0.1.13"},"publishConfig":{"access":"public"},"typesVersions":{"*":{"core":["./dist/core.d.ts"]}},"_hasShrinkwrap":false,"//typesVersions":"[BL-231] `exports` maps are invisible to TypeScript's legacy node10 moduleResolution, which CJS consumers (memory-core: module=CommonJS) are pinned to. typesVersions is the supported way to expose the ./core subpath's types to them. Runtime resolution still goes through `exports` above, for both Node and esbuild.","_npmOperationalInternal":{"tmp":"tmp/sox-ingest_0.1.0_1783616908746_0.7344550821119491","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"name":"@adhd/sox-ingest","version":"0.1.1","description":"Write-path single-item transforms for the memory domain — content-hash (SHA-256), extractive summary (lead-N sentences, zero LLM), deterministic tag extraction, and chunking/normalization. Pure stateless functions; no storage deps.","license":"MIT","private":false,"publishConfig":{"access":"public"},"engines":{"node":">=20"},"type":"module","main":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","default":"./dist/index.js"},"./core":{"types":"./dist/core.d.ts","default":"./dist/core.js"},"./package.json":"./package.json"},"//typesVersions":"[BL-231] `exports` maps are invisible to TypeScript's legacy node10 moduleResolution, which CJS consumers (memory-core: module=CommonJS) are pinned to. typesVersions is the supported way to expose the ./core subpath's types to them. Runtime resolution still goes through `exports` above, for both Node and esbuild.","typesVersions":{"*":{"core":["./dist/core.d.ts"]}},"sox":{"area":"data","group":"ingest","concerns":["content-hash (SHA-256 of normalized content — used for graph-store dedup)","extractive summary (lead-N sentences — first summaryMaxSentences sentences, no scoring; zero LLM; content under 100 chars is returned unchanged via content.trim())","deterministic tag extraction (noun phrases, high-frequency terms, tagMaxCount)","chunking (maxChars / overlapChars sliding window)","per-chunk contentHash for dedup at the chunk level"],"invariants":["zero-LLM, zero-I/O, synchronous — ingest() is a pure function, always safe to call in the write path without latency budget concerns (it lives in core.ts and reaches no chunker; the AstChunker WASM grammar load IS disk I/O but is not on ingest()'s path)","deterministic + byte-reproducible — same input always produces the same hash, summary, and tags (no random or time-based components)","[BL-231] dist/core.js is the CJS-SAFE entry — its module graph must never contain a top-level await. dist/index.js (the root) is ESM-ONLY and cannot be require()d: it re-exports AstChunker, whose module-scope `await Parser.init()` is required to keep chunk()/estimate() synchronous. CommonJS consumers import @adhd/sox-ingest/core. Guarded by tools/test-bl231-cjs-boundary.mjs"],"entrypoints":["dist/index.js","dist/core.js"]},"dependencies":{"tree-sitter-wasms":"0.1.13","web-tree-sitter":"0.25.10"},"keywords":["ingest","data","transform","summary","tags","typescript"],"repository":{"type":"git","url":"git+https://github.com/PseudoSky/adhd.git"},"homepage":"https://github.com/PseudoSky/adhd","_id":"@adhd/sox-ingest@0.1.1","bugs":{"url":"https://github.com/PseudoSky/adhd/issues"},"_integrity":"sha512-jX0h4fDw+Tc7Dtyz1escfZFBA66Sjra5jB0rySBlPxxkNozKv+LVIKZVt8BenmcWNVADbHRhiiLUXkt2BfApgQ==","_resolved":"/private/var/folders/yg/cfczgtx54bzfh74lx2_mv0z80000gp/T/6e4ef5b947f1cd4cd1198e48664edaae/adhd-sox-ingest-0.1.1.tgz","_from":"file:adhd-sox-ingest-0.1.1.tgz","_nodeVersion":"24.11.1","_npmVersion":"11.6.2","dist":{"integrity":"sha512-jX0h4fDw+Tc7Dtyz1escfZFBA66Sjra5jB0rySBlPxxkNozKv+LVIKZVt8BenmcWNVADbHRhiiLUXkt2BfApgQ==","shasum":"8f9dffac18bba57e58951187a4965df69996db7a","tarball":"https://registry.npmjs.org/@adhd/sox-ingest/-/sox-ingest-0.1.1.tgz","fileCount":32,"unpackedSize":112331,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQDV9INsq2y+WrMEVxjGAybjQL3TyVzEobDBz3VJ/+duZwIhAJpYuJifmm10XmeVqKSNRC4ROmQzAZl4DaoK63i9+YBW"}]},"_npmUser":{"name":"pseudosky","email":"skywinston.sk@gmail.com"},"directories":{},"maintainers":[{"name":"pseudosky","email":"skywinston.sk@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/sox-ingest_0.1.1_1788566481968_0.8359116847162489"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-09T17:08:28.647Z","modified":"2026-09-05T00:01:22.364Z","0.1.0":"2026-07-09T17:08:28.883Z","0.1.1":"2026-09-05T00:01:22.176Z"},"license":"MIT","description":"Write-path single-item transforms for the memory domain — content-hash (SHA-256), extractive summary (lead-N sentences, zero LLM), deterministic tag extraction, and chunking/normalization. Pure stateless functions; no storage deps.","maintainers":[{"name":"pseudosky","email":"skywinston.sk@gmail.com"}],"readme":"# @adhd/sox-ingest\n\nWrite-path transforms for turning raw text into retrievable memory: content-hashing (SHA-256\ndedup), extractive summarization (lead-sentence, zero LLM), deterministic tag extraction, and two\nlevels of chunking — a simple sliding-window/sentence splitter, and a full format-aware chunker\nregistry (AST-aware code chunking via tree-sitter, heading-aware chunking for Markdown/MDX/RST/\nAsciiDoc, and a mixed-format chunker that nests the two). Every function is pure, synchronous, and\ndeterministic — no network calls, no disk I/O, no randomness, same input always produces the same\noutput.\n\n```bash\npnpm add @adhd/sox-ingest\n```\n\n## Quick start\n\n```typescript\nimport { ingest } from '@adhd/sox-ingest';\n\nconst result = ingest('Bi-temporal edges supersede facts rather than overwriting them. ' +\n  'This keeps the full history of every revision available for audit.');\n\nconsole.log(result.contentHash); // 64-char SHA-256 hex digest, stable across repeated calls\nconsole.log(result.summary);     // lead-sentence extractive summary (no LLM)\nconsole.log(result.tags);        // deterministic high-frequency terms, e.g. ['temporal', 'edges', ...]\n```\n\nAsk for chunks by passing `chunk` options — useful when `content` is too large to embed or send to\na model in one piece:\n\n```typescript\nconst chunked = ingest(longDocument, {\n  chunk: { maxChars: 2000, overlapChars: 200 },\n});\nfor (const chunk of chunked.chunks ?? []) {\n  console.log(chunk.index, chunk.charOffset, chunk.contentHash);\n}\n```\n\n## Two entry points — pick the one that matches your module system\n\n```typescript\nimport { ingest, hexSha256, splitIntoChunksSentence } from '@adhd/sox-ingest';       // ESM only\nimport { ingest, hexSha256, splitIntoChunksSentence } from '@adhd/sox-ingest/core';  // ESM or CommonJS\n```\n\nThe package root (`@adhd/sox-ingest`) also pulls in the format-aware chunkers (`AstChunker`,\n`HeadingChunker`, `MixedFormatChunker`), and `AstChunker`'s module preloads its tree-sitter WASM\ngrammars with a module-scope `await` — which makes the whole root **ESM-only**; Node refuses to\n`require()` a module graph containing a top-level await. If your project is CommonJS (or you only\nneed `ingest`/`hexSha256`/`splitIntoChunksSentence` and want to avoid pulling in tree-sitter at\nall), import the `@adhd/sox-ingest/core` subpath instead — it re-exports the same three functions\nand types from a module with no top-level await, safe to `require()`.\n\n## API reference\n\n### Core transforms (`@adhd/sox-ingest` or `@adhd/sox-ingest/core`)\n\n```typescript\ninterface IngestOpts {\n  summaryMaxSentences?: number;   // default 3\n  tagMaxCount?: number;            // default 10\n  chunk?: { maxChars?: number; overlapChars?: number }; // default 2000 / 200 — omit to skip chunking\n}\ninterface IngestChunk {\n  index: number;\n  content: string;\n  contentHash: string;   // SHA-256 of this chunk's raw content\n  charOffset: number;\n}\ninterface IngestResult {\n  contentHash: string;   // SHA-256 of normalized (trimmed, whitespace-collapsed) content\n  summary: string;\n  tags: string[];\n  chunks?: IngestChunk[]; // present only when opts.chunk was supplied\n}\n\nfunction ingest(content: string, opts?: IngestOpts): IngestResult;\nfunction hexSha256(data: string): string;\nfunction splitIntoChunksSentence(text: string, chunkTokens: number): string[];\n```\n\n- **`hexSha256`** is a raw hasher — it does not normalize input. `ingest()` normalizes\n  (trim + collapse whitespace) before hashing; call `hexSha256(content.trim().toLowerCase())`\n  yourself if you need dedup-hash parity with a lowercase-normalizing store instead.\n- **Summaries under 100 characters are returned unchanged** (`content.trim()`) rather than run\n  through sentence splitting — there's nothing to extract from a fragment that short.\n- **`splitIntoChunksSentence(text, chunkTokens)`** differs from `ingest()`'s own `chunk` option: it\n  splits at `chunkTokens * 4` characters preferring sentence boundaries and produces **no overlap**\n  between chunks, versus `ingest()`'s fixed-size sliding window **with** overlap. Pick\n  `splitIntoChunksSentence` when you want chunks that never cut a sentence in half; pick\n  `ingest({ chunk })` when you want per-chunk content hashes and consistent overlap for\n  retrieval-context stitching.\n\n### Format-aware chunkers (`@adhd/sox-ingest`, ESM only)\n\nEvery chunker implements the same synchronous interface and always attaches a `SourceMap` so a\nchunk can be traced back to its exact line range in the source document:\n\n```typescript\ninterface SourceMap {\n  sourceStartLine: number;\n  sourceEndLine: number;\n  sourceUrl?: string;\n  sourceSha?: string;\n}\ninterface Chunk {\n  text: string;\n  sourceMap: SourceMap;\n  metadata: {\n    chunkerId: string;\n    language?: string;\n    heading?: string;         // e.g. \"Installation > Prerequisites\"\n    parentDocId?: string;\n    chunkIndex: number;\n    isHeadingRoot?: boolean;\n  };\n  embedding?: Float32Array;   // empty at chunking time; populate after vectorizing\n}\ninterface ChunkerOptions {\n  sourceUrl?: string;\n  sourceSha?: string;\n  parentDocId?: string;\n  minFunctionLines?: number;   // AstChunker: merge shorter declarations into the previous chunk (default 3)\n  maxHeadingDepth?: number;    // HeadingChunker: split down to this heading depth (default 6)\n}\ninterface Chunker {\n  readonly id: string;\n  readonly supportedLanguages: string[];\n  chunk(document: string, options?: ChunkerOptions): Chunk[];\n  estimate(document: string): number;\n}\n```\n\n#### `AstChunker` — syntax-aware code chunking\n\nParses real source with `web-tree-sitter` (WASM grammars, no native build step) and splits on\ntop-level declaration boundaries — a function or class body is never split across two chunks:\n\n```typescript\nimport { AstChunker } from '@adhd/sox-ingest';\n\nconst chunker = new AstChunker('typescript'); // 'typescript' | 'python' | 'java' | 'csharp'\nconst chunks = chunker.chunk(`\nfunction foo() {\n  return 1;\n}\n\nfunction bar() {\n  return 2;\n}\n`);\nconsole.log(chunks.length, chunks[0]?.metadata.chunkerId); // 2 'ast:treesitter:typescript'\n```\n\nDeclarations shorter than `minFunctionLines` (default 3) are merged into the preceding chunk\ninstead of becoming their own singleton chunk.\n\n#### `HeadingChunker` — Markdown / MDX / RST / AsciiDoc\n\nSplits a document on heading boundaries and stamps a `heading` breadcrumb (e.g.\n`\"Installation > Prerequisites\"`) onto each chunk's metadata:\n\n```typescript\nimport { HeadingChunker } from '@adhd/sox-ingest';\n\nconst chunker = new HeadingChunker('markdown'); // 'markdown' | 'mdx' | 'rst' | 'asciidoc'\nconst chunks = chunker.chunk('# Setup\\n\\nRun `pnpm install`.\\n\\n## Prerequisites\\n\\nNode 20+.');\n```\n\n#### `MixedFormatChunker` — headings with embedded code fences\n\nRuns the heading chunker first (parent sections), then the AST chunker over each fenced code block\nfound within a section (child chunks) — so a Markdown doc with embedded TypeScript examples\nproduces both prose chunks and syntax-aware code chunks, correctly nested:\n\n```typescript\nimport { MixedFormatChunker, extractFencedCodeBlocks, mapFenceLanguage } from '@adhd/sox-ingest';\n\nconst chunker = new MixedFormatChunker('markdown'); // 'markdown' | 'mdx' | 'rst' | 'asciidoc'\nconst chunks = chunker.chunk(document);\n\n// Lower-level helpers used internally, also useful standalone:\nmapFenceLanguage('ts');   // 'typescript' — normalizes fence tags/aliases to an AstChunker language\nmapFenceLanguage('bash'); // null — unsupported for AST chunking, left as a prose chunk\nextractFencedCodeBlocks(document, 'markdown'); // FencedCodeBlock[] — raw fence extraction, no AST parse\n```\n\n#### `ChunkerRegistry` — language-keyed lookup\n\nThe package root auto-registers one instance of every chunker above (by language) into\n`globalChunkerRegistry` at import time — no runtime reflection, static registration:\n\n```typescript\nimport { globalChunkerRegistry } from '@adhd/sox-ingest';\n\nconst chunkers = globalChunkerRegistry.getForLanguage('markdown');\n// → both 'heading:markdown' and 'mixed:markdown' — pick whichever fits your pipeline\nconsole.log(globalChunkerRegistry.list());\n```\n\nBuild your own registry (e.g. to register only a subset, or a custom chunker) with\n`new ChunkerRegistry()` — `register(id, factory, languages)`, `get(id)`, `getForLanguage(lang)`,\n`list()`, and `seal()` to lock it against further registration.\n\nA chunker throws `PermanentChunkingError` for a genuinely unsupported input (e.g. an unrecognized\nheading syntax) and `TransientChunkingError` for a retryable failure — catch the two separately\nrather than treating every chunking failure the same way.\n\n`ChunkStaleReason` (`'source_updated' | 'chunker_upgraded' | 'ttl_expired'`) and\n`StaleChunkConfig` (`{ staleThresholdDays }`) are exported types for a consumer that wants to track\nwhen a previously-chunked document needs re-chunking — this package does not itself schedule or\nrun that check.\n\n## Invariants\n\n- **Zero-LLM, zero-I/O, synchronous.** `ingest()` never touches the network or the filesystem and\n  never awaits anything — safe to call on every write with no latency budget concerns.\n- **Deterministic and byte-reproducible.** The same input always produces the same hash, summary,\n  tags, and chunk boundaries — no random or time-based components anywhere in this package.\n- **No storage dependency.** This package has zero database/adapter dependencies of its own; it\n  does not inherit or participate in any store's write-concurrency model. A caller (such as\n  `@adhd/sox-memory-core`, which re-exports `hexSha256`/`splitIntoChunksSentence` from the `/core`\n  subpath for its own CommonJS build) is responsible for persisting whatever `ingest()` returns.\n\n## License\n\nMIT\n","readmeFilename":"README.md","homepage":"https://github.com/PseudoSky/adhd","keywords":["ingest","data","transform","summary","tags","typescript"],"repository":{"type":"git","url":"git+https://github.com/PseudoSky/adhd.git"},"bugs":{"url":"https://github.com/PseudoSky/adhd/issues"}}