{"_id":"@boole/truncate-tokens","name":"@boole/truncate-tokens","dist-tags":{"latest":"2.0.0"},"versions":{"2.0.0":{"name":"@boole/truncate-tokens","version":"2.0.0","description":"Truncate strings to a token budget without cutting mid-token or mid-character. Provider-agnostic with adapter support.","type":"module","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js","require":"./dist/index.cjs"},"./openai":{"types":"./dist/openai.d.ts","import":"./dist/openai.js","require":"./dist/openai.cjs"},"./anthropic":{"types":"./dist/anthropic.d.ts","import":"./dist/anthropic.js","require":"./dist/anthropic.cjs"}},"main":"./dist/index.cjs","module":"./dist/index.js","types":"./dist/index.d.ts","scripts":{"build":"tsup","test":"vitest run","typecheck":"tsc --noEmit","prepublishOnly":"npm run typecheck && npm run build && npm run test"},"engines":{"node":">=18"},"peerDependencies":{"gpt-tokenizer":">=2.0.0"},"peerDependenciesMeta":{"gpt-tokenizer":{"optional":true}},"devDependencies":{"gpt-tokenizer":"^2.8.1","tsup":"^8.3.5","typescript":"^5.7.2","vitest":"^2.1.8"},"keywords":["tokens","truncate","llm","tokenizer","openai","anthropic","claude","gpt","tiktoken"],"license":"BSD-2-Clause","repository":{"type":"git","url":"git+https://github.com/jordanplows/truncate-tokens.git"},"_id":"@boole/truncate-tokens@2.0.0","gitHead":"d23c3815408f64f0713ded1ec5a4217b157621a5","bugs":{"url":"https://github.com/jordanplows/truncate-tokens/issues"},"homepage":"https://github.com/jordanplows/truncate-tokens#readme","_nodeVersion":"22.23.2","_npmVersion":"10.9.8","dist":{"integrity":"sha512-5NIsumMJvnwBNtXwORAUMOssdK8o3oqXeqWJ+bqzOKA5WMY23KxtF8B/rwwUfGSpDY839+5FdkWxLmAWz/wvZA==","shasum":"40cd4a9f0f993a488643cd5768f2317cf8d5936a","tarball":"https://registry.npmjs.org/@boole/truncate-tokens/-/truncate-tokens-2.0.0.tgz","fileCount":35,"unpackedSize":37064,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIDEv2xn9UHkcqf+7k3ghdc2Y1xhFwi8NgyBsm4iG7r+pAiBY4CrQl1YhcDfk4CoxNGR3ty+8xVq5xNkBe+D9VHZhQQ=="}]},"_npmUser":{"name":"jordan.plows","email":"plowstjordan@gmail.com"},"directories":{},"maintainers":[{"name":"jordan.plows","email":"plowstjordan@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/truncate-tokens_2.0.0_1787679553031_0.9589486248095695"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-25T17:39:12.812Z","2.0.0":"2026-08-25T17:39:13.173Z","modified":"2026-08-25T17:39:13.467Z"},"maintainers":[{"name":"jordan.plows","email":"plowstjordan@gmail.com"}],"description":"Truncate strings to a token budget without cutting mid-token or mid-character. Provider-agnostic with adapter support.","homepage":"https://github.com/jordanplows/truncate-tokens#readme","keywords":["tokens","truncate","llm","tokenizer","openai","anthropic","claude","gpt","tiktoken"],"repository":{"type":"git","url":"git+https://github.com/jordanplows/truncate-tokens.git"},"bugs":{"url":"https://github.com/jordanplows/truncate-tokens/issues"},"license":"BSD-2-Clause","readme":"# truncate-tokens\n\nTruncate strings to an LLM token budget without cutting mid-token or mid-character. Provider-agnostic — works with any tokenizer via a small adapter interface.\n\n## Why not just use gpt-tokenizer directly?\n\n- **Provider-agnostic**: Works with OpenAI, Claude, or any custom tokenizer — not locked to tiktoken.\n- **Counting-only path**: Claude doesn't expose a reversible decode. This package handles that with a binary-search strategy over a token counter.\n- **Correctness edge cases centralized**: Surrogate pair safety, multi-byte characters spanning token boundaries, suffix budget accounting — all handled in one place.\n\n## Install\n\n```bash\nnpm install truncate-tokens\n```\n\nFor OpenAI/tiktoken support, also install the optional peer dependency:\n\n```bash\nnpm install gpt-tokenizer\n```\n\n## Quick Start\n\n### OpenAI (encode/decode codec)\n\n```ts\nimport { truncateToTokenLimit } from 'truncate-tokens/openai';\nimport { createOpenAICodec } from 'truncate-tokens/openai';\n\nconst codec = createOpenAICodec('cl100k_base');\nconst result = await truncateToTokenLimit(\n  'Your very long text here...',\n  100, // max tokens\n  codec,\n  { suffix: '...' }\n);\n```\n\n### Anthropic / Claude (count-only)\n\n```ts\nimport { truncateToTokenLimitByCount } from 'truncate-tokens/anthropic';\nimport { createAnthropicCounter } from 'truncate-tokens/anthropic';\n\n// Wrap your own counting function (e.g. from @anthropic-ai/sdk)\nconst counter = createAnthropicCounter(async (text) => {\n  const response = await anthropic.messages.countTokens({\n    model: 'claude-sonnet-4-20250514',\n    messages: [{ role: 'user', content: text }],\n  });\n  return response.input_tokens;\n});\n\nconst result = await truncateToTokenLimitByCount(\n  'Your very long text here...',\n  100,\n  counter,\n  { suffix: '...' }\n);\n```\n\n### Custom tokenizers\n\nImplement `TokenCodec` (if you have encode + decode) or `TokenCounter` (if you only have a count function) directly — no built-in adapter needed:\n\n```ts\nimport { truncateToTokenLimit, truncateToTokenLimitByCount } from 'truncate-tokens';\nimport type { TokenCodec, TokenCounter } from 'truncate-tokens';\n\nconst myCodec: TokenCodec = {\n  encode(text) { /* ... */ },\n  decode(tokens) { /* ... */ },\n};\n\nconst myCounter: TokenCounter = {\n  count(text) { /* ... */ },\n};\n```\n\n## API\n\n### `truncateToTokenLimit(text, maxTokens, codec, options?)`\n\nFast, exact path for tokenizers with full encode/decode. Encodes → slices tokens → decodes the surviving slice in one call.\n\n### `truncateToTokenLimitByCount(text, maxTokens, counter, options?)`\n\nFor tokenizers that only expose counting. Uses binary search on character length with a linear correction pass to guarantee budget compliance.\n\n### `countTokens(text, codecOrCounter)`\n\nReturns the token count of a string using either a codec or counter.\n\n### Options\n\n| Option | Type | Default | Description |\n|--------|------|---------|-------------|\n| `suffix` | `string` | `undefined` | Appended to truncated output (e.g. `\"...\"`) |\n| `includeSuffixInBudget` | `boolean` | `true` | Deduct the suffix's token cost from maxTokens |\n\n## License\n\nMIT\n# truncate-tokens\n","readmeFilename":"README.md","_rev":"1-4d7970ca3932bdc861a960eeb72ad5be"}