{"_id":"@bunkojp/text-sentiment","name":"@bunkojp/text-sentiment","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@bunkojp/text-sentiment","version":"0.1.0","license":"CC0-1.0","type":"module","main":"dist/index.cjs","module":"dist/index.js","types":"dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js","require":"./dist/index.cjs"},"./data/sentiment-ja":{"types":"./dist/data/sentiment-ja.d.ts","import":"./dist/data/sentiment-ja.js","require":"./dist/data/sentiment-ja.cjs"},"./data/sentiment-en":{"types":"./dist/data/sentiment-en.d.ts","import":"./dist/data/sentiment-en.js","require":"./dist/data/sentiment-en.cjs"},"./data/toxic-ja":{"types":"./dist/data/toxic-ja.d.ts","import":"./dist/data/toxic-ja.js","require":"./dist/data/toxic-ja.cjs"},"./data/toxic-en":{"types":"./dist/data/toxic-en.d.ts","import":"./dist/data/toxic-en.js","require":"./dist/data/toxic-en.cjs"}},"packageManager":"bun@1.3.11","sideEffects":false,"scripts":{"build":"vite build","clean":"rimraf dist","lint":"eslint .","lint:fix":"eslint . --fix","format":"prettier --write .","typecheck":"tsc -p tsconfig.json --noEmit","test:cov":"vitest run --coverage","test":"vitest --run","build:data":"bun run scripts/build-data.ts && bun run scripts/embed-data.ts","demo":"vite --config demo/vite.config.ts"},"devDependencies":{"@eslint-community/eslint-plugin-eslint-comments":"^4.7.1","@eslint/js":"^10.0.1","@types/node":"^25.5.0","@types/react":"^19.2.14","@types/react-dom":"^19.2.3","@typescript-eslint/eslint-plugin":"^8.57.1","@typescript-eslint/parser":"^8.57.1","@vitest/coverage-v8":"^4.1.0","ajv":"^8.18.0","cuss":"^2.2.0","eslint":"^10.0.3","eslint-config-prettier":"^10.1.8","eslint-plugin-import":"^2.32.0","eslint-plugin-jsdoc":"^62.8.0","prettier":"^3.8.1","react":"^19.2.4","react-dom":"^19.2.4","rimraf":"^6.1.3","typescript":"^5.9.3","typescript-eslint":"^8.57.1","vite":"^8.0.0","vite-plugin-dts":"^4","vitest":"^4.1.0"},"dependencies":{"@msgpack/msgpack":"^3.1.3","fflate":"^0.8.2"},"gitHead":"d9120b815d27bd6016f5da38ebb6edfbee123e17","_id":"@bunkojp/text-sentiment@0.1.0","description":"Multilingual sentiment analysis with Scunthorpe-safe tokenization.","_nodeVersion":"22.18.0","_npmVersion":"11.11.0","dist":{"integrity":"sha512-zr+QMZioKHmsobvC5M+J2EpZVWnHdqawutB1H4uDM63roKsZpCgL5HoDiOt2+aI9+kasozag8nWJaYmdX9Z7Aw==","shasum":"927991d14c6e42efba3ec883d304fb1535b690d1","tarball":"https://registry.npmjs.org/@bunkojp/text-sentiment/-/text-sentiment-0.1.0.tgz","fileCount":27,"unpackedSize":337640,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIA4jNAdjZb5bNJlrEvGrmB+1IyjlfpJx83Vh5JmEsNQWAiAMNrYr2cqmTXWpnhgHh2mplKZp++2HNotj3BPts7luhw=="}]},"_npmUser":{"name":"trkbt10","email":"triple.quartet+npm@gmail.com"},"directories":{},"maintainers":[{"name":"trkbt10","email":"triple.quartet+npm@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/text-sentiment_0.1.0_1775198729983_0.07190300533850813"},"_hasShrinkwrap":false}},"time":{"created":"2026-04-03T06:45:29.908Z","0.1.0":"2026-04-03T06:45:30.139Z","modified":"2026-04-03T06:45:30.325Z"},"maintainers":[{"name":"trkbt10","email":"triple.quartet+npm@gmail.com"}],"description":"Multilingual sentiment analysis with Scunthorpe-safe tokenization.","license":"CC0-1.0","readme":"# @bunkojp/text-sentiment\n\nMultilingual sentiment analysis with Scunthorpe-safe tokenization.\n\n- **Multiple strategies** — Naive Bayes (primary), TF-IDF, NCD, and weighted ensemble\n- **Toxic content detection** — offensive word matching that never flags safe compound words\n- **No server required** — runs entirely in-process (Node, Bun, or browser)\n- **Data-driven** — all lexicons loaded from pre-built compressed binaries, no hardcoded word lists in source\n\n\n## Overview\n\n### Architecture\n\n```\nanalyze(text, language, options?)\n  |\n  +-- tokenizeText(text, language)       <- Single Source of Truth\n  |     |\n  |     +-- JA: mikan.js-style segmenter (character-class boundaries)\n  |     +-- EN: whitespace split + punctuation strip\n  |\n  +-- tokens -> classifyByNaiveBayes()   <- pre-trained model from binary\n  +-- tokens -> classifyByTfidf()\n  +-- tokens -> classifyByNcd()\n  +-- tokens -> toxic detection (exact token match)\n  +-- tokens -> category breakdown (lexicon lookup)\n```\n\n### Scunthorpe Problem Prevention\n\nThe Japanese segmenter splits text at character-class transitions (kanji / hiragana / katakana / latin). Katakana compound words stay as single tokens, so offensive substrings embedded within them are never produced as independent tokens.\n\nFor English, standard whitespace tokenization already keeps compound words like \"Scunthorpe\" and \"cocktail\" intact.\n\n### Data Format\n\nAll lexicon data is stored as compressed MessagePack binaries in `src/data/`:\n\n| Optimization | Effect |\n|---|---|\n| Columnar layout | words[], scores[], categories[] stored separately |\n| Score delta encoding | sorted scores have small consecutive differences |\n| NB palette indexing | log-likelihood triplets compressed to palette + uint8 index |\n| Category enum | string categories mapped to uint8 |\n| deflate level 9 | maximum compression |\n\n### Supported Languages\n\n| Language | Sentiment Lexicon | Toxic Lexicon |\n|---|---|---|\n| Japanese (ja) | 11,293 words (Tohoku Univ. via oseti) | 748 words |\n| English (en) | 8,219 words (AFINN + VADER) | 1,540 words (cuss) |\n\nAdding a new language requires only running the build script with new corpus URLs — no code changes.\n\n\n## Getting Started\n\n### Quick Start (Node / Bun)\n\n```ts\nimport { readFileSync } from \"node:fs\";\nimport { analyze, registerLexiconFromBinary, registerToxicLexiconFromBinary } from \"@bunkojp/text-sentiment\";\n\n// Load lexicon binaries (once at startup)\nregisterLexiconFromBinary(\"ja\", new Uint8Array(readFileSync(\"node_modules/@bunkojp/text-sentiment/src/data/sentiment-ja.bin\")));\nregisterLexiconFromBinary(\"en\", new Uint8Array(readFileSync(\"node_modules/@bunkojp/text-sentiment/src/data/sentiment-en.bin\")));\nregisterToxicLexiconFromBinary(\"ja\", new Uint8Array(readFileSync(\"node_modules/@bunkojp/text-sentiment/src/data/toxic-ja.bin\")));\nregisterToxicLexiconFromBinary(\"en\", new Uint8Array(readFileSync(\"node_modules/@bunkojp/text-sentiment/src/data/toxic-en.bin\")));\n\n// Analyze\nconst result = analyze(\"This movie is absolutely wonderful.\", \"en\");\nconsole.log(result.sentiment.label);      // \"positive\"\nconsole.log(result.sentiment.confidence);  // 0.95\n```\n\n### Quick Start (Browser)\n\n```ts\nimport { analyze, registerLexiconFromBinary, registerToxicLexiconFromBinary } from \"@bunkojp/text-sentiment\";\n\n// Fetch and register binaries\nconst data = await fetch(\"/sentiment-en.bin\").then(r => r.arrayBuffer());\nregisterLexiconFromBinary(\"en\", new Uint8Array(data));\n\nconst result = analyze(\"Terrible experience.\", \"en\");\nconsole.log(result.sentiment.label); // \"negative\"\n```\n\n\n## Usage\n\n### Basic Sentiment Analysis\n\n```ts\n// Default: Naive Bayes\nanalyze(\"素晴らしい作品です\", \"ja\")\n// { sentiment: { label: \"positive\", confidence: 0.94, scores: {...} }, tokens: [...] }\n\n// Select strategy\nanalyze(\"Great product\", \"en\", { strategy: \"tfidf\" })\nanalyze(\"Great product\", \"en\", { strategy: \"ncd\" })\n```\n\n### Ensemble\n\nWeighted combination of all three strategies:\n\n```ts\nanalyze(\"素晴らしい作品です\", \"ja\", { ensemble: {} })\n// Default weights: naive-bayes 0.6, tfidf 0.25, ncd 0.15\n\n// Custom weights\nanalyze(\"text\", \"en\", {\n  ensemble: {\n    strategies: [\"naive-bayes\", \"tfidf\"],\n    weights: { \"naive-bayes\": 0.7, \"tfidf\": 0.3 },\n  },\n})\n```\n\n### Toxic Content Detection\n\n```ts\nconst r = analyze(\"text\", \"ja\", { toxic: true });\nr.toxic?.toxic     // boolean\nr.toxic?.matches   // [{ word, severity, category }]\n```\n\n### Per-Category Breakdown\n\n```ts\nconst r = analyze(\"text\", \"ja\", { categories: true });\nr.categories\n// { general: { label, confidence, scores }, quality: {...}, ... }\n```\n\nCategories: `general`, `quality`, `service`, `price`, `usability`, `emotion`, `appearance`\n\n### All Options Combined\n\n```ts\nanalyze(\"text\", \"ja\", {\n  strategy: \"naive-bayes\",\n  ensemble: { weights: { \"naive-bayes\": 0.6, tfidf: 0.25, ncd: 0.15 } },\n  categories: true,\n  toxic: true,\n  smoothing: 1,\n  neutralThreshold: 0.05,\n})\n```\n\n### Low-Level API\n\nFor direct access to individual classifiers:\n\n```ts\nimport { tokenizeText, classifyByNaiveBayes, getLexicon } from \"@bunkojp/text-sentiment\";\n\nconst tokens = tokenizeText(\"text\", \"ja\");\nconst lexicon = getLexicon(\"ja\");\nconst result = classifyByNaiveBayes(tokens, lexicon);\n```\n\n\n## Installation\n\n```bash\nnpm install @bunkojp/text-sentiment\n# or\nbun add @bunkojp/text-sentiment\n```\n\n### Building from Source\n\n```bash\ngit clone https://github.com/bunko-jp/text-sentiment.git\ncd text-sentiment\nbun install\nbun run build:data   # Download corpora and build lexicon binaries\nbun run build        # Build library\nbun run test         # Run tests\n```\n\n### Rebuilding Lexicon Data\n\nThe lexicon binaries in `src/data/` are pre-built and included in the package. To rebuild from external corpora:\n\n```bash\nbun run build:data\n```\n\nThis downloads from:\n- Tohoku University sentiment dictionary (via [oseti](https://github.com/ikegami-yukino/oseti))\n- [AFINN-165](https://github.com/fnielsen/afinn) + [VADER](https://github.com/cjhutto/vaderSentiment)\n- [inappropriate-words-ja](https://github.com/MosasoM/inappropriate-words-ja) + [LDNOOBW V2](https://github.com/LDNOOBWV2/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words_V2)\n- [words/cuss](https://github.com/words/cuss)\n\nSee [THIRD-PARTY-LICENSES](THIRD-PARTY-LICENSES) for full license details.\n\n### Demo\n\n```bash\nbun run demo\n# Opens http://localhost:5173 with a React-based interactive demo\n```\n\n\n## License\n\nCC0-1.0 - see [LICENSE](LICENSE) for details.\n","readmeFilename":"README.md","_rev":"1-81e0adf162813a68cffc4f8b8f356a34"}