{"_id":"@ainova-mya/unicode","_rev":"4-37b16afdc355828340529a09b9474908","name":"@ainova-mya/unicode","dist-tags":{"latest":"1.0.3"},"versions":{"1.0.0":{"name":"@ainova-mya/unicode","version":"1.0.0","keywords":[],"author":{"url":" TAKASHI ","name":"San Linn Htet"},"license":"MIT","_id":"@ainova-mya/unicode@1.0.0","maintainers":[{"name":"ainova-mya","email":"takashilinn.personal@gmail.com"}],"dist":{"shasum":"300ed32bfcbc5238c7ca491c535036524668d820","tarball":"https://registry.npmjs.org/@ainova-mya/unicode/-/unicode-1.0.0.tgz","fileCount":29,"integrity":"sha512-DTKgqcgurDlLy6Uay8ASlknJk/kc7n0GFy7Y53FpsnnVi4hRPfDbWuCDvgIKKU0/wIe8UjxPbHz5qQO2Wg6G+g==","signatures":[{"sig":"MEUCIQDUKAVUVQGMGVG+jMDLkY91ybXI4iT3fkbsqz7OpjsxxAIgKOwvwUBLJD9xIhZAENP1+s2UFCX2bNT9zXcpTihVNW8=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":59075},"main":"./dist/index.cjs","type":"module","types":"./dist/index.d.ts","scripts":{"dev":"npx tsx ./src/index.ts","build":"tsup"},"_npmUser":{"name":"ainova-mya","email":"takashilinn.personal@gmail.com"},"_npmVersion":"10.9.8","description":"Myanmar Unicode processing library","directories":{},"_nodeVersion":"22.23.2","dependencies":{"tslib":"^2.8.1"},"_hasShrinkwrap":false,"devDependencies":{"tsup":"^8.5.1","eslint":"^8.38.0","typescript":"^5.6.2","@typescript-eslint/parser":"^5.59.0","@typescript-eslint/eslint-plugin":"^5.59.0"},"_npmOperationalInternal":{"tmp":"tmp/unicode_1.0.0_1786091557474_0.16415795468204353","host":"s3://npm-registry-packages-npm-production"}},"1.0.1":{"name":"@ainova-mya/unicode","version":"1.0.1","keywords":[],"author":{"url":" TAKASHI ","name":"San Linn Htet"},"license":"MIT","_id":"@ainova-mya/unicode@1.0.1","maintainers":[{"name":"ainova-mya","email":"takashilinn.personal@gmail.com"}],"homepage":"https://github.com/SanLinnHtet-dev/mya-ainova/tree/main/packages/mya-unicode","bugs":{"url":"https://github.com/SanLinnHtet-dev/mya-ainova/issues"},"dist":{"shasum":"1e2518068ca7c72e0dcffccf514637fb525f6d72","tarball":"https://registry.npmjs.org/@ainova-mya/unicode/-/unicode-1.0.1.tgz","fileCount":29,"integrity":"sha512-dhql3ZVGSIMOt3PRgY9rDxPWeRObpY35NmJX11kkM8+/F6HmQ9UHxliwNKvFh5MjMJsSSfPqcH9qKOqe3+J4GA==","signatures":[{"sig":"MEQCICiH33VG/AZUqypQDqo3VI5lLfSKHNhh+wusuKvgH4AQAiBixIzV6gmloCEOw+pegjvmgQH+yR+e2w46qT1sYfACoQ==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":59415},"main":"./dist/index.cjs","type":"module","types":"./dist/index.d.ts","gitHead":"68de570481e1c45279814c888798a45f76ad3b8e","scripts":{"dev":"npx tsx ./src/index.ts","build":"tsup"},"_npmUser":{"name":"ainova-mya","email":"takashilinn.personal@gmail.com"},"repository":{"url":"git+https://github.com/SanLinnHtet-dev/mya-ainova.git","type":"git","directory":"packages/mya-unicode"},"_npmVersion":"10.9.8","description":"Myanmar Unicode processing library","directories":{},"_nodeVersion":"22.23.2","dependencies":{"tslib":"^2.8.1"},"_hasShrinkwrap":false,"devDependencies":{"tsup":"^8.5.1","eslint":"^8.38.0","typescript":"^5.6.2","@typescript-eslint/parser":"^5.59.0","@typescript-eslint/eslint-plugin":"^5.59.0"},"_npmOperationalInternal":{"tmp":"tmp/unicode_1.0.1_1786529487825_0.7291308743476841","host":"s3://npm-registry-packages-npm-production"}},"1.0.2":{"name":"@ainova-mya/unicode","version":"1.0.2","keywords":[],"author":{"url":" TAKASHI ","name":"San Linn Htet"},"license":"MIT","_id":"@ainova-mya/unicode@1.0.2","maintainers":[{"name":"ainova-mya","email":"takashilinn.personal@gmail.com"}],"homepage":"https://github.com/SanLinnHtet-dev/mya-ainova/tree/main/packages/mya-unicode","bugs":{"url":"https://github.com/SanLinnHtet-dev/mya-ainova/issues"},"dist":{"shasum":"7be56996988c2bd9924bac687f1c1b04c0c8a1b0","tarball":"https://registry.npmjs.org/@ainova-mya/unicode/-/unicode-1.0.2.tgz","fileCount":29,"integrity":"sha512-ZySa7nTcBzPLFGyJGjXAmboebEMvNnB/5yo2r8wbd4dAr8WSGvK6t8srKeKxlSHx8HVnqa1GGoRDxWNOz35ULw==","signatures":[{"sig":"MEYCIQCUiBTHeu+91Wb7vGvWLDRFp2rmq8Zkxf8+L/KD2IesLwIhANOwJ03k0U+SRJyQVvEKWRg9WlJqlhY0EudTf2k94ag5","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":74445},"main":"./dist/index.cjs","type":"module","types":"./dist/index.d.ts","gitHead":"68de570481e1c45279814c888798a45f76ad3b8e","scripts":{"dev":"npx tsx ./src/index.ts","build":"tsup"},"_npmUser":{"name":"ainova-mya","email":"takashilinn.personal@gmail.com"},"repository":{"url":"git+https://github.com/SanLinnHtet-dev/mya-ainova.git","type":"git","directory":"packages/mya-unicode"},"_npmVersion":"10.9.8","description":"Myanmar Unicode processing library","directories":{},"_nodeVersion":"22.23.2","dependencies":{"tslib":"^2.8.1"},"_hasShrinkwrap":false,"devDependencies":{"tsup":"^8.5.1","eslint":"^8.38.0","typescript":"^5.6.2","@typescript-eslint/parser":"^5.59.0","@typescript-eslint/eslint-plugin":"^5.59.0"},"_npmOperationalInternal":{"tmp":"tmp/unicode_1.0.2_1786530088846_0.9015132867716369","host":"s3://npm-registry-packages-npm-production"}},"1.0.3":{"name":"@ainova-mya/unicode","version":"1.0.3","description":"Myanmar Unicode processing library","author":{"name":"San Linn Htet","url":" TAKASHI "},"repository":{"type":"git","url":"git+https://github.com/SanLinnHtet-dev/mya-ainova.git","directory":"packages/mya-unicode"},"bugs":{"url":"https://github.com/SanLinnHtet-dev/mya-ainova/issues"},"homepage":"https://github.com/SanLinnHtet-dev/mya-ainova/tree/main/packages/mya-unicode","type":"module","main":"./dist/index.cjs","types":"./dist/index.d.ts","scripts":{"dev":"npx tsx ./src/index.ts","build":"tsup"},"keywords":[],"license":"MIT","dependencies":{"tslib":"^2.8.1"},"devDependencies":{"@typescript-eslint/eslint-plugin":"^5.59.0","@typescript-eslint/parser":"^5.59.0","eslint":"^8.38.0","tsup":"^8.5.1","typescript":"^5.6.2"},"_id":"@ainova-mya/unicode@1.0.3","gitHead":"68de570481e1c45279814c888798a45f76ad3b8e","_nodeVersion":"22.23.2","_npmVersion":"10.9.8","dist":{"integrity":"sha512-gBWL3DCuaelRxMoFhsuFDXTiZ5HGnGXsiusyrr1S9hwrMx/kBbPFKRhS4SWOQYZomFtsMRoYrtFQqdXJA1TGrg==","shasum":"5f5f04271c9ed7e651fc9ccd8a91a04bc7f5283c","tarball":"https://registry.npmjs.org/@ainova-mya/unicode/-/unicode-1.0.3.tgz","fileCount":29,"unpackedSize":75282,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCICdGt20vE+3DQ+9AnX8nwzfdf3XlBcTXnT4UqsAvntgaAiBNt4gD/7Y99VFP5E8Yh/o2032Fqj1F0GP/R+N8W2/kYA=="}]},"_npmUser":{"name":"ainova-mya","email":"takashilinn.personal@gmail.com"},"directories":{},"maintainers":[{"name":"ainova-mya","email":"takashilinn.personal@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/unicode_1.0.3_1786532385472_0.38715645687274636"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-07T08:32:37.235Z","modified":"2026-08-12T10:59:46.090Z","1.0.0":"2026-08-07T08:32:37.645Z","1.0.1":"2026-08-12T10:11:27.993Z","1.0.2":"2026-08-12T10:21:29.000Z","1.0.3":"2026-08-12T10:59:45.790Z"},"bugs":{"url":"https://github.com/SanLinnHtet-dev/mya-ainova/issues"},"author":{"name":"San Linn Htet","url":" TAKASHI "},"license":"MIT","homepage":"https://github.com/SanLinnHtet-dev/mya-ainova/tree/main/packages/mya-unicode","keywords":[],"repository":{"type":"git","url":"git+https://github.com/SanLinnHtet-dev/mya-ainova.git","directory":"packages/mya-unicode"},"description":"Myanmar Unicode processing library","maintainers":[{"name":"ainova-mya","email":"takashilinn.personal@gmail.com"}],"readme":"# @ainova-mya/unicode\n\nမြန်မာစာ Unicode processing အတွက် JavaScript / TypeScript library ဖြစ်ပါတယ်။\n\nMyanmar Unicode processing library for JavaScript and TypeScript.\n\n---\n\n## Table of Contents\n\n- [English](#english)\n  - [Overview](#overview)\n  - [Features](#features)\n    - [Unicode Normalization](#unicode-normalization)\n    - [Whitespace Cleaning](#whitespace-cleaning)\n    - [Unicode Repair](#unicode-repair)\n    - [Validation](#validation)\n    - [Segmentation](#segmentation)\n  - [Installation](#installation)\n  - [Usage](#usage)\n  - [Processing Pipeline](#processing-pipeline)\n  - [Available APIs](#available-apis)\n  - [API Overview](#api-overview)\n  - [TypeScript Support](#typescript-support)\n  - [Use Cases](#use-cases)\n  - [Project Status](#project-status)\n  - [License](#license)\n- [မြန်မာဘာသာ](#မြန်မာဘာသာ)\n  - [အကြောင်းအရာ](#အကြောင်းအရာ)\n  - [လုပ်ဆောင်ချက်များ](#လုပ်ဆောင်ချက်များ)\n  - [ထည့်သွင်းခြင်း](#ထည့်သွင်းခြင်း)\n  - [အသုံးပြုပုံ](#အသုံးပြုပုံ)\n  - [Text Processing Pipeline](#text-processing-pipeline-1)\n  - [အသုံးပြုနိုင်သော API များ](#အသုံးပြုနိုင်သော-api-များ)\n  - [API အကျဉ်းချုပ်](#api-အကျဉ်းချုပ်)\n  - [TypeScript Support](#typescript-support-1)\n  - [အသုံးပြုနိုင်သောနေရာများ](#အသုံးပြုနိုင်သောနေရာများ)\n  - [Project Status](#project-status-1)\n  - [License](#license-1)\n\n---\n\n# မြန်မာဘာသာ\n\n## အကြောင်းအရာ\n\n**@ainova-mya/unicode** သည် မြန်မာ Unicode စာသားများကို JavaScript နှင့် TypeScript တွင် စီမံခန့်ခွဲရန်အတွက် ဖန်တီးထားသော library ဖြစ်ပါတယ်။\n\nမြန်မာစာသားများကို:\n\n- Unicode normalization\n- Text cleaning\n- Unicode validation\n- Unicode repair\n- Text segmentation\n- Tokenization\n\nတို့အတွက် အသုံးပြုနိုင်ပါတယ်။\n\nဒီ project ရဲ့ ရည်ရွယ်ချက်က မြန်မာစာ processing အတွက် ယုံကြည်စိတ်ချရသော အခြေခံ layer တစ်ခုကို တည်ဆောက်ပေးရန် ဖြစ်ပါတယ်။\n\nအသုံးပြုနိုင်သောနေရာများမှာ:\n\n- OCR စာသားပြုပြင်ခြင်း\n- Myanmar Unicode normalization\n- Unicode validation\n- Text cleaning\n- စာသားခွဲခြမ်းခြင်း\n- Document processing\n- AI pipeline များ\n- Search system များ\n- RAG system များ\n- NLP system များ\n\n---\n\n## လုပ်ဆောင်ချက်များ\n\n### Unicode Normalization\n\nမြန်မာ Unicode စာသားများကို တူညီသော format တစ်ခုအဖြစ် ပြောင်းလဲပေးပါသည်။\n\n```ts\nimport { normalizeUnicode } from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ';\n\nconst normalized = normalizeUnicode(text);\n\nconsole.log(normalized);\n```\n\n---\n\n### Whitespace Cleanup\n\nအပို space များ၊ zero-width character များနှင့် line ending များကို ပြင်ဆင်ပေးပါသည်။\n\n```ts\nimport { normalizeWhitespace } from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ   မြန်မာစာ';\n\nconst cleaned = normalizeWhitespace(text);\n\nconsole.log(cleaned);\n```\n\n---\n\n### Unicode Repair\n\nOCR မှ ထွက်လာသော စာသားများနှင့် text cleaning အတွက် အခြေခံ Unicode repair utilities များ ပါဝင်ပါတယ်။\n\nပါဝင်သောအရာများမှာ:\n\n- ထပ်နေသော Myanmar mark များ ဖယ်ရှားခြင်း\n- ပုံစံတူ character များ ပြင်ဆင်ခြင်း\n\n```ts\nimport { fixConfusable, removeDuplicateMarks } from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ';\n\nconst fixed = fixConfusable(text);\n\nconst cleaned = removeDuplicateMarks(fixed);\n\nconsole.log(cleaned);\n```\n\n---\n\n### Validation\n\nMyanmar Unicode စာသားများကို စစ်ဆေးနိုင်ပါတယ်။\n\n```ts\nimport {\n  containsMyanmar,\n  containsLatin,\n  isMyanmarOnly,\n  isValidUnicode,\n} from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ';\n\nconsole.log(containsMyanmar(text));\nconsole.log(containsLatin(text));\nconsole.log(isMyanmarOnly(text));\nconsole.log(isValidUnicode(text));\n```\n\nအသုံးပြုနိုင်သော API များ:\n\n- `containsMyanmar()`\n- `containsLatin()`\n- `isMyanmarOnly()`\n- `isValidUnicode()`\n\n---\n\n### Segmentation\n\nမြန်မာစာသားများကို အခြေခံအဆင့် ခွဲခြမ်းနိုင်ပါတယ်။\n\nပါဝင်သော API များ:\n\n- `segmentLine()`\n- `segmentParagraph()`\n- `segmentSentence()`\n- `segmentSyllable()`\n- `segmentWord()`\n- `tokenize()`\n\nဥပမာ:\n\n```ts\nimport {\n  segmentSentence,\n  segmentSyllable,\n  tokenize,\n} from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ။ မြန်မာစာကို လေ့လာနေပါတယ်။';\n\nconst sentences = segmentSentence(text);\n\nconst syllables = segmentSyllable(text);\n\nconst tokens = tokenize(text);\n\nconsole.log(sentences);\nconsole.log(syllables);\nconsole.log(tokens);\n```\n\n---\n\n## ထည့်သွင်းခြင်း\n\n### npm ဖြင့်\n\n```bash\nnpm install @ainova-mya/unicode\n```\n\n### pnpm ဖြင့်\n\n```bash\npnpm add @ainova-mya/unicode\n```\n\n### yarn ဖြင့်\n\n```bash\nyarn add @ainova-mya/unicode\n```\n\n---\n\n## အသုံးပြုပုံ\n\n### TypeScript\n\n```ts\nimport {\n  normalizeUnicode,\n  normalizeWhitespace,\n  tokenize,\n} from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ';\n\nconst clean = normalizeWhitespace(normalizeUnicode(text));\n\nconsole.log(clean);\n\nconst result = tokenize(clean);\n\nconsole.log(result);\n```\n\n### JavaScript\n\n```js\nimport {\n  normalizeUnicode,\n  normalizeWhitespace,\n  tokenize,\n} from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ';\n\nconst clean = normalizeWhitespace(normalizeUnicode(text));\n\nconsole.log(clean);\n\nconst result = tokenize(clean);\n\nconsole.log(result);\n```\n\n---\n\n## Text Processing Pipeline\n\nLibrary ထဲက function များကို ပေါင်းစပ်ပြီး မြန်မာစာ processing pipeline တစ်ခုအဖြစ် အသုံးပြုနိုင်ပါတယ်။\n\n```ts\nimport {\n  normalizeUnicode,\n  normalizeWhitespace,\n  fixConfusable,\n  removeDuplicateMarks,\n  tokenize,\n} from '@ainova-mya/unicode';\n\nconst input = 'မင်္ဂလာပါ';\n\nconst normalized = normalizeUnicode(input);\n\nconst cleaned = normalizeWhitespace(normalized);\n\nconst repaired = fixConfusable(cleaned);\n\nconst finalText = removeDuplicateMarks(repaired);\n\nconst tokens = tokenize(finalText);\n\nconsole.log(finalText);\nconsole.log(tokens);\n```\n\nအခြေခံ processing flow:\n\n```text\nInput Text\n    ↓\nUnicode Normalization\n    ↓\nWhitespace Cleaning\n    ↓\nUnicode Repair\n    ↓\nDuplicate Mark Removal\n    ↓\nSegmentation / Tokenization\n    ↓\nProcessed Myanmar Text\n```\n\n---\n\n## အသုံးပြုနိုင်သော API များ\n\n### Normalization\n\n```ts\nnormalizeUnicode(text);\nnormalizeWhitespace(text);\n```\n\n### Unicode Repair\n\n```ts\nfixConfusable(text);\nremoveDuplicateMarks(text);\n```\n\n### Validation\n\n```ts\ncontainsMyanmar(text);\ncontainsLatin(text);\nisMyanmarOnly(text);\nisValidUnicode(text);\n```\n\n### Segmentation\n\n```ts\nsegmentLine(text);\nsegmentParagraph(text);\nsegmentSentence(text);\nsegmentSyllable(text);\nsegmentWord(text);\ntokenize(text);\n```\n\n---\n\n## API အကျဉ်းချုပ်\n\n| API                      | လုပ်ဆောင်ချက်                                  |\n| ------------------------ | ---------------------------------------------- |\n| `normalizeUnicode()`     | Myanmar Unicode စာသားကို normalize လုပ်ခြင်း   |\n| `normalizeWhitespace()`  | Space နှင့် line ending များကို သန့်ရှင်းခြင်း |\n| `fixConfusable()`        | ပုံစံတူ character များကို ပြင်ဆင်ခြင်း         |\n| `removeDuplicateMarks()` | ထပ်နေသော Myanmar mark များ ဖယ်ရှားခြင်း        |\n| `containsMyanmar()`      | Myanmar character ပါ/မပါ စစ်ခြင်း              |\n| `containsLatin()`        | Latin character ပါ/မပါ စစ်ခြင်း                |\n| `isMyanmarOnly()`        | Myanmar စာသားသာ ဖြစ်/မဖြစ် စစ်ခြင်း            |\n| `isValidUnicode()`       | Unicode စာသား မှန်ကန်မှု စစ်ခြင်း              |\n| `segmentLine()`          | Line အလိုက် ခွဲခြင်း                           |\n| `segmentParagraph()`     | Paragraph အလိုက် ခွဲခြင်း                      |\n| `segmentSentence()`      | Sentence အလိုက် ခွဲခြင်း                       |\n| `segmentSyllable()`      | Syllable အလိုက် ခွဲခြင်း                       |\n| `segmentWord()`          | Word အလိုက် ခွဲခြင်း                           |\n| `tokenize()`             | Token များအဖြစ် ခွဲခြင်း                       |\n\n---\n\n## TypeScript Support\n\nဒီ library ကို TypeScript project များတွင် အသုံးပြုနိုင်ပြီး TypeScript types များလည်း support လုပ်ထားပါတယ်။\n\n```ts\nimport { normalizeUnicode, tokenize } from '@ainova-mya/unicode';\n\nconst text: string = 'မင်္ဂလာပါ';\n\nconst normalized: string = normalizeUnicode(text);\n\nconst tokens = tokenize(normalized);\n\nconsole.log(tokens);\n```\n\n---\n\n## အသုံးပြုနိုင်သောနေရာများ\n\n### OCR Processing\n\nOCR မှရရှိလာသော မြန်မာစာသားများကို ပြန်လည်သန့်ရှင်းရန် processing layer အဖြစ် အသုံးပြုနိုင်ပါတယ်။\n\n```text\nOCR Output\n    ↓\nUnicode Normalization\n    ↓\nWhitespace Cleaning\n    ↓\nUnicode Repair\n    ↓\nMyanmar Text\n```\n\n---\n\n### Search\n\nMyanmar text ကို search engine ထဲသို့ index မလုပ်မီ preprocessing layer အဖြစ် အသုံးပြုနိုင်ပါတယ်။\n\n```text\nRaw Text\n    ↓\nNormalize\n    ↓\nClean\n    ↓\nSegment / Tokenize\n    ↓\nSearch Index\n```\n\n---\n\n### AI / RAG\n\nEmbedding သို့မဟုတ် retrieval မလုပ်မီ မြန်မာစာသားများကို normalize နှင့် clean လုပ်နိုင်ပါတယ်။\n\n```text\nDocument\n    ↓\nText Extraction\n    ↓\nMyanmar Unicode Processing\n    ↓\nText Segmentation\n    ↓\nChunking\n    ↓\nEmbedding\n    ↓\nVector Database\n```\n\n---\n\n### Document Processing\n\nDocument processing pipeline များတွင် foundation layer အဖြစ် အသုံးပြုနိုင်ပါတယ်။\n\n```text\nPDF / DOCX / Image\n        ↓\n    Text Extraction\n        ↓\nMyanmar Unicode Processing\n        ↓\n    Text Segmentation\n        ↓\n    Structured Text\n```\n\n---\n\n# Project Status\n\n**Current Version:** `v1.0.0`\n\nVersion 1 တွင် အဓိကအားဖြင့်:\n\n- Unicode foundation\n- Text cleaning\n- Unicode validation\n- Unicode repair\n- Basic text segmentation\n- Basic tokenization\n\nတို့ကို အဓိကထားပြီး တည်ဆောက်ထားပါတယ်။\n\n### Future Plans\n\nနောက်လာမည့် version များတွင် အောက်ပါ features များကို ထပ်မံထည့်သွင်းရန် ရည်ရွယ်ထားပါတယ်။\n\n- Advanced Myanmar word segmentation\n- Dictionary-based NLP\n- Advanced Unicode repair engine\n- Spell checking\n- OCR optimization\n- Myanmar tokenizer improvements\n- Advanced Myanmar NLP utilities\n\n---\n\n# License\n\nMIT License\n\nThis project is open source and available for personal and commercial use.\n\nဒီ project ကို MIT License အောက်တွင် open source အဖြစ် ဖြန့်ချိထားပါတယ်။\n\nကိုယ်ပိုင် project များနှင့် commercial project များတွင် အသုံးပြုနိုင်ပါတယ်။\n\n---\n\n# English\n\n## Overview\n\n**@ainova-mya/unicode** is a Myanmar Unicode processing library for JavaScript and TypeScript.\n\nIt provides utilities for processing, cleaning, validating, normalizing, and segmenting Myanmar Unicode text.\n\nThe goal of this project is to provide a reliable foundation layer for Myanmar language text processing.\n\nIt can be used as a foundation for:\n\n- OCR text processing\n- Myanmar text normalization\n- Unicode validation\n- Text cleaning\n- Text segmentation\n- Document processing\n- AI pipelines\n- Search systems\n- RAG systems\n- Natural Language Processing (NLP)\n\n---\n\n## Features\n\n### Unicode Normalization\n\nNormalize Myanmar Unicode text into a consistent format.\n\n```ts\nimport { normalizeUnicode } from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ';\n\nconst normalized = normalizeUnicode(text);\n\nconsole.log(normalized);\n```\n\n---\n\n### Whitespace Cleaning\n\nRemove unnecessary whitespace, zero-width characters, and normalize line endings.\n\n```ts\nimport { normalizeWhitespace } from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ   မြန်မာစာ';\n\nconst cleaned = normalizeWhitespace(text);\n\nconsole.log(cleaned);\n```\n\n---\n\n### Unicode Repair\n\nProvides basic Unicode repair utilities for OCR output and text cleaning.\n\nCurrent repair features include:\n\n- Duplicate Myanmar mark removal\n- Confusable character correction\n\n```ts\nimport { fixConfusable, removeDuplicateMarks } from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ';\n\nconst fixed = fixConfusable(text);\n\nconst cleaned = removeDuplicateMarks(fixed);\n\nconsole.log(cleaned);\n```\n\n---\n\n### Validation\n\nCheck Myanmar Unicode content and identify different types of text.\n\n```ts\nimport {\n  containsMyanmar,\n  containsLatin,\n  isMyanmarOnly,\n  isValidUnicode,\n} from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ';\n\nconsole.log(containsMyanmar(text));\nconsole.log(containsLatin(text));\nconsole.log(isMyanmarOnly(text));\nconsole.log(isValidUnicode(text));\n```\n\nAvailable validation APIs:\n\n- `containsMyanmar()`\n- `containsLatin()`\n- `isMyanmarOnly()`\n- `isValidUnicode()`\n\n---\n\n### Segmentation\n\nProvides basic Myanmar text segmentation utilities.\n\nAvailable segmentation APIs:\n\n- `segmentLine()`\n- `segmentParagraph()`\n- `segmentSentence()`\n- `segmentSyllable()`\n- `segmentWord()`\n- `tokenize()`\n\nExample:\n\n```ts\nimport {\n  segmentSentence,\n  segmentSyllable,\n  tokenize,\n} from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ။ မြန်မာစာကို လေ့လာနေပါတယ်။';\n\nconst sentences = segmentSentence(text);\n\nconst syllables = segmentSyllable(text);\n\nconst tokens = tokenize(text);\n\nconsole.log(sentences);\nconsole.log(syllables);\nconsole.log(tokens);\n```\n\n---\n\n## Installation\n\n### npm\n\n```bash\nnpm install @ainova-mya/unicode\n```\n\n### pnpm\n\n```bash\npnpm add @ainova-mya/unicode\n```\n\n### yarn\n\n```bash\nyarn add @ainova-mya/unicode\n```\n\n---\n\n## Usage\n\n### TypeScript\n\n```ts\nimport {\n  normalizeUnicode,\n  normalizeWhitespace,\n  tokenize,\n} from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ';\n\nconst clean = normalizeWhitespace(normalizeUnicode(text));\n\nconsole.log(clean);\n\nconst result = tokenize(clean);\n\nconsole.log(result);\n```\n\n### JavaScript\n\n```js\nimport {\n  normalizeUnicode,\n  normalizeWhitespace,\n  tokenize,\n} from '@ainova-mya/unicode';\n\nconst text = 'မင်္ဂလာပါ';\n\nconst clean = normalizeWhitespace(normalizeUnicode(text));\n\nconsole.log(clean);\n\nconst result = tokenize(clean);\n\nconsole.log(result);\n```\n\n---\n\n## Processing Pipeline\n\nA typical Myanmar text processing pipeline can be built by combining the provided utilities.\n\n```ts\nimport {\n  normalizeUnicode,\n  normalizeWhitespace,\n  fixConfusable,\n  removeDuplicateMarks,\n  tokenize,\n} from '@ainova-mya/unicode';\n\nconst input = 'မင်္ဂလာပါ';\n\nconst normalized = normalizeUnicode(input);\n\nconst cleaned = normalizeWhitespace(normalized);\n\nconst repaired = fixConfusable(cleaned);\n\nconst finalText = removeDuplicateMarks(repaired);\n\nconst tokens = tokenize(finalText);\n\nconsole.log(finalText);\nconsole.log(tokens);\n```\n\nThe general processing flow is:\n\n```text\nInput Text\n    ↓\nUnicode Normalization\n    ↓\nWhitespace Cleaning\n    ↓\nUnicode Repair\n    ↓\nDuplicate Mark Removal\n    ↓\nSegmentation / Tokenization\n    ↓\nProcessed Myanmar Text\n```\n\n---\n\n## Available APIs\n\n### Normalization\n\n```ts\nnormalizeUnicode(text);\nnormalizeWhitespace(text);\n```\n\n### Unicode Repair\n\n```ts\nfixConfusable(text);\nremoveDuplicateMarks(text);\n```\n\n### Validation\n\n```ts\ncontainsMyanmar(text);\ncontainsLatin(text);\nisMyanmarOnly(text);\nisValidUnicode(text);\n```\n\n### Segmentation\n\n```ts\nsegmentLine(text);\nsegmentParagraph(text);\nsegmentSentence(text);\nsegmentSyllable(text);\nsegmentWord(text);\ntokenize(text);\n```\n\n---\n\n## API Overview\n\n| API                      | Description                                      |\n| ------------------------ | ------------------------------------------------ |\n| `normalizeUnicode()`     | Normalize Myanmar Unicode text                   |\n| `normalizeWhitespace()`  | Clean whitespace and line endings                |\n| `fixConfusable()`        | Repair confusable characters                     |\n| `removeDuplicateMarks()` | Remove duplicate Myanmar marks                   |\n| `containsMyanmar()`      | Check whether text contains Myanmar characters   |\n| `containsLatin()`        | Check whether text contains Latin characters     |\n| `isMyanmarOnly()`        | Check whether text contains only Myanmar content |\n| `isValidUnicode()`       | Validate Unicode text                            |\n| `segmentLine()`          | Segment text by lines                            |\n| `segmentParagraph()`     | Segment text by paragraphs                       |\n| `segmentSentence()`      | Segment text by sentences                        |\n| `segmentSyllable()`      | Segment Myanmar text into syllables              |\n| `segmentWord()`          | Segment Myanmar text into words                  |\n| `tokenize()`             | Tokenize Myanmar text                            |\n\n---\n\n## TypeScript Support\n\nThe library is written with TypeScript support and provides TypeScript types.\n\n```ts\nimport { normalizeUnicode, tokenize } from '@ainova-mya/unicode';\n\nconst text: string = 'မင်္ဂလာပါ';\n\nconst normalized: string = normalizeUnicode(text);\n\nconst tokens = tokenize(normalized);\n\nconsole.log(tokens);\n```\n\n---\n\n## Use Cases\n\n### OCR Processing\n\nUseful as a post-processing layer for OCR output.\n\n```text\nOCR Output\n    ↓\nUnicode Normalization\n    ↓\nWhitespace Cleaning\n    ↓\nUnicode Repair\n    ↓\nMyanmar Text\n```\n\n### Search\n\nThe library can be used as a preprocessing layer before indexing Myanmar text into search engines.\n\n```text\nRaw Text\n    ↓\nNormalize\n    ↓\nClean\n    ↓\nSegment / Tokenize\n    ↓\nSearch Index\n```\n\n### AI / RAG\n\nMyanmar text can be normalized and cleaned before embedding or retrieval.\n\n```text\nDocument\n    ↓\nText Extraction\n    ↓\nMyanmar Unicode Processing\n    ↓\nText Segmentation\n    ↓\nChunking\n    ↓\nEmbedding\n    ↓\nVector Database\n```\n\n### Document Processing\n\nIt can also be used as a foundation layer in document processing pipelines.\n\n```text\nPDF / DOCX / Image\n        ↓\n    Text Extraction\n        ↓\nMyanmar Unicode Processing\n        ↓\n    Text Segmentation\n        ↓\n    Structured Text\n```\n","readmeFilename":"README.md"}