{"_id":"@captemulation/grapheme-splitter","_rev":"2-56daacd7b912e2b679f2985590275034","name":"@captemulation/grapheme-splitter","description":"A JavaScipt library that breaks strings into their individual user-perceived characters. It supports emojis!","dist-tags":{"latest":"1.0.0"},"versions":{"1.0.0":{"name":"@captemulation/grapheme-splitter","version":"1.0.0","description":"A JavaScipt library that breaks strings into their individual user-perceived characters. It supports emojis!","homepage":"https://github.com/orling/grapheme-splitter","author":{"name":"Orlin Georgiev"},"contributors":[{"name":"Lucas Tadeu Teixeira","email":"lucas@fastmail.nl","url":"https://lucas.is"}],"main":"index.js","license":"MIT","keywords":["utf-8","strings","emoji","split"],"scripts":{"test":"tape tests/grapheme_splitter_tests.js"},"repository":{"type":"git","url":"git+https://github.com/orling/grapheme-splitter.git"},"bugs":{"url":"https://github.com/orling/grapheme-splitter/issues"},"dependencies":{},"devDependencies":{"tape":"^4.6.3"},"engines":{"npm":"~7.3.0"},"gitHead":"42176f0f4b379cb647f3837289592a3eb7f7b0f5","_id":"@captemulation/grapheme-splitter@1.0.0","_shasum":"2c9888ada8ddf8ead46a68f46faae82fdf6e3482","_from":".","_npmVersion":"3.10.10","_nodeVersion":"6.9.4","_npmUser":{"name":"captemulation","email":"john.holmes.dean@gmail.com"},"dist":{"shasum":"2c9888ada8ddf8ead46a68f46faae82fdf6e3482","tarball":"https://registry.npmjs.org/@captemulation/grapheme-splitter/-/grapheme-splitter-1.0.0.tgz","integrity":"sha512-Aqa4MuQN+zM/nlScBMlIXRtJwZ874twE5j+HQq14mHaH7iV1Y+jJdOPNTzpZu64UDb5qTgIJld+HadUdO+OtJA==","signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEQCIE9WOetiRfAUkZ1ZYqWAGt6zKocijqbr/TFUEdLvJeFVAiAhSCdGs3M7JwZx+di4E9X9lS0qWi1+w3I2xyuWIATFGQ=="}]},"maintainers":[{"name":"captemulation","email":"john.holmes.dean@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/grapheme-splitter-1.0.0.tgz_1497903476574_0.7255026970524341"}}},"readme":"# Background\n\nIn JavaScript there is not always a one-to-one relationship between string characters and what a user would call a separate visual \"letter\". Some symbols are represented by several characters. This can cause issues when splitting strings and inadvertently cutting a multi-char letter in half, or when you need the actual number of letters in a string.\n\nFor example, emoji characters like \"🌷\",\"🎁\",\"💩\",\"😜\" and \"👍\" are represented by two JavaScript characters each (high surrogate and low surrogate). That is, \n\n```javascript\n\"🌷\".length == 2\n```\n\nWhat's more, some languages often include combining marks - characters that are used to modify the letters before them. Common examples are the German letter ü and the Spanish letter ñ. Sometimes they can be represented alternatively both as a single character and as a letter + combining mark, with both forms equally valid:\n    \n```javascript\nvar two = \"ñ\"; // unnormalized two-char n+◌̃  , i.e. \"\\u006E\\u0303\";\nvar one = \"ñ\"; // normalized single-char, i.e. \"\\u00F1\"\nconsole.log(one!=two); // prints 'true'\n```\n\nUnicode normalization, as performed by the popular punycode.js library or ECMAScript 6's String.normalize, can **sometimes** fix those differences and turn two-char sequences into single characters. But it is **not** enough in all cases. Some languages like Hindi make extensive use of combining marks on their letters, that have no dedicated single-codepoint Unicode sequences, due to the sheer number of possible combinations.\nFor example, the Hindi word \"अनुच्छेद\" is comprised of 5 letters and 3 combining marks:\n\nअ + न + ु + च + ् + छ + े + द\n\nwhich is in fact just 5 user-perceived letters:\n\nअ + नु + च् + छे + द\n\nand which Unicode normalization would not combine properly.\nThere are also the unusual letter+combining mark combinations which have no dedicated Unicode codepoint. The string Z͑ͫ̓ͪ̂ͫ̽͏̴̙̤̞͉͚̯̞̠͍A̴̵̜̰͔ͫ͗͢L̠ͨͧͩ͘G̴̻͈͍͔̹̑͗̎̅͛́Ǫ̵̹̻̝̳͂̌̌͘ obviously has 5 separate letters, but is in fact comprised of 58 JavaScript characters, most of which are combining marks.\n\nEnter the grapheme-splitter.js library. It can be used to properly split JavaScript strings into what a human user would call separate letters (or \"extended grapheme clusters\" in Unicode terminology), no matter what their internal representation is. It is an implementation of the Unicode UAX-29 standard. \n\n# Installation\n\nTo install `grapheme-splitter` to your project, use the NPM command below:\n\n```\n$ npm install --save grapheme-splitter\n```\n\n# Tests\n\nTo run the tests on `grapheme-splitter`, use the command below:\n\n```\n$ npm test\n```\n\n# Usage\n\nJust initialize and use:\n\n```javascript\nvar splitter = new GraphemeSplitter();\n\n// split the string to an array of grapheme clusters (one string each)\nvar graphemes = splitter.splitGraphemes(string);\n\n// or do this if you just need their number\nvar graphemeCount = splitter.countGraphemes(string);\n```\n\n# Examples\n\n```javascript\nvar splitter = new GraphemeSplitter();\n\n// plain latin alphabet - nothing spectacular\nsplitter.splitGraphemes(\"abcd\"); // returns [\"a\", \"b\", \"c\", \"d\"]\n\n// two-char emojis and four-char country flag\nsplitter.splitGraphemes(\"🌷🎁💩😜👍🇺🇸\"); // returns [\"🌷\",\"🎁\",\"💩\",\"😜\",\"👍\",\"🇺🇸\"]\n\n// diacritics as combining marks, 10 JavaScript chars\nsplitter.splitGraphemes(\"Ĺo͂ře᷒m̅\"); // returns [\"Ĺ\",\"o͂\",\"ř\",\"e᷒\",\"m̅\"]\n\n// individual Korean characters (Jamo), 4 JavaScript chars\nsplitter.splitGraphemes(\"뎌쉐\"); // returns [\"뎌\",\"쉐\"]\n\n// Hindi text with combining marks, 8 JavaScript chars\nsplitter.splitGraphemes(\"अनुच्छेद\"); // returns [\"अ\",\"नु\",\"च्\",\"छे\",\"द\"]\n\n// demonic multiple combining marks, 75 JavaScript chars\nsplitter.splitGraphemes(\"Z͑ͫ̓ͪ̂ͫ̽͏̴̙̤̞͉͚̯̞̠͍A̴̵̜̰͔ͫ͗͢L̠ͨͧͩ͘G̴̻͈͍͔̹̑͗̎̅͛́Ǫ̵̹̻̝̳͂̌̌͘!͖̬̰̙̗̿̋ͥͥ̂ͣ̐́́͜͞\"); // returns [\"Z͑ͫ̓ͪ̂ͫ̽͏̴̙̤̞͉͚̯̞̠͍\",\"A̴̵̜̰͔ͫ͗͢\",\"L̠ͨͧͩ͘\",\"G̴̻͈͍͔̹̑͗̎̅͛́\",\"Ǫ̵̹̻̝̳͂̌̌͘\",\"!͖̬̰̙̗̿̋ͥͥ̂ͣ̐́́͜͞\"]\n```\n\n# Acknowledgements\n\nThis library is heavily influenced by Devon Govett's excellent grapheme-breaker CoffeeScript library at https://github.com/devongovett/grapheme-breaker with an emphasis on ease of integration and pure JavaScript implementation.\n\n\n\n","maintainers":[{"name":"captemulation","email":"john.holmes.dean@gmail.com"}],"time":{"modified":"2022-06-12T15:40:55.415Z","created":"2017-06-19T20:17:57.704Z","1.0.0":"2017-06-19T20:17:57.704Z"},"homepage":"https://github.com/orling/grapheme-splitter","keywords":["utf-8","strings","emoji","split"],"repository":{"type":"git","url":"git+https://github.com/orling/grapheme-splitter.git"},"contributors":[{"name":"Lucas Tadeu Teixeira","email":"lucas@fastmail.nl","url":"https://lucas.is"}],"author":{"name":"Orlin Georgiev"},"bugs":{"url":"https://github.com/orling/grapheme-splitter/issues"},"license":"MIT","readmeFilename":"README.md"}