{"_id":"@agulbra/uts58","name":"@agulbra/uts58","dist-tags":{"latest":"0.2.3"},"versions":{"0.2.3":{"name":"@agulbra/uts58","version":"0.2.3","description":"UTS #58 web-link extraction for JavaScript.","type":"module","main":"src/index.js","types":"src/index.d.ts","exports":{".":{"types":"./src/index.d.ts","default":"./src/index.js"},"./iana":{"types":"./src/index.d.ts","default":"./src/index-iana.js"},"./core":{"types":"./src/core.d.ts","default":"./src/core.js"}},"sideEffects":false,"engines":{"node":">=18"},"scripts":{"test":"node --test test/*.test.js","maketables":"node tools/maketables.js","maketlds":"node tools/maketlds.js","prepack":"cp ../LICENSE LICENSE","postpack":"rm -f LICENSE"},"keywords":["uts58","url","extractor","idn","twitter-text"],"author":{"name":"Arnt Gulbrandsen","email":"arnt@gulbrandsen.priv.no"},"license":"BSD-2-Clause","publishConfig":{"access":"public"},"repository":{"type":"git","url":"git+https://github.com/arnt/uts58.git","directory":"javascript"},"bugs":{"url":"https://github.com/arnt/uts58/issues"},"homepage":"https://github.com/arnt/uts58#readme","dependencies":{"punycode":"^2.3.1"},"_id":"@agulbra/uts58@0.2.3","_integrity":"sha512-2mTNkQHIr0zOCLcKU1FlF3dslAdSHZjIRWY+RkhthZ75KZRZELJHivmkdpevHgjmzBVNhf/2yUAvpwG4HsELCw==","_resolved":"/home/arnt/src/uts58/javascript/agulbra-uts58-0.2.3.tgz","_from":"file:/home/arnt/src/uts58/javascript/agulbra-uts58-0.2.3.tgz","_nodeVersion":"24.16.0","_npmVersion":"11.16.0","dist":{"integrity":"sha512-2mTNkQHIr0zOCLcKU1FlF3dslAdSHZjIRWY+RkhthZ75KZRZELJHivmkdpevHgjmzBVNhf/2yUAvpwG4HsELCw==","shasum":"dbdb7a24a358e1fc03acb115100ffef5535efc46","tarball":"https://registry.npmjs.org/@agulbra/uts58/-/uts58-0.2.3.tgz","fileCount":14,"unpackedSize":72236,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIB+eulFOnL5FJvIF5O34ULYu9CenHad0bs5kO9dh3LtfAiBiSE3Ktu2OsozY4xnySYZrxFpxW4Yy0ldqGCfjFvxRFw=="}]},"_npmUser":{"name":"agulbra","email":"arnt@gulbrandsen.priv.no"},"directories":{},"maintainers":[{"name":"agulbra","email":"arnt@gulbrandsen.priv.no"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/uts58_0.2.3_1782390517730_0.8684139324682272"},"_hasShrinkwrap":false}},"time":{"created":"2026-06-25T12:28:37.578Z","0.2.3":"2026-06-25T12:28:37.886Z","modified":"2026-06-25T12:28:38.093Z"},"maintainers":[{"name":"agulbra","email":"arnt@gulbrandsen.priv.no"}],"description":"UTS #58 web-link extraction for JavaScript.","homepage":"https://github.com/arnt/uts58#readme","keywords":["uts58","url","extractor","idn","twitter-text"],"repository":{"type":"git","url":"git+https://github.com/arnt/uts58.git","directory":"javascript"},"author":{"name":"Arnt Gulbrandsen","email":"arnt@gulbrandsen.priv.no"},"bugs":{"url":"https://github.com/arnt/uts58/issues"},"license":"BSD-2-Clause","readme":"# uts58\n\nA JavaScript implementation of [UTS58](https://www.unicode.org/reports/tr58/),\nthe Unicode spec for finding links in running text. Given a chunk of text, it\nreturns the URLs in it along with their UTF-16 offsets. This is a port of\nthe [Ruby `uts58` gem](https://github.com/arnt/uts58); the test suite is the\nsame, with very minor differences.\n\nTested extensively on relevant OSes: [![CI](https://github.com/arnt/uts58/actions/workflows/javascript.yml/badge.svg)](https://github.com/arnt/uts58/actions/workflows/javascript.yml)\n\n## Install\n\n```sh\nnpm install @agulbra/uts58\n```\n\nESM-only; Node 18+ (uses Unicode property escapes and lookbehinds).\n\n## Usage\n\n```js\nimport { extractUrls, extractUrlsWithIndices } from '@agulbra/uts58';\n\nextractUrlsWithIndices('see https://example.com/ for details');\n// => [{ url: 'https://example.com/', indices: [4, 24] }]\n\nextractUrls('see https://example.com/ for details');\n// => ['https://example.com/']\n```\n\nThe API mirrors `twitter-text`'s `Extractor#extractUrlsWithIndices` closely;\nit was written to provide what Mastodon-style consumers need. The two\ntop-level functions above also strip partly overlapping matches via\n`Extractor#removeOverlappingEntities`: candidates are sorted by start\noffset and anything that begins inside an earlier survivor's span is\ndropped. Length doesn't enter into it — the earlier-starting candidate\nwins even if a later one is longer. Import `Extractor` directly if you'd\nrather merge with other extractors (mentions, hashtags, …) and resolve\noverlap across all of them yourself.\n\nUnlike `twitter-text`, the functions take no options object. What counts as\na link is fixed by [UTS58](https://www.unicode.org/reports/tr58/); there is\nno `extractUrlsWithoutProtocol`-style switch, because the spec already says\nhow scheme-less input is handled (see below).\n\nInput without a scheme is recognised, and `https://` is prepended in the\nreturned `url`:\n\n```js\nextractUrlsWithIndices('blogspot.com is still a thing');\n// => [{ url: 'https://blogspot.com', indices: [0, 12] }]\n```\n\nIDNs are decoded to use UTF-8 in the output, for better readability:\n\n```js\nextractUrls('xn-----ctdbabcfhu9c2b9l1acccr4c.xn--mgbah1a3hjkrd')[0];\n// => 'https://تجربة-القبول-الشامل.موريتانيا'\n```\n\n(Admittedly that output isn't very readable if you can't read Arabic. But\nthe input wasn't readable to anyone, no matter what languages they can read.)\n\nTrailing punctuation, balanced brackets, ports, paths, queries and\nfragments are handled per the spec. Indices in the output are UTF-16\ncode unit offsets — the units used by `String.prototype.slice`,\n`String#length`, the DOM, and text editors — so `text.slice(start,\nend)` returns the matched substring directly, even across characters\noutside the BMP. (The Ruby gem reports codepoint offsets instead, which\nare idiomatic for Ruby strings; `example.com/🐪/#camel` is a good test\nof the difference, since the emoji is one codepoint but two UTF-16\nunits.)\n\n## Email addresses\n\n```js\nimport {\n  extractEmailAddresses,\n  extractEmailAddressesWithIndices,\n  extractEntities,\n  extractEntitiesWithIndices,\n} from '@agulbra/uts58';\n\nextractEmailAddressesWithIndices('contact info@grå.org today');\n// => [{ email: 'info@grå.org', url: 'mailto:info@grå.org', indices: [8, 20] }]\n\nextractEmailAddresses('contact info@grå.org today');\n// => ['info@grå.org']\n```\n\nEach result carries both the bare `email` and a `mailto:` `url`, so it\ndrops straight into anything that already renders a `url` entity. The\ndomain is IDN-decoded the same way as in `extractUrls`, and a leading\n`mailto:` in the input is absorbed into `indices` per UTS58 5.2.\n\nA plain address overlaps the bare domain that `extractUrls` would find\nafter the `@`. `extractEntitiesWithIndices` runs both extractors,\nsorts by start offset, and removes overlaps — the earlier-starting\nemail wins over the domain inside it:\n\n```js\nextractEntities('mail arnt@grå.org or see blogspot.com');\n// => ['mailto:arnt@grå.org', 'https://blogspot.com']\n```\n\n## Choosing the public-suffix check\n\nTo decide whether `something.example` is a plausible link, the extractor\nchecks the host against a public-suffix table. Which table is a bundle-time\nchoice — pick the entry point that fits, and your bundler ships only that\ntable:\n\n| import | table | gzipped |\n| --- | --- | --- |\n| `@agulbra/uts58` (default) | Public Suffix List, ICANN section | ~5 KB |\n| `@agulbra/uts58/iana` | IANA root-zone TLDs | ~5 KB |\n| `@agulbra/uts58/core` | none — you supply the check | 0 KB |\n\nAll three expose the same API. The two tables are about the same size and\nagree on nearly every host — the difference is how strict the check is for\nthe handful of TLDs that only register at the second level. `@agulbra/uts58/iana`\nasks only \"is the rightmost label a real TLD\": enough to tell `blogspot.jp`\nfrom `blogspot.exe` and to reject typos like `example.cmo`, but it treats a\nbare `foo.za` as plausible. The default `@agulbra/uts58` carries the PSL, which knows\nSouth Africa registers under `co.za` / `org.za` and so rejects a bare\n`foo.za`. If that distinction doesn't matter to you, the tables are\ninterchangeable.\n\nNeither reproduces the PSL's wildcard/exception rules: the question here is\n\"could this be a link\", not \"where is the exact registrable boundary\", so a\nflat membership test is all it does. The PSL table is folded accordingly —\n`!` exceptions dropped, `*.foo` collapsed to `foo`, and any suffix made\nredundant by a shorter one removed (with `no` present, `møre-og-romsdal.no`\nis dropped). That folding is what keeps it down to ~5 KB.\n\n`@agulbra/uts58/core` bundles no table at all. Bring your own check — over a suffix\nset, or wrapping a library you already depend on:\n\n```js\nimport { Extractor } from '@agulbra/uts58/core';\nimport { parse } from 'tldts';\n\nconst ex = new Extractor({\n  isPlausibleHost: (host) => {\n    const p = parse(host);\n    return !!p.domain && p.isIcann && p.publicSuffix !== 'invalid';\n  },\n});\nex.extractUrls('see example.com here');\n```\n\nThat route is also how you get exact PSL semantics back, at the cost of a\ndependency you choose rather than one this package forces on you. (`@agulbra/uts58`\nitself depends only on `punycode`.)\n\n## Suggested test cases and notable behaviour\n\nA few sharp edges worth covering in your own tests if you're swapping\n`twitter-text` out, or just using this from scratch.\n\n**The `href` and the visible text are not the same string.** For\n`see example.com here`, the `indices` span\n`example.com`, but `url` is the longer `https://example.com`.\nUse `url` for the `href` attribute, slice the original text by\n`indices` for the visible content. A test that compares\n`text.slice(start, end) === url` will fail on every scheme-less input,\nand on every IDN where A-labels were decoded.\n\n**No options object.** There is no `extractUrlsWithoutProtocol: false`\nswitch. If you want only scheme-bearing URLs, filter on the matched\nsubstring (not on `url`, which always carries `https://`):\n\n```js\nextractUrlsWithIndices(text)\n  .filter((r) => /^https?:\\/\\//i.test(text.slice(...r.indices)));\n```\n\n**No `autoLink`.** This package extracts; it doesn't render. There's\nno equivalent of `twitter-text`'s `autoLink`, `autoLinkUrlsCustom`,\n`htmlEscape`, etc. — building HTML is the caller's job, which keeps\nescaping decisions where they belong.\n\n**No mentions, hashtags, cashtags, replies.** UTS58 doesn't define\nthem, so this package doesn't either. If you need them, run\n`twitter-text` (or another extractor) for those alongside this one and\nmerge with `Extractor#removeOverlappingEntities`.\n\n**Overlap resolution is start-wins, not longest-wins.** Worth a test\nwhen you merge entities from multiple extractors. `ask\nalice@example.com/02074960909 for details` shows why. The raw\nextractors find both email `alice@example.com` at `[4, 21]` and url\n`https://example.com/02074960909` at `[10, 33]`. Start-wins keeps\n`alice@example.com`, which is what a reader would call right, at least\none who reads 02074960909 as a phone number.\nLongest-wins would keep the longer `https://example.com/02074960909`.\n\n**`maxLength` measures the matched input span.** Not the returned URL.\nA 12-codepoint cap keeps `blogspot.com` (12) and drops\n`https://example.com`, even though the input span and `url` happen to\nbe identical there. The asymmetry shows up the other way for\n`example.com` (11 input cp, 19 in `url`) — the cap of 12 keeps it.\n\nEven though most of the API counts in terms of UTF-16, like the String\nclass, maxLength uses codepoints. The reason is that most of the API\nis for software developers, but maxLength is for end-users, and\ncodepoints are closer to what end-users see. \"☺\" and \"😀\" are both one\ncodepoint long. (This isn't quite perfect: \"é\" may be either one or\ntwo codepoints long and \"🇳🇴\" is two. Life is hard.)\n\n**`mailto:` is absorbed into `indices`.** Per UTS58 5.2, the input\n`mailto:abcd@example.com` returns an entity whose span covers the\nwhole 23-codepoint run, not just the address. The `email` field still\nholds the bare address. If your link-rendering code assumes the span\nstarts at the local-part, mailto inputs will surprise it.\n\n## What's not here\n\n- **Link validation.** Recognised URLs are not fetched, normalised\n  beyond IDN decoding, or their hostnames checked in the DNS. There is\n  no attempt at checking for possible attacks (`Мышκин.рф` and\n  `Мышкин.рф` are both detected, note the Greek kappa in the middle of\n  the prince's name).\n\n## Regenerating the generated tables\n\n`src/constants.js` (the link-termination and bracket tables) is packed from\nthe Ruby reference's `constants.rb`:\n\n```sh\nnpm run maketables\n```\n\nThat reads `../ruby/lib/uts58/constants.rb` by default; pass a path to point\nit elsewhere. `constants.rb` is itself generated on the Ruby side, where\n`tools/maketables.rb` downloads the UTS58 data files (`LinkTerm.txt` and\n`LinkBracket.txt`) straight from unicode.org — so the source of truth is\nUnicode's published data, not a copy kept in this repo.\n\n`src/tlds-iana.js` and `src/suffixes-psl.js` are the public-suffix tables.\nWith no arguments they're fetched from their canonical sources\n([IANA](https://data.iana.org/TLD/tlds-alpha-by-domain.txt),\n[PSL](https://publicsuffix.org/list/public_suffix_list.dat)); pass local\npaths to regenerate offline:\n\n```sh\nnpm run maketlds\nnpm run maketlds -- /path/to/tlds-alpha-by-domain.txt /path/to/public_suffix_list.dat\n```\n\n## License\n\nBSD-2-Clause. See `LICENSE`.\n\nFWIW, I wrote this as part of my work at ICANN and will maintain it as\npart of the same work. (I resolve problems relating to Unicode in\ndomains, email addresses and similar, so more people, more\ncommunities, can use the internet in the way they prefer.)\n","readmeFilename":"README.md","_rev":"1-0db07985b51b2778e521a9d806f76130"}