{"_id":"@akshatowo/pdf-search","_rev":"2-d1d9e38782914b6a81711717739257ca","name":"@akshatowo/pdf-search","dist-tags":{"latest":"0.1.1"},"versions":{"0.1.0":{"name":"@akshatowo/pdf-search","version":"0.1.0","keywords":["cli","pdf","search"],"author":{"name":"Akshat Singh","email":"akshatsinghdelhi@gmail.com"},"license":"MIT","_id":"@akshatowo/pdf-search@0.1.0","maintainers":[{"name":"akshatowo","email":"akshatsinghdelhi@gmail.com"}],"homepage":"https://github.com/akshat-OwO/pdf-search","bugs":{"url":"https://github.com/akshat-OwO/pdf-search/issues"},"bin":{"pdf-search":"dist/index.mjs"},"dist":{"shasum":"3e290d04a3f850c4cfa7a38be9b44e491645da30","tarball":"https://registry.npmjs.org/@akshatowo/pdf-search/-/pdf-search-0.1.0.tgz","fileCount":4,"integrity":"sha512-rmw5tTKQQ1i1sV767UewGUbCwtqGK7cES7GOCcFf2ss+OaBIOG8N27IlG3njkinIT3XMGQEsAiFtmVMLOB16+w==","signatures":[{"sig":"MEYCIQDL3wejB2b1FLS2Kk1wdvJX8jPGSFIMiBgCqgbixmnYsQIhAJOBn2QjpDNaSGaG18y1sGY4EQjpTbWMXV/1rSijnwS2","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":26897},"type":"module","_from":"file:akshatowo-pdf-search-0.1.0.tgz","exports":{".":"./dist/index.mjs","./package.json":"./package.json"},"scripts":{"dev":"vp pack --watch","test":"vp test","build":"vp pack","check":"vp check"},"_npmUser":{"name":"akshatowo","email":"akshatsinghdelhi@gmail.com"},"_resolved":"/private/var/folders/yv/c82gq2pd5vl1h1z28h8bv3p00000gn/T/00ab0de011bb85d148c2ab34b537a38a/akshatowo-pdf-search-0.1.0.tgz","_integrity":"sha512-rmw5tTKQQ1i1sV767UewGUbCwtqGK7cES7GOCcFf2ss+OaBIOG8N27IlG3njkinIT3XMGQEsAiFtmVMLOB16+w==","repository":{"url":"git+https://github.com/akshat-OwO/pdf-search.git","type":"git"},"_npmVersion":"11.9.0","description":"CLI for searching text in PDF files with page-numbered results.","directories":{},"_nodeVersion":"24.14.0","dependencies":{"pdfjs-dist":"^5.5.207"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"bumpp":"^11.0.1","pdf-lib":"^1.17.1","vite-plus":"^0.1.11","typescript":"^5.9.3","@types/node":"^25.5.0","@typescript/native-preview":"7.0.0-dev.20260316.1"},"_npmOperationalInternal":{"tmp":"tmp/pdf-search_0.1.0_1774375804366_0.13207319309524745","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"name":"@akshatowo/pdf-search","version":"0.1.1","description":"CLI for searching text in PDF files with page-numbered results.","keywords":["cli","pdf","search"],"homepage":"https://github.com/akshat-OwO/pdf-search","bugs":{"url":"https://github.com/akshat-OwO/pdf-search/issues"},"license":"MIT","author":{"name":"Akshat Singh","email":"akshatsinghdelhi@gmail.com"},"repository":{"type":"git","url":"git+https://github.com/akshat-OwO/pdf-search.git"},"bin":{"pdf-search":"dist/index.mjs"},"type":"module","exports":{".":"./dist/index.mjs","./package.json":"./package.json"},"publishConfig":{"access":"public"},"scripts":{"build":"vp pack","dev":"vp pack --watch","test":"vp test","check":"vp check","prepublishOnly":"vp run build","prepare":"vp config"},"dependencies":{"pdfjs-dist":"^5.5.207"},"devDependencies":{"@types/node":"^25.5.0","@typescript/native-preview":"7.0.0-dev.20260316.1","bumpp":"^11.0.1","pdf-lib":"^1.17.1","typescript":"^5.9.3","vite-plus":"^0.1.11"},"engines":{"node":">=18"},"packageManager":"pnpm@10.32.1","pnpm":{"overrides":{"vite":"npm:@voidzero-dev/vite-plus-core@latest","vitest":"npm:@voidzero-dev/vite-plus-test@latest"}},"gitHead":"370e1ce5f49150d81946676d33e48e0b8d895f22","_id":"@akshatowo/pdf-search@0.1.1","_nodeVersion":"24.14.0","_npmVersion":"11.9.0","dist":{"integrity":"sha512-6AMd+k90BODJIflMpDtahwblMfFjIvsLTyCQIbKoJzn3bAV89uwoSABuAid2KgLNm2GpX+dbI9dCkB3cnDxNRA==","shasum":"15c53c4805d4b939c3d206d3c706a82e75667944","tarball":"https://registry.npmjs.org/@akshatowo/pdf-search/-/pdf-search-0.1.1.tgz","fileCount":4,"unpackedSize":35043,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@akshatowo%2fpdf-search@0.1.1","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQD6vB5R+zWf1Bx6FXObr43os4ARqTtDL0W0xvY64DjM1QIhAOm4y2lPOMErjh3MlKV4jMV0Z9AmhJc8qKOCE/v179+1"}]},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:d6d379a6-ec9d-4133-bb66-8a3d8fd5624c"}},"directories":{},"maintainers":[{"name":"akshatowo","email":"akshatsinghdelhi@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/pdf-search_0.1.1_1774415014135_0.032889204300075026"},"_hasShrinkwrap":false}},"time":{"created":"2026-03-24T18:10:04.275Z","modified":"2026-03-25T05:03:34.579Z","0.1.0":"2026-03-24T18:10:04.530Z","0.1.1":"2026-03-25T05:03:34.281Z"},"bugs":{"url":"https://github.com/akshat-OwO/pdf-search/issues"},"author":{"name":"Akshat Singh","email":"akshatsinghdelhi@gmail.com"},"license":"MIT","homepage":"https://github.com/akshat-OwO/pdf-search","keywords":["cli","pdf","search"],"repository":{"type":"git","url":"git+https://github.com/akshat-OwO/pdf-search.git"},"description":"CLI for searching text in PDF files with page-numbered results.","maintainers":[{"name":"akshatowo","email":"akshatsinghdelhi@gmail.com"}],"readme":"# pdf-search\n\n`pdf-search` is a Node CLI for searching text in a PDF and printing matches with\ntheir page numbers. You can also print the extracted text for a single page by\nnumber without scanning every page (PDF.js resolves the page directly).\n\n## Install\n\nFrom npm:\n\n```bash\npnpm add -g @akshatowo/pdf-search\n# or: npm install -g @akshatowo/pdf-search\n```\n\nFrom a clone (contributors):\n\n```bash\npnpm install\n```\n\n## Usage\n\n```bash\npdf-search <pdfPathOrUrl> <query> [options]\npdf-search <pdfPathOrUrl> --and <term> [--and <term> ...] [--or <term> ...] [options]\npdf-search --page <number> <pdfPathOrUrl>\n```\n\n`<pdfPathOrUrl>` may be a filesystem path, a `file://` URL, or an `http://` / `https://` URL to a PDF. Each CLI invocation that loads a remote URL downloads the file again (there is no cross-call cache). The programmatic API behaves the same: `loadPdfDocument` and `getPdfPageText` fetch on every call unless you pass local bytes or a path yourself.\n\n### Options\n\n- `-p, --page <number>`: print extracted text for one page (1-based); do not combine with search arguments or `--context` / `--context-chars` / `--concurrency`\n- `--page-format <mode>`: only with `--page`. **`compact`** (default): single-line style text like search extraction. **`layout`**: keep line breaks from the PDF text runs. **`json`**: one JSON object per line with `page` and `text` (the `text` field uses **layout** formatting)\n- `--and <term>`: require a page to contain this term; repeat for all required terms\n- `--or <term>`: require a page to contain at least one optional term; repeat as needed\n- `-c, --context`: show a short snippet around each match\n- `--context-chars <number>`: control how much surrounding text is shown\n- `--concurrency <number>`: for **local** PDFs (and the bundled CLI), how many **worker threads** share the scan—this is where you usually see a wall-clock speedup on multi-core machines. For **`http(s)` URLs**, the file is scanned **in the main process** (to avoid copying the whole download into every worker), so this flag has **little effect on runtime**; matches and ordering stay the same.\n- `--fetch-timeout-ms <n>`: for `http(s)` PDFs only, abort if the response is not received in time; default is 120000 when omitted. Use `0` to disable the timeout\n- `--max-fetch-bytes <n>`: for `http(s)` PDFs only, reject responses larger than _n_ bytes (checked while streaming)\n- `-h, --help`: print usage help\n\nSearch terms are matched as case-insensitive substrings. Normal runs suppress\nrecoverable PDF parser warnings, and progress is shown on stderr while the file\nis being scanned.\n\n### Security note\n\nIf you pass URLs from untrusted users (for example inside a web application), fetching them can carry **SSRF** risk (internal addresses, redirect chains) and **resource** risk (very large downloads). Prefer allowlists, trusted hosts, or downloading out-of-band. Remote fetches use a default timeout; use `--max-fetch-bytes` (CLI) or `maxFetchBytes` (API) when you need an explicit body size limit.\n\n## Examples\n\nSearch a PDF and print page-level match counts:\n\n```bash\npdf-search \"./docs/guide.pdf\" \"worker threads\"\n```\n\nRequire pages to contain both terms:\n\n```bash\npdf-search \"./docs/guide.pdf\" --and \"worker threads\" --and \"memory pressure\"\n```\n\nRequire pages to contain `worker threads` and at least one of two related terms:\n\n```bash\npdf-search \"./docs/guide.pdf\" --and \"worker threads\" --or \"benchmarking\" --or \"throughput\"\n```\n\nSearch a PDF and show surrounding text for each hit:\n\n```bash\npdf-search \"./docs/guide.pdf\" \"worker threads\" --context\n```\n\nIncrease snippet length and worker-thread concurrency:\n\n```bash\npdf-search \"./docs/guide.pdf\" \"worker threads\" --context --context-chars 80 --concurrency 6\n```\n\nDump normalized text for page 3 only (stdout is suitable for piping):\n\n```bash\npdf-search --page 3 \"./docs/guide.pdf\"\n```\n\nPreserve line structure, or emit JSON for scripts:\n\n```bash\npdf-search --page 3 --page-format layout \"./docs/guide.pdf\"\npdf-search --page 3 --page-format json \"./docs/guide.pdf\"\n```\n\n`--concurrency` defaults to a bounded worker-thread count based on CPU cores, and\nis capped by the total page count to avoid oversubscription. That default matters\nmost for **local** PDFs when the bundled CLI can use worker threads. For\n**remote** `http(s)` PDFs, scanning stays in-process, so changing `--concurrency`\nusually does not improve wall-clock time much.\n\n## Example Output\n\nDefault mode:\n\n```text\nPDF: guide.pdf\nQuery: \"worker threads\"\nPages scanned: 42\nMatches found: 3\n\nPage 4: 1 match\nPage 18: 2 matches\n```\n\nMulti-term mode:\n\n```text\nPDF: guide.pdf\nQuery: all of \"worker threads\"; any of \"benchmarking\", \"throughput\"\nPages scanned: 42\nMatches found: 4\n\nPage 18: 2 matches\nPage 27: 2 matches\n```\n\nContext mode:\n\n```text\nPDF: guide.pdf\nQuery: \"worker threads\"\nPages scanned: 42\nMatches found: 3\n\nPage 4\n  1. ...processing pool uses worker threads to keep the search responsive...\n\nPage 18\n  1. ...a bounded worker threads strategy reduces memory pressure...\n  2. ...benchmarking worker threads across pages improves throughput...\n```\n\n## Development\n\nRun the test suite:\n\n```bash\nvp test\n```\n\nBuild the CLI:\n\n```bash\nvp pack\n```\n\nRun checks:\n\n```bash\nvp check\n```\n\nRun a synthetic concurrency benchmark:\n\n```bash\nvp run bench\n```\n","readmeFilename":"README.md"}