{"_id":"@aliste-sdk/observability","_rev":"4-24267bd32da74ed4808fa4a5cd21d1f4","name":"@aliste-sdk/observability","dist-tags":{"latest":"1.2.1"},"versions":{"1.0.0":{"name":"@aliste-sdk/observability","version":"1.0.0","keywords":["observability","opentelemetry","otel","prometheus","metrics","tracing","aliste"],"author":{"name":"Aliste Technologies","email":"tech@alistetechnologies.com"},"license":"MIT","_id":"@aliste-sdk/observability@1.0.0","maintainers":[{"name":"aliste-tech","email":"tech@alistetechnologies.com"},{"name":"gouravgupta02","email":"gouravguptagg02@gmail.com"}],"homepage":"https://github.com/alistetechnologies/aliste-sdk#readme","bugs":{"url":"https://github.com/alistetechnologies/aliste-sdk/issues"},"dist":{"shasum":"eda68ee817f78c1bfa3c0dda22cb466e5ec8c5e9","tarball":"https://registry.npmjs.org/@aliste-sdk/observability/-/observability-1.0.0.tgz","fileCount":38,"integrity":"sha512-PYIFqDq3Xu2DuGwoH0f2Ce75WZCnyHLvFjnVL1Tnk/CaDdMEmv6jJekwPpdnro6Q+k1w3qsk18GWzyCyJQ8Zyw==","signatures":[{"sig":"MEUCIQDR8VxSoPjXVlnkjr0qJxE/O7T8LCXuy0xYH9oIwl/y/QIgOmXZJPtBUnGOmHtppgys4gMrLXJ6khQkHuI4oVKbGh0=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":48230},"main":"dist/index.js","types":"dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"73efb16c71abf89cdce550776b4aeaf685a3dbe1","scripts":{"build":"tsc","clean":"rm -rf dist","build:watch":"tsc --watch","prepublishOnly":"npm run clean && npm run build"},"_npmUser":{"name":"gouravgupta02","email":"gouravguptagg02@gmail.com"},"repository":{"url":"git+https://github.com/alistetechnologies/aliste-sdk.git","type":"git","directory":"packages/observability"},"_npmVersion":"11.11.0","description":"Production-grade observability SDK — tracing (OpenTelemetry) and metrics (Prometheus)","directories":{},"_nodeVersion":"24.14.1","dependencies":{"prom-client":"^15.1.3","@opentelemetry/api":"^1.9.0","@opentelemetry/sdk-node":"^0.57.0","@opentelemetry/resources":"^1.30.0","@opentelemetry/sdk-trace-base":"^1.30.0","@opentelemetry/semantic-conventions":"^1.28.0","@opentelemetry/exporter-trace-otlp-http":"^0.57.0","@opentelemetry/auto-instrumentations-node":"^0.57.0"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"express":"^4.21.0","typescript":"^5.7.0","@types/node":"^22.0.0","@types/express":"^5.0.0"},"peerDependencies":{"express":"^4.18.0 || ^5.0.0"},"_npmOperationalInternal":{"tmp":"tmp/observability_1.0.0_1776937603783_0.004395567203164807","host":"s3://npm-registry-packages-npm-production"}},"1.1.0":{"name":"@aliste-sdk/observability","version":"1.1.0","keywords":["observability","opentelemetry","otel","prometheus","metrics","tracing","aliste"],"author":{"name":"Aliste Technologies","email":"tech@alistetechnologies.com"},"license":"MIT","_id":"@aliste-sdk/observability@1.1.0","maintainers":[{"name":"aliste-tech","email":"tech@alistetechnologies.com"},{"name":"gouravgupta02","email":"gouravguptagg02@gmail.com"}],"homepage":"https://github.com/alistetechnologies/aliste-sdk#readme","bugs":{"url":"https://github.com/alistetechnologies/aliste-sdk/issues"},"dist":{"shasum":"2d48421617844f69cf51734a70385c486b1fe9a7","tarball":"https://registry.npmjs.org/@aliste-sdk/observability/-/observability-1.1.0.tgz","fileCount":42,"integrity":"sha512-BmcD/NxOfInWxPU1Jp/y+Ptt585fNrP+zPbdmvsl/hIu9pC6IsSod4RIxPRf8YJ0LEeWdTprZDvWL/9WKk9QZA==","signatures":[{"sig":"MEQCIHK/NL2e0cPrAlxEKRL/vLUW9QrC8kXdPBBVxknTRZbsAiBu/WPfxK4lr+3dkiItE5Tw+ghnAyzZlmeQMbB/ApjwGg==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":58389},"main":"dist/index.js","types":"dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"4f5d4a9a90a585f100e5bd44cde40ebb19eba72e","scripts":{"build":"tsc","clean":"rm -rf dist","build:watch":"tsc --watch","prepublishOnly":"npm run clean && npm run build"},"_npmUser":{"name":"gouravgupta02","email":"gouravguptagg02@gmail.com"},"repository":{"url":"git+https://github.com/alistetechnologies/aliste-sdk.git","type":"git","directory":"packages/observability"},"_npmVersion":"11.11.0","description":"Production-grade observability SDK — tracing (OpenTelemetry) and metrics (Prometheus)","directories":{},"_nodeVersion":"24.14.1","dependencies":{"prom-client":"^15.1.3","@opentelemetry/api":"^1.9.0","@opentelemetry/sdk-node":"^0.57.0","@opentelemetry/resources":"^1.30.0","@opentelemetry/sdk-trace-base":"^1.30.0","@opentelemetry/semantic-conventions":"^1.28.0","@opentelemetry/exporter-trace-otlp-http":"^0.57.0","@opentelemetry/auto-instrumentations-node":"^0.57.0"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"express":"^4.21.0","typescript":"^5.7.0","@types/node":"^22.0.0","@types/express":"^5.0.0"},"peerDependencies":{"express":"^4.18.0 || ^5.0.0"},"_npmOperationalInternal":{"tmp":"tmp/observability_1.1.0_1777463267665_0.6876790563424058","host":"s3://npm-registry-packages-npm-production"}},"1.2.0":{"name":"@aliste-sdk/observability","version":"1.2.0","keywords":["observability","opentelemetry","otel","prometheus","metrics","tracing","aliste"],"author":{"name":"Aliste Technologies","email":"tech@alistetechnologies.com"},"license":"MIT","_id":"@aliste-sdk/observability@1.2.0","maintainers":[{"name":"aliste-tech","email":"tech@alistetechnologies.com"},{"name":"gouravgupta02","email":"gouravguptagg02@gmail.com"}],"homepage":"https://github.com/alistetechnologies/aliste-sdk#readme","bugs":{"url":"https://github.com/alistetechnologies/aliste-sdk/issues"},"dist":{"shasum":"d2bda0022aed897ce5cd45c76b85d5826865c4de","tarball":"https://registry.npmjs.org/@aliste-sdk/observability/-/observability-1.2.0.tgz","fileCount":42,"integrity":"sha512-zO+ow19QSPqbOproJ1xX+MnUYJ0w/Cg9m/dBgOC+rrrvUUZm+RhkGTiyHU/JzI2iyRywoR/IrsXX5bR5VAMYZg==","signatures":[{"sig":"MEUCIDimfItIiiP1pYv1zi+XeCMr4LXbtSuOTSkM5kp30mwiAiEAkRUbU8iLBZY29LATjh1q45N4gZqLNnGLlLPKkXSmw4E=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":60660},"main":"dist/index.js","types":"dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"4f5d4a9a90a585f100e5bd44cde40ebb19eba72e","scripts":{"build":"tsc","clean":"rm -rf dist","build:watch":"tsc --watch","prepublishOnly":"npm run clean && npm run build"},"_npmUser":{"name":"gouravgupta02","email":"gouravguptagg02@gmail.com"},"repository":{"url":"git+https://github.com/alistetechnologies/aliste-sdk.git","type":"git","directory":"packages/observability"},"_npmVersion":"11.11.0","description":"Production-grade observability SDK — tracing (OpenTelemetry) and metrics (Prometheus)","directories":{},"_nodeVersion":"24.14.1","dependencies":{"prom-client":"^15.1.3","@opentelemetry/api":"^1.9.0","@opentelemetry/sdk-node":"^0.57.0","@opentelemetry/resources":"^1.30.0","@opentelemetry/sdk-trace-base":"^1.30.0","@opentelemetry/semantic-conventions":"^1.28.0","@opentelemetry/exporter-trace-otlp-http":"^0.57.0","@opentelemetry/auto-instrumentations-node":"^0.57.0"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"express":"^4.21.0","typescript":"^5.7.0","@types/node":"^22.0.0","@types/express":"^5.0.0"},"peerDependencies":{"express":"^4.18.0 || ^5.0.0"},"_npmOperationalInternal":{"tmp":"tmp/observability_1.2.0_1777961309448_0.20750845758267755","host":"s3://npm-registry-packages-npm-production"}},"1.2.1":{"name":"@aliste-sdk/observability","version":"1.2.1","description":"Production-grade observability SDK — tracing (OpenTelemetry) and metrics (Prometheus)","main":"dist/index.js","types":"dist/index.d.ts","scripts":{"build":"tsc","build:watch":"tsc --watch","clean":"rm -rf dist","prepublishOnly":"npm run clean && npm run build"},"keywords":["observability","opentelemetry","otel","prometheus","metrics","tracing","aliste"],"author":{"name":"Aliste Technologies","email":"tech@alistetechnologies.com"},"license":"MIT","publishConfig":{"access":"public"},"repository":{"type":"git","url":"git+https://github.com/alistetechnologies/aliste-sdk.git","directory":"packages/observability"},"dependencies":{"@opentelemetry/api":"1.9.0","@opentelemetry/auto-instrumentations-node":"0.57.0","@opentelemetry/exporter-trace-otlp-http":"0.57.0","@opentelemetry/resources":"1.30.0","@opentelemetry/sdk-node":"0.57.0","@opentelemetry/sdk-trace-base":"1.30.0","@opentelemetry/semantic-conventions":"1.28.0","prom-client":"15.1.3"},"peerDependencies":{"express":"^4.18.0 || ^5.0.0"},"devDependencies":{"@types/express":"^5.0.0","@types/node":"^22.0.0","express":"^4.21.0","typescript":"^5.7.0"},"engines":{"node":">=18.0.0"},"gitHead":"268bce5c577e918798f3ec3742b3a6f1eb777204","_id":"@aliste-sdk/observability@1.2.1","bugs":{"url":"https://github.com/alistetechnologies/aliste-sdk/issues"},"homepage":"https://github.com/alistetechnologies/aliste-sdk#readme","_nodeVersion":"24.14.1","_npmVersion":"11.11.0","dist":{"integrity":"sha512-JcqMSMQJ8o+/4zk+S4kD7QwmahWx7W0Iq+9aosxvCdv9ztCJTUGsCLKEigc6HPwm5tldmmsbzBpt7GGOtRKuvA==","shasum":"f64d2cde41827215e0c27039967ead76e6d5d344","tarball":"https://registry.npmjs.org/@aliste-sdk/observability/-/observability-1.2.1.tgz","fileCount":42,"unpackedSize":60652,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQDIWNNOEAfi5rAj/ZOqiIf5S+hspXSOGt/okNpMpqZi4QIgPoMMg/mIhuutNP2KcvVYqe5AcbpvHDFFP9fhPrSxnV0="}]},"_npmUser":{"name":"gouravgupta02","email":"gouravguptagg02@gmail.com"},"directories":{},"maintainers":[{"name":"aliste-tech","email":"tech@alistetechnologies.com"},{"name":"gouravgupta02","email":"gouravguptagg02@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/observability_1.2.1_1777965898192_0.9542515793645896"},"_hasShrinkwrap":false}},"time":{"created":"2026-04-23T09:46:43.311Z","modified":"2026-05-05T07:24:58.445Z","1.0.0":"2026-04-23T09:46:43.927Z","1.1.0":"2026-04-29T11:47:47.804Z","1.2.0":"2026-05-05T06:08:29.598Z","1.2.1":"2026-05-05T07:24:58.332Z"},"bugs":{"url":"https://github.com/alistetechnologies/aliste-sdk/issues"},"author":{"name":"Aliste Technologies","email":"tech@alistetechnologies.com"},"license":"MIT","homepage":"https://github.com/alistetechnologies/aliste-sdk#readme","keywords":["observability","opentelemetry","otel","prometheus","metrics","tracing","aliste"],"repository":{"type":"git","url":"git+https://github.com/alistetechnologies/aliste-sdk.git","directory":"packages/observability"},"description":"Production-grade observability SDK — tracing (OpenTelemetry) and metrics (Prometheus)","maintainers":[{"name":"aliste-tech","email":"tech@alistetechnologies.com"},{"name":"gouravgupta02","email":"gouravguptagg02@gmail.com"}],"readme":"# @aliste-sdk/observability\n\nProduction-grade observability for Node.js + Express services. Drop in OpenTelemetry tracing and Prometheus metrics with a two-call setup.\n\n---\n\n## Quick Start\n\n> **Assumes:** Node.js ≥ 18, an Express application, and an OTLP-compatible collector (e.g. [OpenTelemetry Collector](https://opentelemetry.io/docs/collector/)) reachable at `OTEL_EXPORTER_OTLP_ENDPOINT`.\n\n### CommonJS (`require`)\n\nThe SDK compiles to CommonJS and works directly with `require()`. In CJS there is no import hoisting, so `require()` calls execute in the exact order they appear — no dynamic import trick needed.\n\nThe two calls have different homes:\n- `initTracing()` — in the **entry point** (`server.js` / `bin/www`), before any other `require()`\n- `initMetrics(app)` — inside **`app.js`**, before route definitions\n\n```js\n// server.js — process entry point\nconst { initTracing } = require(\"@aliste-sdk/observability\");\n\ninitTracing({\n  serviceName: process.env.OTEL_SERVICE_NAME,\n  serviceVersion: process.env.SERVICE_VERSION,\n  deploymentEnvironment: process.env.NODE_ENV,\n  otlpEndpoint: process.env.OTEL_EXPORTER_OTLP_ENDPOINT,\n  samplerRatio: Number(process.env.OTEL_TRACES_SAMPLER_ARG ?? \"1.0\"),\n});\n\n// express, http, mongoose, etc. all load here — after OTel patches are applied\nconst app = require(\"./app\");\nconst http = require(\"http\");\n\nhttp.createServer(app).listen(Number(process.env.PORT ?? 3000));\n```\n\n```js\n// app.js\nconst express = require(\"express\");\nconst { initMetrics } = require(\"@aliste-sdk/observability\");\n\nconst app = express();\n\n// Must come before any route definitions\ninitMetrics(app, {\n  serviceName: process.env.OTEL_SERVICE_NAME,\n  serviceVersion: process.env.SERVICE_VERSION,\n  deploymentEnvironment: process.env.NODE_ENV,\n});\n\napp.get(\"/users/:id\", (req, res) => { ... });\n// ... rest of routes\n\nmodule.exports = app;\n```\n\n### ESM / TypeScript (`import`)\n\nIn ESM, static `import` statements are hoisted before any code runs. Use a dynamic `import()` to control load order.\n\n```ts\n// index.ts — process entry point\nimport { initTracing } from '@aliste-sdk/observability';\n\ninitTracing();\n\n// Dynamic import runs AFTER initTracing() patches are applied\nimport('./app').then(({ default: app }) => app.listen(3000));\n```\n\n```ts\n// app.ts\nimport express from 'express';\nimport { initMetrics } from '@aliste-sdk/observability';\n\nconst app = express();\ninitMetrics(app);\n\nexport default app;\n```\n\n```bash\nOTEL_SERVICE_NAME=my-service \\\nOTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 \\\nNODE_ENV=production \\\nnode dist/index.js\n```\n\n---\n\n## Table of Contents\n\n1. [Architecture Overview](#architecture-overview)\n2. [Features](#features)\n3. [Installation](#installation)\n4. [Initialization — Critical Order](#initialization--critical-order)\n5. [Metrics Usage](#metrics-usage)\n6. [Environment Variables](#environment-variables)\n7. [Service Identity and Labels](#service-identity-and-labels)\n8. [Route Normalization](#route-normalization)\n9. [Double Initialization Safety](#double-initialization-safety)\n10. [Security Considerations](#security-considerations)\n11. [Example Usage](#example-usage)\n12. [Troubleshooting](#troubleshooting)\n\n---\n\n## Architecture Overview\n\nThe SDK exposes two completely separate initialization flows because tracing and metrics have fundamentally different initialization timing requirements.\n\n### Why `initTracing()` must run before any other imports\n\nOpenTelemetry auto-instrumentation works by monkey-patching modules **at require/import time**. When `@opentelemetry/auto-instrumentations-node` loads, it intercepts `require()` calls and wraps the target module (e.g. `express`, `http`, `pg`) with span-generating code. If Express is already loaded into memory before OTel starts, those patches never apply — you get a running process with zero traces.\n\n`initTracing()` is synchronous and has no side effects on the rest of your application. It must be the literal first statement in your process entry file, before any `import` that could trigger the loading of an instrumentable module.\n\n### Why `initMetrics()` is separate\n\n`initMetrics(app)` requires an Express `Application` instance to register the `/metrics` endpoint and attach the request-tracking middleware. That instance does not exist until after Express is imported and instantiated. Prometheus metrics do not require pre-import timing — they simply need to be wired up before the server starts accepting traffic.\n\n### Lifecycle separation\n\n```\nProcess start\n    │\n    ▼\ninitTracing()           ← must be first; patches http, express, etc.\n    │\n    ▼\nimport express\ncreate app\n    │\n    ▼\ninitMetrics(app)        ← registers /metrics + attaches middleware\n    │\n    ▼\napp.listen()\n```\n\n---\n\n## Features\n\n- **OpenTelemetry tracing** via `NodeSDK` with `getNodeAutoInstrumentations`\n- **Prometheus metrics** via `prom-client` using an **isolated custom `Registry`** (never the global singleton)\n- **Express middleware** that tracks request duration, active requests, response size, and errors\n- **Route normalization** that strips numeric IDs, UUIDs, phone numbers, and alphanumeric device IDs from label values before they reach Prometheus\n- **Double-initialization safety** — both `initTracing` and `createCollectors` are idempotent\n- **Environment-driven configuration** with optional programmatic overrides\n- **Graceful shutdown** on `SIGTERM` via `sdk.shutdown()`\n\n---\n\n## Installation\n\n```bash\nnpm install @aliste-sdk/observability\n```\n\nExpress is a peer dependency. Install it separately if not already present:\n\n```bash\nnpm install express\nnpm install --save-dev @types/express\n```\n\n---\n\n## Initialization — Critical Order\n\nThe only hard rule is: **`initTracing()` must run before any `require()` or `import` that loads an instrumentable module** (express, http, mongoose, pg, etc.). How you achieve that depends on whether your codebase uses CommonJS or ESM.\n\n### CommonJS — straightforward\n\n`require()` is synchronous and sequential — no hoisting, no dynamic import tricks needed.\n\n**`initTracing()`** goes in the entry point (`server.js` / `bin/www`), as the very first statement before any other `require()`.\n\n**`initMetrics(app)`** goes inside `app.js`, before route definitions. This is the critical constraint: Express processes middleware in the order it was registered. If `initMetrics` is called after routes are defined — even by just a few lines — the metrics middleware ends up behind route handlers in the stack. Route handlers send responses without calling `next()`, so the middleware never runs and no metrics are recorded.\n\n```js\n// server.js — only initTracing here\nconst { initTracing } = require('@aliste-sdk/observability');\n\ninitTracing({\n  serviceName: 'my-service',\n  serviceVersion: '2.1.0',\n  otlpEndpoint: 'http://otel-collector:4318',\n  samplerRatio: 0.1,\n});\n\nconst app = require('./app'); // express loads here, after OTel patches\nconst http = require('http');\nhttp.createServer(app).listen(3000);\n```\n\n```js\n// app.js — initMetrics before routes\nconst express = require('express');\nconst { initMetrics } = require('@aliste-sdk/observability');\n\nconst app = express();\n\ninitMetrics(app, { serviceName: 'my-service' }); // ← must be before routes\n\napp.get('/users/:id', (req, res) => { ... });    // ← routes after\n// ...\nmodule.exports = app;\n```\n\n### ESM / TypeScript — requires a dynamic import\n\nIn ESM, static `import` declarations are **hoisted** to the top of the module by the JavaScript engine before any code executes. This means even if you write `initTracing()` before an `import app from './app'` line, the runtime loads `app.ts` (and therefore express) first. The only way to control load order in ESM is a dynamic `import()`.\n\n```ts\n// index.ts\nimport { initTracing } from '@aliste-sdk/observability';\n\ninitTracing({ serviceName: 'my-service', samplerRatio: 0.1 });\n\n// Dynamic import: app module is resolved only after this point\nimport('./app').then(({ default: app }) => {\n  app.listen(3000);\n});\n```\n\n### Incorrect usage — tracing will be broken\n\n```ts\n// WRONG (ESM) — static import is hoisted; express loads before initTracing() runs\nimport { initTracing } from '@aliste-sdk/observability';\nimport app from './app'; // hoisted — express already loaded here\n\ninitTracing(); // too late\n```\n\n```js\n// WRONG (CJS) — initTracing called after app loads\nconst app = require('./app'); // express already loaded here\n\nconst { initTracing } = require('@aliste-sdk/observability');\ninitTracing(); // too late: patches never applied\n```\n\n```js\n// WRONG (CJS) — initMetrics called after routes are defined\nconst app = require('./app'); // all routes already registered inside app.js\n\nconst { initMetrics } = require('@aliste-sdk/observability');\ninitMetrics(app); // middleware is behind routes in the stack — never fires\n```\n\nWhen tracing is initialized too late, `NodeSDK.start()` still runs without error, but the auto-instrumentation hooks were never registered, so no spans are generated for HTTP requests.\n\n---\n\n## Metrics Usage\n\nCall `initMetrics(app)` after creating your Express application. It does two things:\n\n1. Registers `metricsMiddleware()` globally via `app.use()` — this must happen before any route definitions to ensure every request is tracked.\n2. Mounts a `GET /metrics` endpoint that serializes the registry in Prometheus text format.\n\n```ts\n// app.ts\nimport express from 'express';\nimport { initMetrics } from '@aliste-sdk/observability';\n\nconst app = express();\napp.use(express.json());\n\ninitMetrics(app, {\n  serviceName: 'my-service',\n  serviceVersion: '2.1.0',\n  deploymentEnvironment: 'production',\n  metricsPath: '/metrics',         // default\n  collectDefaultMetrics: true,     // default — enables Node.js process metrics\n  defaultMetricsInterval: 5000,    // default — ms between GC/event loop samples\n  // Restrict /metrics to your Prometheus server's IP or internal subnet.\n  // Falls back to METRICS_ALLOWED_IPS env var (comma-separated).\n  // may leave empty only in development\n  allowedIPs: ['10.0.0.0/8', '192.168.1.50'],\n});\n\n// Route definitions go AFTER initMetrics so they are covered by the middleware.\napp.get('/users/:id', ...);\n```\n\n### Custom Registry (not global)\n\nThe SDK creates a fresh `prom-client` `Registry` instance rather than using `prom-client`'s global default registry. This means:\n\n- Metrics from different packages or test runs cannot bleed into each other.\n- The `/metrics` endpoint only serializes this SDK's registry.\n- If you define your own application-level metrics, register them against a separate registry — not the global one.\n\n### Metrics collected\n\n| Metric | Type | Labels |\n|---|---|---|\n| `http_request_total` | Counter | `method`, `route`, `status_code` |\n| `http_request_duration_seconds` | Histogram | `method`, `route`, `status_code` |\n| `active_http_requests` | Gauge | — |\n| `http_errors_total` | Counter | `method`, `route`, `status_code` |\n| `http_response_size_bytes` | Histogram | `method`, `route`, `status_code` |\n| `http_request_size_bytes` | Histogram | `method`, `route` |\n\nPlus all standard Node.js process metrics from `collectDefaultMetrics` (heap, GC, event loop lag, file descriptors, etc.).\n\n### Middleware behavior\n\n`metricsMiddleware()` uses `res.on('finish')` to record all metrics after the response is fully sent. The request duration timer starts on the incoming request, not on `finish`. Active request count is incremented on arrival and decremented on finish. Response size is read from the `content-length` response header.\n\n**Request size** (`http_request_size_bytes`) is measured as follows:\n1. If the `content-length` request header is present, its value is used directly.\n2. Otherwise, if a parsed `req.body` is available (e.g. via `express.json()`), the size is estimated as `Buffer.byteLength(JSON.stringify(req.body))`.\n3. If neither is available (e.g. GET requests with no body), `0` is observed.\n\nIf `initMetrics()` has not been called before `metricsMiddleware()` runs, the middleware silently passes through — it will not throw.\n\n---\n\n## Environment Variables\n\nThe SDK reads directly from `process.env`. It does **not** load or parse `.env` files.\n\n### Variables read by this SDK\n\n| Variable | Used by | Default | Description |\n|---|---|---|---|\n| `OTEL_SERVICE_NAME` | tracing, metrics | `unknown-service` | Service name attached to every trace and metric |\n| `OTEL_EXPORTER_OTLP_ENDPOINT` | tracing | `http://localhost:4318` | Base URL of the OTLP collector. The SDK appends `/v1/traces` automatically. |\n| `OTEL_TRACES_SAMPLER_ARG` | tracing | `1.0` | Sampling ratio as a float between `0` and `1`. Values outside this range are clamped. |\n| `OTEL_DEBUG` | tracing | `false` | Set to `\"true\"` to enable verbose OTel diagnostic output to stdout |\n| `SERVICE_VERSION` | tracing, metrics | `0.0.0` | Application version attached to every trace and metric |\n| `NODE_ENV` | tracing, metrics | `development` | Deployment environment (`production`, `staging`, `development`) |\n| `INSTANCE_ID` | tracing, metrics | `${HOSTNAME}-${PID}` | Unique instance identifier. Auto-generated from `HOSTNAME` and `process.pid` if not set. |\n| `HOSTNAME` | tracing, metrics | `local` | Used as part of the auto-generated instance ID when `INSTANCE_ID` is not set |\n| `METRICS_ALLOWED_IPS` | metrics | *(none — open)* | Comma-separated list of IPs or CIDR ranges allowed to scrape `/metrics`. Example: `\"10.0.0.1,10.0.0.0/8\"`. Takes effect only when `allowedIPs` is not set in the config object. If empty, a warning is logged. |\n\n### Variables honored by the underlying OTel SDK (not this SDK's config resolver)\n\n| Variable | Notes |\n|---|---|\n| `OTEL_RESOURCE_ATTRIBUTES` | Additional resource attributes in `key=value,key=value` format. Merged by `NodeSDK` at startup. |\n\n### Variables NOT honored\n\n| Variable | Why |\n|---|---|\n| `OTEL_TRACES_SAMPLER` | This SDK always configures a `ParentBasedSampler` wrapping `TraceIdRatioBasedSampler`. The env var is ignored because an explicit sampler is passed to `NodeSDK`. |\n\n### Precedence\n\nFor every configuration value: **config object argument > environment variable > built-in default**.\n\n---\n\n## Service Identity and Labels\n\nFour identity fields are resolved at initialization time and attached to both telemetry systems.\n\n### OpenTelemetry Resource attributes\n\nAttached to the `Resource` passed to `NodeSDK`. Every span exported by this process carries these attributes:\n\n| OTel attribute | Source field |\n|---|---|\n| `service.name` | `serviceName` |\n| `service.version` | `serviceVersion` |\n| `deployment.environment` | `deploymentEnvironment` |\n| `service.instance.id` | `instanceId` |\n\n### Prometheus default labels\n\nAttached via `registry.setDefaultLabels()`. Every metric scraped from `/metrics` carries these labels:\n\n| Prometheus label key | Source field |\n|---|---|\n| `service` | `serviceName` |\n| `version` | `serviceVersion` |\n| `environment` | `deploymentEnvironment` |\n| `instance` | `instanceId` |\n\nNote that Prometheus label keys use flat names (no dots) because dots are not valid in Prometheus label names. The OTel attributes use the standard dotted semantic convention names.\n\n---\n\n## Route Normalization\n\n### Why high-cardinality labels are dangerous\n\nPrometheus stores one time series per unique label combination. If a label like `route` contains raw user-supplied values such as `/users/12345` or `/orders/550e8400-e29b-41d4-a716-446655440000`, every unique ID creates a new time series. A service with millions of users would produce millions of time series, causing Prometheus memory exhaustion, slow queries, and alerting failures. This is called **cardinality explosion**.\n\n### How `normalizeRoute` works\n\nThe utility in [src/utils/route.ts](src/utils/route.ts) applies a single general-purpose regex to every path segment, then strips the query string:\n\n```\n/(?:\\d+|\\+\\d{7,}|[a-zA-Z0-9+.\\-]*\\d{4,}[a-zA-Z0-9]*)(?=\\/|$)/g\n```\n\nA segment is replaced with `:id` when it is:\n\n- **All digits** — `123`, `42`\n- **E.164 phone number** — `+919106521492` (leading `+` followed by 7+ digits)\n- **Alphanumeric with 4+ consecutive digits** — `H322556`, `DEV0042`, `550e8400-e29b-41d4-a716-446655440000`\n\nThe 4-digit threshold preserves short version-style segments like `v3`, `v12`, or `house2` (fewer than 4 consecutive digits). Query strings are stripped before the regex runs.\n\n### Examples\n\n| Raw `req.originalUrl` | Normalized route |\n|---|---|\n| `/users/123` | `/users/:id` |\n| `/users/123?expand=profile` | `/users/:id` |\n| `/users/507f1f77bcf86cd799439011` | `/users/:id` |\n| `/users/550e8400-e29b-41d4-a716-446655440000` | `/users/:id` |\n| `/devices/H322556` | `/devices/:id` |\n| `/sms/+919106521492` | `/sms/:id` |\n| `/api/v1/orders` | `/api/v1/orders` |\n| `/health` | `/health` |\n\nThe middleware uses `req.originalUrl` (the full unmodified path) with `req.path` as a fallback.\n\n---\n\n## Double Initialization Safety\n\n### Tracing\n\n```ts\n// src/tracing/tracer.ts\nlet sdk: NodeSDK | null = null;\n\nexport function initTracing(config?: TracingConfig): void {\n  if (sdk !== null) return;  // ← no-op on second call\n  ...\n}\n```\n\n### Metrics\n\n```ts\n// src/metrics/collectors.ts\nlet collectors: MetricCollectors | null = null;\n\nexport function createCollectors(config: ResolvedMetricsConfig): MetricCollectors {\n  if (collectors !== null) return collectors;  // ← returns existing instance\n  ...\n}\n```\n\nBoth guards use module-level singleton variables. This matters because:\n\n- In Node.js module caching is per-process, so calling `initTracing()` twice in the same process would attempt to start two `NodeSDK` instances and register duplicate SIGTERM handlers.\n- Calling `createCollectors()` twice would attempt to register duplicate metric names in the same registry, causing a `prom-client` error.\n- The singleton pattern makes the SDK safe to call from library code, middleware, or test setup without worrying about call ordering.\n\n---\n\n## Security Considerations\n\n### `/metrics` endpoint exposure\n\nThe `/metrics` endpoint exposes internal runtime data: memory usage, GC pause durations, event loop lag, request rates, error rates, and response times. This information is valuable to an attacker for understanding service behavior and load patterns.\n\n**The SDK includes a built-in IP allowlist for `/metrics`.**\n\nConfigure it via the `allowedIPs` option or the `METRICS_ALLOWED_IPS` environment variable:\n\n```ts\ninitMetrics(app, {\n  allowedIPs: ['10.0.0.0/8', '192.168.1.50'],  // exact IPs or CIDR ranges\n});\n```\n\n```bash\nMETRICS_ALLOWED_IPS=\"10.0.0.0/8,192.168.1.50\" node server.js\n```\n\nWhen `allowedIPs` is non-empty, any request to `/metrics` from an IP not in the list receives `HTTP 403 Forbidden`. Both exact IPv4 addresses and CIDR ranges are supported. IPv6-mapped IPv4 addresses (e.g. `::ffff:10.0.0.1`) are normalized before comparison.\n\nIf `allowedIPs` is empty, the SDK logs a warning at startup:\n\n```\n[observability] /metrics is open to all IPs. Set METRICS_ALLOWED_IPS to restrict access.\n```\n\nFor defense in depth, combine the IP allowlist with:\n\n- **Network policy / firewall** — Allow scrape traffic only from your Prometheus server's IP range.\n- **Reverse proxy rule** — In Nginx or your ingress controller, block `/metrics` from public-facing traffic.\n\nNever expose `/metrics` directly on a public-facing port.\n\n---\n\n## Example Usage\n\nThe [examples/](examples/) directory contains working examples for both module systems.\n\n### CommonJS\n\n| File | Purpose |\n|---|---|\n| [examples/server.js](examples/server.js) | Entry point — calls `initTracing()`, then `require('./app-cjs')` |\n| [examples/app-cjs.js](examples/app-cjs.js) | Express app — calls `initMetrics(app)` before route definitions |\n\n### TypeScript / ESM\n\n| File | Purpose |\n|---|---|\n| [examples/index.ts](examples/index.ts) | Entry point — calls `initTracing()`, then dynamic `import('./app')` |\n| [examples/app.ts](examples/app.ts) | Express app with `initMetrics(app)` and sample routes |\n\nAll examples define the same routes, which exercise the route normalizer:\n\n| Defined route | Incoming request | Metric label |\n|---|---|---|\n| `GET /users/:id` | `GET /users/42` | `GET /users/:id` |\n| `GET /users/:id` | `GET /users/507f1f77bcf86cd799439011` | `GET /users/:id` |\n| `POST /orders` | `POST /orders` | `POST /orders` |\n\n---\n\n## Troubleshooting\n\n### Traces are not appearing in the collector\n\n**Cause:** `initTracing()` was called after Express or another instrumentable module was already loaded.\n\n**How to confirm:** Add `debug: true` to the `initTracing` config (or set `OTEL_DEBUG=true`). If you see no span creation logs for incoming HTTP requests, auto-instrumentation did not apply.\n\n**Fix:** Ensure `initTracing()` is the very first statement in the process entry file and that the app module is loaded via a dynamic `import()`.\n\n---\n\n### Metrics are missing from `/metrics`\n\n**Cause 1:** `initMetrics(app)` was never called.\n\n**Cause 2:** `metricsMiddleware()` was called directly without first calling `initMetrics()`. The middleware checks `getCollectors()` and silently passes through if collectors have not been created — no error is thrown, but no metrics are recorded.\n\n**Fix:** Always call `initMetrics(app)` before starting the server. Verify the `/metrics` endpoint responds with HTTP 200 and `Content-Type: text/plain`.\n\n---\n\n### Duplicate metric registration error from prom-client\n\n**Cause:** Two separate Node.js module instances of `@aliste-sdk/observability` are loaded in the same process (e.g. due to `npm link`, a broken monorepo hoisting configuration, or a nested `node_modules`). Each module instance has its own `collectors` module-level variable, so the singleton guard does not protect across instances.\n\n**How to confirm:** Run `npm ls @aliste-sdk/observability` in the consuming project. If you see more than one version or path, there are duplicate instances.\n\n**Fix:** Ensure the package is deduplicated. In a monorepo, hoist the package to the root `node_modules`.\n\n---\n\n### `/metrics` returns HTTP 403\n\n**Cause:** The scraping client's IP is not in the `allowedIPs` list (or `METRICS_ALLOWED_IPS` env var).\n\n**How to confirm:** Check the IP of your Prometheus server or the machine making the scrape request. Compare it against the configured allowlist.\n\n**Fix:** Add the scraper's IP or its subnet to `allowedIPs`:\n\n```ts\ninitMetrics(app, {\n  allowedIPs: ['10.0.0.0/8', '192.168.1.50'],\n});\n```\n\nOr via the environment variable:\n\n```bash\nMETRICS_ALLOWED_IPS=\"10.0.0.0/8,192.168.1.50\"\n```\n\nTo allow all IPs temporarily during development, set `allowedIPs: []` (or leave `METRICS_ALLOWED_IPS` unset). A warning will be logged.\n\n---\n\n### OTLP exporter connection refused\n\n**Cause:** `OTEL_EXPORTER_OTLP_ENDPOINT` points to an unreachable collector.\n\n**Behavior:** The SDK does not crash on export failure. Spans are batched and silently dropped after retry exhaustion.\n\n**Fix:** Verify the collector is reachable from the service at the configured endpoint. The SDK appends `/v1/traces` to the base URL automatically — do not include it in `OTEL_EXPORTER_OTLP_ENDPOINT`.\n\n---\n\n## License\n\n[MIT](../../LICENSE)\n","readmeFilename":"README.md"}