{"_id":"@ankurkumarz/pi-gpu-sre-ops","name":"@ankurkumarz/pi-gpu-sre-ops","dist-tags":{"latest":"0.2.0"},"versions":{"0.2.0":{"name":"@ankurkumarz/pi-gpu-sre-ops","version":"0.2.0","description":"Runtime-configurable Pi package for safe GPU inference SRE, token efficiency analysis, and proposal-only optimization workflows.","type":"module","private":false,"bin":{"gpuagent":"bin/gpuagent.js"},"keywords":["pi-package","pi-agent","gpu","sre","llm","inference","vllm","sglang","triton","observability","token-optimization"],"pi":{"extensions":["./extensions"],"skills":["./skills"],"prompts":["./prompts"]},"scripts":{"pack:check":"npm pack --dry-run","test:syntax":"node --check extensions/index.ts && node --check extensions/token-inference-optimizer.ts","test:pi":"pi -e ./extensions/index.ts","publish:check":"npm run test:syntax && npm run pack:check"},"peerDependencies":{"@earendil-works/pi-coding-agent":"*","typebox":"*"},"publishConfig":{"access":"public"},"license":"MIT","_id":"@ankurkumarz/pi-gpu-sre-ops@0.2.0","_nodeVersion":"24.13.0","_npmVersion":"11.6.2","dist":{"integrity":"sha512-eo8bg4Wlwif3Z+WXUwyFdsQM3BecFLAUQ+jeVMiYUqPZ9EGrgr0PYuAHYC+S4iiJuzWrXnCgBLfJHkhqmRz04w==","shasum":"ae76a73459742ddf79cfdab574f6ccf0d4932f42","tarball":"https://registry.npmjs.org/@ankurkumarz/pi-gpu-sre-ops/-/pi-gpu-sre-ops-0.2.0.tgz","fileCount":22,"unpackedSize":42694,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQDO+dGXuVjzJcNlWrArI3r617dAz1oVRsY3gIipKBDafAIgSzNVPVNz7N6+ni+e3+gW6zZUttQFrnwV9US7LA1rEI8="}]},"_npmUser":{"name":"ankurkumarz","email":"ankurkumar78@gmail.com"},"directories":{},"maintainers":[{"name":"ankurkumarz","email":"ankurkumar78@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/pi-gpu-sre-ops_0.2.0_1788916122366_0.6388548757308705"},"_hasShrinkwrap":false}},"time":{"created":"2026-09-09T01:08:42.251Z","0.2.0":"2026-09-09T01:08:42.524Z","modified":"2026-09-09T01:08:42.889Z"},"maintainers":[{"name":"ankurkumarz","email":"ankurkumar78@gmail.com"}],"description":"Runtime-configurable Pi package for safe GPU inference SRE, token efficiency analysis, and proposal-only optimization workflows.","keywords":["pi-package","pi-agent","gpu","sre","llm","inference","vllm","sglang","triton","observability","token-optimization"],"license":"MIT","readme":"# @ankurkumarz/pi-gpu-sre-ops\n\nA safe, runtime-configurable Pi package for GPU inference SRE, incident triage, token-efficiency analysis, and proposal-only optimization workflows.\n\n## What 0.1.0 includes\n\n- `gpu-incident-triage` skill for evidence-based incident investigation.\n- `token-inference-optimization` skill for prefill, decode, cache, scheduler, and capacity analysis.\n- Runtime-neutral profile contract supporting vLLM, SGLang, Triton, TensorRT-LLM, TGI, and custom runtimes.\n- Read-only tools for runtime profile lookup, inference health, token efficiency, and optimization proposals.\n- Mock mode by default—no live request is made unless an authenticated Ops Gateway is configured.\n- Proposal-only optimization plans with canary, approval, and rollback requirements.\n- Guardrails that block destructive shell patterns and disable general shell execution in production.\n\n## Security boundary\n\nThis public package contains generic skills, schemas, mock data, and connector contracts only. Do **not** include client endpoints, credentials, kubeconfigs, raw production logs, raw prompts, private runbooks, or customer data.\n\nA production Ops Gateway must enforce tenant/service authorization, RBAC, secret handling, query validation, redaction, audit logging, approval validation, and configuration rollout controls. The Pi package must never be the sole production security boundary.\n\n## Install\n\n```bash\npi install npm:@ankurkumarz/pi-gpu-sre-ops@0.1.0\n```\n\nFor a project-local installation:\n\n```bash\npi install -l npm:@ankurkumarz/pi-gpu-sre-ops@0.1.0\n```\n\n## Test locally\n\n```bash\ncp .env.example .env\npi -e ./extensions/index.ts\n```\n\nWithout `OPS_GATEWAY_URL` and `OPS_GATEWAY_TOKEN`, all operational tools return clearly marked fixture data.\n\nExample prompt:\n\n```text\nUse gpu-incident-triage for incident INC-1234 in tenant personal-lab,\nservice llm-inference, environment dev, from 09:00 to 10:00 PDT.\nThen use token-inference-optimization and draft a proposal-only canary plan.\n```\n\n## Validate package contents\n\n```bash\nnpm install\nnpm run publish:check\n```\n\nReview the `npm pack --dry-run` file list before publishing. It must not contain `.env`, credentials, customer information, private runbooks, or generated local artifacts.\n\n## Publish\n\n```bash\nnpm login\nnpm whoami\nnpm publish --access public\n```\n\nAfter any published version, npm does not allow republishing the same name/version. Increment the version before the next release:\n\n```bash\nnpm version patch\nnpm publish --access public\n```\n\n## Client configuration model\n\nKeep real, Git-versioned client runtime profiles in a private control-plane repository. Validate YAML with JSON Schema and policy-as-code in CI, then publish signed/traceable normalized profiles to an authenticated Ops Gateway. The agent should receive only the profile resolved for its authorized incident, tenant, service, and environment.\n","readmeFilename":"README.md","_rev":"1-d4afd3275dbb389514239602f0639aa4"}