{"_id":"@awais0198/evalgate","name":"@awais0198/evalgate","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@awais0198/evalgate","version":"0.1.0","description":"An open evaluation and audit layer for AI coding agents — scores tool-call diffs, calibrates against human overrides, and gives non-engineers a way to review flagged changes.","type":"module","license":"MIT","author":{"name":"Awais Ahmad"},"homepage":"https://github.com/awais0198/evalgate","repository":{"type":"git","url":"git+https://github.com/awais0198/evalgate.git"},"keywords":["claude-code","ai-agents","llm-evaluation","llm-as-judge","code-review","evaluation-harness"],"bin":{"evalgate":"dist/cli.js"},"main":"./dist/cli.js","scripts":{"build":"tsc -p tsconfig.json","dev":"tsx src/cli.ts","capture":"tsx src/cli.ts capture","score":"tsx src/cli.ts score","init":"tsx src/cli.ts init","test":"node --import tsx --test tests/**/*.test.ts","prepublishOnly":"npm run build && npm test"},"dependencies":{"better-sqlite3":"^11.3.0","commander":"^12.1.0","diff":"^7.0.0","yaml":"^2.5.1"},"devDependencies":{"@types/better-sqlite3":"^7.6.11","@types/diff":"^6.0.0","@types/node":"^22.5.0","tsx":"^4.19.0","typescript":"^5.5.4"},"engines":{"node":">=18"},"publishConfig":{"access":"public"},"_id":"@awais0198/evalgate@0.1.0","gitHead":"ae5e5264f54b189494a8a1007308231d89331532","types":"./dist/cli.d.ts","bugs":{"url":"https://github.com/awais0198/evalgate/issues"},"_nodeVersion":"22.22.0","_npmVersion":"10.9.4","dist":{"integrity":"sha512-2yAhTxOAX7tJdiqwucr5nzBD3+8ZwDQpoEq44YNdFJpfF76KKdjiazl+g/aGFajI0JA281jV5ntusZT7LNRseA==","shasum":"96eaedc3bccf90e1e3d89fb17f1a6c60a8040110","tarball":"https://registry.npmjs.org/@awais0198/evalgate/-/evalgate-0.1.0.tgz","fileCount":49,"unpackedSize":110243,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQCP/UpI+UzsgjPu2Njlg1bscXlq1EaU1pTt8n6+HZ0viAIgJauX/P0vx4H6PxWIXA2UyHWQ9TjcTrUIQ3lEOF8MG6M="}]},"_npmUser":{"name":"awais0198","email":"awais.builds@gmail.com"},"directories":{},"maintainers":[{"name":"awais0198","email":"awais.builds@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/evalgate_0.1.0_1784742160442_0.20714421324006427"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-22T17:42:40.330Z","0.1.0":"2026-07-22T17:42:40.590Z","modified":"2026-07-22T17:42:40.779Z"},"maintainers":[{"name":"awais0198","email":"awais.builds@gmail.com"}],"description":"An open evaluation and audit layer for AI coding agents — scores tool-call diffs, calibrates against human overrides, and gives non-engineers a way to review flagged changes.","homepage":"https://github.com/awais0198/evalgate","keywords":["claude-code","ai-agents","llm-evaluation","llm-as-judge","code-review","evaluation-harness"],"repository":{"type":"git","url":"git+https://github.com/awais0198/evalgate.git"},"author":{"name":"Awais Ahmad"},"bugs":{"url":"https://github.com/awais0198/evalgate/issues"},"license":"MIT","readme":"# evalgate 🚦\n\n**The vibe check for your AI coding agent.**\n\nClaude Code, Cursor, Copilot — they all ship diffs, run commands, and call\ntools all day. None of them tell you if what just happened was actually\n*good*. evalgate is the audit layer that does: it scores every session,\ntracks quality over time, and lets a non-engineer (PM, QA, security lead)\nreview the sketchy stuff directly — no ticket required.\n\nNo cap: it's local-first, agent-agnostic, and works with **zero API keys**\nout of the box.\n\n---\n\n## Why do I need this?\n\nBecause right now you're doing one of two things: reading every AI-written\ndiff line by line (slow, doesn't scale, defeats the point of delegating),\nor trusting it blind (fast, but eventually something ships that shouldn't\nhave). evalgate is the middle option.\n\n- **Catches the stuff that actually matters, automatically.** Risky shell\n  commands, missing test coverage, scope creep, a diff that doesn't\n  actually match what was asked for — flagged without you having to go\n  looking for it.\n- **Gets smarter about *your* standards over time.** Every time you\n  override a score, that decision becomes context for future judging. It's\n  not a fixed rubric grading you — it converges toward how your team\n  actually reviews code.\n- **Non-engineers can review without an engineering ticket.** A PM, QA\n  lead, or security reviewer can open the dashboard, see exactly what\n  changed and why it was flagged, and approve or reject it with a reason —\n  no Slack thread, no \"can someone explain this diff to me.\"\n- **Costs nothing to run.** No hosted service, no database to stand up, no\n  API key to buy credits for — it rides on the Claude Code plan you\n  already have and stores everything in a local SQLite file.\n- **Not locked to one agent.** Claude Code gets the richest signal via\n  hooks, but the git adapter scores every commit regardless of what\n  produced it — so it still works the day you (or your team) switch tools.\n\nBasically: it's the difference between finding out something went wrong\nwhen a user hits it, versus finding out at the end of the session, with a\nreason attached.\n\n---\n\n## Quickstart\n\nPublished on npm as a scoped package (the plain name was already taken) —\nthe `evalgate` command itself is unaffected, only the install line changes:\n\n```bash\nnpm install -g @awais0198/evalgate\n```\n\nBuilding from source instead (contributing, or trying an unpublished change):\n\n```bash\nnpm install && npm run build\nnpm link            # makes the `evalgate` command available everywhere\n\ncd /path/to/the/repo/you/want/watched\nevalgate init\n```\n\nThen wire it into Claude Code — one JSON block:\n\n```json\n{\n  \"hooks\": {\n    \"SessionStart\": [{ \"hooks\": [{ \"type\": \"command\", \"command\": \"evalgate session-start\" }] }],\n    \"PreToolUse\": [{ \"matcher\": \"Edit\", \"hooks\": [{ \"type\": \"command\", \"command\": \"evalgate snapshot\" }] }],\n    \"PostToolUse\": [{ \"matcher\": \"Edit|Write|Bash\", \"hooks\": [{ \"type\": \"command\", \"command\": \"evalgate capture\" }] }],\n    \"SessionEnd\": [{ \"hooks\": [{ \"type\": \"command\", \"command\": \"evalgate score --session=${session_id}\" }] }]\n  }\n}\n```\n\n`SessionStart` is the one hook that looks forward instead of back: if the\nprevious session in this repo was flagged, its reasoning gets surfaced\nstraight into the new session's context — so the agent gets a shot at not\nrepeating whatever got dinged last time, not just a human seeing it later\nin the dashboard. Silent when the last session wasn't flagged.\n\nSave that as `.claude/settings.json` in the repo you initialized (or merge\nit into an existing one). Now just work normally — nothing else to run.\n\nYou'll see a one-line result after **every** Edit/Write/Bash call, not just\nat the end — this is static checks only (free, instant), so it doesn't add\nany latency to your editing loop:\n\n```\n[evalgate] Edit src/foo.ts — ok (3 distinct file(s) touched this session)\n[evalgate] ⚠ Bash rm -rf /tmp/build — 1 risky command(s) detected: rm -rf\n```\n\nSet `EVALGATE_QUIET=1` if you only want the session-end summary. At the\nend of every session you'll also see the fuller, judged report:\n\n```\n[evalgate] session abc123: 0.87 (ok)\n  - no-risky-commands: 1.00  No risky shell commands detected in this session.\n  - test-coverage-touched: 1.00  Test files were touched alongside source changes.\n  - scope-contained: 1.00  3 distinct file(s) touched this session.\n  - intent-alignment: 0.90  Changes match the apparent refactor goal.\n  - code-quality: 0.85  Clean, no dead code left behind.\n```\n\nBelow 0.6 and it flags itself, so the stuff worth a second look actually\ngets one.\n\n---\n\n## How it works\n\n1. **Capture** 📸 — a Claude Code hook fires on every tool call and pipes\n   the event to `evalgate capture`, which normalizes it into an\n   agent-agnostic record and stores it locally in `.evalgate/evalgate.db`\n   (SQLite — zero config, nothing to stand up, no cloud dependency). On a\n   different agent? The [git adapter](#using-any-other-agent-cursor-copilot-etc)\n   covers you instead.\n2. **Score** 🧠 — at session end, evalgate runs everything recorded\n   against a rubric: deterministic static checks (risky commands, missing\n   test coverage, scope creep — instant, no LLM needed) *plus* an\n   LLM-as-judge pass for the stuff that genuinely needs judgment (does the\n   diff match the likely intent, is the code actually good or just\n   technically working).\n3. **Calibrate** 🎯 — every time a human overrides a score in the\n   dashboard, it's logged with a summary of what that session touched.\n   Future judge calls retrieve the most *similar* past overrides — not\n   just the most recent — so scoring gets closer to how your team actually\n   reviews code the longer you use it.\n4. **Review** 👀 — the dashboard lets non-engineers see and override\n   flagged sessions directly, with a reason, no engineering ticket in the\n   loop.\n\n---\n\n## Where do I actually see results?\n\nThree places, depending on how deep you want to look:\n\n- **Right in your terminal**, at the end of every Claude Code session —\n  the score block shown above, printed automatically.\n- **The dashboard** — the real place to live. Run:\n\n  ```bash\n  cd dashboard && npm install && npm run build && cd ..\n  evalgate dashboard\n  ```\n\n  Opens at `http://localhost:4787`: full session list (flagged ones\n  visible up top), click into any session for the diffs it scored and the\n  reasoning behind each criterion, plus a one-click override form. The\n  **trends** tab plots overall score over time and per-criterion movement,\n  so you can see whether quality is drifting before it becomes a pattern\n  you only notice in hindsight.\n- **The raw SQLite file**, if you want to query it yourself —\n  `.evalgate/evalgate.db` in whatever repo you initialized:\n\n  ```bash\n  sqlite3 .evalgate/evalgate.db \\\n    \"select sessionId, overallScore, flagged, scoredAt from scores order by scoredAt desc limit 10;\"\n  ```\n\n---\n\n## LLM judging — genuinely zero config\n\nThis is the part most tools make you pay for upfront. evalgate doesn't.\n\nIf the `claude` CLI is installed and logged in — which it already is if\nyou run Claude Code — evalgate quietly shells out to `claude -p` for the\ntwo judgment-based criteria (`intent-alignment`, `code-quality`). It rides\non your existing Claude Code plan. **No separate API key, no prepaid\ncredits, nothing to buy.**\n\n```bash\nwhich claude   # prints a path? judging is already on\n```\n\n**Only if you need it** — running evalgate somewhere the CLI isn't\ninstalled, like a bare CI runner — you can opt into a standalone API key\ninstead:\n\n```bash\ncp .env.example .env\n# set EVALGATE_ANTHROPIC_API_KEY in .env\n```\n\nIf a key is set, it takes priority over the CLI. If neither is available,\nevalgate doesn't break — the two judged criteria just score a neutral 0.5\nwith an explanation instead of a real judgment, so scoring degrades\ngracefully instead of failing loud.\n\n---\n\n## Using any other agent (Cursor, Copilot, etc.) {#using-any-other-agent-cursor-copilot-etc}\n\nClaude Code's hooks give the richest per-tool-call signal, but not every\nagent exposes one. The universal fallback scores each **commit**,\nregardless of what produced it — works for literally anything that commits\nto git, agent or human:\n\n```bash\ncp hooks/post-commit.sample .git/hooks/post-commit\nchmod +x .git/hooks/post-commit\n```\n\nEvery commit is now automatically captured and scored. No agent-specific\nintegration required — it doesn't care what wrote the code.\n\n### GitHub PR comments — also zero config\n\nIf the current branch has an open PR and the `gh` CLI is installed and\nlogged in, every `git-capture` run (so every commit, via the post-commit\nhook above) also posts the score as a PR comment automatically — same\nphilosophy as judging: no GitHub token to generate, it uses `gh`'s\nexisting auth. It edits its own last comment rather than spamming a new\none per commit, so the PR thread stays readable.\n\nRunning scoring as a separate CI step instead? Comment explicitly:\n\n```bash\nevalgate score --session=$(git rev-parse HEAD)\nevalgate pr-comment\n```\n\nNo open PR, or no `gh` CLI? It's silently skipped — this is a bonus on top\nof local scoring, never a requirement for it.\n\n---\n\n## Making it grade the way you want\n\nWant a different bar for what counts as \"good\"? Rubrics are just YAML —\ndrop a `.evalgate/rubric.yaml` in your repo and it beats the bundled\ndefault automatically. Tune your own criteria and weights, no fork needed.\n\nTwo presets ship in `rubric/presets/` so you don't have to start from\nscratch:\n\n```bash\ncp rubric/presets/security-strict.yaml .evalgate/rubric.yaml   # infra, auth, payments — one risky command tanks the score\ncp rubric/presets/fast-prototype.yaml .evalgate/rubric.yaml    # spikes/internal tools — speed over test coverage & narrow diffs\n```\n\n## License\n\nMIT\n","readmeFilename":"README.md","_rev":"1-7eb46564ba5330342eebd3b032e7b4b5"}