{"_id":"@agentic-evals/mcp","name":"@agentic-evals/mcp","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@agentic-evals/mcp","version":"0.1.0","description":"Agentic Evaluation Library — MCP server for deterministic and non-deterministic UI/UX evaluations with swarm orchestration","type":"module","license":"MIT","main":"dist/mcp/server.js","bin":{"agentic-evals":"dist/mcp/server.js"},"publishConfig":{"access":"public"},"scripts":{"build":"tsc","dev":"tsc --watch","start":"node dist/mcp/server.js","lint":"eslint . --ext .ts","test":"vitest"},"dependencies":{"@anthropic-ai/sdk":"^0.39.0","@modelcontextprotocol/sdk":"^1.12.1","axe-core":"^4.10.0","glob":"^11.0.0","nanoid":"^5.1.0","playwright":"^1.50.0","zod":"^3.24.0"},"devDependencies":{"@types/node":"^22.0.0","typescript":"^5.7.0","vitest":"^3.0.0"},"engines":{"node":">=20.0.0"},"_id":"@agentic-evals/mcp@0.1.0","gitHead":"563c6f1b9775d5d6fe14fb581c861ba2f12bdd61","types":"./dist/mcp/server.d.ts","_nodeVersion":"22.18.0","_npmVersion":"10.9.3","dist":{"integrity":"sha512-90TBpqlTebxE9lp7GthnTMP+9WEo+KW4cjqT6rCQooH5ZvvJsxxRroIA6T6sE1U9eoaxvHdn5CCd+SjouvJLRQ==","shasum":"db8295c60b48fd7078a62aa25b96a1af737a1c08","tarball":"https://registry.npmjs.org/@agentic-evals/mcp/-/mcp-0.1.0.tgz","fileCount":222,"unpackedSize":1020677,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQCA12AfL0no/pYvZGJoSx9HXgsqg+BUdJ+Wy9h/QoWi9QIhAIcSYJ8EPrU4fGNegO2Qro8Tks5oDVhmqRgrTWpQkQ17"}]},"_npmUser":{"name":"unthink","email":"alex@unthinkmedia.com"},"directories":{},"maintainers":[{"name":"unthink","email":"alex@unthinkmedia.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/mcp_0.1.0_1775595884795_0.6664621967594362"},"_hasShrinkwrap":false}},"time":{"created":"2026-04-07T21:04:44.716Z","0.1.0":"2026-04-07T21:04:44.995Z","modified":"2026-04-07T21:04:45.215Z"},"maintainers":[{"name":"unthink","email":"alex@unthinkmedia.com"}],"description":"Agentic Evaluation Library — MCP server for deterministic and non-deterministic UI/UX evaluations with swarm orchestration","license":"MIT","readme":"# @agentic-evals/mcp\n\nMCP server for running deterministic and non-deterministic UI/UX evaluations using multi-agent swarms. Combines automated tooling (Lighthouse, axe-core, ESLint) with LLM-judged rubric evaluations, synthesizes findings through a deliberation protocol, and outputs prioritized GitHub issues.\n\n## Quick Start\n\n```bash\nnpm install\nnpm run build\nnpm start          # starts MCP server on stdio\n```\n\n### Use with VS Code / Copilot\n\nAdd to `.vscode/mcp.json`:\n\n```json\n{\n  \"servers\": {\n    \"agentic-evals\": {\n      \"type\": \"stdio\",\n      \"command\": \"node\",\n      \"args\": [\"dist/mcp/server.js\"]\n    }\n  }\n}\n```\n\nThen ask Copilot: *\"Run a quick review of http://localhost:3000\"*\n\n## Architecture\n\n```\nUser / Copilot\n  │  MCP tool call (run_swarm, run_eval, …)\n  ▼\nMCP Server (17 tools, 2 resources, 3 prompts)\n  │\n  ▼\nSwarmOrchestrator.execute()\n  │\n  ├─ Phase 1: FAN OUT (parallel batches)\n  │   ├─ Deterministic agents (Lighthouse, axe, ESLint, Stylelint, Prettier)\n  │   └─ Non-deterministic agents (LLM + rubric + evidence)\n  │\n  ├─ Phase 1.5: RE-EVALUATION (confidence-gated)\n  │   └─ High-severity + low-confidence findings → targeted evidence\n  │       (zoom, hover, focus states, element-level capture)\n  │\n  ├─ Phase 1.75: CROSS-RUN MEMORY\n  │   └─ Annotate findings: new | persistent | regression | resolved\n  │\n  └─ Phase 2: DELIBERATION (LLM synthesis)\n      ├─ Merge overlapping findings\n      ├─ Challenge across agents\n      ├─ Prioritize (severity × effort → P0–P3)\n      └─ Actionize into concrete fixes\n  │\n  ▼\nGitHub Issue Creator → [P0] [Domain] Title with evidence & labels\n```\n\n## Wizard Flow\n\nThe wizard provides a session-based pipeline that guides the evaluation from start to finish. Each step auto-populates the next — the agent never needs to figure out what to call next.\n\n```\n📋 Plan  →  ✅ Approve  →  🔍 Run  →  🤔 Deliberate  →  🎯 Act  →  📝 Confirm  →  🎉 Done\n```\n\n| Step | What Happens | User Action |\n|------|-------------|-------------|\n| **Plan** | `wizard_start` assembles an eval team based on project signals | Review team, add/remove evals |\n| **Approve** | `wizard_advance(approve)` locks the team and launches the swarm | Confirm the roster |\n| **Run** | Swarm executes in parallel, returns phased narrative | — (automatic) |\n| **Deliberate** | Agent follows deliberation protocol to merge/prioritize findings | Review prioritized findings |\n| **Act** | `wizard_advance(submit_findings)` formats findings as issues | Choose: issues, actionize, or fix |\n| **Confirm** | `wizard_advance(choose_action)` shows dry-run preview | Confirm to create |\n| **Done** | `wizard_advance(confirm)` creates GitHub issues/PRs | — |\n\nSession state persists across tool calls — findings, issues, and evidence are never lost between steps.\n\n## MCP Tools\n\n### Wizard Flow (Recommended)\n\n| Tool | Description |\n|------|-------------|\n| `wizard_start` | Start a guided evaluation wizard — creates a session, auto-assembles team, walks through each phase |\n| `wizard_advance` | Advance the wizard to the next phase (approve → run → deliberate → act → confirm) |\n| `wizard_status` | Check wizard session status or list all active sessions |\n\n### Individual Tools\n\n| Tool | Description |\n|------|-------------|\n| `list_evals` | List available evaluation plugins |\n| `scan_project` | Discover project tooling (linters, tests, formatters) and register as plugins |\n| `run_eval` | Run a single evaluation plugin against a URL or project |\n| `run_suite` | Run multiple evaluations as a suite |\n| `run_swarm` | Execute a full multi-agent swarm evaluation |\n| `list_presets` | List swarm presets (quick-scan, full-review, deep-dive, a11y-focus) |\n| `capture_evidence` | Capture screenshots, DOM, and styles at multiple viewports |\n| `capture_user_flow` | Execute a user flow (click, fill, navigate, hover) and capture evidence at each step |\n| `reevaluate_findings` | Re-evaluate high-severity/low-confidence findings with targeted evidence (hover, focus, zoom) |\n| `get_run_history` | View evaluation run history, trend data, regressions, and resolutions |\n| `get_rubric` | Load a rubric with knowledge context |\n| `list_rubrics` | List all available rubrics |\n| `create_issues` | Convert findings into GitHub issues |\n| `review_page` | One-shot page review (evidence + eval + report) |\n| `register_plugin` | Register a custom eval plugin at runtime |\n| `scaffold_deterministic_eval` | Create a custom deterministic eval in `.evals/` (plugin + command + config) |\n| `scaffold_non_deterministic_eval` | Create a custom non-deterministic eval in `.evals/` (rubric + plugin + config) |\n\n## Swarm Presets\n\n| Preset | Agents | Rounds | Use Case |\n|--------|--------|--------|----------|\n| `quick-scan` | 3 | 1 | Fast smoke test |\n| `full-review` | 6 | 2 | Standard review |\n| `deep-dive` | 8 | 3 | Comprehensive audit |\n| `a11y-focus` | 4 | 2 | Accessibility-focused |\n\n## Knowledge & Rubrics\n\nNon-deterministic evals are powered by rubrics in [`knowledge/rubrics/`](knowledge/rubrics/). Override per-project by placing rubrics in `.evals/rubrics/` in the target project.\n\nAvailable rubrics: accessibility, content/copy, information architecture, performance perception, responsive design, UX heuristics, visual design.\n\nSupporting knowledge: [`knowledge/principles/`](knowledge/principles/) (Gestalt, Fitts' law, cognitive load, color theory, typography) and [`knowledge/standards/`](knowledge/standards/) (WCAG 2.2, design systems).\n\nSee [`knowledge/README.md`](knowledge/README.md) for details on the rubric resolution system.\n\n### Per-project `.evals/` structure\n\nRun `scan_project` or call `initProject()` to scaffold a fully customizable `.evals/` directory:\n\n```\n.evals/\n├── config.json                      # enable/disable evals, adjust settings\n├── commands.json                    # custom CLI-based evals\n├── deliberation-protocol.md         # override deliberation prompts\n├── issue-template.md                # override GitHub issue format\n├── rubrics/                         # override or extend rubric scoring\n│   ├── visual-design.md\n│   └── ...\n├── knowledge/\n│   ├── principles/                  # override or add design principles\n│   │   ├── gestalt.md\n│   │   ├── brand-guidelines.md      # (your own)\n│   │   └── ...\n│   └── standards/                   # override or add design standards\n│       ├── wcag-2.2.md\n│       ├── your-design-system.md    # (your own)\n│       └── ...\n└── plugins/                         # override or extend eval plugins\n    ├── deterministic/\n    ├── non-deterministic/\n    └── README.md\n```\n\n**Resolution order:** project `.evals/` overrides → library defaults. Project files always win on name collision, and you can add new files that don't exist in the library.\n\n## Custom Evals\n\n### Command-based (deterministic)\n\nAdd `.evals/commands.json` to your project:\n\n```json\n{\n  \"commands\": [\n    {\n      \"name\": \"storybook-build\",\n      \"command\": \"npm run build-storybook\",\n      \"kind\": \"build\",\n      \"failOn\": \"exit-code\"\n    }\n  ]\n}\n```\n\n### Plugin-based (non-deterministic)\n\nSee [`examples/example-plugin.ts`](examples/example-plugin.ts) for a template.\n\n### Guided creation (via Copilot)\n\nUse the `create-eval` prompt — Copilot will interview you, then scaffold everything:\n\n> *\"Create a non-deterministic eval for brand compliance\"*\n\nOr call the tools directly: `scaffold_deterministic_eval` / `scaffold_non_deterministic_eval`.\n\n## Development\n\n```bash\nnpm run dev        # watch mode\nnpm run lint\nnpm test           # vitest\n```\n\n## License\n\nMIT\n","readmeFilename":"README.md","_rev":"1-b9ca2e23e1f636812ac5a58defb233ce"}