{"_id":"@aryan_mangod/toolscore","name":"@aryan_mangod/toolscore","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@aryan_mangod/toolscore","version":"0.1.0","description":"See how real models actually use your MCP tools. Selection accuracy, parameter accuracy, confusion pairs, graded A to F.","license":"MIT","author":{"name":"Aryan-MP"},"type":"module","homepage":"https://github.com/Aryan-MP/toolscore#readme","repository":{"type":"git","url":"git+https://github.com/Aryan-MP/toolscore.git"},"bugs":{"url":"https://github.com/Aryan-MP/toolscore/issues"},"bin":{"toolscore":"dist/cli.js"},"scripts":{"build":"tsc","dev":"tsx src/cli.ts","demo":"tsx src/cli.ts scan --stdio \"tsx demo/flawed-server.ts\" --provider mock","compile":"bun build ./src/cli.ts --compile --outfile ./dist-bin/toolscore","compile:all":"bun run compile:linux-x64 && bun run compile:linux-arm64 && bun run compile:macos-x64 && bun run compile:macos-arm64 && bun run compile:windows-x64","compile:linux-x64":"bun build ./src/cli.ts --compile --target=bun-linux-x64 --outfile ./dist-bin/toolscore-linux-x64","compile:linux-arm64":"bun build ./src/cli.ts --compile --target=bun-linux-arm64 --outfile ./dist-bin/toolscore-linux-arm64","compile:macos-x64":"bun build ./src/cli.ts --compile --target=bun-darwin-x64 --outfile ./dist-bin/toolscore-macos-x64","compile:macos-arm64":"bun build ./src/cli.ts --compile --target=bun-darwin-arm64 --outfile ./dist-bin/toolscore-macos-arm64","compile:windows-x64":"bun build ./src/cli.ts --compile --target=bun-windows-x64 --outfile ./dist-bin/toolscore-windows-x64.exe"},"keywords":["mcp","model-context-protocol","ai-agents","tool-calling","evaluation","benchmark"],"dependencies":{"@anthropic-ai/sdk":"^0.112.1","@clack/prompts":"^1.7.0","@modelcontextprotocol/sdk":"^1.29.0","commander":"^15.0.0","picocolors":"^1.1.1","zod":"^4.4.3"},"devDependencies":{"@types/node":"^26.1.1","tsx":"^4.23.1","typescript":"^7.0.2"},"engines":{"node":">=22.12.0"},"_id":"@aryan_mangod/toolscore@0.1.0","gitHead":"7c9e9e01351aa101ab880eb1a707e082cf75fe85","_nodeVersion":"22.23.1","_npmVersion":"10.9.8","dist":{"integrity":"sha512-y7DFVJushWeoHILjg+dXsxaxFgBnaL9Bw3A89VZBWOeSqu4LNvETpltKR9TxcsLzKVD4QHtufyrIJdteDX977Q==","shasum":"8c1f1d033ed17fd669cb4227f80bd408acc15a41","tarball":"https://registry.npmjs.org/@aryan_mangod/toolscore/-/toolscore-0.1.0.tgz","fileCount":54,"unpackedSize":121543,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@aryan_mangod%2ftoolscore@0.1.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQDxyT8h42qfS9wm+1lJLeyGjc0dIswbH1aP3MLYU3JOGQIgcEOZbgcgGPV50++NF/8nRwxPYGujqLAEh19sIEED3ao="}]},"_npmUser":{"name":"aryan_mangod","email":"aryanmp2003@gmail.com"},"directories":{},"maintainers":[{"name":"aryan_mangod","email":"aryanmp2003@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/toolscore_0.1.0_1784711751343_0.007417274049616118"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-22T09:15:51.118Z","0.1.0":"2026-07-22T09:15:51.536Z","modified":"2026-07-22T09:15:52.150Z"},"maintainers":[{"name":"aryan_mangod","email":"aryanmp2003@gmail.com"}],"description":"See how real models actually use your MCP tools. Selection accuracy, parameter accuracy, confusion pairs, graded A to F.","homepage":"https://github.com/Aryan-MP/toolscore#readme","keywords":["mcp","model-context-protocol","ai-agents","tool-calling","evaluation","benchmark"],"repository":{"type":"git","url":"git+https://github.com/Aryan-MP/toolscore.git"},"author":{"name":"Aryan-MP"},"bugs":{"url":"https://github.com/Aryan-MP/toolscore/issues"},"license":"MIT","readme":"<p align=\"center\">\n  <picture>\n    <source media=\"(prefers-color-scheme: dark)\" srcset=\"assets/logo-dark.svg\">\n    <source media=\"(prefers-color-scheme: light)\" srcset=\"assets/logo-light.svg\">\n    <img alt=\"toolscore\" src=\"assets/logo-light.svg\" width=\"360\">\n  </picture>\n</p>\n\n<p align=\"center\">\n  <a href=\"https://www.npmjs.com/package/@aryan_mangod/toolscore\"><img alt=\"npm version\" src=\"https://img.shields.io/npm/v/%40aryan_mangod%2Ftoolscore.svg\"></a>\n  <a href=\"https://github.com/Aryan-MP/toolscore/releases\"><img alt=\"GitHub release\" src=\"https://img.shields.io/github/v/release/Aryan-MP/toolscore?sort=semver\"></a>\n  <a href=\"https://github.com/Aryan-MP/toolscore/blob/main/LICENSE\"><img alt=\"License: MIT\" src=\"https://img.shields.io/badge/license-MIT-informational.svg\"></a>\n</p>\n\n<p align=\"center\"><b>See how real models actually use your MCP tools.</b> Selection accuracy, parameter accuracy, confusion pairs, graded A to F.</p>\n\nStatic linters check that your tool descriptions *exist*. toolscore checks that they *work*: it generates realistic tasks for your toolset (with ground-truth labels built in by construction), runs them against a real model, and tells you which tools the model confuses, which parameters it gets wrong, and what to fix.\n\n## Why\n\nAgent toolsets fail silently. Research on production agents shows tool-selection accuracy collapses as toolsets grow (down to roughly 13% on large tool menus), and a wrong parameter cascades into a wrong final answer most of the time. Audits of the MCP ecosystem grade the overwhelming majority of published tools **F**, mostly for descriptions a model cannot disambiguate.\n\nThe teams who fix this today do it by hand: building custom eval harnesses, iterating on tool descriptions until selection accuracy climbs. One documented case took tool selection from 60% to 100% purely by rewriting descriptions. toolscore turns that loop into one command.\n\n## Install\n\n```bash\n# No Node.js required. Standalone binary, checksum-verified at install time.\ncurl -fsSL https://raw.githubusercontent.com/Aryan-MP/toolscore/main/install.sh | bash\n\n# Or, if you are already in a Node/npm project\nnpm install -g @aryan_mangod/toolscore\n```\n\n## Quick start\n\n```bash\n# Optional: interactive setup, saves defaults to .toolscorerc.json\nnpx @aryan_mangod/toolscore init\n\n# Scan your MCP server (BYOK, uses your ANTHROPIC_API_KEY)\nnpx @aryan_mangod/toolscore scan --stdio \"node ./my-mcp-server.js\"\n\n# No API key? Use your Claude subscription via the Claude Code CLI\nnpx @aryan_mangod/toolscore scan --stdio \"node ./my-mcp-server.js\" --provider claude-code\n\n# Or try the pipeline against the bundled flawed demo server, fully offline\nnpm run demo\n\n# Then let a model rewrite the descriptions it got wrong, and prove the delta\nnpx @aryan_mangod/toolscore fix --stdio \"node ./my-mcp-server.js\" --provider claude-code\n```\n\n`scan` and `fix` show live progress (connecting, generating tasks, running, scoring) and print nothing but clean JSON on stdout when run with `--json`, so they are safe to pipe in CI.\n\nReal output from a real model (Claude, via `--provider claude-code`) scanning the bundled demo server:\n\n```\n  toolscore · tsx demo/flawed-server.ts\n  7 tools · 21 tasks · claude-code/default-prompted\n\n    Grade: B+  (89/100)\n\n  ├─ Tool selection       86%\n  └─ Parameter accuracy   94%  (when the right tool was chosen)\n\n  Lexical baseline: 0%  (bag-of-words matcher on the same tasks)\n\n  Confusion pairs  (expected → actually called)\n    find_content → search_docs  ×1\n    add_document → create_note  ×1\n    get_user → fetch_account    ×1\n```\n\nThe demo server has three deliberately planted flaws: three near-identical search tools, a vague note/document split, and two user-lookup tools with interchangeable descriptions. The model stumbled on **exactly those three**.\n\n## How it works\n\n1. **Discover.** Connects to your MCP server over stdio and lists its tools.\n2. **Generate.** Creates realistic user tasks from each tool, so the correct answer is known by construction: pick tool T, write a request that objectively requires T with concrete parameter values embedded in the text. Extra tasks deliberately target your most confusable tool pairs (ranked by description similarity), hunting where failures actually live instead of sampling randomly.\n3. **Run.** Presents the **full** toolset and each task to the model, records which tool it picks and with what arguments.\n4. **Score.** Selection accuracy (60% weight) plus parameter accuracy among correct selections (40% weight) produce a 0 to 100 score and an A to F grade, plus a confusion-pair diagnosis pointing at the exact tools to fix.\n\n### The lexical baseline: keeping the benchmark honest\n\nA tool benchmark can silently flatter you: if generated tasks reuse your tools' vocabulary (\"search the docs for…\" leading to `search_docs`), any model passes by keyword matching and the grade means nothing. toolscore defends against this two ways:\n\n- Task generation is **anti-leakage by instruction**: tasks describe the user's goal in the user's own words and must never echo a tool's name or distinctive vocabulary.\n- Every scan also runs a **bag-of-words matcher** over the same tasks. If that dumb baseline scores high, the report warns you that the grade is inflated. (On the demo scan above: model 86%, baseline 0%. The tasks genuinely require understanding intent.)\n\n### `toolscore fix`\n\n`fix` runs a scan, has the model rewrite the descriptions of every tool involved in a failure, then **re-runs the identical task set** against the rewritten toolset and shows the before and after delta, including whether the resolved confusion pairs actually resolved, and whether the new descriptions cheated by echoing task vocabulary (the lexical baseline is compared too). You get copy-pasteable descriptions with a verified score improvement, or an honest report that they did not help.\n\n### Reproducibility\n\nGenerated task sets are cached in `.toolscore/`, keyed by a hash of your toolset's schemas. Same server, same task set, same model produces comparable grades run to run, which is what makes regression gating in CI meaningful. Pass `--regenerate` after changing your tools.\n\n## CLI reference\n\n```\ntoolscore init                        interactive setup, saves defaults to .toolscorerc.json\ntoolscore scan --stdio \"<command>\"    benchmark the toolset, print the graded report\ntoolscore fix  --stdio \"<command>\"    propose description rewrites and verify them\n\n  --stdio <command>      command that starts your MCP server (stdio transport)   [required unless set by `init`]\n  --provider <name>      anthropic | claude-code | mock       (default: anthropic)\n  --model <id>           model id for the provider            (default: provider's default)\n  --tasks-per-tool <n>   tasks generated per tool             (default: 4)\n  --regenerate           rebuild the cached task set\n  --json                 machine-readable output for scripting/CI\n```\n\nProviders:\n\n- **`anthropic`**: raw Claude API via `ANTHROPIC_API_KEY`, native tool-use. The reference provider.\n- **`claude-code`**: drives the [Claude Code](https://claude.com/claude-code) CLI, billed to your Claude subscription. **No API key needed.** Tool selection is prompted (toolset shown as JSON) rather than the native tool-use API, so grades are comparable across `claude-code` runs but not directly against raw-API runs; the report labels the model `*-prompted` accordingly.\n- **`mock`**: keyless bag-of-words matcher. Powers the offline demo and the lexical baseline.\n\nExit codes: `scan` returns `0` for grade D or better, `1` for grade F, `2` on error. `fix` returns `0` when the score improved or held, `1` when rewrites made it worse, `2` on error. Usable in CI as-is.\n\n## Project structure\n\n```\nsrc/\n  cli.ts             command-line entry point (scan, fix, init)\n  ui.ts              terminal UI: spinners, intro/outro framing, brand mark\n  config.ts          .toolscorerc.json read/write\n  mcp/client.ts      stdio MCP connection and tool discovery\n  providers/         model backends behind one interface (anthropic, claude-code, mock)\n  taskgen/           task generation, confusable-pair targeting, cached task sets\n  runner.ts          executes tasks against the provider, checks calls\n  scoring.ts         metrics, weights, A-F grading, confusion pairs\n  fix.ts             description-rewrite proposals and verification toolset patching\n  report.ts          terminal reports (scan and fix before/after)\ndemo/\n  flawed-server.ts   deliberately bad MCP server: overlapping tools, vague\n                     descriptions, so you can see a failing grade in 60 seconds\nassets/\n  logo-light.svg, logo-dark.svg    full lockup for light/dark backgrounds\n  mark-light.svg, mark-dark.svg    icon only, for favicons and small placements\n```\n\n## Design principles\n\n- **Zero-config.** Point it at a server, get a grade. No test files, no YAML, no scorer code.\n- **Prescriptive.** A score without a diagnosis is trivia. Confusion pairs and parameter mismatches name the exact thing to fix.\n- **BYOK.** Your keys, your data, your machine. Nothing leaves your environment.\n- **Provider-pluggable.** A model backend is one small interface: raw APIs, CLI harnesses, or local models all fit.\n- **Self-skeptical.** A benchmark that cannot detect its own inflation is marketing. The lexical baseline runs on every scan and calls out leaky task sets in the report.\n\n## Roadmap\n\n- **Providers:** `ollama` (local, keyless), OpenAI, Gemini\n- **`toolscore ci`**: baseline comparison and regression gating on every commit\n- **Distractor scaling**: accuracy vs. toolset-size curves, to answer which tools you should even expose\n- **Embedding-based pair targeting**: replace lexical similarity for sharper adversarial tasks\n- **HTML report**: shareable scan reports\n\n## Development\n\n```bash\nnpm install\nnpm run build     # compile TypeScript to dist/\nnpm run demo      # end-to-end scan of the flawed demo server, no API key needed\nnpx tsx src/cli.ts scan --stdio \"...\" --provider anthropic\n```\n\nRequires Node 22.12 or later.\n\n### Building standalone binaries\n\nReleases are compiled with [Bun](https://bun.sh) to self-contained native executables (no Node.js runtime needed to run them) for Linux, macOS, and Windows, x64 and arm64:\n\n```bash\nbun run compile          # current platform, into dist-bin/\nbun run compile:all      # cross-compile every target from one machine\n```\n\nCI builds every target on tag push, publishes `checksums.txt` alongside the binaries in the GitHub release, and `install.sh` verifies the checksum before installing. See `.github/workflows/release.yml` for details.\n\n## Contributing\n\nIssues and pull requests are welcome. See [CONTRIBUTING.md](CONTRIBUTING.md) for how the project is organized and how to run it locally.\n\n## License\n\nMIT. See [LICENSE](LICENSE).\n","readmeFilename":"README.md","_rev":"1-e5302eff89fced7fc5dfb551659b9fef"}