{"_id":"@abdur-raheem/tardi","name":"@abdur-raheem/tardi","dist-tags":{"latest":"1.0.0"},"versions":{"1.0.0":{"name":"@abdur-raheem/tardi","version":"1.0.0","publishConfig":{"access":"public"},"main":"dist/index.js","bin":{"tardi":"dist/index.js"},"scripts":{"build":"tsc","postbuild":"node dist/schema-generator.js","test":"vitest run src/","release":"npm run build && changeset publish"},"keywords":["ai-testing","llm-evaluations","agentic-testing","llmops","ai-agents","developer-tools","cli","test-framework","llm-as-a-judge"],"author":{"name":"Tardi Contributors"},"license":"ISC","description":"A deterministic testing framework for LLM agents and non-deterministic AI pipelines, minimizing token cost via tiered assertions.","dependencies":{"@ai-sdk/anthropic":"^4.0.5","@ai-sdk/google":"^4.0.5","@ai-sdk/openai":"^4.0.5","@clack/prompts":"^1.6.0","@napi-rs/keyring":"^1.3.0","ai":"^7.0.10","ajv":"^8.20.0","boxen":"^8.0.1","chalk":"^5.6.2","ci-info":"^4.4.0","cli-table3":"^0.6.5","commander":"^15.0.0","diff":"^9.0.0","dotenv":"^17.4.2","glob":"^13.0.6","ora":"^9.4.1","p-limit":"^7.3.0","pngjs":"^7.0.0","ts-node":"^10.9.2","typescript":"^6.0.3","yaml":"^2.9.0","zod":"^4.4.3","zod-to-json-schema":"^3.25.2","zod-validation-error":"^5.0.0"},"devDependencies":{"@changesets/cli":"^2.31.0","@types/node":"^26.1.0","@types/pngjs":"^6.0.5","eslint":"^10.6.0","prettier":"^3.9.4","vitest":"^4.1.9"},"repository":{"type":"git","url":"git+https://github.com/Nyx-abu/tardi.git"},"engines":{"node":">=18.0.0"},"_id":"@abdur-raheem/tardi@1.0.0","gitHead":"9c7d23ad84dc57c5cd4d6be7977b068bb2dd5c55","bugs":{"url":"https://github.com/Nyx-abu/tardi/issues"},"homepage":"https://github.com/Nyx-abu/tardi#readme","_nodeVersion":"22.15.0","_npmVersion":"10.6.0","dist":{"integrity":"sha512-SWMH2R50ocV+Tzx5jfjiMjPZw1qSDett1Zi3LGkRpubTITWPdyPfzTcfaXhiEchavKLk6O2dFMkNJRDSoy6pow==","shasum":"c4e08fe2fc5bb89c8b2492a2714215a14e88ce84","tarball":"https://registry.npmjs.org/@abdur-raheem/tardi/-/tardi-1.0.0.tgz","fileCount":16,"unpackedSize":154546,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIEujwL73c8k2geiRsmFi7rQdMx6qGfLCYlQsaeTEE9EvAiBeMidTMWGJ1M/JXI6U1w38Q1lzLXB/0P54WhyIZUwdhA=="}]},"_npmUser":{"name":"abdur-raheem","email":"shotonishi26@gmail.com"},"directories":{},"maintainers":[{"name":"abdur-raheem","email":"shotonishi26@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/tardi_1.0.0_1783086550284_0.07427591356603758"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-03T13:49:10.150Z","1.0.0":"2026-07-03T13:49:10.424Z","modified":"2026-07-03T13:49:10.598Z"},"maintainers":[{"name":"abdur-raheem","email":"shotonishi26@gmail.com"}],"description":"A deterministic testing framework for LLM agents and non-deterministic AI pipelines, minimizing token cost via tiered assertions.","homepage":"https://github.com/Nyx-abu/tardi#readme","keywords":["ai-testing","llm-evaluations","agentic-testing","llmops","ai-agents","developer-tools","cli","test-framework","llm-as-a-judge"],"repository":{"type":"git","url":"git+https://github.com/Nyx-abu/tardi.git"},"author":{"name":"Tardi Contributors"},"bugs":{"url":"https://github.com/Nyx-abu/tardi/issues"},"license":"ISC","readme":"<p align=\"center\">\n  <img src=\"assets/tardi-logo.svg\" alt=\"Tardi Logo\" width=\"650\" />\n</p>\n\n<h1 align=\"center\">Tardi</h1>\n\n<p align=\"center\">\n  <strong>Deterministic testing for LLM agents and non-deterministic pipelines.</strong>\n</p>\n\n<p align=\"center\">\n  <a href=\"https://npmjs.com/package/tardi-cli\"><img src=\"https://img.shields.io/npm/v/tardi-cli.svg\" alt=\"NPM Version\" /></a>\n  <a href=\"LICENSE\"><img src=\"https://img.shields.io/badge/License-ISC-blue.svg\" alt=\"License\" /></a>\n  <a href=\"https://github.com/Nyx-abu/tardi/actions\"><img src=\"https://img.shields.io/github/actions/workflow/status/Nyx-abu/tardi/ci.yml?branch=master\" alt=\"Build Status\" /></a>\n</p>\n\n<p align=\"center\">\n  <a href=\"#about\">About</a> •\n  <a href=\"#features\">Features</a> •\n  <a href=\"#the-tardi-pipeline\">The Pipeline</a> •\n  <a href=\"#installation\">Installation</a> •\n  <a href=\"#quickstart\">Quickstart</a> •\n  <a href=\"#examples-output\">Examples & Output</a>\n</p>\n\n---\n\nTardi is an open-source testing framework engineered specifically for evaluating agentic workflows, autonomous LLM scripts, and non-deterministic applications. \n\n## About\n\nTesting AI agents presents a unique challenge: traditional testing frameworks cannot evaluate non-deterministic natural language outputs, and raw \"LLM-as-a-judge\" evaluation pipelines are prohibitively expensive and prone to hallucination at scale.\n\n**Tardi solves this by implementing a tiered assertion gauntlet.** Instead of blindly sending every agent execution trace to an evaluation model, Tardi enforces strict deterministic constraints first. If your agent crashes, hangs in an infinite loop, or returns malformed JSON, Tardi fails the test immediately—preventing unnecessary LLM API calls and accelerating your feedback loop.\n\n## Features\n\n- **Tiered Evaluation Engine:** Catch process crashes, timeouts, and schema mismatches deterministically before triggering expensive LLM judges.\n- **Concurrency & Rate-Limiting:** Execute test suites in parallel with intelligent chunking to maximize throughput without exceeding API rate limits.\n- **Provider Agnostic:** Built on the standard Vercel AI SDK, Tardi supports hot-swappable evaluation models from OpenAI, Google, Anthropic, and local endpoints.\n- **Interactive REPL & Natural Language CLI:** Includes a zero-friction CLI environment for rapid test synthesis, execution, and debugging, powered by an onboard NLP intent parser.\n- **Secure Credential Management:** Safely stores provider API keys in your native OS keychain during local development.\n\n## The Tardi Pipeline\n\nWhen you execute an iteration, Tardi evaluates the agent's output through a strict, cost-saving pipeline. Deterministic checks always run first.\n\n```mermaid\ngraph TD\n    Start([Run Agent]) --> ProcessCheck{Process Crashed?}\n    ProcessCheck -- Yes --> FailCrash[FAIL: CRASH]\n    ProcessCheck -- No --> TimeoutCheck{Exceeded Timeout?}\n    \n    TimeoutCheck -- Yes --> FailTimeout[FAIL: TIMEOUT]\n    TimeoutCheck -- No --> RegexCheck{Matches Regex?}\n    \n    RegexCheck -- No --> FailRegex[FAIL: REGEX_MISMATCH]\n    RegexCheck -- Yes --> SchemaCheck{Valid JSON Schema?}\n    \n    SchemaCheck -- No --> FailSchema[FAIL: SCHEMA_MISMATCH]\n    SchemaCheck -- Yes --> LLMJudge{LLM Judge Approves?}\n    \n    LLMJudge -- No --> FailJudge[FAIL: LLM_JUDGE_FAIL]\n    LLMJudge -- Yes --> Pass([PASS])\n    \n    style FailCrash fill:#fbd4d4,stroke:#f87171,color:#000\n    style FailTimeout fill:#fbd4d4,stroke:#f87171,color:#000\n    style FailRegex fill:#fbd4d4,stroke:#f87171,color:#000\n    style FailSchema fill:#fbd4d4,stroke:#f87171,color:#000\n    style FailJudge fill:#fbd4d4,stroke:#f87171,color:#000\n    style Pass fill:#dcfce7,stroke:#4ade80,color:#000\n```\n\n## Installation\n\nInstall the CLI globally via npm to use the `tardi` command anywhere:\n\n```bash\nnpm install -g tardi-cli\n```\n\n## Quickstart\n\n### 1. Auto-Synthesize from a GitHub Repository\nTardi can automatically clone an agent repository, detect its entry point, and synthesize a test gauntlet based on golden runs.\n\n```bash\ntardi github https://github.com/Nyx-abu/demo-agent.git\n```\n\n<p align=\"center\">\n  <img src=\"assets/demo-github.svg\" alt=\"Tardi Auto-Synthesis\" width=\"800\" />\n</p>\n\n### 2. Manual Initialization\nIf you prefer to configure your suite manually:\n\n```bash\ntardi init\ntardi auth login google\ntardi run tests/\n```\n\n## Examples & Output\n\nTardi utilizes simple YAML files (`*.tardi.yaml`) to define test suites. Below are detailed examples of how Tardi behaves under different failure and success conditions.\n\n### Scenario A: Deterministic JSON Schema Validation\n\nA common requirement for AI agents is to return structured JSON. Tardi catches malformed JSON or missing fields *deterministically*, bypassing the LLM Judge entirely.\n\n**Test Configuration:**\n```yaml\n# tests/json-schema.tardi.yaml\nname: Agent JSON Output Test\ncommand: node src/agent.js\niterations: 5\nconcurrency: 2\n\nassertions:\n  jsonSchema:\n    type: object\n    required: [\"status\", \"result\"]\n```\n\n**Expected Tardi Output (Schema Mismatch):**\nIf the agent forgets to include the `result` field, Tardi immediately aborts the pipeline and outputs a schema mismatch error:\n\n```text\n ╭────────────────────────  Tardi Execution Summary  ────────────────────────╮\n │   ┌─────────────┬──────────┐                                              │\n │   │ Metric      │ Value    │                                              │\n │   ├─────────────┼──────────┤                                              │\n │   │ Total Runs  │ 5        │                                              │\n │   │ Passed      │ 4        │                                              │\n │   │ Failed      │ 1        │                                              │\n │   │ Flakiness   │ 20.00%   │                                              │\n │   └─────────────┴──────────┘                                              │\n │                                                                           │\n │   Failures:                                                               │\n │   - Iteration 3 [SCHEMA_MISMATCH]: Output failed JSON Schema validation.  │\n │     Missing required property: 'result'.                                  │\n ╰───────────────────────────────────────────────────────────────────────────╯\n```\n\n### Scenario B: Semantic Evaluation with LLM-as-a-Judge\n\nIf your agent passes all deterministic constraints, Tardi forwards the raw output to an LLM evaluator. The LLM evaluator uses your custom rubric to score the output.\n\n**Test Configuration:**\n```yaml\n# tests/evaluator.tardi.yaml\nname: Semantic Agent Test\ncommand: node src/summarize.js\niterations: 3\n\nassertions:\n  regex: \"summary:\"\n\nevaluator:\n  provider: google\n  model: gemini-2.5-flash\n  prompt: |\n    Evaluate if the agent successfully summarized the input document.\n    Return only 'PASS' or 'FAIL'.\n```\n\n**Expected Tardi Output (LLM Judge Failure):**\n\n<p align=\"center\">\n  <img src=\"assets/demo-terminal.svg\" alt=\"Tardi Execution Summary\" width=\"800\" />\n</p>\n\n### Scenario C: Uncaught Exceptions & Process Crashes\n\nWhen testing complex autonomous scripts, agents often crash mid-execution. Tardi intercepts process crashes and records the exit codes, allowing you to debug stability issues over multiple iterations.\n\n**Expected Tardi Output (Crash):**\n```text\n ╭────────────────────────  Tardi Execution Summary  ────────────────────────╮\n │   Failures:                                                               │\n │   - Iteration 2 [CRASH]: Agent process exited with code 1.                │\n │     Stderr: Error: Uncaught exception in src/agent.js:42                  │\n ╰───────────────────────────────────────────────────────────────────────────╯\n```\n\n## CI/CD Usage\n\nTardi detects continuous integration environments automatically and disables interactive prompts. Inject API keys directly via environment variables for your pipeline:\n\n```bash\nGOOGLE_GENERATIVE_AI_API_KEY=\"your-key-here\" tardi run tests/\n```\n\n## Contributing\n\nPull requests are welcome. For major architectural changes, please open an issue first to discuss your proposed modifications.\n\nPlease see our [Contributing Guide](CONTRIBUTING.md) and our [Code of Conduct](CODE_OF_CONDUCT.md).\n\n## License\n\nThis project is licensed under the [ISC License](LICENSE).\n","readmeFilename":"README.md","_rev":"1-8c4590259b45a581c3aefaa99e5e6a25"}