{"_id":"@bonnie-mcconnell/liminal","name":"@bonnie-mcconnell/liminal","dist-tags":{"latest":"0.4.2"},"versions":{"0.4.2":{"name":"@bonnie-mcconnell/liminal","version":"0.4.2","description":"Tool-use orchestration for the Anthropic API - typed errors, SHA-256 content-hash caching, DAG scheduling, typed event stream, and structured observability.","keywords":["llm","agent","anthropic","claude","tool-use","orchestration","typescript","dag","scheduler","events","observability"],"type":"module","license":"MIT","author":{"name":"Bonnie McConnell"},"repository":{"type":"git","url":"git+https://github.com/bonnie-mcconnell/liminal.git"},"main":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"import":"./dist/index.js","types":"./dist/index.d.ts"}},"scripts":{"build":"tsc","dev":"tsc --watch","test":"vitest run","test:watch":"vitest","test:coverage":"vitest run --coverage","typecheck":"tsc --project tsconfig.test.json --noEmit","lint":"eslint src tests examples benchmarks","format":"prettier --write src tests examples benchmarks","format:check":"prettier --check src tests examples benchmarks","demo":"tsx examples/research-agent.ts","demo:dry":"tsx examples/research-agent.ts --dry-run","bench":"tsx benchmarks/parallel-vs-sequential.ts","bench:custom":"tsx benchmarks/parallel-vs-sequential.ts --calls 4 --delay 300"},"dependencies":{"@anthropic-ai/sdk":"^0.90.0","zod":"^3.23.0","zod-to-json-schema":"^3.23.0"},"devDependencies":{"@types/node":"^20.14.0","@typescript-eslint/eslint-plugin":"^8.0.0","@typescript-eslint/parser":"^8.0.0","@vitest/coverage-v8":"^1.6.0","eslint":"^8.57.0","prettier":"^3.3.0","tsx":"^4.15.0","typescript":"^5.6.0","vitest":"^1.6.0"},"engines":{"node":">=20.0.0"},"publishConfig":{"access":"public"},"gitHead":"5ebdbb2b6505ad1438f44349de77daa2de5aa907","_id":"@bonnie-mcconnell/liminal@0.4.2","bugs":{"url":"https://github.com/bonnie-mcconnell/liminal/issues"},"homepage":"https://github.com/bonnie-mcconnell/liminal#readme","_nodeVersion":"24.13.0","_npmVersion":"11.6.2","dist":{"integrity":"sha512-ZqsvXN03jk7agmOpVOjIykTFAqT3WmU5pRqlUHDDDNqwmKaQHWPAbP6bN/J6CsNnLpO012d6fcu7FhKodoomuw==","shasum":"a45be88df6444ba4d2841c298dcf1989bd2377d1","tarball":"https://registry.npmjs.org/@bonnie-mcconnell/liminal/-/liminal-0.4.2.tgz","fileCount":103,"unpackedSize":239982,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQDfy8vKZS1jSVB0Mj4I4gj2rNVg48bE8SROmhEF3wtO/AIgGhaAOCIpnchfEbXJG7kzAypi+CC7K+R8VbFm1RU+CgI="}]},"_npmUser":{"name":"bonnie-mcconnell","email":"bonniep.mcconnell@gmail.com"},"directories":{},"maintainers":[{"name":"bonnie-mcconnell","email":"bonniep.mcconnell@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/liminal_0.4.2_1776678396021_0.8885758241153103"},"_hasShrinkwrap":false}},"time":{"created":"2026-04-20T09:46:35.855Z","0.4.2":"2026-04-20T09:46:36.185Z","modified":"2026-04-20T09:46:36.723Z"},"maintainers":[{"name":"bonnie-mcconnell","email":"bonniep.mcconnell@gmail.com"}],"description":"Tool-use orchestration for the Anthropic API - typed errors, SHA-256 content-hash caching, DAG scheduling, typed event stream, and structured observability.","homepage":"https://github.com/bonnie-mcconnell/liminal#readme","keywords":["llm","agent","anthropic","claude","tool-use","orchestration","typescript","dag","scheduler","events","observability"],"repository":{"type":"git","url":"git+https://github.com/bonnie-mcconnell/liminal.git"},"author":{"name":"Bonnie McConnell"},"bugs":{"url":"https://github.com/bonnie-mcconnell/liminal/issues"},"license":"MIT","readme":"# Liminal\n\n![CI](https://github.com/bonnie-mcconnell/liminal/actions/workflows/ci.yml/badge.svg)\n\nA TypeScript library that manages the tool-use loop in LLM agents: runs independent calls in parallel, sequences dependent ones via a DAG, retries failures with backoff, caches results by content hash, and keeps tool errors from crashing the run. 323 tests. Three runtime dependencies.\n\n```\nRun run_4a9f2b1c8d3e  ·  2.01s  ·  3,204 tokens\n├─ Step 1  201ms  ·  280 tokens  ·  3 calls (1 level)\n│  ├─ web_search(\"typescript strict mode benefits\")  →  success  198ms  [parallel]\n│  ├─ web_search(\"typescript adoption statistics\")   →  success  201ms  [parallel]\n│  └─ calculator(\"500 * 0.62 * 0.40\")               →  cache hit  0ms   [parallel]\n├─ Step 2  1,203ms  ·  180 tokens  ·  2 calls (1 level)\n│  ├─ file_reader(\"examples/context.md\")             →  success  4ms    [parallel]\n│  └─ fetch(\"GET https://api.github.com/...\\\")        →  success  891ms  [parallel]\n└─ Final answer  (487 tokens out)\n```\n\nBoth web searches and the calculator ran simultaneously in step 1. The file read and HTTP fetch ran simultaneously in step 2.\n\n## Why I built this\n\nI was building an agent that made two web searches per turn - independent queries, no reason one had to wait for the other. Every implementation I found ran them sequentially, because the model issues tool calls as a list and the obvious thing is to execute them one by one. On a turn with three independent 200ms calls, that's 600ms. The three calls should take 200ms.\n\nFixing it properly meant the agent loop needed to know which calls were independent and which had real data dependencies. Once I started thinking about that, the loop was also obviously doing too many other things: calling the model, managing execution order, handling retries, writing to cache. Those are separate problems and they don't belong tangled together.\n\nSo I pulled them apart. The scheduler (Kahn's algorithm) groups calls into execution levels. The executor handles timeout, retry, and cache per-call. The loop just calls the model, hands tool calls to the scheduler, runs each level via `Promise.allSettled`, and feeds results back.\n\nThe other thing that bothered me: most agent implementations let tool errors throw. One failed tool call crashes the entire run. The right behaviour is to return a typed error result to the model and let it decide - try different parameters, use a different tool, or answer from what it already has. The executor never throws.\n\n## Measured performance\n\n```\n$ npm run bench\n\n────────────────────────────────────────────────────────────\n  liminal - parallel vs sequential benchmark\n────────────────────────────────────────────────────────────\n  3 tool calls × 200ms delay each\n  5 runs per configuration\n\n  Strategy      Mean        Stddev      Samples\n  ────────────────────────────────────────────────────────\n  Sequential    624ms       9ms         617ms  626ms  620ms  641ms  617ms\n  Parallel      211ms       4ms         215ms  206ms  214ms  215ms  207ms\n\n  Speedup:              2.95×\n  Theoretical maximum:  3.00×\n  Efficiency:           98.4% of theoretical\n```\n\nScheduling overhead is under 11ms. The ~9ms stddev in sequential times is OS scheduler noise on Windows, not the library.\n\n## How it works\n\n```\nAgent\n  └─ for each iteration:\n       ├─ check abort signal\n       ├─ check budget (tokens, steps)\n       ├─ call Anthropic API\n       ├─ if no tool calls → done\n       ├─ resolveScheduledCalls: apply toolDependencies graph → ScheduledCall[]\n       ├─ Scheduler (Kahn's): topological sort → execution levels\n       ├─ for each level: Promise.allSettled(executor.execute(call, signal))\n       └─ feed results back to model\n\nToolExecutor (per call)\n  ├─ check AbortSignal\n  ├─ check ResultCache (SHA-256 content-addressable key)\n  ├─ validate input (Zod)\n  ├─ Promise.race(execute(), timeout, abortSignal)\n  ├─ retry with exponential backoff + jitter\n  ├─ validate output (Zod)\n  └─ write to ResultCache\n```\n\nThe scheduler groups calls into levels using Kahn's algorithm - everything in a level runs via `Promise.allSettled` (not `Promise.all`, so one failure doesn't cancel siblings), and cycles throw `CyclicDependencyError` at scheduling time rather than producing mysterious ordering at runtime.\n\nThe result cache is SHA-256 content-addressable. `{a:1,b:2}` and `{b:2,a:1}` hit the same entry because object keys are canonicalised before hashing. Each tool gets its own LRU store so a high-traffic tool can't evict results belonging to others.\n\nThe executor never throws. Every failure path - not found, bad input, timeout, execution error, retry exhaustion - returns a typed `ToolResult` with a machine-readable error code the model can reason about.\n\n## Design decisions\n\nThe two things I spent the most time on were the scheduler and the cache key.\n\nFor the scheduler, I chose Kahn's algorithm over DFS topological sort because cycle detection is implicit - any node with nonzero in-degree after the sweep is part of a cycle. DFS gives you the same ordering but requires a separate visited-set check bolted on afterwards.\n\nFor the cache key, the non-obvious part is canonicalisation. `JSON.stringify({a:1,b:2})` and `JSON.stringify({b:2,a:1})` produce different strings, so the same logical tool input misses the cache depending on property insertion order. Sorting object keys recursively before hashing fixes this. I use SHA-256 rather than a faster hash (djb2, FNV) because non-cryptographic hashes can cluster on structured JSON - similar inputs produce similar digests, which raises the practical collision rate above the birthday-bound theoretical rate. The 64-bit prefix gives P(collision) ≈ 2.7×10⁻⁸ for 10⁶ inputs, which is acceptable for a cache where a false positive serves stale data rather than causing corruption.\n\nTwo interface decisions: `CachePolicy` is a discriminated union rather than a flat object, so accessing `ttlMs` on a `\"no-cache\"` policy is a compile error instead of a silent runtime bug. Cache capacity is configured at construction rather than per write - early versions took `maxEntries` on every `set()` call, which is a leaky interface that any Redis backend would have to accept a parameter it can't use.\n\nThe dependency graph is validated at construction time because a misspelled tool name previously produced no error and no sequencing - it silently dropped the dependency and the ordering bug only showed up at runtime. Retry jitter is there because without it, clients that all fail at the same moment retry at the same moment, hitting a recovering service with the same burst that just took it down.\n\n## What I'd do differently\n\nThe `toolDependencies` graph is declared statically on the agent, not per-call. If you want `summarise_results` to depend on `web_search`, you declare it globally - it applies every turn, even turns where `web_search` isn't called (which the resolver handles by ignoring absent dependencies). The right interface is probably per-invocation dependency hints from the model, but the Anthropic API doesn't expose that in a structured way yet. This isn't just a waiting-for-the-API situation - it reflects a real design constraint: static declaration is explicit and testable, but it prevents context-dependent sequencing that a smarter graph would support.\n\nThe cache key uses the first 16 hex chars (64 bits) of the SHA-256 digest. 64 bits gives P(collision) ≈ 2.7×10⁻⁸ at 10⁶ distinct inputs - fine for tool-call caching where a false positive serves stale data rather than causing corruption. But the truncation is a choice with a real tradeoff: the full 64-char digest would eliminate collision risk entirely at the cost of a larger key footprint per entry. For a distributed Redis cache processing millions of calls per day, the full digest is the right call. I'd make this configurable at `ResultCache` construction time.\n\nFor tools that implement cooperative cancellation (accepting `signal?: AbortSignal` in their `execute` function), in-flight work stops immediately when a timeout or abort fires - no background resource consumption, no duplicate side effects on retry. The built-in tools all do this: `fetchTool` and `webSearchTool` forward the signal to `fetch()`, and `fileReaderTool` checks it at each I/O boundary. Custom tools that ignore the signal still work correctly via the external `Promise.race`, but their timed-out execution continues in the background until it settles.\n\n## Installation\n\n```bash\nnpm install @bonnie-mcconnell/liminal\n```\n\nNode 20+ required.\n\n**To run the demo:**\n\n```bash\nnpm run demo:dry    # inspect the task and tools without an API key\nnpm run demo        # live run - requires ANTHROPIC_API_KEY\nnpm run bench       # measure parallel vs sequential performance\n```\n\n## Quick start\n\n```typescript\nimport { Agent, ToolRegistry, calculatorTool, webSearchTool, renderTrace } from \"@bonnie-mcconnell/liminal\";\n\nconst registry = new ToolRegistry().register(calculatorTool).register(webSearchTool);\n\nconst agent = new Agent(registry, {\n  model: \"claude-haiku-4-5-20251001\", // use opus-4-6 for harder tasks\n  budget: { maxTotalTokens: 10_000 },\n});\n\nconst result = await agent.run(\n  \"Search for TypeScript adoption trends, then calculate: \" +\n    \"if 40% of 500 engineers use TypeScript, how many is that?\",\n);\n\nif (result.status === \"success\") {\n  console.log(result.output);\n  console.log(renderTrace(result.trace));\n}\n```\n\n## Built-in tools\n\n| Tool | What it does | Caching |\n|---|---|---|\n| `calculatorTool` | Evaluates math expressions via a recursive-descent parser - no `eval()` | Content-hash, 24h TTL |\n| `webSearchTool` | Web search via the Brave API (labeled mock results when no API key is set) | Content-hash, 10min TTL |\n| `fileReaderTool` | Reads files relative to cwd - rejects absolute paths and directory traversal | Content-hash, 30s TTL |\n| `fetchTool` | HTTP requests (GET/POST/PUT/PATCH/DELETE/HEAD) with body truncation | No-cache (side effects) |\n\n## Tool dependencies\n\nBy default, all tool calls in a single model turn run concurrently. When one tool genuinely needs the output of another, declare it in `toolDependencies`. All names must be registered in the registry - the constructor throws immediately if any are unknown.\n\n```typescript\nconst agent = new Agent(registry, {\n  model: \"claude-haiku-4-5-20251001\",\n  toolDependencies: {\n    // summarise_results always runs after web_search completes\n    summarise_results: [\"web_search\"],\n    // analyse_data runs after both\n    analyse_data: [\"web_search\", \"summarise_results\"],\n  },\n});\n```\n\nDependencies on tools not called in a given turn are silently ignored. Cycles throw `CyclicDependencyError` immediately.\n\n## Cancellation\n\n```typescript\nconst runPromise = agent.run(longTask);\n\nsetTimeout(() => agent.abort(), 10_000);\n\nconst result = await runPromise;\nif (result.status === \"error\") {\n  console.log(result.error.code);   // \"PLANNER_ERROR\"\n  console.log(result.trace.steps.length); // steps completed before cancel\n}\n```\n\n`run()` always resolves - it never rejects. Calling `abort()` before `run()` is valid; calling it when no run is active is a no-op.\n\n## Live event stream\n\n```typescript\nconst agent = new Agent(registry, config, {\n  onEvent(event) {\n    switch (event.type) {\n      case \"dispatched\":\n        console.log(`→ ${event.toolName} (attempt ${event.attempt})`);\n        break;\n      case \"retrying\":\n        console.warn(`  ↻ retrying in ${event.delayMs}ms`);\n        break;\n      case \"succeeded\":\n        metrics.histogram(\"tool.duration_ms\", event.durationMs, { tool: event.toolName });\n        break;\n      case \"failed\":\n        logger.error(\"tool failed\", { tool: event.toolName, code: event.error.code });\n        break;\n    }\n  },\n});\n```\n\nComplete lifecycle per call:\n\n```\ndispatched → succeeded                                            # first attempt success\ncache_hit                                                         # no dispatch\ndispatched → attempt_failed → retrying → dispatched → succeeded  # retry success\ndispatched → attempt_failed → ... → failed                       # exhausted\nfailed                                                            # pre-dispatch (not found, invalid input)\n```\n\n## Custom tools\n\n```typescript\nimport { z } from \"zod\";\nimport { ToolTimeoutError } from \"@bonnie-mcconnell/liminal\";\nimport type { ToolDefinition } from \"@bonnie-mcconnell/liminal\";\n\nconst weatherTool: ToolDefinition = {\n  name: \"get_weather\",\n  description:\n    \"Returns current weather for a city. \" +\n    \"Use when the task requires weather or temperature data.\",\n  inputSchema: z.object({\n    city: z.string().describe(\"City name, e.g. 'Auckland' or 'London, UK'\"),\n    units: z.enum([\"celsius\", \"fahrenheit\"]).default(\"celsius\"),\n  }),\n  outputSchema: z.object({\n    temperature: z.number(),\n    conditions: z.string(),\n  }),\n  execute: async ({ city, units }, signal) => fetchWeather(city, units, signal),\n  summarize: ({ city, units }) => `${city} (${units})`,\n  policy: {\n    timeoutMs: 10_000,\n    retry: {\n      maxAttempts: 3,\n      backoff: \"exponential\",\n      baseDelayMs: 500,\n      maxDelayMs: 10_000,\n      jitterMs: 200,\n      shouldRetry: (err) => err instanceof ToolTimeoutError,\n    },\n    cache: { strategy: \"content-hash\", ttlMs: 5 * 60_000, vary: [], maxEntries: 256 },\n  },\n};\n```\n\n## Sharing a cache across runs\n\n```typescript\nimport { Agent, ToolRegistry, ResultCache, type Cache, calculatorTool } from \"@bonnie-mcconnell/liminal\";\n\nconst cache: Cache = new ResultCache();\nconst registry = new ToolRegistry().register(calculatorTool);\n\nconst agent1 = new Agent(registry, { model: \"claude-haiku-4-5-20251001\" }, { cache });\nconst agent2 = new Agent(registry, { model: \"claude-haiku-4-5-20251001\" }, { cache });\n\nconst stats = cache.stats(\"calculator\");\nconsole.log(`Hit rate: ${((stats?.hitRate ?? 0) * 100).toFixed(1)}%`);\n```\n\n## Error handling\n\n```typescript\nimport { BudgetExceededError, MaxIterationsError } from \"@bonnie-mcconnell/liminal\";\n\nconst result = await agent.run(task);\n\nif (result.status === \"error\") {\n  if (result.error instanceof BudgetExceededError) {\n    // result.error.budgetType → \"tokens\" | \"steps\"\n    // result.error.limit, result.error.used\n  } else if (result.error instanceof MaxIterationsError) {\n    // Model is looping - check tool descriptions and prompt design\n  }\n  // result.trace is always present, even on failure.\n  console.log(`${result.trace.steps.length} steps completed`);\n}\n```\n\n## Observability\n\nEvery significant event is written as newline-delimited JSON:\n\n```\n{\"ts\":\"...\",\"level\":\"info\",\"runId\":\"run_4a9f2b\",\"event\":\"agent.started\",\"data\":{\"model\":\"claude-haiku-4-5-20251001\"}}\n{\"ts\":\"...\",\"level\":\"debug\",\"runId\":\"run_4a9f2b\",\"event\":\"tool.dispatched\",\"data\":{\"callId\":\"c1\",\"toolName\":\"web_search\",\"attempt\":1}}\n{\"ts\":\"...\",\"level\":\"warn\",\"runId\":\"run_4a9f2b\",\"event\":\"tool.retrying\",\"data\":{\"callId\":\"c1\",\"attempt\":2,\"delayMs\":623}}\n{\"ts\":\"...\",\"level\":\"info\",\"runId\":\"run_4a9f2b\",\"event\":\"tool.succeeded\",\"data\":{\"durationMs\":780,\"cacheHit\":false}}\n{\"ts\":\"...\",\"level\":\"info\",\"runId\":\"run_4a9f2b\",\"event\":\"agent.completed\",\"data\":{\"totalTokens\":2841,\"steps\":3}}\n```\n\n`LOG_LEVEL=debug` shows the execution plan, cache checks, and every dispatch. `LOG_LEVEL=warn` shows only retries and failures. NDJSON is ingested without configuration by Datadog, CloudWatch, and Splunk.\n\n## Tests\n\n323 tests across 15 files.\n\n```bash\nnpm test               # unit + integration\nnpm run test:coverage  # with lcov report\n```\n\nThe integration tests replace the Anthropic SDK with a deterministic mock - no credentials needed, fully reproducible. Covered failure modes: timeouts, input validation errors, retry exhaustion, dependency cycles, budget limits, partial parallel failures, `toolDependencies` sequencing, construction-time validation, and `abort()` pre-run and mid-run cancellation.\n\nUnit coverage: 99% statements, 93% branches.\n\n## Extending it\n\n**Distributed cache.** The `Cache` interface is three methods: `configure`, `get`, `set`. Implement it against Redis and inject at construction. The executor and agent are unchanged - they don't know or care what's behind the interface. The SHA-256 key scheme works across processes because it's deterministic: the same logical input always produces the same 16-char hex key regardless of where it was generated.\n\n**Model-agnostic.** The Anthropic SDK lives in one file (`agent.ts`). The only thing that would change to support OpenAI or Gemini is the API call and response parsing inside `Agent.run()`. Everything downstream - the scheduler, executor, cache, event stream - operates on `ToolCall[]` and `ToolResult[]`, which are your types, not the SDK's.\n\n**Trace persistence.** `ExecutionTrace` is a plain object with no circular references. Store it by `runId` and you get run replay, prompt A/B testing against historical inputs, and per-task cost attribution.\n\n**Streaming tool results.** Currently `execute()` awaits the complete result. Making it return `AsyncIterable<ToolEvent>` - where `succeeded` is the terminal event - would let long-running tools stream partial progress. The scheduler and cache are unaffected; the blast radius is `ToolExecutor` and the agent loop's result-collection logic.\n\n## Demo\n\n```bash\nexport ANTHROPIC_API_KEY=sk-...\nnpm run demo        # runs a multi-step research task\nnpm run demo:dry    # prints the task and tools without calling the API\n\n# Real web search (mock data used otherwise):\nexport BRAVE_SEARCH_API_KEY=BSA...\nnpm run demo\n```\n\n## Structure\n\n```\nsrc/\n├── core/          agent, executor, registry, scheduler, cache, lru, defaults\n├── errors/        typed error hierarchy (LiminalError subclasses)\n├── tools/         calculator, web_search, file_reader, fetch\n├── observability/ logger (NDJSON), trace renderer, EventEmitter\n├── types/         ToolDefinition, AgentResult, ExecutionTrace, ToolEvent, policies\n└── index.ts       public API - everything not exported here is an internal detail\n\nbenchmarks/\n└── parallel-vs-sequential.ts   measures scheduler speedup with real wall-clock numbers\n\ntests/\n├── unit/          one file per source module (14 suites)\n└── integration/   full agent loop with deterministic mock Anthropic client\n```","readmeFilename":"README.md","_rev":"1-54f6cf82ca1045dbbb2d60a949b89d73"}