{"_id":"@dory-agentic/dory-agentic-sdk","name":"@dory-agentic/dory-agentic-sdk","dist-tags":{"latest":"0.2.0"},"versions":{"0.2.0":{"name":"@dory-agentic/dory-agentic-sdk","version":"0.2.0","description":"Dory Agentic SDK: LFS context management with AI SDK tool loops for long-running agents","type":"module","license":"MIT","repository":{"type":"git","url":"git+https://github.com/ahmedmigo/dory-agentic-sdk.git"},"bugs":{"url":"https://github.com/ahmedmigo/dory-agentic-sdk/issues"},"homepage":"https://github.com/ahmedmigo/dory-agentic-sdk#readme","keywords":["ai-sdk","agent","context-window","lfs","long-running","tool-loop","memory","llm"],"engines":{"node":">=18"},"publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"main":"dist/index.js","types":"dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"build":"tsc -p tsconfig.build.json","test":"npm run test:unit","test:unit":"vitest run --exclude \"tests/lfs-ollama.integration.test.ts\"","test:ollama":"vitest run tests/lfs-ollama.integration.test.ts","test:all":"npm run test:unit && npm run test:ollama","typecheck":"tsc --noEmit","prepublishOnly":"npm run typecheck && npm run test:unit && npm run build"},"dependencies":{"@ai-sdk/mcp":"^1.0.45","@ai-sdk/openai":"^3.0.67","@ai-sdk/openai-compatible":"^2.0.48","ai":"^6.0.193","zod":"^4.4.3"},"devDependencies":{"@types/node":"^25.9.1","typescript":"^6.0.3","vitest":"^4.1.7"},"_id":"@dory-agentic/dory-agentic-sdk@0.2.0","gitHead":"9e51ff06a3e5e9681c310aa98b6dde4b9841e162","_nodeVersion":"22.11.0","_npmVersion":"11.3.0","dist":{"integrity":"sha512-47YDSWPKLt5wurrgjODSrOpO5XWG1ggTTEazNdvMAu7CTKoNaJEVpZ0KC6TsGmuL7okhtCjVlPIyoZWCX+vgPw==","shasum":"a66655e5e98a220de27a186c5daca8ae6c21fc3a","tarball":"https://registry.npmjs.org/@dory-agentic/dory-agentic-sdk/-/dory-agentic-sdk-0.2.0.tgz","fileCount":90,"unpackedSize":193743,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIGSosPydzfvOjyFLuThPPUSxZob27HFvag1JhdZSdeELAiAOOWRdXzNMhNGQe9yQlp36vYMaKa4dWOzZFK0HWSArUA=="}]},"_npmUser":{"name":"agenaidy","email":"ahmad.migo@gmail.com"},"directories":{},"maintainers":[{"name":"agenaidy","email":"ahmad.migo@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/dory-agentic-sdk_0.2.0_1780409869306_0.47210078124386756"},"_hasShrinkwrap":false}},"time":{"created":"2026-06-02T14:17:49.185Z","0.2.0":"2026-06-02T14:17:49.458Z","modified":"2026-06-02T14:17:49.670Z"},"maintainers":[{"name":"agenaidy","email":"ahmad.migo@gmail.com"}],"description":"Dory Agentic SDK: LFS context management with AI SDK tool loops for long-running agents","homepage":"https://github.com/ahmedmigo/dory-agentic-sdk#readme","keywords":["ai-sdk","agent","context-window","lfs","long-running","tool-loop","memory","llm"],"repository":{"type":"git","url":"git+https://github.com/ahmedmigo/dory-agentic-sdk.git"},"bugs":{"url":"https://github.com/ahmedmigo/dory-agentic-sdk/issues"},"license":"MIT","readme":"<p align=\"center\">\n  <strong>Dory Agentic SDK</strong><br/>\n  <em>Long-running agents that never forget the plot.</em>\n</p>\n\n<p align=\"center\">\n  <code>@dory-agentic/dory-agentic-sdk</code>\n</p>\n\n---\n\n## The story of a session that would not end\n\nImagine you pair with an agent on a real project. Not a five-message demo, but a **long session**. You explore a repo, run tools, fix tests, argue about architecture, ship a patch, then ask a question that depends on something you said **two hours ago**.\n\nMost agent stacks treat context like a fixed-size backpack. Every new turn adds weight. Eventually something has to fall out: old instructions, a failing test log, the variable name you agreed on in message twelve. The model does not get angry. It simply **stops seeing** what mattered.\n\n**Dory Agentic SDK** exists for that moment.\n\nWe built **Long-Form Session (LFS)**, a virtual memory layer that sits between your agent loop and the model. The full conversation still lives in a session store, but **what the model sees on each call is curated dynamically**: expanded when there is room, compressed when pressure rises, and restorable when the agent needs the exact words again.\n\nYou keep the long story. The model keeps a window it can actually think inside.\n\n### Proof at million-token scale\n\nOn **AgencyBench V2** (GAIR/AgencyBench, ACL 2026), six scenarios were run with **~1M tokens** of agent history each (Backend and Code domains, `gemini-2.5-flash-lite`). Native Gemini hit **context truncation** on every run. With LFS v4.6:\n\n| Metric | Native Gemini | LFS v4.6 |\n|--------|---------------|----------|\n| Avg rubric score | 93.7 / 100 | **99.7 / 100** (+6.0) |\n| Avg latency | 25.6s | **19.0s** (~26% faster) |\n| Success rate | 6 / 6 | 6 / 6 |\n| Tokens per call (typical) | ~1.42M (truncated) | ~720K (managed) |\n\nStandout case: **Code scenario 3** scored **80/100** natively vs **100/100** with LFS (+20), after **524 parity swaps** and a **memory reflection** that hydrated **9 archived chunks** back into active context. Backend and Code category averages moved from **96.3** and **91.0** to **99.3** and **100.0** respectively.\n\nThat is the story in numbers: the session stays long, the window stays honest, and the agent can still find the needle.\n\n---\n\n## The long session problem\n\nAfter dozens of turns, three forces collide:\n\n1. **Volume**: code, tool output, and reasoning traces pile up faster than anyone summarizes them.\n2. **A hard ceiling**: every model has a context limit; you cannot paste infinity into one call.\n3. **Silent loss**: naive truncation drops the wrong messages: not always the oldest, not always the least important.\n\n![The long session problem: chat timeline grows while the context window fills to 100%](docs/images/infographic-01-the-problem.png)\n\nThat is the gap we address. Not by pretending the window is bigger, but by **managing it like memory**.\n\n---\n\n## What changes with LFS\n\n| Without dynamic context | With Dory Agentic SDK |\n|-------------------------|------------------------|\n| One growing blob sent every call | A **session store** plus a **prepared view** per step |\n| Old turns vanish unpredictably | Low-utility chunks **archive** with labeled placeholders |\n| Agent cannot recover detail | **`remind`** tool hydrates exact text when needed |\n| Same policy at turn 3 and turn 300 | **Preflight routing**: light path when calm, full MMU when pressured |\n\n![Before and after: lost context vs organized active and archived memory](docs/images/infographic-03-before-after.png)\n\n### Benefits for builders\n\n- **Hours-long sessions**: coding agents, support bots, research assistants that accumulate real history.\n- **Dynamic context window**: fill ratio is monitored every model round; eviction tracks relevance, recency, and agent credit scores, not arbitrary cutoffs.\n- **Tool-loop native**: built on the AI SDK `ToolLoopAgent`; LFS runs in `prepareStep` before each inference, so every tool call sees a fresh, budgeted prompt.\n- **Production-minded**: path-safe skill loading, MCP lifecycle disposal, SSRF guards, session cache bounds.\n- **Two lanes**: `standard` for baseline comparison; `optimized` for the full memory manager.\n\n---\n\n## How the agent scores messages (and what gets forgotten)\n\nLFS does not delete history. It **ranks** it every turn so the least useful material leaves the active window first, while the full session remains in storage.\n\n![Scoring and long-term memory: utility, forgetScore, and archive path](docs/images/infographic-04-scoring-and-long-memory.png)\n\n### Step 1: Chunk and index\n\nLarge user turns and tool output are split into **chunks**. Each chunk is embedded in a lightweight vector index so the current user query can retrieve semantically related pieces.\n\n### Step 2: Utility score (what to keep in focus)\n\nFor each chunk, LFS computes **utility** from the current query:\n\n```\nU = β · S_rel + (1 − β) · W_agent\n```\n\n| Symbol | Meaning |\n|--------|---------|\n| **S_rel** | Semantic similarity between the chunk and the active query (vector retrieval) |\n| **W_agent** | Agent-assigned credit weight (what the model marked as important via the scoring manifest) |\n| **β** | Blend factor (default 0.5): balance retrieval vs agent judgment |\n\n**forgetScore** is derived as `1 − U`. Low forgetScore means “keep near the model.” High forgetScore means “safe to archive.”\n\n### Step 3: Temporal decay (what drifts toward archive)\n\nEven strong chunks age unless the agent keeps using them:\n\n- Each turn, **decay** increases forgetScore for conversational messages (mass-aware: larger messages decay faster under pressure).\n- **CODE_ASSET** chunks get a short **shield** when referenced within the last few turns, then fade exponentially as the task moves on.\n- **TOOL_LOG** chunks are penalized first under pressure (noisy logs leave before core instructions).\n- **Hydrated** chunks (brought back via `remind` or reflection) get a **grace period** with forgetScore pinned to 0.\n\nWhen forgetScore reaches **1.0**, the message is a candidate for **displacement**: content moves to **long-term memory** and the active window shows a compact, labeled placeholder.\n\n### Step 4: Parity swap into long-term memory\n\nWhen projected tokens exceed the **memory cap** (default **70%** of `contextMax`, tunable from 50% to 99%):\n\n1. Chunks are sorted by utility (lowest first).\n2. Lowest-utility **active** chunks are marked `isArchived` and copied to **`longTermMemory`**.\n3. The model still sees **where** something went (ids, summaries, manifest), not a silent hole.\n4. The agent can call **`remind`** or trigger a **reflection** to pull exact text back (AgencyBench run: 9 chunks hydrated in one Code scenario).\n\n**Credit budget** scales with window size: under pressure the model receives a bounded number of scoring credits so it must declare what it relied on, which feeds **W_agent** on later turns.\n\n---\n\n## How it works (every model call)\n\nThink of each turn as a **flight check**:\n\n1. **Merge** new messages into the session.\n2. **Preflight**: if projected tokens stay below ~50% of the window, take the fast **PASS_THRU** lane; otherwise engage **ELASTIC_SWAP** and the memory manager.\n3. **Chunk** large user content and index it for similarity search.\n4. **Score and swap**: parity swap ranks chunks by utility (semantic match + agent weights + decay); the least useful leave the window first.\n5. **Budget credits**: under pressure, the model receives a scoring budget and manifest metadata so it cites what it actually used.\n6. **Format**: only then does the prompt go to the model, wrapped, labeled, and honest about what is archived.\n\n![LFS flow: session store → preflight → chunk → vector match → parity swap → dynamic context](docs/images/infographic-02-lfs-flow.png)\n\n```mermaid\nflowchart LR\n  subgraph store [Session]\n    History[fullHistory]\n    LTM[longTermMemory]\n    Weights[agentWeights]\n  end\n\n  subgraph prepare [prepareStep on every tool call]\n    Router[preflightRouter]\n    Chunk[chunkManager]\n    Vector[vectorDb]\n    Parity[paritySwapper]\n    Engine[memoryEngine]\n    Format[formatForModel]\n  end\n\n  History --> Router\n  Router --> Chunk --> Vector --> Parity --> Engine --> Format\n  Parity --> LTM\n  Format --> Model[(Language model)]\n```\n\nThe agent loop (tools, streaming, cancellation) stays in the AI SDK. **LFS only decides what the model reads.**\n\n---\n\n## A day in the life (narrative walkthrough)\n\n**Turn 1 to 5: Plenty of space.**  \nPreflight returns `PASS_THRU`. The model sees the conversation almost as-is. LFS is quiet; you pay almost no overhead.\n\n**Turn 40: A fat log arrives.**  \nTool output pushes projected tokens past the threshold. LFS switches to `ELASTIC_SWAP`, chunks the new material, and starts scoring older messages. Anything irrelevant drifts toward archive summaries: still named, still addressable, no longer eating tokens.\n\n**Turn 41: The agent needs the exact error line.**  \nIt calls `remind` with a message or chunk id. The SDK pulls the raw text back from long-term storage into active context for that step. Surgical, not a full replay of history.\n\n**Turn 200: Still the same session id.**  \n`createLfsAgent` and `LfsContextAdapter` share one session store. Memory cap fraction (default **70%**, tunable **50% to 99%**) tells LFS how aggressively to keep the window below the physical limit before the provider truncates.\n\nThat is the innovation: **context as a managed resource**, not a static string.\n\n---\n\n## Install\n\n```bash\nnpm install @dory-agentic/dory-agentic-sdk\n```\n\nMonorepo / local development:\n\n```bash\nnpm install file:../dory-agentic-sdk\n```\n\n```json\n{\n  \"dependencies\": {\n    \"@dory-agentic/dory-agentic-sdk\": \"file:../dory-agentic-sdk\"\n  }\n}\n```\n\n---\n\n## Quick start\n\n### Full agent (tool loop + LFS)\n\n```typescript\nimport { createLfsAgent, createOllamaProvider } from \"@dory-agentic/dory-agentic-sdk\";\n\nconst ollama = createOllamaProvider();\nconst model = ollama(\"qwen2.5-coder:7b\");\n\nconst agent = await createLfsAgent({\n  sessionId: \"project-alpha\",\n  lane: \"optimized\",\n  model,\n  contextMax: 32768,\n  memoryCapFraction: 0.7,\n});\n\n// Use agent.toolLoopAgent with your AI SDK patterns, then:\nawait agent.dispose();\n```\n\n### Prepare context only (bring your own loop)\n\n```typescript\nimport { LfsContextAdapter } from \"@dory-agentic/dory-agentic-sdk\";\n\nconst adapter = new LfsContextAdapter();\nconst prepared = await adapter.prepareStep({\n  sessionId: \"project-alpha\",\n  lane: \"optimized\",\n  incomingMessages: history,\n  replaceSessionHistory: true,\n  contextMax: 32768,\n  memoryCapFraction: 0.7,\n  useRemindTool: true,\n});\n\n// prepared.messagesForModel → send to your model\n// prepared.isMmuActive, prepared.executionLane → observability\n```\n\n### One-shot helper\n\n```typescript\nimport { runLfsTurn } from \"@dory-agentic/dory-agentic-sdk\";\n\nconst result = await runLfsTurn({\n  sessionId: \"bench-1\",\n  lane: \"optimized\",\n  incomingMessages: messages,\n  replaceSessionHistory: true,\n  contextMax: 8192,\n});\n```\n\n---\n\n## Core concepts\n\n| Concept | Meaning |\n|---------|---------|\n| **Session store** | Canonical history, archived chunks, forget scores, agent weights |\n| **Utility / forgetScore** | Rank for retention; archive when pressure forces lowest-U chunks out |\n| **PASS_THRU** | Low pressure; minimal memory management |\n| **ELASTIC_SWAP** | High pressure; chunking, parity swap, credit budget active |\n| **Memory cap fraction** | How full LFS allows the window to get before aggressive eviction (50% to 99%) |\n| **longTermMemory** | Archived chunks keyed by id; placeholders stay in the active manifest |\n| **remind** | Tool to hydrate archived content back into active context |\n| **Lanes** | `standard` (baseline) vs `optimized` (full LFS) |\n\n---\n\n## Configuration\n\n| Variable | Purpose |\n|----------|---------|\n| `OLLAMA_BASE_URL` | Base URL for Ollama integration tests |\n| `DORY_AGENTIC_SKILLS_ROOT` | Allowed directory root for `skillLoader` |\n| `DORY_AGENTIC_MCP_FETCH_ALLOWLIST` | Comma-separated host allowlist for fetch MCP URLs |\n| `MCP_FETCH_URL_ALLOWLIST` | Alias for the above |\n\n---\n\n## Development\n\n```bash\nnpm install\nnpm run typecheck\nnpm run test:unit      # fast, no network\nnpm run test:ollama    # optional; needs local Ollama\nnpm run build\n```\n\n---\n\n## Architecture (repository layout)\n\n```\nsrc/\n  agent/          createLfsAgent, runLfsTurn, ToolLoopAgent wiring\n  lfs/            preflight, chunking, vector index, parity swap, memory engine\n  history/        manifest formatting for the model\n  tools/          remind, skillLoader, registry\n  mcp/            MCP client helpers\n  providers/      Ollama and provider registry\n```\n\n---\n\n## Related packages\n\n- **`@dory-agentic/dory-agentic-sdk`** (this repo): LFS virtual memory for AI SDK agent loops.\n- **`@dorycode-ai/sdk`**: separate HTTP client for the DoryCode server.\n\n---\n\n## License\n\nMIT. See [package.json](package.json).\n\n---\n\n<p align=\"center\">\n  <sub>Built for agents that outlive the context window.</sub>\n</p>\n","readmeFilename":"README.md","_rev":"1-14e7a5fbf5e7fbf160a9bc30c4a4b067"}