{"_id":"@cheapestinference/openclaw-ratelimit-retry","name":"@cheapestinference/openclaw-ratelimit-retry","dist-tags":{"latest":"1.0.0"},"versions":{"1.0.0":{"name":"@cheapestinference/openclaw-ratelimit-retry","version":"1.0.0","description":"Automatically retry agent conversations that fail due to provider rate limits","type":"module","license":"MIT","repository":{"type":"git","url":"git+https://github.com/cheapestinference/openclaw-plugin-ratelimit-retry.git"},"keywords":["openclaw","plugin","retry","rate-limit","429","budget","ratelimit"],"openclaw":{"extensions":["./index.ts"]},"_id":"@cheapestinference/openclaw-ratelimit-retry@1.0.0","gitHead":"50e56ca5f55fb1314fec53fa060d65bfcfa906ff","bugs":{"url":"https://github.com/cheapestinference/openclaw-plugin-ratelimit-retry/issues"},"homepage":"https://github.com/cheapestinference/openclaw-plugin-ratelimit-retry#readme","_nodeVersion":"22.22.0","_npmVersion":"10.9.4","dist":{"integrity":"sha512-YqDKwA9EqNMami5sR4CT3Lw2Vt22DhO3fpwnkzDu90+dh754juiIVCJqyz+iG8zGIqlQqMxltam/mIqN74kLtw==","shasum":"b894d15f7adf7899c6404838f36899e98b582ee4","tarball":"https://registry.npmjs.org/@cheapestinference/openclaw-ratelimit-retry/-/openclaw-ratelimit-retry-1.0.0.tgz","fileCount":8,"unpackedSize":50178,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIDlDH2dBWutR6wtdR51cg/j6YPgvJTeHZpzmOgGpNXn3AiBiKWRK2ZPSrDtMuZ9MxgfabTc2+n9cAtYO9WVUf4mWjg=="}]},"_npmUser":{"name":"cheapestinference","email":"admin@cheapestinference.com"},"directories":{},"maintainers":[{"name":"cheapestinference","email":"admin@cheapestinference.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/openclaw-ratelimit-retry_1.0.0_1773347670595_0.8653501414485927"},"_hasShrinkwrap":false}},"time":{"created":"2026-03-12T20:34:30.512Z","1.0.0":"2026-03-12T20:34:30.753Z","modified":"2026-03-12T20:34:30.951Z"},"maintainers":[{"name":"cheapestinference","email":"admin@cheapestinference.com"}],"description":"Automatically retry agent conversations that fail due to provider rate limits","homepage":"https://github.com/cheapestinference/openclaw-plugin-ratelimit-retry#readme","keywords":["openclaw","plugin","retry","rate-limit","429","budget","ratelimit"],"repository":{"type":"git","url":"git+https://github.com/cheapestinference/openclaw-plugin-ratelimit-retry.git"},"bugs":{"url":"https://github.com/cheapestinference/openclaw-plugin-ratelimit-retry/issues"},"license":"MIT","readme":"# ratelimit-retry\n\nAn OpenClaw plugin that automatically retries agent conversations killed by provider rate limits.\n\n## Problem\n\nWhen your LLM provider hits a rate limit or budget cap (HTTP 429), every running agent task dies mid-conversation. Nothing resumes them. If you close the dashboard, those conversations are gone. You have to manually find and re-trigger each one after the budget resets.\n\n## Solution\n\nThis plugin hooks into OpenClaw's `agent_end` event, detects retriable errors (429s, rate limits, budget exhaustion), and parks the failed session in a persistent queue on disk. A background service waits for the provider's budget window to reset, then sends `chat.send` to the original session -- resuming the conversation with its full transcript context, as if the user had typed a message.\n\n## Installation\n\n```bash\nopenclaw plugins install @cheapestinference/openclaw-ratelimit-retry\n```\n\nOr copy manually to your extensions directory:\n\n```bash\ncp -r openclaw-plugin-ratelimit-retry ~/.openclaw/extensions/ratelimit-retry\n```\n\nEnable it in OpenClaw config:\n\n```bash\nopenclaw config set plugins.ratelimit-retry.budgetWindowHours 5\nopenclaw config set plugins.ratelimit-retry.maxRetryAttempts 3\n```\n\nNo `npm install` needed. The plugin has zero runtime dependencies.\n\n### Complete example\n\n```yaml\n# ~/.openclaw/config.yaml\nplugins:\n  ratelimit-retry:\n    budgetWindowHours: 5\n    maxRetryAttempts: 3\n    checkIntervalMinutes: 5\n    retryMessage: \"Continue where you left off. The previous attempt failed due to a rate limit that has now reset.\"\n```\n\n## How It Works\n\n```\nAgent run fails (429)\n  |\n  v\nagent_end hook fires\n  |-- Non-retriable error? --> ignore\n  |-- Retriable error?     --> queue to disk\n                                 |\n                                 v\n                  Background timer (every 5 min)\n                    |\n                    |-- Budget window not reset? --> wait\n                    |-- Budget window reset?     --> chat.send to session\n                                                       |\n                                                       |--> Ack received: wait for result\n                                                       |     |--> agent_end success: remove from queue\n                                                       |     |--> agent_end 429: re-queued automatically\n                                                       |--> Send failed: wait for next window\n```\n\nThe retry uses `chat.send` with the original `sessionKey`, which means the gateway loads the complete JSONL transcript and the agent resumes with full context. This is equivalent to the user typing a message in the chat.\n\nThe model is **fire-and-forget with re-detection**: `chat.send` returns an immediate ack (`{ ok, runId, status: \"started\" }`), not the final result. If the retried run fails again with a 429, the `agent_end` hook fires again and the session is re-queued with an incremented attempt counter. This loop continues until the retry succeeds or `maxRetryAttempts` is reached.\n\n## Configuration\n\n| Option | Type | Default | Description |\n|--------|------|---------|-------------|\n| `budgetWindowHours` | `number` | `5` | Budget reset window in hours, aligned to UTC clock boundaries |\n| `maxRetryAttempts` | `number` | `3` | Max retries per session before abandoning |\n| `checkIntervalMinutes` | `number` | `5` | How often the background service checks for pending retries |\n| `retryMessage` | `string` | `\"Continue where you left off...\"` | Message sent to the session to resume the conversation |\n\n## How the Retry Timing Works\n\nMany LLM providers (including LiteLLM) reset budget counters on fixed UTC-aligned windows. With a 5-hour window, the boundaries are:\n\n```\n00:00  05:00  10:00  15:00  20:00  (next day) 00:00\n  |------|------|------|------|------|\n```\n\nWhen an error is queued, the plugin calculates the next boundary after the current time and adds a **1-minute margin** (retries at `HH:01:00` instead of `HH:00:00`) to avoid racing the provider's reset.\n\n**When 24 is not evenly divisible by `windowHours`**: the math still works. If `windowHours` is 7, boundaries fall at 0, 7, 14, 21, and the next one would be 28 -- which overflows to 04:00 the next day. The plugin handles day overflow correctly.\n\n## Error Classification\n\nNon-retriable patterns are checked first. If an error matches a non-retriable pattern, it is never retried, even if it also matches a retriable pattern.\n\n### Retriable (queued for retry)\n\n| Pattern | Catches |\n|---------|---------|\n| `429` | `\"Error code: 429 - ...\"` |\n| `rate limit`, `rate_limit` | `\"RateLimitError: ...\"` |\n| `too many requests` | HTTP 429 reason phrases |\n| `budget` | `\"Budget exceeded for ...\"` |\n| `quota exceeded` | Provider quota messages |\n| `resource exhausted` | gRPC-style exhaustion errors |\n| `tokens per minute`, `tpm` | TPM limit messages |\n\n### Non-retriable (ignored)\n\n| Pattern | Reason |\n|---------|--------|\n| `401`, `402`, `403`, `404` | HTTP client errors -- won't succeed on retry |\n| `invalid api key`, `unauthorized` | Auth errors -- fix your credentials |\n| `invalid request`, `malformed` | Bad request format -- won't succeed on retry |\n| `model not found` | Model doesn't exist |\n| `context length`, `prompt too large` | Context overflow -- message is too long |\n| `insufficient credits` | Billing issue -- requires user action |\n\n## Edge Cases\n\n- **Server restarts**: the queue is persisted to `{stateDir}/ratelimit-retry/queue.json` and reloaded on startup.\n- **Same session errors multiple times**: deduplicated by `sessionKey`. The existing entry is updated with incremented attempts and a recalculated `retryAfter`.\n- **Retry fails with 429 again**: `agent_end` fires again, re-queuing with incremented attempts. Natural loop until success or `maxRetryAttempts`.\n- **Gateway unreachable during retry**: connection error is caught, entry's `retryAfter` is pushed to the next budget window to avoid hammering a down gateway every tick.\n- **Max attempts exceeded**: entry is removed from queue and a warning is logged.\n- **Sub-agent sessions**: handled identically -- `sessionKey` format `agent:X:subagent:Y` works the same way.\n- **Timer fires during active retry**: a `retryInProgress` guard prevents overlapping batches.\n- **Queue file corrupted**: JSON parse errors are caught; service starts with an empty queue and logs a warning.\n- **Queue overflow**: capped at 100 entries. Oldest entries are evicted when full.\n- **Atomic writes**: queue is written to a uniquely-named `.tmp` file first, then renamed, to prevent corruption on crashes or concurrent writes.\n\n## Limitations\n\n- **Fire-and-forget window**: after `chat.send` returns its ack, there is a brief period where the retried run is in progress. If it fails with 429 again immediately, there is a small window before the `agent_end` hook fires and re-queues it. This is by design -- the re-detection loop handles it.\n- **`chat.send` requires a non-empty message**: the retry always sends the configured `retryMessage`. It cannot send an empty message to silently resume.\n- **No partial-run recovery**: the plugin resumes the conversation from the last completed turn. It does not replay partial streaming output that was interrupted.\n- **Single-instance only**: the queue is a local JSON file with no locking. Running multiple OpenClaw instances sharing the same `~/.openclaw/` directory is not supported.\n- **No backpressure on the provider**: the plugin retries all ready sessions in sequence. If you have many queued sessions, they all fire at the start of the next window.\n\n## License\n\n[MIT](LICENSE)\n\n## Contributing\n\nContributions are welcome. Please open an issue first to discuss what you would like to change.\n\n```bash\ngit clone https://github.com/cheapestinference/openclaw-plugin-ratelimit-retry\ncd openclaw-plugin-ratelimit-retry\n# No build step. OpenClaw loads .ts files directly via Jiti.\n```\n","readmeFilename":"README.md","_rev":"1-de2061c6fcea156bb8591663c23a2f8d"}