{"_id":"@aid-on/llm-queue-dispatcher","name":"@aid-on/llm-queue-dispatcher","dist-tags":{"latest":"1.0.2"},"versions":{"1.0.2":{"name":"@aid-on/llm-queue-dispatcher","version":"1.0.2","description":"高度なスコアリングアルゴリズムを備えたLLMリクエスト用キューディスパッチャー - LLM Queue Dispatcher for LLM requests with advanced scoring algorithms","main":"dist/index.js","types":"dist/index.d.ts","module":"dist/index.mjs","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.mjs","require":"./dist/index.js"}},"scripts":{"build":"tsup","dev":"tsup --watch","type-check":"tsc --noEmit","lint":"eslint src --ext .ts","test":"vitest run","test:coverage":"vitest --coverage","demo:dev":"vite","demo:build":"vite build","demo:preview":"vite preview","prepublishOnly":"npm run build && npm run type-check"},"keywords":["llm","queue","rate-limit","intelligent","scoring","priority","typescript"],"author":{"name":"aid-on"},"license":"MIT","engines":{"node":">=16.0.0"},"publishConfig":{"access":"public"},"repository":{"type":"git","url":"git+https://github.com/Aid-On/llm-queue-dispatcher.git"},"homepage":"https://Aid-On.github.io/llm-queue-dispatcher/","bugs":{"url":"https://github.com/Aid-On/llm-queue-dispatcher/issues"},"devDependencies":{"@types/node":"^20.11.0","@typescript-eslint/eslint-plugin":"^6.19.0","@typescript-eslint/parser":"^6.19.0","@vitest/coverage-v8":"^1.2.0","eslint":"^8.56.0","tsup":"^8.0.1","typescript":"^5.3.3","vite":"^5.4.19","vitest":"^1.2.0"},"peerDependencies":{"@aid-on/llm-throttle":"^1.0.1"},"_id":"@aid-on/llm-queue-dispatcher@1.0.2","gitHead":"5ff1d82e02ab7a5d0c3f915fd9e19f9336aad7ac","_nodeVersion":"20.10.0","_npmVersion":"10.2.3","dist":{"integrity":"sha512-TnH9D8wMbrS99JX6ZFpVTRpxbRzfivQZulvhjTZti63+s3hiyW6HKLNVjMEs30FgeiUaEaBlQe1kRugLIsWchw==","shasum":"9540b3bc900dba8f08f37b82807b4f39b5d4bd2b","tarball":"https://registry.npmjs.org/@aid-on/llm-queue-dispatcher/-/llm-queue-dispatcher-1.0.2.tgz","fileCount":10,"unpackedSize":240694,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCICe0v/jKWIVUy+aSlcBG7FqnSWHy4O7XUy35oZqu1iCDAiEA/FvB5vdpIr5/Du/iuR4qXQJVneGKSYk4yR8dNgi3NwQ="}]},"_npmUser":{"name":"aid-on","email":"hiromi.motodera@aid-on.org"},"directories":{},"maintainers":[{"name":"aid-on","email":"hiromi.motodera@aid-on.org"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/llm-queue-dispatcher_1.0.2_1754216653454_0.010094484456041775"},"_hasShrinkwrap":false}},"time":{"created":"2025-08-03T10:24:13.357Z","1.0.2":"2025-08-03T10:24:13.626Z","modified":"2025-08-03T10:24:13.945Z"},"maintainers":[{"name":"aid-on","email":"hiromi.motodera@aid-on.org"}],"description":"高度なスコアリングアルゴリズムを備えたLLMリクエスト用キューディスパッチャー - LLM Queue Dispatcher for LLM requests with advanced scoring algorithms","homepage":"https://Aid-On.github.io/llm-queue-dispatcher/","keywords":["llm","queue","rate-limit","intelligent","scoring","priority","typescript"],"repository":{"type":"git","url":"git+https://github.com/Aid-On/llm-queue-dispatcher.git"},"author":{"name":"aid-on"},"bugs":{"url":"https://github.com/Aid-On/llm-queue-dispatcher/issues"},"license":"MIT","readme":"# @aid-on/llm-queue-dispatcher\n\n> 🧠 LLM Queue Dispatcher for LLM requests with advanced scoring algorithms and rate-limit awareness\n\n[![npm version](https://badge.fury.io/js/%40aid-on%2Fllm-queue-dispatcher.svg)](https://www.npmjs.com/package/@aid-on/llm-queue-dispatcher)\n[![TypeScript](https://img.shields.io/badge/TypeScript-Ready-blue.svg)](https://www.typescriptlang.org/)\n[![MIT License](https://img.shields.io/badge/license-MIT-green.svg)](./LICENSE)\n\n## Overview\n\n`@aid-on/llm-queue-dispatcher` is a sophisticated queueing system designed specifically for LLM (Large Language Model) request processing. It uses advanced scoring algorithms to intelligently select the most optimal requests for processing based on multiple factors including priority, token efficiency, wait time, and rate limiting constraints.\n\n🚀 **[Try the Interactive Demo](https://aid-on.github.io/llm-queue-dispatcher/)** - Experience the queue in action with real-time visualizations!\n\n## Features\n\n- 🎯 **Multi-dimensional Scoring**: Considers priority, efficiency, wait time, retry count, token fit, and processing time\n- ⚡ **Rate Limiter Integration**: Seamlessly works with `@aid-on/llm-throttle` for rate-aware processing\n- 🔄 **Prefetching Support**: Optional message prefetching for improved throughput\n- 🗂️ **Abstract Storage**: Pluggable storage adapters (in-memory included, extend for SQS, Redis, etc.)\n- 📊 **Rich Metrics**: Comprehensive monitoring and performance metrics\n- 🏭 **Factory Functions**: Pre-configured queue types for common use cases\n- 🧪 **TypeScript First**: Fully typed with excellent IDE support\n\n## Installation\n\n```bash\nnpm install @aid-on/llm-queue-dispatcher\n```\n\n## Quick Start\n\n### Basic Usage\n\n```typescript\nimport { createInMemoryLLMQueueDispatcher, Priority } from '@aid-on/llm-queue-dispatcher';\nimport { createLLMThrottle } from '@aid-on/llm-throttle';\n\n// Create queue and rate limiter\nconst queue = createInMemoryLLMQueueDispatcher();\nconst rateLimiter = createLLMThrottle({ rpm: 60, tpm: 10000 });\n\n// Enqueue requests\nawait queue.enqueue({\n  id: 'req-1',\n  payload: { prompt: 'Hello, world!' },\n  priority: Priority.HIGH,\n  tokenInfo: { estimated: 150 },\n  createdAt: new Date()\n});\n\n// Process messages\nconst processable = await queue.dequeue(rateLimiter);\nif (processable) {\n  try {\n    // Your LLM processing logic here\n    const result = await processLLMRequest(processable.message.payload);\n    await processable.markAsProcessed();\n  } catch (error) {\n    await processable.markAsFailed(error);\n  }\n}\n```\n\n### Custom Storage\n\n```typescript\nimport { createLLMQueueDispatcher } from '@aid-on/llm-queue-dispatcher';\n\n// Implement your storage adapter\nclass SQSStorage implements QueueStorageAdapter {\n  async enqueue(message) { /* SQS implementation */ }\n  async dequeue(limit, timeout) { /* SQS implementation */ }\n  // ... other methods\n}\n\nconst queue = createLLMQueueDispatcher(new SQSStorage(), {\n  enablePrefetch: true,\n  bufferSize: 100\n});\n```\n\n## API Reference\n\n### Core Classes\n\n#### `LLMQueueDispatcher<T>`\n\nThe main queue class that handles intelligent message selection and processing.\n\n```typescript\nclass LLMQueueDispatcher<T = LLMPayload> {\n  constructor(storage: QueueStorageAdapter<LLMRequest<T>>, config?: LLMQueueDispatcherConfig)\n  \n  async enqueue(request: LLMRequest<T>): Promise<void>\n  async batchEnqueue(requests: LLMRequest<T>[]): Promise<void>\n  async dequeue(rateLimiter: LLMThrottle): Promise<ProcessableMessage<T> | null>\n  async getQueueMetrics(): Promise<QueueMetrics>\n  async purge(): Promise<void>\n  stop(): void\n}\n```\n\n#### Configuration\n\n```typescript\ninterface LLMQueueDispatcherConfig {\n  bufferSize?: number;                    // Prefetch buffer size (default: 50)\n  enablePrefetch?: boolean;               // Enable message prefetching (default: false)\n  prefetchInterval?: number;              // Prefetch interval in ms (default: 5000)\n  maxCandidatesToEvaluate?: number;       // Max messages to score (default: 20)\n  minScoreThreshold?: number;             // Minimum score to process (default: 0.1)\n  scoring?: ScoringConfig;                // Custom scoring configuration\n  metricsRetentionMs?: number;            // Metrics retention time\n  logger?: Logger;                        // Custom logger\n}\n```\n\n### Factory Functions\n\n#### `createInMemoryQueue<T>(config?)`\nCreates a queue with in-memory storage (ideal for testing and development).\n\n#### `createPrefetchingQueue<T>(storage, config?)`\nCreates a queue with prefetching enabled for high-throughput scenarios.\n\n#### `createSimplePriorityQueue<T>(storage, config?)`\nCreates a queue that primarily uses priority-based selection.\n\n#### `createThroughputOptimizedQueue<T>(storage, config?)`\nCreates a queue optimized for maximum throughput and TPM efficiency.\n\n#### `createFairQueue<T>(storage, config?)`\nCreates a queue that balances fairness with priority (FIFO-like behavior).\n\n### Message Types\n\n```typescript\ninterface LLMRequest<T = LLMPayload> {\n  id: string;\n  payload: T;\n  priority: Priority;\n  tokenInfo: TokenInfo;\n  expectedProcessingTime?: number;\n  metadata?: Record<string, unknown>;\n  createdAt: Date;\n}\n\nenum Priority {\n  URGENT = 0,\n  HIGH = 1,\n  NORMAL = 2,\n  LOW = 3\n}\n```\n\n## Advanced Usage\n\n### Custom Scoring\n\n```typescript\nconst queue = createLLMQueueDispatcher(storage, {\n  scoring: {\n    weights: {\n      priority: 0.3,\n      efficiency: 0.25,\n      waitTime: 0.25,\n      retry: 0.1,\n      tokenFit: 0.1,\n      processingTime: 0.0\n    },\n    customScorers: [{\n      name: 'deadline',\n      weight: 0.2,\n      calculate: (message, context) => {\n        const deadline = message.payload.metadata?.deadline as number;\n        if (!deadline) return 0.5;\n        const timeLeft = deadline - context.currentTime;\n        return Math.max(0, Math.min(1, timeLeft / 3600000)); // 1 hour max\n      }\n    }]\n  }\n});\n```\n\n### Storage Adapter Implementation\n\n```typescript\nclass RedisStorage implements QueueStorageAdapter<LLMRequest> {\n  constructor(private redis: Redis) {}\n  \n  async enqueue(message: LLMRequest): Promise<QueueMessage<LLMRequest>> {\n    const queueMessage = {\n      id: generateId(),\n      payload: message,\n      attributes: {\n        messageId: generateId(),\n        receiptHandle: generateHandle(),\n        enqueuedAt: new Date(),\n        receiveCount: 0\n      }\n    };\n    \n    await this.redis.lpush('queue', JSON.stringify(queueMessage));\n    return queueMessage;\n  }\n  \n  async dequeue(limit: number, visibilityTimeout: number): Promise<QueueMessage<LLMRequest>[]> {\n    // Redis implementation with visibility timeout logic\n    // ...\n  }\n  \n  // Implement other required methods...\n}\n```\n\n### Monitoring and Metrics\n\n```typescript\nconst metrics = await queue.getQueueMetrics();\n\nconsole.log(`Total messages: ${metrics.queue.totalMessages}`);\nconsole.log(`In-flight: ${metrics.processing.activeRequests}`);\nconsole.log(`Throughput: ${metrics.queue.throughput.messagesPerMinute} msg/min`);\nconsole.log(`Buffer utilization: ${metrics.performance.bufferUtilization * 100}%`);\n```\n\n## Integration with LLM Throttle\n\nThe queue is designed to work seamlessly with `@aid-on/llm-throttle`:\n\n```typescript\nimport { createLLMThrottle } from '@aid-on/llm-throttle';\n\nconst rateLimiter = createLLMThrottle({\n  rpm: 60,        // 60 requests per minute\n  tpm: 10000,     // 10,000 tokens per minute\n  burstTPM: 15000 // Allow bursts up to 15,000 TPM\n});\n\n// The queue automatically considers rate limits when selecting messages\nconst processable = await queue.dequeue(rateLimiter);\n```\n\n## Error Handling\n\n```typescript\nconst processable = await queue.dequeue(rateLimiter);\nif (processable) {\n  try {\n    const result = await processLLMRequest(processable.message.payload);\n    await processable.markAsProcessed();\n  } catch (error) {\n    if (error.code === 'RATE_LIMITED') {\n      // Extend visibility timeout to retry later\n      await processable.updateVisibility(300); // 5 minutes\n    } else {\n      // Mark as failed for other errors\n      await processable.markAsFailed(error);\n    }\n  }\n}\n```\n\n## Demo\n\nExplore the interactive demo to see the intelligent queue in action:\n\n```bash\n# Clone the repository\ngit clone https://github.com/aid-on-libs/llm-queue-dispatcher.git\ncd llm-queue-dispatcher\n\n# Install dependencies\nnpm install\n\n# Start the demo server\nnpm run demo:dev\n```\n\nThe demo features:\n- **Real-time Queue Visualization**: See messages moving through pending, processing, and completed states\n- **Interactive Rate Limiting**: Adjust RPM/TPM limits and see their effect on message selection\n- **Scoring Algorithm Demonstration**: View detailed scoring breakdowns for each message\n- **Multiple Queue Types**: Compare different pre-configured queue strategies\n- **Live Metrics Dashboard**: Monitor throughput, efficiency, and performance metrics\n\n## Testing\n\nThe library includes comprehensive test utilities:\n\n```typescript\nimport { createInMemoryLLMQueueDispatcher } from '@aid-on/llm-queue-dispatcher';\n\n// In-memory queue is perfect for testing\nconst testQueue = createInMemoryQueue();\n\n// Test your queue processing logic\nawait testQueue.enqueue(mockRequest);\nconst processable = await testQueue.dequeue(mockRateLimiter);\nexpect(processable).toBeDefined();\n```\n\n## Performance Tips\n\n1. **Enable Prefetching**: For high-throughput scenarios, enable prefetching to reduce dequeue latency\n2. **Tune Buffer Size**: Larger buffers provide more selection options but use more memory\n3. **Optimize Scoring**: Adjust scoring weights based on your specific requirements\n4. **Monitor Metrics**: Use the built-in metrics to identify bottlenecks\n5. **Custom Storage**: Implement storage adapters optimized for your infrastructure\n\n## License\n\nMIT © aid-on\n\n## Contributing\n\nContributions are welcome! Please check our [GitHub repository](https://github.com/aid-on-libs/llm-queue-dispatcher) for issues and contribution guidelines.\n\n---\n\nFor more examples and detailed documentation, visit our [GitHub Pages site](https://aid-on-libs.github.io/llm-queue-dispatcher/).","readmeFilename":"README.md","_rev":"1-f189784d544fa816ccd5aafd604ea898"}