{"_id":"@abhi-arya1/autoscaled","_rev":"2-14d282abe5cac702146e139786ffe123","name":"@abhi-arya1/autoscaled","dist-tags":{"latest":"0.1.0"},"versions":{"0.0.1":{"name":"@abhi-arya1/autoscaled","version":"0.0.1","author":{"name":"Abhigyan Arya","email":"abhigyaa@uci.edu"},"license":"ISC","_id":"@abhi-arya1/autoscaled@0.0.1","maintainers":[{"name":"abhi-arya1","email":"abhi@opennote.me"}],"homepage":"https://github.com/abhi-arya1/autoscaled","bugs":{"url":"https://github.com/abhi-arya1/autoscaled/issues"},"dist":{"shasum":"603b07b05b455f763b00afd8953c53aacfaddce2","tarball":"https://registry.npmjs.org/@abhi-arya1/autoscaled/-/autoscaled-0.0.1.tgz","fileCount":11,"integrity":"sha512-DZxQxxoVamhaPEXmoI6v0WA5kRU5hFFRdtwMb4VOkhySP8DzvU5M4gu8yNKyWQevpwp1tJfLaE1l/rWHQljh/Q==","signatures":[{"sig":"MEQCIGQB0585aBX0/CPo74s7F4y/Bn7sBNE2Y9IKtLc61WAIAiAK0/u2C3AyG1gVJOzrMx5PpDiDrVja7lJbbMPINH4fqA==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":73512},"main":"./src/index.ts","type":"module","types":"./src/index.ts","module":"./src/index.ts","exports":{".":{"types":"./src/index.ts","import":"./src/index.ts"}},"gitHead":"2080a69bfd92f67d281babf07338776fff0599ac","_npmUser":{"name":"abhi-arya1","email":"abhi@opennote.me"},"repository":{"url":"git://github.com/abhi-arya1/autoscaled.git","type":"git"},"_npmVersion":"11.6.2","description":"Build autoscaling container servers with Cloudflare Workers","directories":{},"_nodeVersion":"24.11.1","dependencies":{"nanoid":"^5.1.6","@cloudflare/containers":"^0.0.31"},"_hasShrinkwrap":false,"devDependencies":{"@types/bun":"latest"},"peerDependencies":{"typescript":"^5","@cloudflare/workers-types":"^4.20251213.0"},"_npmOperationalInternal":{"tmp":"tmp/autoscaled_0.0.1_1765989439779_0.4065767746190627","host":"s3://npm-registry-packages-npm-production"}},"0.1.0":{"name":"@abhi-arya1/autoscaled","version":"0.1.0","repository":{"type":"git","url":"git://github.com/abhi-arya1/autoscaled.git"},"homepage":"https://github.com/abhi-arya1/autoscaled","author":{"name":"Abhigyan Arya","email":"abhigyaa@uci.edu"},"license":"ISC","description":"Build autoscaling container servers with Cloudflare Workers","main":"./src/index.ts","module":"./src/index.ts","types":"./src/index.ts","type":"module","exports":{".":{"types":"./src/index.ts","import":"./src/index.ts"}},"devDependencies":{"@types/bun":"latest"},"peerDependencies":{"typescript":"^5","@cloudflare/workers-types":"^4.20251213.0"},"dependencies":{"@cloudflare/containers":"^0.0.31","nanoid":"^5.1.6"},"gitHead":"8627cabc00e8c92a95f8c797ca74c7ae0c6d616a","_id":"@abhi-arya1/autoscaled@0.1.0","bugs":{"url":"https://github.com/abhi-arya1/autoscaled/issues"},"_nodeVersion":"24.11.1","_npmVersion":"11.6.2","dist":{"integrity":"sha512-Ag+Xoa9/Wb5ctkZFE6hEdvJvsS3ccn4q9bBlaI+8b+Zptn7LfBqupJz6Xv1fp/si8H/PWkhNaYlyrvkXoF5JYA==","shasum":"0c735af1b1bff492dd89b02c1e9453483cd607e1","tarball":"https://registry.npmjs.org/@abhi-arya1/autoscaled/-/autoscaled-0.1.0.tgz","fileCount":11,"unpackedSize":73396,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIG+Z6RxC6kL4jZenNqFK7gxN5OyWFhgspDfrElW3xPAQAiEArvcgOud13klPUBSoHyGe+ebZrAb26orB66CNc47YOas="}]},"_npmUser":{"name":"abhi-arya1","email":"abhi@opennote.me"},"directories":{},"maintainers":[{"name":"abhi-arya1","email":"abhi@opennote.me"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/autoscaled_0.1.0_1765989524202_0.14202405449624877"},"_hasShrinkwrap":false}},"time":{"created":"2025-12-17T16:37:19.667Z","modified":"2025-12-17T16:38:44.548Z","0.0.1":"2025-12-17T16:37:19.911Z","0.1.0":"2025-12-17T16:38:44.349Z"},"bugs":{"url":"https://github.com/abhi-arya1/autoscaled/issues"},"author":{"name":"Abhigyan Arya","email":"abhigyaa@uci.edu"},"license":"ISC","homepage":"https://github.com/abhi-arya1/autoscaled","repository":{"type":"git","url":"git://github.com/abhi-arya1/autoscaled.git"},"description":"Build autoscaling container servers with Cloudflare Workers","maintainers":[{"name":"abhi-arya1","email":"abhi@opennote.me"}],"readme":"# AutoscaleD\n\nAutomatically scale Cloudflare Containers based on compute, load, status, and more, for distributed applications on Cloudflare's edge network.\n\n## Installation\n\n```shell\nnpm install @abhi-arya1/autoscaled\n```\n\n## Why Use AutoscaleD?\n\n**AutoscaleD** provides an automatic load balancing service that sits in front of any Cloudflare Container and automatically controls the number of running containers based on actual demand. It can scale to 0, ensuring you have no unused compute after a point, while intelligently scaling up when needed to minimize latency for users worldwide.\n\n## Usage\n\nHere's an example of how to use AutoscaleD:\n\nWrite your code:\n\n```ts\n// Define your Container\nexport class MyContainer extends Container<Env> {\n    // Port the container listens on (default: 8080)\n    defaultPort = 8080;\n\n    // Port 81 for monitor (see Monitoring section below)\n    requiredPorts = [8080, 81];\n\n    // Set this to anything greater than heartbeatInterval, since this will no longer be used, and the Autoscaler will manage sleep/wakeup.\n    sleepAfter = \"2m\";\n    envVars = {\n        MESSAGE: \"I was passed in via the container class!\",\n    };\n}\n\n// Import and Define your Autoscaler\n\nimport { Autoscaler, routeContainerRequest } from \"@abhi-arya1/autoscaled\";\n\nexport class MyAutoscaler extends Autoscaler<Env> {\n    config = {\n        // See Instance Types for available options\n        // https://developers.cloudflare.com/containers/platform-details/limits/\n        instance: \"standard-1\",\n        maxInstances: 5,\n        minInstances: 1,\n    };\n\n    constructor(ctx: DurableObjectState, env: Env) {\n        super(ctx, env, env.MY_CONTAINER);\n    }\n}\n\nexport default {\n    // Set up your fetch handler to use configured server\n    // `env.AUTOSCALER` is the Wrangler binding to your Autoscaler class, such as MyAutoscaler above\n    async fetch(request: Request, env: Env): Promise<Response> {\n        return (\n            (await routeContainerRequest(request, env.AUTOSCALER)) ||\n            new Response(\"Not Found\", { status: 404 })\n        );\n    },\n} satisfies ExportedHandler<Env>;\n```\n\nAnd configure your `wrangler.toml`:\n\n```toml\nname = \"my-worker\"\nmain = \"index.ts\"\n\n[[durable_objects.bindings]]\nname = \"MyAutoscaler\"\nclass_name = \"MyAutoscaler\"\n\n[[migrations]]\ntag = \"v1\" # Should be unique for each entry\nnew_classes = [\"MyAutoscaler\"]\n```\n\nAny request to your worker will now be routed to the least loaded container through the autoscaler, with minimal latency impact.\n\n## Customizing `Autoscaler`\n\n`Autoscaler` is a class that extends `DurableObject`. You can override any of the `DurableObject` methods on `Autoscaler` to add custom behavior.\n\nYou can override the `config` property to customize the autoscaler's behavior.\n\n```ts\nexport interface AutoscalerConfig {\n    /**\n     * The instance type, to account for compute constraints of a container\n     * @default \"standard-1\"\n     */\n    instance: InstanceType;\n    /**\n     * The maximum number of containers that can run\n     * @default 10\n     */\n    maxInstances: number;\n    /**\n     * The minimum number of containers that should be running\n     * @default 0\n     */\n    minInstances?: number;\n    /**\n     * The maximum number of requests that a container can handle at once before it is considered overloaded\n     * This is optional and will be ignored if not provided.\n     */\n    maxRequestsPerInstance?: number;\n    /**\n     * The capacity threshold (0-1) at which to trigger optimistic scale-up\n     * For example, 0.7 means scale up when instances reach 70% of maxRequestsPerInstance\n     * This prevents race conditions by scaling up before hitting capacity\n     * @default 0.7\n     */\n    scaleUpCapacityThreshold?: number;\n    /**\n     * The interval at which the autoscaler will check for stale or unhealthy instances\n     * @default 30_000 (milliseconds)\n     */\n    heartbeatInterval?: number;\n    /**\n     * The threshold for considering an instance stale (i.e. last request was more than this threshold ago)\n     * @default 120_000 (milliseconds, i.e. 2 minutes)\n     */\n    staleThreshold?: number;\n    /**\n     * The specific CPU percentage threshold for scaling up a new instance\n     */\n    scaleThesholdCPU?: number;\n    /**\n     * The specific memory usage percentage threshold for scaling up a new instance\n     */\n    scaleThesholdMemoryMiB?: number;\n    /**\n     * The specific disk usage percentage threshold for scaling up a new instance\n     */\n    scaleThesholdDiskGB?: number;\n    /**\n     * The overall threshold for scaling up a new instance\n     * You can either use this to control all at once, or use the specific thresholds above to control each metric individually.\n     * @default 75 (Percent)\n     */\n    scaleThreshold?: number;\n    /**\n     * Milliseconds to wait after scale-up before scaling again\n     * @default 60_000 (1 minute)\n     */\n    scaleUpCooldown?: number;\n    /**\n     * Milliseconds to wait after scale-down before scaling again\n     * @default 120_000 (2 minutes)\n     */\n    scaleDownCooldown?: number;\n    /**\n     * General threshold for scaling down (defaults to scaleThreshold - 45% for hysteresis)\n     */\n    scaleDownThreshold?: number;\n    /**\n     * CPU percentage threshold for scaling down\n     */\n    scaleDownThresholdCPU?: number;\n    /**\n     * Memory usage percentage threshold for scaling down\n     */\n    scaleDownThresholdMemory?: number;\n    /**\n     * Disk usage percentage threshold for scaling down\n     */\n    scaleDownThresholdDisk?: number;\n    /**\n     * Number of consecutive health check failures before marking unhealthy\n     * @default 3\n     */\n    healthCheckRetries?: number;\n    /**\n     * Maximum time to wait for instance draining (milliseconds)\n     * @default 60_000 (1 minute)\n     */\n    drainTimeout?: number;\n    /**\n     * The endpoint to monitor the autoscaler's health\n     * @default \"/healthz\"\n     */\n    monitoringEndpoint?: string;\n    /**\n     * The endpoint to fetch container compute monitoring from\n     * @default \"http://localhost:81/monitorz\"\n     */\n    monitorzURL?: string;\n}\n```\n\n## Monitoring Containers\n\nYou might've seen the `monitorzURL` option in the config, or the fact that port 81 is required in the container class.\n\nSince Cloudflare Containers are still in beta, and do not have a dedicated metrics endpoint (that I could find), the [monitor](../monitor/) package implements a lightweight HTTP server in Go that can be used to monitor resources for scaling.\n\n> [!NOTE] Platform Targets\n> The monitor is by default built for `GOOS=linux GOARCH=amd64` in this package. Feel free to rebuild it for your own platform.\n\nIn your Container `Dockerfile`, you can copy the monitor and then\n\n```dockerfile\nCOPY monitor /monitor\n\nEXPOSE 81 # Add your ports to this list\n\n# Use monitor in exec mode: monitor runs on port 81 and execs the server\nCMD [\"/monitor\", ...] # your executable goes here\n```\n\nThe full documentation for the monitor is available [here](../monitor/README.md).\n\nNote that the configuration option is offered if you do not want to use the provided monitor, and instead your own endpoint. You can set the `monitorzURL` option to your own path in the container. This has to be of the form `GET http://<your-endpoint>/<your-path>`.\n\nIt must return a JSON object with the following fields for usage in the Autoscaler.\n\n```ts\nexport type MonitorzData = {\n    // Percentages on a 0-100 scale\n    cpu_usage: number;\n    memory_usage: number;\n    disk_usage: number;\n};\n```\n\n# How does it work?\n\nAutoscaleD is built as a Cloudflare Durable Object that acts as an intelligent load balancer and autoscaler for Cloudflare Containers. It maintains state about all running container instances and makes scaling decisions based on metrics, request load, and health status.\n\n> I was a little lazy with this write-up so everything below this point has been helped in some part by Claude Code.\n\n## Architecture\n\nThe Autoscaler consists of four core components that work together:\n\n- **`AutoscalerState`**: Manages persistent state in SQL (Durable Object storage) tracking all instances, their metrics, health status, and capacity limits\n- **`Router`**: Selects the best instance for each incoming request based on load balancing criteria\n- **`Scaler`**: Evaluates when to scale up or down based on configured thresholds and metrics\n- **`InstanceManager`**: Handles the container lifecycle (creation, destruction, health checks, metrics collection)\n\n## Request Flow\n\nWhen a request arrives at your worker:\n\n1. **Routing**: The `Router` selects the least-loaded healthy instance that isn't draining and has capacity available (compute and request based)\n2. **Health Check**: If the selected instance is unhealthy, the Autoscaler attempts to create a replacement instance or route to another healthy instance\n3. **Request Execution**: The request is forwarded to the selected container instance\n4. **Load Tracking**: Active request counts are incremented before the request and decremented after completion\n5. **Optimistic Scale-Up**: If an instance crosses its capacity thresholds, a new instance is created proactively in the background to prevent overload\n\n## Periodic Heartbeat (Alarm Handler)\n\nThe Autoscaler runs a periodic heartbeat (default: every 30 seconds) that performs several maintenance tasks:\n\n1. **Cleanup**: Removes stale instances that no longer exist in the container namespace\n2. **Metrics Collection**: Fetches CPU, memory, and disk usage from each instance's monitoring endpoint (`/monitorz`)\n3. **Health Checks**: Performs health checks on all instances and marks them unhealthy after consecutive failures\n4. **Scale-Up Evaluation**: Checks if any instance is crossing CPU/memory/disk thresholds (transitioning from below to above) and creates new instances if needed\n5. **Scale-Down Evaluation**: Checks if all instances are below scale-down thresholds and initiates draining for excess instances\n6. **Drain Processing**: Monitors draining instances and removes them once they have no active requests or the drain timeout expires\n\n## Scaling Mechanisms\n\n### Scale-Up Triggers\n\nThe Autoscaler can scale up based on two mechanisms:\n\n1. **Metrics-Based Scaling**: When any instance crosses configured CPU, memory, or disk thresholds (via `scaleThreshold` or specific thresholds like `scaleThesholdCPU`)\n    - Uses \"just crossed\" detection to prevent duplicate scale-ups\n    - Tracks each instance's `threshold_crossed_at` timestamp\n    - Only triggers scale-up if the instance hasn't crossed within the cooldown period (60s)\n    - Prevents duplicate instance placements from the same threshold breach\n\n2. **Request-Based Scaling**: When the average requests per instance exceeds `maxRequestsPerInstance` (if configured)\n    - Uses optimistic scaling with `scaleUpCapacityThreshold` (default: 70%)\n    - Scales proactively before hitting maximum capacity to prevent overload\n\nScale-ups respect:\n\n- `maxInstances` limit\n- `scaleUpCooldown` period (default: 60 seconds) to prevent rapid scaling\n\n### Scale-Down Process\n\nScale-down uses a hysteresis pattern to prevent flapping:\n\n1. **Evaluation**: Only scales down when ALL healthy instances are below the scale-down thresholds (typically 45% below scale-up thresholds)\n2. **Selection**: Chooses instances with the fewest active requests and oldest heartbeat times\n3. **Draining**: Marks selected instances as \"draining\" and stops routing new requests to them\n4. **Removal**: Waits for active requests to complete (up to `drainTimeout`, default: 60 seconds) before destroying the instance\n\nScale-downs respect:\n\n- `minInstances` limit\n- `scaleDownCooldown` period (default: 120 seconds)\n\n## Instance Lifecycle\n\n### Creation\n\nWhen a new instance is needed:\n\n1. A slot is reserved in the capacity tracking system (prevents exceeding `maxInstances`)\n2. A new container is created with a unique name (using `nanoid`)\n3. The container is started and waits for ports to be ready\n4. The instance is registered in the state database with initial metrics\n5. Health check is performed to verify the instance is ready\n\n### Health Monitoring\n\nEach instance is continuously monitored:\n\n- **Health Checks**: Performed periodically via the `monitoringEndpoint` (default: `/healthz`)\n- **Failure Tracking**: Instances are marked unhealthy after `healthCheckRetries` consecutive failures (default: 3)\n- **Metrics Collection**: CPU, memory, and disk usage are fetched from the `monitorzURL` endpoint\n- **Keep-Alive**: Healthy instances receive periodic keep-alive requests to prevent sleep\n\n### Removal\n\nInstances are removed when:\n\n1. They become stale (no longer exist in the container namespace)\n2. They're selected for scale-down and successfully drain\n3. They're replaced due to being unhealthy\n4. They exceed the drain timeout while draining\n\n## State Management\n\nThe Autoscaler uses SQL (via Durable Object storage) to maintain:\n\n- **Instance Records**: Name, creation time, active request count, current metrics (CPU/memory/disk), health status, draining status, and threshold crossing timestamps\n- **Capacity Tracking**: Current and maximum instance counts (prevents race conditions)\n- **Scaling State**: Timestamps of last scale-up and scale-down (for cooldown enforcement)\n- **Threshold Tracking**: Per-instance `threshold_crossed_at` timestamps to prevent duplicate scale-ups from compute metrics\n\nThis persistent state ensures the Autoscaler can recover from restarts and maintain consistency across concurrent operations.\n\n## Initialization\n\nWhen the Autoscaler Durable Object is first created or restarted:\n\n1. **Database Migration**: Creates necessary tables if they don't exist\n2. **Stale Cleanup**: Removes any instances from the database that no longer exist\n3. **Capacity Sync**: Synchronizes the capacity counter with actual instance count\n4. **Warm-Up**: Creates `minInstances` containers to ensure minimum capacity is available\n5. **Alarm Scheduling**: Schedules the first heartbeat alarm\n\nThis ensures the Autoscaler starts in a consistent state and maintains the configured minimum instances.\n\n# Limitations\n\nThis autoscaler is still not a native solution, just a wrapper around functionality that Cloudflare already provides. I know that one is slated on the roadmap, but I wanted to give a shot at making one a reality anyways.\n\nThanks for reading and let me know what you think!\n","readmeFilename":"README.md"}