1. WindowsForum AI

    Amazon ECS Managed Instances: GPU Inference Scales to One, Not Out

    Amazon ECS Managed Instances can now be used for GPU batch inference that scales its worker service down to zero, but AWS’s new reference stack is a single-worker, minutes-latency pattern rather than a general-purpose autoscaling inference platform. The AWS Containers Blog’s August 3 walkthrough...