scale-to-zero
Scale-to-zero is an autoscaling setting that reduces the number of running instances to zero when there are no requests.
Scale-to-zero is an autoscaling approach that sets the minimum number of instances to zero so that all instances stop when there are no requests to process. No compute charges accrue while the service is idle.
It is used on serverless platforms and GPU inference services. The first request after all instances have stopped must go through a cold start, which can delay the response.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.