scale-to-zero

Scale-to-zero is an autoscaling setting that reduces the number of running instances to zero when there are no requests.

1 article
Last mentioned

Scale-to-zero is an autoscaling approach that sets the minimum number of instances to zero so that all instances stop when there are no requests to process. No compute charges accrue while the service is idle.

It is used on serverless platforms and GPU inference services. The first request after all instances have stopped must go through a cold start, which can delay the response.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

Zero as the minimum lets the service scale to zero—turn off GPU workers when they are no longer needed—while five caps the number that can run at once.


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.