Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Should you keep a microservice always on or let it scale to zero? Use scale-to-zero when traffic is intermittent and requests can tolerate startup delay; keep some capacity ready when a slow first response would materially hurt users. Warm capacity costs more, and it reduces rather than guarantees away latency. The right choice depends on your service’s traffic, startup work, latency target, and billing configuration.

What a cold start means

A cold start is the work required to prepare a new execution environment or container before it can handle a request. Depending on the platform and application, that can include provisioning runtime capacity, loading code and dependencies, and initializing connections or other resources.

When a service has scaled to zero, a new request may trigger that work and wait for it to finish. Later requests may reuse a ready environment, but the duration and frequency of cold starts vary with the service and its platform. They are not a fixed delay you can assume for every request.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scale to zero or keep capacity ready?

Operating choice Potential benefit Trade-off Often suits
Scale to zero Can reduce charges for idle runtime capacity when the chosen platform and billing mode stop charging for it. A request arriving after the service has scaled down may wait for provisioning and initialization. Intermittent or low-frequency traffic where startup delay is acceptable.
Keep minimum or provisioned capacity ready Can reduce initialization-related delay for requests that fit within the ready capacity. Ready capacity incurs charges, and bursts beyond it or other runtime work can still affect latency. Interactive services with a meaningful first-response latency objective.

“Always on” is shorthand, not a single cloud setting. Providers expose different controls, with different scaling and billing behavior. For example, Google Cloud Run minimum instances, AWS Lambda provisioned concurrency, and Azure Functions hosting plans do not operate identically.

How the major platforms handle warm capacity

Google Cloud Run

Cloud Run scales instances in response to incoming load. Google documents minimum instances as a way to keep instances available and reduce latency, including when a service would otherwise scale from zero. Google describes the choice as a trade-off between cold-start latency and pending-request latency; keeping minimum instances available incurs charges. The actual billing depends on whether the service uses request-based or instance-based billing, so there is no single idle price that applies to every Cloud Run service. See Cloud Run minimum instances, instance autoscaling, and What is Cloud Run.

For functions on Cloud Run, Google recommends minimum instances for latency-sensitive workloads and notes that load-time initialization affects startup latency. Keep initialization focused on what the first request actually needs. See Google’s functions best practices.

AWS Lambda

AWS provisioned concurrency pre-initializes execution environments to reduce cold-start latency and has additional charges. AWS describes it as useful for reducing cold-start latency and designed to make functions available with double-digit millisecond response times; that is a design intent, not a latency SLA.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not confuse provisioned concurrency with reserved concurrency. Reserved concurrency sets a concurrency limit and reserves capacity, but does not pre-initialize environments. AWS says asynchronous workloads often have less need for provisioned concurrency than interactive workloads.

AWS’s Lambda execution-environment lifecycle documentation says cold starts typically occur in under 1% of invocations and that their duration ranges from under 100 ms to over 1 second. These are AWS’s general statements, not a guarantee for a particular function or a benchmark that applies to Cloud Run, Azure Functions, or all workloads.

Microsoft Azure Functions

Azure Functions behavior depends on the hosting plan. The Consumption plan can scale to zero, which can mean startup latency when an invocation arrives. The Premium plan supports always-ready instances, while the Dedicated plan can run continuously on prescribed instances. Compare the relevant plan’s behavior rather than treating Azure Functions as having one universal cold-start mode. Details are in Microsoft’s Scale and Hosting documentation.

How to choose for your service

Decide from observed workload behavior rather than the label “always on.” Use these checks to compare the latency benefit with the cost of ready capacity:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Set a latency objective. Decide what first-request and tail latency your users can tolerate. A warm-capacity setting is useful only if it helps meet the objective that matters to the service.
  • Examine traffic shape. Note how often requests arrive, how long idle gaps last, how bursty traffic is, and how much concurrency the service needs. A small ready pool may help steady traffic but may not cover a burst.
  • Measure startup work. Identify what happens before the first request can be served. Dependency loading and connection setup can extend startup, so avoid doing nonessential initialization on the critical path.
  • Estimate ready capacity needs. Determine how much capacity must be available to cover the latency-sensitive traffic. Requests beyond that configured capacity may still encounter scaling or other runtime delays.
  • Check the actual billing mode. Compare idle and active charges for the provider, region, plan, and configuration you use. Scaling to zero does not necessarily mean every related charge disappears, and warm-capacity charges differ by service and billing mode.

For an intermittent workload that can tolerate a slower first response, try scale-to-zero and verify both the resulting latency and whether the selected billing mode reduces idle charges. For an interactive service where the first response matters, consider a minimum or provisioned capacity setting sized to observed demand. In either case, compare measured latency percentiles and total spend for the actual workload, region, concurrency, and plan; an exact cost cannot be inferred without those inputs.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What warm capacity does—and does not—solve

Keeping capacity ready can reduce delay caused by initializing an environment, but it does not make every request instantaneous. Traffic can exceed the ready capacity, and other runtime effects can still add latency. Likewise, scale-to-zero does not mean every request will be cold: the delay concerns requests that arrive when new capacity must be started.

There is no universal winner. Choose based on the latency objective, traffic frequency and burstiness, startup work, concurrency, minimum capacity required, and the specific platform’s idle and active billing behavior.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.