TL;DR: Kubernetes 1.37's native HPA scale-to-zero eliminates a node-draining limitation but exposes a cold-start tax that can make GPU workloads cost more at zero than at one. The fix is not a better autoscaler; it is a tiered hot-cold architecture that matches the metric to the workload.
Key Takeaways: - Scale-to-zero in 1.37 is a Beta feature, but it requires an object or external metric to trigger a return from zero. - A GPU cold start runs through a pipeline of image pull, weight loading, and CUDA initialization, during which the GPU is billed but not serving traffic, so frequent scale-up events erase any idle savings. - The right architecture is hybrid: hot pods for user-facing inference, scale-to-zero pods for queues and batch jobs, never CPU as the driver.
The Month Your GPU Bill Climbed After You Scaled to Zero

Your CFO asked why the GPU bill went up the same month you scaled idle inference pods to zero. The answer is not a billing error. It is a cold-start tax that 1.37 made easier to pay.
Kubernetes 1.37 shipped on August 26, 2026, under the codename "Garhwal," and made HPA scale-to-zero a Beta feature. The intuition feels airtight.
Fewer pods mean lower cost, so zero pods should mean the lowest cost. But a GPU behind an idle warm pod bills at the same hourly rate whether it processes one request or one thousand. The "zero cost at zero pods" assumption is technically true and economically misleading.
What teams see in production is worse than flat spend. The same month they enabled scale-to-zero for long-tail inference models, their total GPU spend held or grew. The culprit is not the zero state. It is the cost of getting out of it, paid dozens of times per day.
If you are also wrestling with AI/ML training pipelines that already bleed on idle, the pattern looks familiar.
To understand why a feature designed to cut waste instead amplified it, you have to revisit the workaround it was supposed to replace.
The Shared-Node Blind Spot That Made Cluster-Autoscaler Useless
Before 1.37, the standard workaround for idle GPU cost was a cluster-autoscaler that removes entire nodes once every pod on them terminates. This only works when a node hosts nothing but the idle workload.
In real production clusters, low-traffic inference services share nodes with busier services. The cluster-autoscaler cannot drain a node that still has a healthy pod on it. The expensive GPU stays reserved, the bill stays high, and teams assume Kubernetes simply cannot do scale-to-zero well. For teams whose cloud security solutions posture also includes strict node-level isolation, this shared tenancy pattern makes the problem worse.
1.37's HPA-level scale-to-zero was built to fix exactly this. It scales individual workloads to zero without waiting for the node to drain.
The pod is gone, the GPU reservation is gone, and the cluster-autoscaler can finally reclaim the node. On paper, this is the cleanest cost lever Kubernetes has shipped in years.
But fixing the node problem created a new one. The moment a pod returns from zero, it has to rebuild everything a warm pod was quietly holding in memory.
What a GPU Cold Start Actually Costs in Wall-Clock Time
A cold GPU pod is not just "a pod starting." It is a multi-stage pipeline.
First, the kubelet pulls a multi-gigabyte CUDA base image. Then the container runtime starts, the model weights load from object storage, and CUDA allocates its context. After that, the KV cache fills.
For large transformer models, this pipeline runs long enough that the GPU sits billed but idle during every cold start. The exact window depends on image caching and model size.
During that window the GPU is billed but not serving traffic. The cost of one cold start can exceed the savings from the idle time it replaces.
If a workload cold-starts often, the billable time across cold starts can exceed the cost of leaving a warm pod running. Scale-to-zero becomes more expensive than running hot. This mirrors how AI/ML training becomes a budget black hole when teams optimize for the wrong phase.
Request-driven inference is the worst fit, because users wait behind a cold start. Queue consumers and batch processors are the best fit, because they can tolerate latency and the cold start hides behind a backlog. The workload shape, not the autoscaler, decides the economics.
Knowing the cold-start cost is necessary but not sufficient. The metric you choose to drive the HPA determines how often you pay that cost.
Inside 1.37's ScaledToZero: Why Your Metric Choice Is the Bill

1.37's HPA can scale a workload to zero, but only when configured with an object metric or external metric. CPU and memory metrics cannot trigger scale-from-zero.
The controller resolves the ambiguity of a zero replica count with a new `ScaledToZero=True` status condition. This tells reconciliation loops that the HPA owns the zero state and should keep evaluating the metric.
This sounds like a small detail. It is the entire bill. An external metric like a Prometheus gauge for requests per second gives precise control over when the pod returns. A poorly tuned threshold can trigger dozens of cold starts per hour on noisy traffic.
An object metric tied to a message queue length is more stable, but only works for async consumers. If your cloud security solutions stack already routes through a sidecar that adds metric jitter, expect even more thrash.
The metric choice is the dial that sets how often your team pays the cold-start tax. Most teams pick it without modeling the cost. The teams that get this right treat metric selection as a financial decision, not a YAML decision.
That modeling is the next step. It produces a decision matrix that looks very different from the "always scale to zero" advice in most tutorials.
The Decision Matrix: When Scale-to-Zero Saves Money and When It Bleeds
Workloads that win at scale-to-zero share three traits. They sit idle for long stretches, and they tolerate multi-minute startup latency. They also have no human or SLA-bound caller waiting on the first request.
Queue consumers, async enrichment, and batch processors fit the profile. Workloads that lose sit on the opposite side. Request-driven inference for high-traffic or latency-sensitive LLMs falls here, because the cold start lands inside the user-visible response window.
The breakeven formula is simple. `cold_start_cost x scale_up_frequency` must be less than `hourly_gpu_rate x idle_hours_saved`.
A model that serves only a handful of requests per day will lose money at scale-to-zero. One cold start costs more than the idle time it saves. A model that serves frequent short bursts will also lose money, because the pods keep scaling down between bursts and the cold-start tax compounds.
The sweet spot is a workload that sits idle for hours at a time and tolerates a multi-minute cold start when it wakes. For everything else, AI/ML training platforms that already run hot need a different lever.
If your traffic shape does not match the sweet spot, scale-to-zero is not a cost optimization. It is a cost transfer from idle time to cold-start time.
Once you know which workloads belong in which column, the architecture writes itself. It is rarely "all scale-to-zero" or "all run hot."
A Hybrid Hot-Cold Architecture That Pays for Itself in 30 Days
The pattern that works has two tiers. Tier 1 holds hot pods with `minReplicas` set to at least 1 for steady-traffic models and any inference endpoint serving user-facing requests. Tier 2 holds HPA-managed scale-to-zero pods for queue-driven batch jobs, async enrichment pipelines, and rarely-called models that can tolerate cold starts.
1apiVersion: autoscaling/v22kind: HorizontalPodAutoscaler3metadata:4 name: long-tail-inference5spec:6 scaleTargetRef:7 apiVersion: apps/v18 kind: Deployment9 name: summarizer-v210 minReplicas: 011 maxReplicas: 412 metrics: - type: External13 external:14 metric:15 name: http_requests_per_second16 selector:17 matchLabels:18 model: summarizer-v219 target:20 type: AverageValue21 averageValue: "0.1"22 behavior:23 scaleUp:24 stabilizationWindowSeconds: 6025 scaleDown:26 stabilizationWindowSeconds: 300
Three rules separate teams whose GPU bills drop from teams whose bills plateau.
First, use an external metric for HTTP-driven scale-from-zero and an object metric for queue depth. Never use CPU. Second, set a stabilization window of at least 5 minutes to absorb burst noise.
Third, pin CUDA base images on the node or use a node image cache. Image pull adds a major share of cold start time for most teams. If your cloud security solutions stack includes image signing, the cache must respect the signature policy or you reintroduce supply-chain risk.
This tiered pattern is the difference between a 1.37 upgrade that saves money and one that quietly inflates it.
The architectural shift is small. The financial impact is not. The teams that get it right see a different category of outcome.
What Changes When You Get the Scale-to-Zero Economics Right
GPU spend becomes predictable. The cold-start tax is bounded by the architecture rather than the traffic pattern. Infrastructure teams stop fighting thrash and start reasoning about cost per request rather than cost per pod.
The `ScaledToZero` condition gives you a single status field to alert on. That is a much cleaner signal than inferring idle state from replica counts.
The net effect is not a percentage saving. It is the elimination of an entire failure mode where scale-to-zero silently costs more than running hot.
Teams that build their cloud security solutions telemetry to read the `ScaledToZero` condition can catch misconfigured scale-up thresholds before they become budget events.
For a FinOps view, see Why LLM Auto-Scaling Is Bleeding Money & Breaking Compliance. It covers the same metric problem from a different angle.
If you run RAG at scale, the Kubernetes cost playbook for AI workloads shows how scale-to-zero fits into a broader idle-cost strategy.
Below are the questions CTOs ask most often when auditing their own scale-to-zero setup.
Frequently Asked Questions
Does Kubernetes 1.37 support scale-to-zero natively?
Yes. Kubernetes 1.37, shipped August 26, 2026, moved HPA scale-to-zero to Beta. A HorizontalPodAutoscaler configured with an object or external metric can now scale a workload to zero replicas and back. It needs no KEDA, Knative, or alpha feature gate.
Why does scale-to-zero cost more than running hot for some GPU workloads?
The cold-start pipeline includes image pull, model weight load, CUDA init, and KV cache allocation. During this, the GPU stays billed but idle.
If a workload cold-starts often, the cumulative billable GPU time across cold starts can exceed the cost of leaving a warm pod running.
Can HPA scale to zero using CPU or memory metrics in Kubernetes 1.37?
No. The scale-from-zero path requires an object metric or external metric. CPU and memory metrics can keep a workload scaled up, but they cannot trigger a return from zero. That is why 1.37 ships guidance to point the HPA at queue depth or HTTP request rate.
What is the ScaledToZero status condition?
It is a new status condition on the HorizontalPodAutoscaler. The controller sets it to True when it has scaled a workload to zero.
It clears up the ambiguity between "HPA scaled to zero" and "an operator paused the workload." It also tells reconciliation loops to keep evaluating the configured metric.
When should I use KEDA instead of native 1.37 scale-to-zero?
Use KEDA when you need scalers beyond what core HPA supports. Examples include Kafka consumer lag, CloudWatch metrics, cron schedules, and scaler chaining. For straightforward external-metric-driven scaling on Prometheus or a queue, native 1.37 HPA removes the dependency and the version-skew risk.
For a cross-check on either path, the KEDA cost traps breakdown walks through the same cold-start pattern from a different toolchain.
About the author
Mayank Singh is a software developer at Levitation Infotech, where he builds web and AI-powered applications across the company’s fintech, healthcare, and enterprise projects.
