Kubernetes HPA Scale to Zero: What v1.37 Changed and When KEDA Still Wins
Kubernetes v1.37 made HPA scale to zero beta and on by default, but only object and external metrics can wake a workload. Here is the wake-up budget, the cost math, and the KEDA decision.
By VVV Ops ·
A batch cluster runs forty queue consumers around the clock and does six hours of real work a day. The other eighteen hours are paid idle, and on GPU nodes that idle is the single largest line on the invoice. Kubernetes HPA scale to zero closes that gap without an add-on: minReplicas: 0 now works in a stock cluster. The feature is narrower than the headline suggests, and the release it landed in is easy to get wrong. This post covers what actually shipped, the wake-up latency you are signing up for, and how we decide whether KEDA still earns its place in the cluster.
What v1.37 actually shipped
Kubernetes v1.37 "Garhwal" was released on 26 August 2026 with 67 enhancements: 16 stable, 23 beta, 27 alpha, and one deprecation. Scaling to zero is one of the beta graduations, and the project's explainer, published on 2 September 2026 by Johannes Würbach, states it plainly: "This feature is now Beta and enabled by default."
That last clause is the part worth reading twice. No feature gate to flip, no controller to install. If you upgrade to v1.37, any HorizontalPodAutoscaler in your cluster can be set to minReplicas: 0 today.
If you see this attributed to v1.36, that is wrong. The alpha HPAScaleToZero gate has existed for years, which is where the confusion comes from, and the HPA API reference still carries the older wording: "minReplicas is allowed to be 0 if the alpha feature gate HPAScaleToZero is enabled and at least one Object or External metric is configured." The second half of that sentence is the constraint that survived graduation, and it is the whole story.
Why CPU and memory cannot wake a sleeping workload
Almost every HPA in production scales on CPU, and not one of them can scale to zero. The reason is structural rather than a missing feature. The Kubernetes explainer puts it this way: "Once the replica count reaches zero, there are no Pods left to measure and no signal that can tell the HPA to scale back up."
Resource metrics are produced by the pods being scaled. Take the pods away and the signal goes with them. Object and external metrics come from somewhere else, so a queue depth or a consumer lag keeps reporting while zero workers are running.
Practically, this means scale to zero is a feature for workloads that pull. Queue consumers, batch processors, scheduled ETL, model inference behind a job queue. If your workload is an HTTP service that scales on CPU, this release changed nothing for you, and forcing it is a bad trade for reasons covered further down.
Wiring an external metric the HPA can read at zero
You need three pieces: a metric that exists independently of the workload, an adapter that exposes it through the External Metrics API, and an HPA that targets it.
The Kubernetes example uses a Prometheus series called queue_consumer_lag. Exposing it through the Prometheus Adapter takes an externalRules entry:
# prometheus-adapter values.yaml
externalRules:
- seriesQuery: '{__name__="queue_consumer_lag",name!=""}'
metricsQuery: sum(<<.Series>>{<<.LabelMatchers>>}) by (name)
resources:
overrides: { namespace: {resource: "namespace"} }
Verify the metric is actually reachable before you touch the HPA. This is the step teams skip, and it is the one that turns a rollout into an incident:
kubectl get --raw \
"/apis/external.metrics.k8s.io/v1beta1/namespaces/default/queue_consumer_lag" | jq .
If that returns an empty item list, stop. An HPA pointed at a metric the API cannot serve will not scale, and once the workload is at zero it will stay there.
The HPA itself is unremarkable, which is the point:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: worker-tasks
namespace: default
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: worker-tasks
minReplicas: 0
maxReplicas: 40
metrics:
- type: External
external:
metric:
name: queue_consumer_lag
selector:
matchLabels:
name: worker_tasks
target:
type: AverageValue
averageValue: "100"
The wake-up budget you are actually signing up for
"Cold start" gets waved at as a single number. It is four numbers, and three of them are yours to control. The Kubernetes docs give the control loop defaults: the interval "is set by the --horizontal-pod-autoscaler-sync-period parameter to the kube-controller-manager (and the default interval is 15 seconds)", the scale-up stabilization window defaults to 0 seconds, and scale-up may add 100% of current replicas or 4 pods per 15-second period, whichever is larger.
| Stage | Where the delay comes from | Default or typical | How to shrink it | |---|---|---|---| | Metric freshness | Prometheus scrape interval plus adapter cache | Your scrape interval, commonly 30s | Scrape the queue exporter on a shorter interval than the rest of the fleet | | HPA observation | --horizontal-pod-autoscaler-sync-period | 15s | Rarely worth tuning cluster-wide | | Scale-up decision | behavior.scaleUp.stabilizationWindowSeconds | 0s | Already immediate; leave it | | Pod start | Scheduling, node provisioning, image pull, app boot | Seconds to minutes | Pre-pull images, keep a warm node, cut init containers |
Note the first row against the fourth. If a node has to be provisioned because the last pod on it went away, the pod start stage dominates everything above it. That interaction between pod autoscaling and node autoscaling is the part teams get wrong, and we covered the node side of it in AWS Auto Scaling with Terraform.
There is a fifth number nobody budgets for: behavior.scaleDown.stabilizationWindowSeconds defaults to 300 seconds. The last replica sticks around for five minutes after the metric drops. A workload whose idle gaps are four minutes long will never reach zero at all, and you will have added complexity for no saving.
What the idle hours are worth
Run the arithmetic before the migration, not after. A 30-day month is 720 hours. Reclaimed idle is (24 - busy hours per day) x 30 x replicas x node hourly rate.
| Workload profile | Busy hours per day | Idle hours per month | Share of the month reclaimed | |---|---|---|---| | Nightly batch | 2 | 660 | 92% | | Business-hours ETL | 8 | 480 | 67% | | Bursty queue consumer | 14 | 300 | 42% |
We are deliberately not printing a dollar figure. GPU and CPU on-demand rates change, and the only number worth taking to a CFO is the one from your own bill and the vendor's current pricing page. Multiply the idle hours by that rate yourself.
Two caveats that decide whether the saving is real. The pods must actually free a node, or you have removed nothing from the invoice. And reserved capacity or savings plans keep billing whether the pods run or not, so scale to zero pays best on on-demand and spot. We work through both in Kubernetes Cost Optimization.
Native HPA or KEDA: how we decide
KEDA has done this job since long before v1.37, and the honest comparison is narrower than "core beats add-on". Here are the documented ScaledObject defaults for KEDA 2.17 against what the platform now gives you.
| Concern | Native HPA in v1.37 | KEDA 2.17 | |---|---|---| | Wake-up signal | Object or external metric only | 70-plus scalers, including SQS, Kafka and Azure Service Bus, with no adapter rule to write | | Floor | minReplicas: 0 | minReplicaCount, default 0, plus idleReplicaCount, which supports only 0 | | Poll interval | 15s controller sync | pollingInterval, default 30 seconds | | Scale-down delay | stabilizationWindowSeconds, default 300s | cooldownPeriod, default 300 seconds | | Metric source failure | Scaling stops; workload stays where it is | fallback with failureThreshold and replicas | | HTTP traffic | No buffering | Requires the separate HTTP add-on | | Operational surface | None beyond a metrics adapter | A controller, a metrics server and CRDs to upgrade |
Our rule: if the signal is already a Prometheus series and you already run the Prometheus Adapter, drop KEDA and use the native HPA. You are maintaining a controller to duplicate a field. If your triggers are cloud queues, drop the adapter and keep KEDA, because writing and maintaining an externalRules entry per SQS queue is worse than running one more controller.
The row that decides genuinely hard cases is metric source failure. The API reference says scaling "is active as long as at least one metric value is available", which means an adapter outage leaves a zeroed workload at zero while the queue fills behind it. Native Kubernetes has no answer to that, and KEDA's fallback block does. If your queue backs up into an SLA breach, that alone is worth the controller.
Where scale to zero costs more than it saves
Three failure modes, in the order we see them.
Request-driven services are the first and the worst. The Kubernetes explainer is explicit: "Kubernetes Services do not buffer requests while no Pods are ready, so HTTP and other request-driven workloads need a separate buffering layer." A zeroed HTTP deployment does not return a slow response. It returns a connection error. Do not scale a user-facing service to zero without an activator proxy in front of it.
Second, monitoring that counts pods. Every dashboard and alert that treats "replicas = 0" as an outage will now fire on healthy behaviour. Move the alert to queue age or oldest unacked message before you change the deployment, not after the first 3am page.
Third, the adapter becomes a tier-one dependency. Your metrics pipeline used to be observability. Once a workload's ability to start depends on it, it is production infrastructure and needs the same review as the rest of the control plane. Our Kubernetes production readiness checklist is a reasonable place to start on that.
A rollout that does not page anyone
We run this in five steps and it takes about a week per workload.
- Pick one queue consumer with a durable queue and no request path. Never start with the busiest one.
- Expose the queue metric and confirm
kubectl get --rawon the external metrics API returns a value. Watch it for a full business cycle to catch gaps in the scrape. - Point an HPA at the external metric with
minReplicas: 1. Confirm it scales sensibly between 1 and N before it is allowed anywhere near 0. - Rewrite the alerts. Queue age and oldest unacked message replace pod count. Add an alert for the metric itself going stale, because that is the failure that leaves you at zero.
- Set
minReplicas: 0, then measure real wake-up latency against the budget table above. If it exceeds what the queue's consumers can tolerate, go back to 1 and keep the money.
Step 4 is the one that gets cut for time, and it is the one that turns a cost win into an incident.
When to Get Help
Scale to zero is a small field with a large blast radius. The work is not writing the YAML, it is knowing which workloads can tolerate a cold start, whether your node autoscaler will actually return the capacity, and what breaks when the metrics pipeline becomes load-bearing. Getting that wrong costs more than the idle capacity ever did.
We do this work on client clusters: audit which workloads qualify, model the saving against your actual bill, and run the migration with the alerting rewritten first. If you are upgrading to v1.37 this quarter and want the cost win without the 3am surprise, get in touch.