Kubernetes In-Place Pod Resize in Production: What GA Actually Changed

In-place pod resize is GA and VPA can drive it without evicting pods. The savings are real, but five documented limits decide which of your workloads are silently skipped.

By VVV Ops ·

Your right-sizing job still works by killing pods. Something watches actual usage, decides a container asked for four times the CPU it uses, and then evicts it so the new numbers take effect on a fresh pod. Everyone accepts the restart as the price of reclaiming waste. That price disappeared. Kubernetes in-place pod resize in production is now a GA feature, on by default since the v1.33 beta, and the Vertical Pod Autoscaler can drive it without evicting anything. The catch is that a resize can be accepted, deferred, or quietly skipped, and most teams wire it up without learning the difference.

The gap between generally available and safe to enable

In-place pod resize graduated to stable in Kubernetes v1.35 on 19 December 2025, after alpha in v1.27 and beta in v1.33. The InPlacePodVerticalScaling feature gate is locked on, so there is nothing to enable on a current cluster. Amazon EKS added 1.35 on 27 January 2026 and 1.36 in June, so if you are inside standard support you already have this.

That is the part everyone reports. What matters operationally is that GA describes the API contract, not your workloads. The kubelet will accept a patch to a pod's resize subresource and then reconcile toward it on a best-effort basis. Your job is to know which of your pods are structurally ineligible, because those pods will sit in a pending condition rather than fail loudly, and a right-sizing loop that assumes success will report savings it never made.

One prerequisite is easy to miss: your kubectl client must be at least v1.32 for the --subresource=resize flag. A CI runner pinned to an older client will fail in a way that reads like a permissions problem.

Five limits that decide whether a resize lands

These come from the upstream limitations list, and they are the first thing we check on a client cluster.

| Limit | What it blocks | Who it hits | |---|---|---| | QoS class is fixed at creation | Guaranteed pods must keep requests equal to limits; Burstable pods cannot have requests meet limits for CPU and memory at once; BestEffort pods cannot gain requests or limits at all | Anyone whose recommender emits requests without matching limits | | Memory decrease is best-effort | With a NotRequired memory policy, if usage already exceeds the target limit the resize is skipped and stays in progress | Java and Node services that hold a large heap | | Static CPU or memory manager policy | Pods on nodes using those policies cannot be resized in place at all | Latency-sensitive and NUMA-pinned workloads | | Swap | Memory requests cannot be resized unless the memory resizePolicy is RestartContainer | Clusters that enabled node swap for burst headroom | | Container type | Non-restartable init containers and ephemeral containers cannot be resized; sidecars can | Service meshes and log shippers, in the good way |

Two more are worth knowing. Only CPU and memory can be resized, so GPU requests are untouched and stay a placement decision rather than a tuning one, which is why we handle them separately in our GPU chargeback playbook. And once a request or limit is set it cannot be removed, only changed to another value.

The QoS rule catches more teams than any other. If your workloads are Guaranteed because a platform policy requires requests to equal limits, every resize has to move both numbers together or it is rejected. Write that into the recommender's output format before you turn anything on, not after your first batch of pending conditions.

Reading resize status instead of trusting the patch

A resize request has three outcomes, and the pod tells you which one you got. The kubelet sets PodResizePending when it cannot grant the request immediately, with reason: Infeasible when the node cannot ever satisfy it (more resources than the node has) and reason: Deferred when it might later, once something else frees up. PodResizeInProgress means the request was accepted and allocated but the change is still being applied, with reason: Error and a message if actuation fails.

Deferred requests are not dead. The kubelet retries them periodically, for example when another pod leaves the node, and when several are waiting it retries by PriorityClass first, then Guaranteed before Burstable, then whichever has waited longest. Having kube-scheduler preempt lower-priority pods to make room for a deferred resize is a separate alpha gate (InPlacePodVerticalScalingSchedulerPreemption, v1.37, off by default), so do not count on it.

Here is the loop we hand to platform teams, using the real example shape from the upstream task doc:

# Ask for more CPU on a running pod
kubectl patch pod resize-demo -n qos-example --subresource resize --patch \
  '{"spec":{"containers":[{"name":"pause", "resources":{"requests":{"cpu":"800m"}, "limits":{"cpu":"800m"}}}]}}'

# Did it actually land, or is it pending?
kubectl get pod resize-demo -n qos-example \
  -o jsonpath='{range .status.conditions[?(@.type=="PodResizePending")]}{.reason}{": "}{.message}{"\n"}{end}'

# What is the container running with right now?
kubectl get pod resize-demo -n qos-example \
  -o jsonpath='{.status.containerStatuses[0].resources}{"\n"}{.status.containerStatuses[0].restartCount}{"\n"}'

Check status.containerStatuses[*].resources for what the container is configured with now, not the spec. On an infeasible resize the status still shows the previous values and restartCount does not change, so a naive diff of spec against reality is how you end up believing a resize succeeded. There is also status.containerStatuses[*].allocatedResources, which holds values the kubelet has confirmed for internal scheduling; upstream says to focus on resources for monitoring and validation, and we agree. Do not build dashboards on allocatedResources.

Restart behaviour is per resource, set in the container spec:

    resizePolicy:
    - resourceName: cpu
      restartPolicy: NotRequired
    - resourceName: memory
      restartPolicy: RestartContainer

NotRequired is the default and applies the change to the running container. RestartContainer restarts it, which upstream notes is often necessary for memory because many runtimes cannot adjust their allocation while running. Our default for JVM and Node services is exactly the split above: resize CPU live, restart on memory. For Go and Rust services we use NotRequired on both. One constraint to respect: if a pod's overall restartPolicy is Never, every container resizePolicy must be NotRequired.

Wiring VPA to resize instead of evict

Patching pods by hand is a demo. This matters because VPA can now do it, and the upstream feature doc records InPlaceOrRecreate as alpha in VPA 1.4.0, beta in 1.5.0, and GA in 1.6.0. The installation doc puts the current default at VPA 1.8.0 with release branches maintained for 1.8, 1.7 and 1.6, and the compatibility table pairs 1.8.x with Kubernetes 1.36 to 1.38, 1.7.x with 1.35 to 1.37, and 1.6.x with 1.34 to 1.36. Every supported combination has in-place updates at GA.

apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: my-vpa
spec:
  updatePolicy:
    updateMode: "InPlaceOrRecreate"

InPlaceOrRecreate tries the resize first and falls back to recreating the pod when it cannot. Upstream lists the fallback triggers, and they are worth reading as a list of things that will surprise you: the resize is infeasible, it stays deferred for more than five minutes, it stays in progress for more than an hour, the update would change the pod's QoS class, or memory limit downscaling is needed with a no-restart policy. So this mode does not eliminate evictions. It reduces them, and it moves the remaining ones into cases you can enumerate.

If you genuinely cannot tolerate eviction, VPA 1.7.0 added an InPlace mode, documented as alpha, which never falls back and instead defers and retries. We would not put an alpha update mode in front of production traffic yet, but it is the right thing to pilot on a stateful workload where eviction is the whole problem.

Two operational details. VPA updates all containers in a pod together, so a single ineligible sidecar drags the whole pod into fallback. And by default VPA respects disruption budgets even for in-place updates; the --in-place-skip-disruption-budget flag, off by default, lets it skip those checks when every container has a NotRequired policy for both CPU and memory. Turn that on only after you have confirmed the policies, because with any RestartContainer policy present the budgets are still enforced and you will have changed nothing.

For visibility, VPA exposes counters worth scraping on day one:

vpa_updater_in_place_updatable_pods_total
vpa_updater_in_place_updated_pods_total
vpa_updater_failed_in_place_update_attempts_total

The ratio of the second to the first is your real eligibility rate. If it sits below half, you have a QoS or policy problem, not a VPA problem.

What the reclaimed requests are actually worth

The savings do not come from resizing. They come from lowering requests, which lets the scheduler pack more pods per node, which lets the cluster autoscaler remove nodes. In-place resize only changes how much disruption you accept to get there, and that matters because disruption is the reason most teams run their right-sizing loop monthly instead of daily.

The arithmetic, with inputs you should replace with your own:

  • 200 application pods, each requesting 1 vCPU, actually using 0.35 vCPU at p95
  • Right-sizing to 0.5 vCPU requests reclaims 100 vCPU of requests
  • At 8 allocatable vCPU per node, that is about 12 nodes of capacity the autoscaler can release
  • At an assumed $0.14 per node-hour and 730 hours a month, 12 nodes is about $1,230 per month

Substitute your own node price from your bill. The point is the shape. Reclaiming half the requested CPU on a 200-pod namespace frees double-digit node counts, and the loop that used to run monthly because restarts were expensive can now run weekly or nightly. In our own engagements the second-order win is larger than the first. Teams stop padding requests defensively once they trust that a wrong number gets corrected without a restart.

For where else the money hides in a Kubernetes bill, our cost optimization framework covers the other three levers.

Vertical or horizontal: how we choose

These are different tools, and the usual mistake is using one where the other belongs. We wrote up the horizontal and scale-to-zero side in Kubernetes HPA scale to zero; this is the complement, not a replacement.

| Situation | Reach for | Why | |---|---|---| | Stateless web service, traffic varies hourly | HPA | Add replicas; per-pod sizing barely moves the bill | | Singleton or leader-elected service | VPA in-place | Cannot add replicas, so the only lever is the pod's own size | | Long-running batch or stream consumer | VPA in-place | Restarts lose in-flight state and replay costs real money | | Requests padded 3x across a whole namespace | VPA in-place, recommend-only first | The waste is in the numbers, not the replica count | | Workload idle for hours at a time | HPA with an external metric, or KEDA | Vertical scaling has no path to zero | | NUMA-pinned or static CPU manager pods | Neither, size by hand | In-place resize is blocked outright |

Do not run HPA on CPU and VPA on CPU for the same workload. They fight, because VPA moves the requests that HPA's utilisation target is measured against. Split by resource or scope VPA to memory only.

What we would enable first

Start in updateMode: "Off" for two weeks and treat VPA purely as a recommender. Export the recommendations, compare them with your p95 usage, and find the workloads where the recommendation would violate the QoS rule. That list is your real work.

Then pick one namespace that is not customer-facing, set resizePolicy explicitly on every container rather than relying on the default, and move it to InPlaceOrRecreate. Watch vpa_updater_failed_in_place_update_attempts_total and the pending conditions for a week. Widen the blast radius only after that ratio looks sane, and consider --in-place-skip-disruption-budget only after that.

Plan the next step now rather than rediscovering it later. Pod-level in-place resize reached beta in v1.36 on 30 April 2026 behind the InPlacePodLevelResourcesVerticalScaling gate, enabled by default, which lets you resize a pod's aggregate .spec.resources budget instead of each container. For multi-container pods where the sidecar and the app trade headroom, that is the model you actually want, and it is worth designing toward even while you deploy container-level resize today.

When to get help

Turning this on is a two-day change. Knowing which of your workloads it will silently skip, and rebuilding a right-sizing loop that reads resize conditions instead of assuming success, takes longer and is where the savings live. If you are running a cluster where requests are padded across the board and nobody wants to be the one who evicts the stateful service, we can do the eligibility audit and hand you the rollout.

Talk to us about your cluster.

Tags: kubernetes in-place pod resize in production, vertical pod autoscaler in place updates, kubernetes right sizing without restarts, kubernetes cost optimization, kubernetes consulting, secure kubernetes deployments