Kube-proxy IPVS to nftables Migration: The 2026 Playbook
Kubernetes has published the releases where kube-proxy IPVS mode stops working. Here is the pre-flight check, the NodePort default that silently breaks callers, and the node-by-node cutover we run for clients.
By VVV Ops ·
If your clusters run kube-proxy in ipvs mode, you now have a removal date rather than a recommendation. Kubernetes deprecated the mode in v1.35, added the gate that will switch it off in v1.37, and has published the two releases where it stops working: v1.40 disables it by default, v1.43 deletes the code. A kube-proxy IPVS to nftables migration is not hard, but the default that changes underneath you is not the one the release notes talk about, and we have watched it take NodePort traffic down on the first node. This is the sequencing, the pre-flight checks, and the cutover we run for clients.
The deprecation clock in real version numbers
KEP-5495 lays out five stages. The Kubernetes documentation states it plainly: "Support for ipvs mode will be disabled by default from Kubernetes v1.40 (you can re-enable it with the KubeProxyIPVS feature gate); ipvs mode will be fully removed in Kubernetes v1.43."
| Release | What happens | Date | |---|---|---| | v1.35 | Mode marked deprecated, kube-proxy logs a warning on start | Released 17 Dec 2025 | | v1.37 | KubeProxyIPVS feature gate added, Default: true | Released 26 Aug 2026 | | v1.40 | Gate flips to Default: false. Without an override, kube-proxy exits with an error listing the valid modes | Projected, not announced | | v1.43 | Gate locked, pkg/proxy/ipvs removed from the tree | Projected, not announced | | v1.46 | Feature gate removed | Projected, not announced |
Kubernetes shipped v1.35 in December 2025, v1.36 in April 2026, and v1.37 in August 2026, which is three minor releases a year. On that cadence v1.40 lands around August 2027. Treat that as arithmetic, not as a date the release team has committed to.
A year away sounds comfortable. It is not, because the project maintains only the three most recent minor branches and gives each roughly a year of patch support. A cluster sitting on v1.37 today runs out of patches in October 2027, and by then v1.40 is the current release. Teams that upgrade once a year will meet the flipped default and the control plane upgrade in the same maintenance window, which is the combination to avoid.
Do it now, on your own schedule, decoupled from any upgrade.
Why the IPVS backend is being removed
The honest reason is maintenance, and the KEP says so: "sig-network currently lacks maintainers who are familar with the ipvs backend code." Bug reports against ipvs have been getting answered with a suggestion to move to nftables for a while.
The deeper reason is that ipvs never bought what people assume it bought. The kernel IPVS API alone cannot express Kubernetes Services, so the ipvs backend uses the iptables API on top of IPVS. If you adopted ipvs to escape iptables, you did not escape it. The Kubernetes documentation's own retrospective is blunt: the mode "was an experiment," it hit its throughput goals, but "the kernel IPVS API turned out to be a bad match for the Kubernetes Services API, and the ipvs backend was never able to implement all of the edge cases of Kubernetes Service functionality correctly."
The nftables mode is the designated replacement for both older backends. The same docs page says it "is essentially a replacement for both the iptables and ipvs modes, with better performance than either of them, and is recommended as a replacement for ipvs." It installs rules that pick a backend Pod at random, and it processes endpoint changes faster than iptables mode does. The docs also tell you where the default is going: "In Kubernetes 1.37, this is iptables, but a future version of Kubernetes will change the default to nftables."
That last sentence is worth acting on separately. If your kube-proxy config has no explicit mode field, a future upgrade will change your dataplane for you. Set the field even on clusters you are not migrating yet.
The scheduler question, and why it matters less than you think
This is where IPVS migrations stall. Someone points out that nftables mode picks endpoints at random, and asks what happens to lc or sh. The full IPVS scheduler list is real: rr, wrr, lc, wlc, lblc, lblcr, sh, dh, sed, nq. The KEP authors saw this coming. Their own checklist for the v1.37 stage includes a migration blog post with "an explanation of why IPVS schedulers aren't actually useful in Kubernetes, which a lot of ipvs users don't seem to realize."
Here is the mechanism. Every node runs its own kube-proxy with its own IPVS state, and no node knows what any other node is doing. Microsoft's AKS documentation puts a number on the consequence: "IPVS load balancing operates in each node independently and is only aware of connections that flow through the local node. This means that while LeastConnection results in a more even load under a higher number of connections, when a low number of connections occur (# connects < 2 * node count), traffic might be unbalanced."
So lc is not least-connection across your fleet. It is least-connection across whatever share of traffic one node happened to see. On a 40-node cluster with modest per-service traffic, that is close to random with extra steps and worse debuggability.
Two schedulers do encode a real intent, and both have a Kubernetes-native replacement.
| IPVS scheduler | What you actually wanted | Where it lives now | |---|---|---| | sh, lblc, lblcr | The same client keeps hitting the same Pod | .spec.sessionAffinity: ClientIP on the Service | | lc, wlc, sed, nq | Don't overload a slow replica | Readiness probes, HPA, and a real load balancer in front | | rr, wrr, dh | Nothing specific, it was the default someone picked | Random selection in nftables mode |
Session affinity is part of the Service API, not a proxy-mode feature, so it survives the migration. Tune the window with .spec.sessionAffinityConfig.clientIP.timeoutSeconds, which defaults to 10800 seconds, or three hours.
If a service genuinely needs load-aware balancing, kube-proxy was never the right layer for it. That belongs in a service mesh or an L7 proxy, and if you are already replacing an ingress controller this year, our Ingress-NGINX retirement migration guide covers the L7 side of that decision. This post stays at L4.
The default that will break you
Every generic "iptables vs nftables" post covers the NodePort change. Almost none of them tell IPVS users it applies to them, because the official migration note is written for iptables mode. The kube-proxy config reference is the source that settles it. On nodePortAddresses: "If unset, this defaults to 'all' in iptables and ipvs mode, and to 'primary' in nftables mode."
Read that again if you run NodePorts. Today, on ipvs, your NodePort services answer on every local IP on every node: the primary address, the secondary interface facing your management network, the address on a secondary ENI, whatever a VIP manager has parked on the box. After the cutover they answer on the node's primary IPv4 and IPv6 addresses only.
Nothing errors. The rules are simply not installed on the other addresses, so anything pointed at a secondary IP gets connection refused, and those callers are usually the ones nobody owns any more: an old monitoring scrape, a partner integration, a hardware load balancer pool that predates your ingress controller.
Inventory before you migrate, not after. Then either fix the callers or pin the old behaviour explicitly:
# KubeProxyConfiguration for a node group mid-migration.
# nodePortAddresses is set explicitly to preserve pre-migration reachability.
apiVersion: kubeproxy.config.k8s.io/v1alpha1
kind: KubeProxyConfiguration
mode: "nftables"
nodePortAddresses:
- "all" # keywords: primary, localhost, all. CIDRs also accepted.
nftables:
masqueradeAll: false
syncPeriod: 30s
minSyncPeriod: 1s
Set it to all to hold current behaviour through the cutover, then tighten it to primary as a separate, reversible change once you have fixed the callers. Two changes, two blast radii, two rollbacks.
Two more differences are worth knowing, though neither is as sharp.
Kernel 5.13 is a hard floor. The docs are unambiguous that nftables mode "requires kernel 5.13 or later." The KEP authors expect every kernel too old for it to be out of LTS by the end of 2026, which is true for the distributions most teams run, and irrelevant if you have one appliance-flavoured node pool stuck on an old vendor image. Check the pool, not the fleet average.
Conntrack behaviour differs. The same page notes that kernels before 6.1 carry a bug that can reset long-lived TCP connections to service IPs. The iptables mode installs a workaround; nftables mode does not install one by default. Check the iptables_ct_state_invalid_dropped_packets_total metric to see whether your cluster depends on it, and pass --conntrack-tcp-be-liberal if it does.
Pre-flight checks
Five commands. Run them against a real cluster before you write the change ticket.
# 1. Kernel floor, per node pool. Anything below 5.13 blocks the migration.
kubectl get nodes -o custom-columns=\
NODE:.metadata.name,KERNEL:.status.nodeInfo.kernelVersion
# 2. Which NodePort services exist at all. Empty output means the
# nodePortAddresses change above cannot hurt you.
kubectl get svc -A --field-selector spec.type=NodePort
# 3. Every non-primary address a node currently answers on.
kubectl get nodes -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{range .status.addresses[*]}{.type}={.address}{" "}{end}{"\n"}{end}'
# 4. Services relying on client stickiness, which must survive the move.
kubectl get svc -A -o json | jq -r '.items[]
| select(.spec.sessionAffinity=="ClientIP")
| "\(.metadata.namespace)/\(.metadata.name)"'
# 5. The scheduler configured today. The ConfigMap is called kube-proxy
# on kubeadm clusters and kube-proxy-config on EKS; adjust the name.
kubectl -n kube-system get cm kube-proxy-config -o yaml | grep -A3 'ipvs:'
Add one check the commands cannot make for you. Ask whether your CNI plugin cares. Cilium and Calico in eBPF mode may replace kube-proxy entirely, in which case this migration does not apply. On AKS, Microsoft documents that IPVS "doesn't support Azure Network Policy," so a cluster on IPVS there already made a policy trade-off worth revisiting during the move.
Running the cutover node by node
The reason this migration is safer than an ingress cutover is a design property most people miss. KEP-3866 states that rolling out the new mode in a live cluster is expected to be safe "even though this will result in different proxy modes running on different nodes, because Kubernetes service proxying is defined in such a way that no node needs to be aware of the implementation details of the service proxy implementation on any other node."
A node in nftables mode and a node in ipvs mode coexist in the same cluster with no shared state to corrupt. That gives you a per-node canary, which is a much better position than the all-or-nothing ConfigMap edit most walkthroughs describe.
On a self-managed cluster, run a second kube-proxy DaemonSet against a second ConfigMap, selected by a node label:
# Illustrative. Clone your existing kube-proxy DaemonSet, change the
# ConfigMap it mounts, and gate both copies on a node label.
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: kube-proxy-nftables
namespace: kube-system
spec:
selector:
matchLabels:
k8s-app: kube-proxy-nftables
template:
metadata:
labels:
k8s-app: kube-proxy-nftables
spec:
nodeSelector:
acme.example.internal/proxy-mode: nftables
hostNetwork: true
# containers, volumes and RBAC copied from the original DaemonSet,
# with the ConfigMap swapped for the nftables one.
Add the inverse nodeSelector to the original DaemonSet so exactly one kube-proxy runs per node, then move nodes across by labelling them:
kubectl label node ip-10-0-1-42 acme.example.internal/proxy-mode=nftables --overwrite
Verify on that node before touching a second one. Kube-proxy in nftables mode creates a table named kube-proxy in the ip family, and another in ip6:
nft list table ip kube-proxy | head -40 # rules present
ipvsadm -Ln # old IPVS entries cleared
Then check the things the rules do not tell you. Curl a ClusterIP and a NodePort from that node. Curl the NodePort from off-node, using whichever address your callers actually use. Watch error rates on services with endpoints on that node for a full traffic cycle, not five minutes.
Rollback is removing the label. Kube-proxy on the original DaemonSet cleans up the stale rules when it restarts on that node, which is the behaviour that makes the per-node approach worth the extra DaemonSet.
Migrate one node, then one node pool, then the rest. If your clusters are not yet in a state where you can label a node and watch its blast radius, that is the actual gap, and our Kubernetes production readiness checklist is the better place to start.
Where managed clusters constrain you
The playbook above assumes you control the kube-proxy DaemonSet. On managed control planes you often do not, and the per-node canary collapses into a cluster-wide switch.
| Platform | How the mode is set | What to check first | |---|---|---| | EKS | aws eks update-addon --addon-name kube-proxy --configuration-values '{"mode":"..."}', or edit the kube-proxy-config ConfigMap in kube-system | Run aws eks describe-addon-configuration --addon-name kube-proxy --addon-version <v> and read .configurationSchema to confirm the addon version accepts nftables | | AKS | az aks update --kube-proxy-config kube-proxy.json with {"enabled": true, "mode": "NFTABLES"} | nftables mode is preview: needs the KubeProxyConfigurationPreview feature flag and API version 2025-09-02-preview or later | | GKE | Not exposed as a kube-proxy setting | Google steers new clusters to Dataplane V2, which replaces kube-proxy rather than reconfiguring it |
AWS documents the mode switch and warns that it "is a disruptive change and should be performed in off-hours," which is correct for a cluster-wide flip. Read the EKS guidance before scheduling. Note that the EKS user guide page for the kube-proxy add-on does not mention nftables at all as of this writing, so verify against your addon version's configuration schema instead of assuming.
Our advice for managed clusters is the same shape, executed differently. Build a throwaway cluster on the same platform and version, run the pre-flight and verification steps there, then flip the real cluster in a window. You lose the per-node canary, so buy the confidence somewhere else.
One thing not to do: leave IPVS in place and plan to set the KubeProxyIPVS gate in v1.40. It works, and it buys three releases of nothing. The code is unmaintained today, and you will run the same migration in 2028 with less time and no one left who remembers why the cluster was on IPVS.
When to get help
Most teams can run this themselves. The pre-flight checks take an afternoon, and on a self-managed cluster the per-node cutover is genuinely low risk.
Call someone if the NodePort inventory comes back long and unowned, if a node pool is stuck below kernel 5.13 on a vendor image, if your CNI does its own service handling and nobody is sure what kube-proxy is still doing, or if this is one of several deprecations queued in the same quarter and you need them sequenced rather than raced.
We run dataplane migrations on production Kubernetes without maintenance windows that make the business nervous. If that is the problem in front of you, get in touch and we will look at your clusters before anyone edits a ConfigMap.