
The Case of the Phantom Webhook Timeout: How a Floating IP blocked entire Nutanix Kubernetes Platform Deployment
Contents
While bootstrapping a Nutanix Kubernetes Platform (NKP) 2.18 management cluster on AHV, the entire deployment ground to a halt at the very first component install. nkp create capi-components never got past cert-manager and because everything downstream (cluster-api-operator, and the rest of the cluster bootstrap) depends on that step, the whole cluster build was blocked.
The error itself gave almost no clue as to why:
Post "https://cert-manager-webhook.cert-manager.svc:443/mutate?timeout=30s":
context deadline exceeded (Client.Timeout exceeded while awaiting headers)
This kept repeating for 30+ minutes straight in the kube-apiserver logs, not a one-off startup race, but a persistent failure. The cert-manager Helm release sat stuck in pending-install. Reinstalling it manually made pods look healthy for a moment, but the next webhook-dependent step failed identically. Something below the application layer was broken.
Ruling Things Out, One by One#
The first instinct with any webhook timeout is to check the obvious suspects. So I did, methodically:
- Pod health: cert-manager, cainjector, and webhook pods were all healthy and Ready.
- Endpoints object: correctly populated with the right pod IP and port.
- Webhook reachability from the same subnet: a debug pod on the pod network could reach the webhook fine, TLS handshake and all.
- kube-proxy’s iptables Service rules: present and correct on every node.
- Calico’s route table: tunl0 routes to each node’s pod CIDR looked right.
- firewalld: inactive everywhere.
- MTU/fragmentation: ruled out; 1450-byte ICMP passed with zero loss.
- rp_filter sysctl: adjusted just in case, no change.
Narrowing It Down#
Since kube-apiserver runs with hostNetwork: true, I replicated that path with a test pod pinned to the control-plane node. It couldn’t reach any Service ClusterIP — not even CoreDNS. Restarting kube-proxy changed nothing. The issue wasn’t cert-manager specifically; it was all hostNetwork traffic leaving the control-plane node.
Simultaneous tcpdump captures on both ends confirmed it: IPIP-encapsulated packets (Calico’s overlay protocol) left the control-plane node correctly, but were silently dropped by the receiving worker’s cali-INPUT firewall chain before ever reaching the pod network.
Root Cause: The Node Didn’t Know Its Own IP#
The control-plane node runs kube-vip in ARP mode, announcing a floating API server VIP (10.161.95.130) on the same interface as its real address (10.161.95.220). Kubelet had no –node-ip flag set, so it auto-selected the VIP as the node’s InternalIP instead of its real one.
That one misidentification cascaded:
- Kubernetes registered the node’s InternalIP as the VIP, not its real address.
- Calico inherited that same wrong IP as the node’s overlay identity.
- Every worker’s cali40all-hosts-net trust list allow-listed the VIP — never the real IP.
- But the kernel actually sent IPIP traffic from the real interface address, since that’s just how routing works.
Calico expected packets from one IP; the wire delivered packets from another. Every IPIP packet from the control-plane node got dropped on arrival, taking down all pod-network and hostNetwork-to-pod traffic from that node, including kube-apiserver’s calls to the cert-manager webhook.
The Fix#
# /var/lib/kubelet/kubeadm-flags.env
KUBELET_KUBEADM_ARGS="--cloud-provider= ... --provider-id=preprovisioned:////10.161.95.220 --node-ip=10.161.95.220"
systemctl restart kubelet
After restarting kubelet, the Node’s InternalIP corrected itself, Calico’s annotation followed suit, and worker ipsets updated automatically. Webhook calls started succeeding immediately.
Running kube-vip alongside Calico on a preprovisioned cluster? Check your –node-ip flag before you deploy, it can save you days of debugging.