The Case of the Phantom Webhook Timeout: How a Floating IP blocked entire Nutanix Kubernetes Platform Deployment

Contents

While bootstrapping a Nutanix Kubernetes Platform (NKP) 2.18 management cluster on AHV, the entire deployment ground to a halt at the very first component install. nkp create capi-components never got past cert-manager and because everything downstream (cluster-api-operator, and the rest of the cluster bootstrap) depends on that step, the whole cluster build was blocked.

The error itself gave almost no clue as to why:

Post "https://cert-manager-webhook.cert-manager.svc:443/mutate?timeout=30s":
  context deadline exceeded (Client.Timeout exceeded while awaiting headers)

This kept repeating for 30+ minutes straight in the kube-apiserver logs, not a one-off startup race, but a persistent failure. The cert-manager Helm release sat stuck in pending-install. Reinstalling it manually made pods look healthy for a moment, but the next webhook-dependent step failed identically. Something below the application layer was broken.

Ruling Things Out, One by One#

The first instinct with any webhook timeout is to check the obvious suspects. So I did, methodically:

  • Pod health: cert-manager, cainjector, and webhook pods were all healthy and Ready.
  • Endpoints object: correctly populated with the right pod IP and port.
  • Webhook reachability from the same subnet: a debug pod on the pod network could reach the webhook fine, TLS handshake and all.
  • kube-proxy’s iptables Service rules: present and correct on every node.
  • Calico’s route table: tunl0 routes to each node’s pod CIDR looked right.
  • firewalld: inactive everywhere.
  • MTU/fragmentation: ruled out; 1450-byte ICMP passed with zero loss.
  • rp_filter sysctl: adjusted just in case, no change.

Narrowing It Down#

Since kube-apiserver runs with hostNetwork: true, I replicated that path with a test pod pinned to the control-plane node. It couldn’t reach any Service ClusterIP — not even CoreDNS. Restarting kube-proxy changed nothing. The issue wasn’t cert-manager specifically; it was all hostNetwork traffic leaving the control-plane node.

Simultaneous tcpdump captures on both ends confirmed it: IPIP-encapsulated packets (Calico’s overlay protocol) left the control-plane node correctly, but were silently dropped by the receiving worker’s cali-INPUT firewall chain before ever reaching the pod network.

Root Cause: The Node Didn’t Know Its Own IP#

The control-plane node runs kube-vip in ARP mode, announcing a floating API server VIP (10.161.95.130) on the same interface as its real address (10.161.95.220). Kubelet had no –node-ip flag set, so it auto-selected the VIP as the node’s InternalIP instead of its real one.

That one misidentification cascaded:

  • Kubernetes registered the node’s InternalIP as the VIP, not its real address.
  • Calico inherited that same wrong IP as the node’s overlay identity.
  • Every worker’s cali40all-hosts-net trust list allow-listed the VIP — never the real IP.
  • But the kernel actually sent IPIP traffic from the real interface address, since that’s just how routing works.

Calico expected packets from one IP; the wire delivered packets from another. Every IPIP packet from the control-plane node got dropped on arrival, taking down all pod-network and hostNetwork-to-pod traffic from that node, including kube-apiserver’s calls to the cert-manager webhook.

The Fix#

# /var/lib/kubelet/kubeadm-flags.env
KUBELET_KUBEADM_ARGS="--cloud-provider= ... --provider-id=preprovisioned:////10.161.95.220 --node-ip=10.161.95.220"
systemctl restart kubelet

After restarting kubelet, the Node’s InternalIP corrected itself, Calico’s annotation followed suit, and worker ipsets updated automatically. Webhook calls started succeeding immediately.

Running kube-vip alongside Calico on a preprovisioned cluster? Check your –node-ip flag before you deploy, it can save you days of debugging.