Skip to main content
Crusoe Support Help Center home page
Crusoe

How-To Enable Cilium socketLB.hostNamespaceOnly on a CMK Cluster

Karan Solanki
Karan Solanki
Updated

Introduction

Cilium, the CNI on Crusoe Managed Kubernetes (CMK), load-balances Kubernetes Services in eBPF at the socket layer. When a pod process calls connect() against a ClusterIP, Cilium's socket-LB rewrites the destination to a backend pod IP before any packet leaves the pod. That is why there is no kube-proxy on CMK nodes.

Socket-LB only works on traffic that originates from a socket on the node. A pod that receives a packet on one interface and forwards it towards a ClusterIP never calls connect() for that flow, so there is nothing to hook.

The ClusterIP leaves the pod untranslated, Cilium classifies it as the world identity, and the SYN is routed out of the physical NIC to nowhere. No drop is recorded and no counter moves; the connection times out. Any pod in the forwarding class is affected: VPN and tunnel ingress proxies such as Tailscale, service-mesh sidecars or ambient redirection layers, and NAT gateway pods.

CMK ships Cilium's default socket-LB coverage, Full. Setting the Helm value socketLB.hostNamespaceOnly=true (ConfigMap key bpf-lb-sock-hostns-only) restricts socket-LB to the host network namespace, so pod-to-ClusterIP traffic is translated by the per-packet eBPF datapath instead, which handles forwarded packets correctly. enable-host-legacy-routing is not required; you keep Cilium's BPF host routing.

This article walks through confirming the current coverage mode, getting the setting applied durably in your cluster's managed Cilium configuration, and rolling it out to nodes.

Prerequisites

  • Kubeconfig With Cluster-Admin Access to the CMK Cluster
  • kubectl Installed Locally
  • helm Installed Locally
  • Ability to Schedule a cilium-agent Restart on Each Node

Instructions

Step 1: Check the Current Socket-LB Coverage

Pick a node and query its cilium-agent:

NODE=<NODE_NAME>
CILIUM_POD=$(kubectl -n kube-system get pods -l k8s-app=cilium --field-selector spec.nodeName=$NODE -o jsonpath='{.items[0].metadata.name}')
kubectl -n kube-system exec $CILIUM_POD -- cilium-dbg status --verbose | grep -i "Socket LB"

Socket LB Coverage: Full is the default. Socket LB Coverage: Hostns-only means the setting is already active on that node.

Check whether the key is present in the cluster configuration:

kubectl -n kube-system get cm cilium-config -o yaml | grep bpf-lb-sock-hostns-only

No output means the key is absent and Cilium is using the upstream default of false.

Step 2: Get the Value into the Cilium Release

socketLB.hostNamespaceOnly is a Helm chart value. For it to be durable it has to live in the Cilium release's values, not only in the rendered cilium-config ConfigMap.

Open a support ticket asking for socketLB.hostNamespaceOnly=true to be set in the CMK-managed Cilium values for your cluster. Include the cluster name and the output from Step 1. The CMK team applies it in the managed values, so future re-renders of the release carry it forward.

Cilium is a Crusoe-managed add-on under the shared responsibility model. If you manage Cilium values on your cluster yourself, coordinate with Crusoe Support before changing them and preserve your existing values (for example with helm upgrade --reuse-values) rather than applying chart defaults. See FAQ: CMK Add-on Lifecycle and Upgrades.

Changing this value only updates cilium-config and it does not roll the agents unless the release sets rollOutCiliumPods=true (the chart default is false). Check your release before the value is applied:

helm get values -n kube-system cilium -a | grep rollOutCiliumPods

If the output is rollOutCiliumPods: true, applying the value restarts every agent in the cluster through the DaemonSet's rolling update rather than on your node-by-node schedule in Step 3. Mention the result in your ticket.

While your ticket is open, you can apply the key directly to the ConfigMap as a stopgap:

kubectl patch cm cilium-config -n kube-system --type merge -p '{"data":{"bpf-lb-sock-hostns-only":"true"}}'

⚠️ Warning: Do not treat patching bpf-lb-sock-hostns-only=true directly into the cilium-config ConfigMap as a permanent fix. A ConfigMap patch is not part of the Helm release values, so it cannot be relied on to survive a re-render of the release. If the key is lost, the failure is silent: agents that restart afterwards come up on Full coverage with no drop or error recorded. A ConfigMap patch is fine as a stopgap while your ticket is open, and it is safe to leave in place alongside the release value since both set the same key to the same value.

ℹ️ Note: There is currently no self-service mechanism for per-cluster Cilium value overrides through the Crusoe Console, CLI, or Terraform, and no published list of which Cilium values CMK manages.

Step 3: Restart cilium-agent on Each Node

cilium-agent reads cilium-config only at startup, so the setting takes effect on a node only after its agent restarts. Nodes added after the setting is in place come up with it automatically.

Restart one node's agent at a time when its workloads can tolerate a brief datapath reload. Set CILIUM_POD for the target node as in Step 1, then delete the pod; the DaemonSet recreates it with the new configuration:

kubectl -n kube-system delete pod $CILIUM_POD

Repeat Step 1 for each node to track which are still on Full.

ℹ️ Note: On GPU node pools running multi-day training jobs you may not be able to restart agents freely. Until every node reports Hostns-only, constrain your forwarding pods (the pods doing the forwarding, not only the backends they front) to a node pool whose agents already have the setting, using a nodeSelector or node affinity. A forwarding pod rescheduled onto a node still on Full will hit the timeout again.

Step 4: Verify

Confirm every node reports Socket LB Coverage: Hostns-only using the Step 1 check, then retry a request through a forwarding pod that previously timed out.

To see the difference in the datapath, look up the cilium-agent on the node where the forwarding pod is running (the pod you deleted in Step 3 no longer exists, and the trace has to run on the forwarding pod's own node), find the forwarding pod's Cilium endpoint ID, and trace from it:

NODE=$(kubectl -n <FORWARDING_POD_NAMESPACE> get pod <FORWARDING_POD_NAME> -o jsonpath='{.spec.nodeName}')
CILIUM_POD=$(kubectl -n kube-system get pods -l k8s-app=cilium --field-selector spec.nodeName=$NODE -o jsonpath='{.items[0].metadata.name}')
kubectl -n kube-system exec $CILIUM_POD -- cilium-dbg endpoint list
kubectl -n kube-system exec $CILIUM_POD -- cilium-dbg monitor --type trace --from <FORWARDING_POD_ENDPOINT_ID>

cilium-dbg monitor streams until you stop it with Ctrl+C, so send a request through the forwarding pod while it runs.

Before the change the flow leaves with the ClusterIP intact and identity world on the physical NIC:

-> network flow 0x0 , identity <POD_IDENTITY>->world state new ifindex ens7 orig-ip 0.0.0.0: <POD_IP>:60998 -> <CLUSTER_IP>:80 tcp SYN

After the change the flow resolves to a backend pod IP with endpoint-to-endpoint identities and the handshake completes:

-> endpoint 726 flow 0x0 , identity <POD_IDENTITY>-><BACKEND_IDENTITY> state new ifindex lxcf3adcf8c3b6a orig-ip <POD_IP>: <POD_IP>:39122 -> <BACKEND_POD_IP>:10080 tcp SYN
-> endpoint 2805 flow 0x6e32c828 , identity <BACKEND_IDENTITY>-><POD_IDENTITY> state reply ifindex lxcb1f9dfe5d9cf orig-ip <BACKEND_POD_IP>: <CLUSTER_IP>:80 -> <POD_IP>:39122 tcp SYN, ACK

Example

A team runs an internal ML platform on a CMK cluster with a CPU node pool for infrastructure and an H100 node pool for training. They expose the platform's Envoy gateway Services to their private network through a Tailscale ingress proxy pod. Requests from the tailnet time out at connect while curl against the same ClusterIP from inside the proxy pod works, and cilium-dbg monitor --type drop shows nothing for the Service CIDR.

Step 1 shows every node on Full and the key absent from cilium-config. They open a support ticket, the CMK team sets socketLB.hostNamespaceOnly=true in the managed values, and the CPU-pool nodes pick it up on their next agent restart. Until the H100 agents can be cycled between training jobs, the team pins both the Envoy gateways and the Tailscale proxy pods to the CPU pool with a nodeSelector, and inbound traffic stays up.

Related Articles

Additional Resources

Related to

Was this article helpful?

0 out of 0 found this helpful

Still need help?

Our support team is ready to assist you with any questions.

Have more questions? Submit a request

Related Articles

Recently Viewed

Comments

0 comments

Article is closed for comments.