TL;DR: pv-media-server, a ReadWriteMany PV backed by smb.csi.k8s.io against a TrueNAS share (//192.168.20.5/media-server), returned stale file handle on stat() from every pod scheduled on it — sonarr, radarr, bazarr, jellyfin — on two different nodes (opal, onix) at the same time. csi-driver-smb keeps exactly one CIFS mount per node (the “globalmount”) and bind-mounts it into each pod’s NodePublishVolume call; it never checks whether that global mount is still alive, so once it goes stale, every pod scheduled on that node inherits the same dead handle, forever, with CreateContainerError on a loop. The mount used serverino (trust server-assigned inode numbers), which is exactly what TRaSH-Guides warns against for CIFS shares backing an *arr stack — server-side inode churn invalidates every cached dentry the kernel client is holding. Fixed with three independent layers: noserverino in the PV’s mountOptions so the client stops trusting inode numbers it can’t verify, a DaemonSet that stats every node’s CIFS globalmount every 30 seconds and umount -ls anything that returns ESTALE (forcing csi-driver-smb to remount on the next pod attach), and a liveness probe that actually stats /data instead of checking whether TCP port 445 is open — the old probe stayed green through the entire incident, because the SMB server was reachable the whole time. Only the mount was dead.
A CSI driver’s job is narrower than it looks: mount the share once per node, then bind-mount it into pods. Nothing in that contract obligates it to notice the mount died. csi-driver-smb doesn’t, and the failure mode it produces — every future pod on a node silently inheriting a corpse of a mount — is invisible to a liveness probe that only checks network reachability. This post is the trace from four crash-looping pods to a three-part fix that treats “the mount is dead” as a first-class, monitored condition instead of an assumption.
Who this is for
You’re running Kubernetes on bare metal or VMs with a shared SMB or NFS volume mounted via a CSI driver — csi-driver-smb, csi-driver-nfs, or similar — and more than one pod or more than one node consuming the same ReadWriteMany PV. You know what a CSI NodeStageVolume/NodePublishVolume split does. What you might not have checked is what your driver does when the staged mount dies underneath it, and whether your liveness probe would actually catch that.
Table of Contents
Open Table of Contents
- The setup: one SMB share, six *arr containers
- The symptom: CreateContainerError, stale file handle, two nodes at once
- Why “two nodes at once” rules out a single bad host
- What csi-driver-smb actually does on NodePublishVolume
- The root cause: serverino trusting inode numbers it can’t verify
- The fix, three layers
- Why the existing liveness probe never caught it
- What’s still not proven
- References
The setup: one SMB share, six *arr containers
Direct answer: Six containers (jellyfin, sonarr, radarr, bazarr, and two disabled — lingarr, lazylibrarian) share one
ReadWriteManyPersistentVolume backed bysmb.csi.k8s.iov1.19.1, pointed at a TrueNAS SMB export. Each pod also ran abusyboxsidecar meant to catch SMB outages, but it only checked TCP reachability of port 445 — not whether the mount itself worked.
The PV is provisioned by hand (not dynamically) as a Pulumi PersistentVolume resource, with a matching PersistentVolumeClaim that every *arr app binds by name:
media_server_share = PersistentVolume(
"pv-media-server",
spec={
"capacity": {"storage": "1000Gi"},
"accessModes": ["ReadWriteMany"],
"storageClassName": "smb",
"mountOptions": [
"dir_mode=0777", "file_mode=0777", "nounix",
"echo_interval=30", "hard", "actimeo=30",
],
"csi": {
"driver": "smb.csi.k8s.io",
"volumeHandle": "smb-server.default.svc.cluster.local/share#media-server#",
"volumeAttributes": {"source": "//192.168.20.5/media-server"},
"nodeStageSecretRef": {"name": "smb-credentials-secret", "namespace": "media-server"},
},
},
)
hard means the kernel CIFS client retries indefinitely on I/O errors rather than surfacing them to the application — the standard choice for a share you don’t want to silently return garbage on a network blip. Each app pod already carried a smb-health-monitor sidecar, added after an earlier SMB outage, that ran nc -z -w 5 192.168.20.5 445 every 30 seconds and wrote a timestamp file a liveness_probe checked. It looked like defense in depth. It wasn’t checking the thing that actually broke.
The symptom: CreateContainerError, stale file handle, two nodes at once
Direct answer:
kubectl describe podon sonarr showed repeatedWarning Failedevents —failed to stat "…/pv-media-server/mount": stat … stale file handle— immediately after a successful image pull, on every container creation attempt, across pod restarts. The same error, same wording, showed up on bazarr and jellyfin on a different node (onix) at the same time. Pulling thecsi-smb-nodeDaemonSet logs on both nodes confirmed the exact kernel error:stale NFS file handle(yes, that string, from a CIFS mount — Linux’s VFS layer reusesESTALEacross filesystem types).
Warning Failed 8s kubelet spec.containers{sonarr}: Error: failed to generate container
"fe130c99…" spec: failed to generate spec: failed to stat
"/var/lib/kubelet/pods/…/volumes/kubernetes.io~csi/pv-media-server/mount":
stat …/mount: stale file handle
kubelet retried the container creation three times in nine seconds — each retry re-pulled the (already-cached) image, then failed at the exact same stat() call. kubectl get pods -n media-server showed sonarr, radarr, bazarr, and jellyfin all in CreateContainerError, all four pods that reference pv-media-server. Every pod that didn’t mount it — prowlarr, qbit-manage, recyclarr — was healthy.
Why “two nodes at once” rules out a single bad host
Direct answer: sonarr was on
opal, bazarr and jellyfin were ononix— two different physical hosts, two different kernel CIFS clients, two differentcsi-smb-nodeDaemonSet pods, both returning the identicalstale NFS file handleerror at overlapping timestamps. A bad NIC, a flaky switch port, or corrupted local state on one node would explain a single-node failure. Two independent kernel CIFS clients going stale within the same few minutes points at the shared thing both clients talk to: the TrueNAS export itself.
ping 192.168.20.5 from inside the cluster returned clean 30–56ms round trips throughout the incident — no packet loss, no route flap. That ruled out a network partition. Whatever invalidated the mount didn’t take the server offline; it just changed something about the filesystem state the SMB server was exposing, and every client holding an open handle from before that change got ESTALE on the next access.
What csi-driver-smb actually does on NodePublishVolume
Direct answer: csi-driver-smb mounts the SMB share exactly once per node, at
NodeStageVolumetime, into a per-volumeglobalmountdirectory under/var/lib/kubelet/plugins/kubernetes.io/csi/smb.csi.k8s.io/<hash>/globalmount. Every pod on that node that references the same volume gets a bind-mount of that one globalmount into its own pod directory viaNodePublishVolume. The driver never re-checks that the globalmount is alive before doing the bind-mount — it just runsmount -o bind. If the globalmount is stale, the bind-mount succeeds (bind-mounting a directory doesn’t require reading through it), and the nextstat()anyone does — kubelet’s own container-creationstat()— is what returns ESTALE.
Confirmed directly by reading mount output on opal as root:
//192.168.20.5/media-server on /var/lib/kubelet/plugins/kubernetes.io/csi/smb.csi.k8s.io/<hash>/globalmount
type cifs (rw,relatime,vers=3.1.1,…,serverino,mapposix,reparse=nfs,…,actimeo=30,closetimeo=1)
//192.168.20.5/media-server on /var/lib/kubelet/pods/<pod-uid>/volumes/kubernetes.io~csi/pv-media-server/mount
type cifs (rw,relatime,vers=3.1.1,…,serverino,mapposix,reparse=nfs,…,actimeo=30,closetimeo=1)
Both entries are the same CIFS session — the second is a bind-mount of the first, both showing the same type cifs because bind-mounts of network filesystems still report the underlying fstype. And in the csi-smb-node container logs, the NodePublishVolume call for a brand-new pod, seconds after the ESTALE errors started, logs a clean Mounting cmd (mount) with arguments (-o bind …) and a successfully response — the driver reported success while handing out a dead mount. The nodeserver.go NodePublishVolume implementation does an existence/mount-point check before the bind, not a liveness check — a stale mount is still, from the kernel’s point of view, “mounted.”
The root cause: serverino trusting inode numbers it can’t verify
Direct answer: the PV’s
mountOptionsdidn’t disableserverino, so the CIFS client defaults to trusting inode numbers the TrueNAS server hands out, rather than generating its own.serverinois the correct default for most workloads — client-generated inodes break hardlink detection across a single mount, which matters for exactly the kind of atomic move*arrapps do between a download directory and a library directory on the same share. But it also means that if the server ever reassigns inode numbers for files the client has cached dentries for — a snapshot rollback, a dataset re-export, an SMB service restart, anything that changes how the backing filesystem enumerates inodes — every one of those cached dentries becomes unresolvable, and the kernel returnsESTALEon the next lookup through them.
This is a documented CIFS hardlink caveat in the TRaSH-Guides ecosystem that *arr app operators run into specifically because of instant-move/hardlink behavior — but the usual framing is “hardlinks silently fail” or “cross-share hardlinks turn into copies,” not “the entire mount goes stale for every consumer.” The stale-handle failure mode is the same root cause — the client trusting server-side inode identity — surfacing worse, at the mount level instead of the file level, because serverino was combined with hard (retry forever, never surface a soft error) and a 30-second actimeo (cache attributes, including inode-derived state, for 30 seconds before re-validating).
I did not get direct access to TrueNAS’s own service or journalctl logs during this incident — SSH to the TrueNAS box used a different key than the one available in the cluster’s admin environment, and by the time that was resolved, the relevant window of logs had rotated. So the specific server-side event that remapped inode numbers is inferred from the mount options and the failure signature, not confirmed from a TrueNAS-side log line. That’s the honest limitation of this post — see What’s still not proven.
The fix, three layers
Direct answer:
noserverinoin the PVmountOptionsstops the client from trusting server-assigned inode numbers at all, removing the specific trigger. A newsmb-stale-mount-healerDaemonSet, running privileged with a bidirectional hostPath mount on/var/lib/kubelet/plugins/kubernetes.io/csi/smb.csi.k8s.io,stats every globalmount every 30 seconds andumount -ls anything that returns an error — the actual auto-recovery csi-driver-smb doesn’t have. And the sidecar’s liveness check was rewritten from a TCP port probe to a realstat /data, so a stale mount actually fails the probe instead of leaving it green.
noserverino is one line in the existing mountOptions list:
"mountOptions": [
"dir_mode=0777", "file_mode=0777", "nounix",
"echo_interval=30", "hard", "actimeo=30",
"noserverino",
],
The mount.cifs man page documents noserverino as: client generates inode numbers instead of using ones provided by the server. It doesn’t eliminate ESTALE as a possibility — a CIFS session can still go stale for other reasons, like a TCP-level reconnect that resets server-side file handle state — but it removes the one failure trigger this incident’s evidence pointed at, and it’s the same recommendation TRaSH-Guides gives for this exact class of app.
The healer DaemonSet is what makes “still went stale for some other reason” survivable instead of fatal, because it turns a permanent outage into a bounded one — worst case, 30 seconds until the next stat, then a forced remount on the next pod attach:
smb_stale_mount_healer = DaemonSet(
"smb-stale-mount-healer",
metadata={"name": "smb-stale-mount-healer", "namespace": "kube-system"},
spec={
"template": {
"spec": {
"containers": [{
"name": "healer",
"image": "busybox:latest",
"command": ["/bin/sh", "-c"],
"args": ["""
while true; do
for d in /var/lib/kubelet/plugins/kubernetes.io/csi/smb.csi.k8s.io/*/globalmount; do
[ -d "$d" ] || continue
if ! timeout 5 stat "$d" >/dev/null 2>&1; then
echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [HEAL] $d is stale, forcing lazy umount"
umount -l "$d" 2>&1
fi
done
sleep 30
done
"""],
"security_context": {"privileged": True},
"volume_mounts": [{
"name": "csi-smb-root",
"mount_path": "/var/lib/kubelet/plugins/kubernetes.io/csi/smb.csi.k8s.io",
"mount_propagation": "Bidirectional",
}],
}],
"volumes": [{
"name": "csi-smb-root",
"hostPath": {"path": "/var/lib/kubelet/plugins/kubernetes.io/csi/smb.csi.k8s.io", "type": "DirectoryOrCreate"},
}],
},
},
},
)
mountPropagation: Bidirectional is what makes umount -l inside the container actually detach the mount on the host, not just inside the container’s own mount namespace — without it, the container would be unmounting a private copy of the mount table and the host-side globalmount, and every pod still bind-mounted from it, would be untouched. umount -l (lazy unmount) detaches the mount point immediately and cleans up the underlying reference once nothing has it open — the right choice here because a -f force unmount can still hang against a truly unresponsive server, and a plain umount fails outright if anything still has the path open.
privileged: true is a real cost, not a formality — this DaemonSet can unmount anything under that host path, on every node, with no further authorization check. It’s scoped as tightly as the mechanism allows (one hostPath, one directory tree, stat+umount only), but the failure mode wasn’t going to be fixed by anything less than something running on the host’s mount namespace.
Why the existing liveness probe never caught it
Direct answer: the original sidecar ran
nc -z -w 5 192.168.20.5 445— a raw TCP connect to the SMB port — every 30 seconds, and considered the mount “alive” if that connect succeeded. Throughout this entire incident, the TCP connect succeeded, because the SMB server was up and accepting connections the whole time; only the previously-established mount’s file handles were invalid. A liveness probe testing network reachability and a liveness probe testing mount health are answering two different questions, and only one of them is the one that actually predictsCreateContainerError.
The fix inverts what’s being tested — stat the thing the application actually depends on, not a proxy for it:
smb_healthcheck_sidecar = {
"args": ["""
STATE=up
while true; do
if timeout 5 stat /data >/dev/null 2>&1; then
date +%s > /healthcheck/alive
[ "$STATE" = "down" ] && echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [RECOVERY] /data mount is healthy again"
STATE=up
else
[ "$STATE" = "up" ] && echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [FAILURE] /data mount unhealthy"
STATE=down
fi
sleep 30
done
"""],
"volume_mounts": [
{"name": "smb-healthcheck", "mount_path": "/healthcheck"},
{"name": "media-server-data", "mount_path": "/data", "read_only": True},
],
}
The sidecar needs its own mount of the same volume to stat it — a network reachability check doesn’t require touching the filesystem, a mount health check inherently does. This is a small but real change in blast radius: the sidecar is now a second consumer of the CIFS mount inside the same pod, so if the healer DaemonSet’s remount ever raced against a pod’s own read/write traffic, both containers would see the same transient state. In practice this hasn’t caused any issue, because both are read-only-safe operations and timeout 5 bounds how long either can block.
timeout 5 in front of stat matters specifically because of the hard mount option: a genuinely unresponsive server (not stale, just slow or partitioned) would otherwise make stat block for as long as the kernel client keeps retrying, and a probe that hangs is worse than one that returns an honest failure.
What’s still not proven
Direct answer: the specific TrueNAS-side event that remapped inode numbers and triggered the stale handles was never identified — SSH access to the TrueNAS box wasn’t available with the key on hand during the incident window, and by the time it was, the relevant service logs had rotated out.
noserverinofixes the class of failure this evidence points at; it doesn’t confirm what actually happened on 2026-09-22, and it wouldn’t protect against every possible cause of a stale CIFS handle.
The strongest counter-argument to this fix: if the real trigger was a TCP-level session drop and reconnect — not an inode remap — noserverino does nothing, because the failure would be in the session layer, not in inode trust. The closetimeo=1 mount option visible in the mount output (a one-second close timeout, likely a csi-driver-smb default rather than something explicitly configured) is short enough that a marginal network blip could plausibly force a reconnect cycle that leaves stale handles regardless of serverino. That’s exactly the scenario the healer DaemonSet is insurance against — it doesn’t care why the mount went stale, only that it did. If noserverino alone had been the fix, the DaemonSet would be redundant; the fact that it’s staying in place is itself the honest admission that the root cause isn’t fully nailed down.
Post-rollout, verified directly: all four *arr pods 2/2 Running with zero restarts, the healer DaemonSet 3/3 ready across every node, stat /data from inside a running sonarr container returning clean inode/mode output, and no stale file handle events in kubectl get events since the fix went in. That’s confirmation the symptom stopped, not confirmation of the exact cause — worth keeping the two claims separate on Monday, when the next storage incident inevitably needs a different diagnosis.
References
- csi-driver-smb
nodeserver.go— the actualNodePublishVolumeimplementation, showing the bind-mount has no liveness check before or after. mount.cifs(8)— documentsserverino/noserverinoandhard/soft, the two option pairs at the center of this incident.- TRaSH-Guides: Hardlinks and Instant Moves — the community source for the
noserverinorecommendation on CIFS shares backing*arrapps, framed around hardlink correctness rather than mount stability, but the same underlying cause. kubernetes-csi/csi-driver-smbcharts — the Helm chart this cluster deploys the driver from (v1.19.1 at time of writing).