TL;DR: CIFS hard mounts (the default) put processes in TASK_UNINTERRUPTIBLE (D-state) when the server disappears — D-state ignores SIGKILL, so containers can’t be removed, pods get stuck Terminating, and if enough accumulate the hung-task watchdog panics the node; NFS-style options (timeo, retrans, nofail) are invalid arguments for the CIFS module, HTTP liveness probes don’t catch the hang because the app stays responsive until it physically touches the mount, and exec probes doing ls /mount/path block for the same reason. The solution is a sidecar that probes TCP 445 every 30 s via nc (no filesystem access), writes a Unix timestamp to a shared emptyDir on success, and a liveness probe on the main container that reads that timestamp and fails if it’s older than 90 s — so the probe can never block on a hung mount; hung_task_panic=1 with panic=5 remains as a backstop if a pod reaches D-state before the sidecar can intervene.
The right abstraction for detecting a hung CIFS mount is not filesystem access — it’s a network-level TCP check against port 445. Any probe that touches the mount itself will block in D-state alongside the workload it’s supposed to rescue. This post documents every standard approach that fails and why it fails at the kernel level, then builds a sidecar pattern that survives NAS outages without touching the filesystem. The echo_interval mount option does bound reconnection time, but it is not sufficient alone: pods that write continuously will enter D-state within that window, so the sidecar must kill the pod before the echo timeout expires.
kill -9 is supposed to kill anything. It won’t kill a process blocked on a hung CIFS mount. When a NAS goes offline and a pod’s containers are mid-I/O, those processes enter TASK_UNINTERRUPTIBLE — the kernel’s D-state — and become invisible to every signal, including SIGKILL. The container runtime can’t remove them. The pod stays Terminating indefinitely. Multiply that across eight pods in a media server namespace and the kernel’s hung-task watchdog fires, panicking the entire node.
This is the situation this cluster ended up in: a k3s homelab with ~8 pods mounting a 1 TB SMB share via the SMB CSI driver. When the NAS went offline, the node froze and required either a physical power-button hold or hung_task_panic=1 to trigger an automatic reboot. The search for a liveness probe that could catch this before it cascaded led through every obvious option — and most of them make the problem worse, not better.
This is a CIFS-specific hang, not to be confused with the unrelated memory-overcommit and NVMe-cascade freezes on the same cluster covered in k3s Node Freeze: Longhorn iSCSI Stalls and Memory Pressure — same symptom (unresponsive node), different mechanism. Why Your Kubernetes Node Freezes indexes this alongside the other storage failure modes found on the same cluster.
This post is for operators who already know their way around Kubernetes liveness probes and CIFS mounts and have hit (or want to avoid) the hung-mount failure mode. It assumes familiarity with pod specs, Kubernetes health checks, and basic Linux process states.
Table of Contents
Open Table of Contents
- Why do D-state processes ignore SIGKILL — and why does that freeze entire nodes?
- What was tried, and why did it fail?
- What works: network-level sidecar with emptyDir heartbeat
- Detection latency: worst-case timeline
- The backstop: when a pod reaches D-state anyway
- Why are network health checks the right abstraction for filesystem resources?
- References
Why do D-state processes ignore SIGKILL — and why does that freeze entire nodes?
CIFS hard mounts put I/O-blocked processes into
TASK_UNINTERRUPTIBLE(D-state), which ignores every signal, includingSIGKILL, and is invisible to the OOM killer. The kernel retries I/O against the unreachable server indefinitely, so the process cannot be forced to exit until the mount recovers or the machine reboots. Kubernetes cannot reap a D-state process, so the pod hangs inTerminating, and if enough pods hit this at once, the kernel’s hung-task watchdog panics the whole node. Thesoftmount option is documented as the escape hatch, but on real kernels it does not reliably prevent D-state hangs under connection loss — the man page’s promise does not match observed behavior. Treat CIFS hard mounts as capable of freezing an entire node, not only the pod that owns the mount, whenever the backing server can disappear.
CIFS mounts are hard by default — the kernel retries I/O indefinitely until the server comes back. A process blocked on a hard mount sits in TASK_UNINTERRUPTIBLE, which means it ignores all signals, including SIGKILL. The OOM killer can’t touch it. kill -9 does nothing. The process will sit there until the mount comes back or the machine reboots.
This is the same behavior as hard NFS mounts. The CIFS module does have a soft option (it’s nominally the default per the man page), but in practice it doesn’t reliably prevent D-state hangs on most kernels — processes still block on I/O to an unreachable server. This is a longstanding kernel bug: the man page says soft mounts won’t hang, but real-world behavior contradicts that under connection-loss scenarios.
What was tried, and why did it fail?
Every standard Kubernetes health-check approach fails against a hung CIFS mount, for different reasons rooted in the kernel. NFS-style mount options (
timeo,retrans,nofail) are not recognized by the CIFS module and cause the mount to fail outright withmount error(22).echo_intervalis a real per-mount option that bounds reconnection time to roughly3 × echo_interval, butecho_retriesno longer exists as a tunable on modern kernels — it is absent from/sys/module/cifs/parameters/on Debian 13 with kernel 6.12. HTTP liveness probes keep passing because the application stays responsive until it touches the mount, so the hang happens after the probe already reported healthy. Exec probes that runlson the mount path block in D-state exactly like the workload they are meant to protect, adding stuck processes instead of catching the failure. None of these operate below the filesystem layer, so none of them can detect the hang before it happens.
timeo, retrans, nofail — NFS options the CIFS module rejects
These are NFS mount options. The CIFS module doesn’t recognize them. The mount fails immediately:
mount error(22): Invalid argument
echo_interval and echo_retries — what does the module actually support?
The CIFS kernel module uses server echo (keepalive) to detect dead connections. echo_interval is a valid per-mount option — it sets the interval in seconds between echo requests (default: 60s) and can be passed via mountOptions in a Kubernetes PV spec. The reconnection timeout is approximately 3 × echo_interval, so echo_interval=5 gives a ~15-second timeout before the kernel considers the server dead and starts returning errors.
The PV in this cluster uses echo_interval=30, giving a 90-second reconnection timeout — intentionally matched to the 90-second heartbeat threshold in the sidecar liveness probe. When the NAS goes offline, both mechanisms converge on the same window: the sidecar kills the pod before the CIFS echo timeout would kick in for idle connections, and the echo timeout bounds how long in-flight I/O can block before the kernel gives up on the connection.
echo_retries is a different story: it was a module-level parameter on older kernels but is absent from /sys/module/cifs/parameters/ on Debian 13 with kernel 6.12:
$ ls /sys/module/cifs/parameters/
CIFSMaxBufSize cifs_max_pending cifs_min_rcv cifs_min_small
dir_cache_timeout disable_legacy_dialects enable_gcm_256
enable_negotiate_signing enable_oplocks require_gcm_256
slow_rsp_threshold
Even with echo_interval tuned down, any I/O already in-flight when the server disappears still blocks until the echo timeout elapses. For pods that continuously write (like a media scanner or a database), this window is enough to enter D-state and accumulate.
Why do HTTP liveness probes pass while the app sleeps into D-state?
The *arr stack (Sonarr, Radarr, Bazarr, etc.) exposes HTTP health endpoints. A liveness probe against /ping will restart the container if it stops responding. But if the SMB share goes down while the app is idle, the HTTP endpoint keeps returning 200. The liveness probe passes. The pod looks healthy. The next time the app tries to write a database record or scan a directory on the mount, it enters D-state. By then the probe can’t help — a D-state process won’t respond to the container restart signal either.
HTTP probes are useful for detecting app crashes. They don’t detect hung mounts.
Exec probes doing ls /mount/path block in D-state too
Same problem. ls on a hung CIFS mount enters D-state. The probe itself hangs. Kubernetes marks the probe as timed out after timeoutSeconds, but the blocked ls process stays alive in D-state, accumulating with each probe interval. The pod eventually gets restarted by the failure threshold, but the old D-state processes can prevent clean container removal.
What works: network-level sidecar with emptyDir heartbeat
The fix is a sidecar container that checks TCP port 445 with
nc -z -w 5 <server> 445every 30 seconds — no filesystem access, so it cannot enter D-state alongside the workload. On success it writes a Unix timestamp to a sharedemptyDirvolume; on failure it logs the event and stops writing. The main container’s liveness probe reads that timestamp file, not the CIFS mount, and fails once it is older than 90 seconds, giving a check that can never block.terminationGracePeriodSeconds: 5matters too: without it, kubelet waits the default 30 seconds for graceful shutdown before sendingSIGKILL, which a D-state process ignores regardless — a short grace period lets the container runtime force-destroy the cgroup sooner. The original HTTP probe moves toreadinessProbe, keeping app-level health checks separate from mount-level health checks.
The key insight: any process that touches a hung CIFS mount will block. The solution is to never touch the mount to check its health. Instead, check at the network level — if the server is unreachable on TCP port 445, the mount will hang on the next I/O.
The sidecar’s TCP-only check loop
A sidecar container runs nc -z -w 5 <server> 445 every 30 seconds. No filesystem access. On success it writes a Unix timestamp to an emptyDir volume shared with the main container. On failure it logs the event and stops writing.
STATE=up
while true; do
if nc -z -w 5 192.168.1.50 445 2>/dev/null; then
date +%s > /healthcheck/alive
if [ "$STATE" = "down" ]; then
echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [RECOVERY] SMB reachable again"
STATE=up
fi
else
if [ "$STATE" = "up" ]; then
echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [FAILURE] SMB unreachable — liveness probe will fail in ~90s"
STATE=down
else
echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [FAILURE] SMB still unreachable"
fi
fi
sleep 30
done
How does the liveness probe read the heartbeat without blocking?
The main container’s liveness probe reads the timestamp from emptyDir — not from the CIFS mount, so it can never block:
livenessProbe:
exec:
command:
- /bin/sh
- -c
- "test $(( $(date +%s) - $(cat /healthcheck/alive 2>/dev/null || echo 0) )) -lt 90"
initialDelaySeconds: 60
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 3
When the NAS goes offline: the sidecar detects it within 30 seconds, stops writing the heartbeat, logs [FAILURE]. After 90 seconds of stale heartbeat the liveness probe fails. After three failures (another 90 seconds) the pod is killed.
The pod spec needs terminationGracePeriodSeconds: 5. Without it, if the app is already in D-state by the time the pod is killed, kubelet waits the full default 30 seconds for graceful shutdown before sending SIGKILL. With 5 seconds, the container runtime moves to force-kill faster. A D-state process still won’t exit on SIGKILL, but the container runtime can forcibly destroy the cgroup and move on.
Wiring the sidecar into the pod spec
The full pod spec addition, in Pulumi Python:
smb_healthcheck_volume = {
"name": "smb-healthcheck",
"emptyDir": {},
}
smb_healthcheck_sidecar = {
"name": "smb-health-monitor",
"image": "busybox:latest",
"image_pull_policy": "IfNotPresent",
"command": ["/bin/sh", "-c"],
"args": [
"""
STATE=up
while true; do
if nc -z -w 5 192.168.1.50 445 2>/dev/null; then
date +%s > /healthcheck/alive
if [ "$STATE" = "down" ]; then
echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [RECOVERY] SMB reachable again"
STATE=up
fi
else
if [ "$STATE" = "up" ]; then
echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [FAILURE] SMB unreachable — liveness probe will fail in ~90s"
STATE=down
else
echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [FAILURE] SMB still unreachable"
fi
fi
sleep 30
done
"""
],
"volume_mounts": [{"name": "smb-healthcheck", "mount_path": "/healthcheck"}],
}
smb_liveness_probe = {
"exec": {
"command": [
"/bin/sh",
"-c",
"test $(( $(date +%s) - $(cat /healthcheck/alive 2>/dev/null || echo 0) )) -lt 90",
]
},
"initial_delay_seconds": 60,
"period_seconds": 30,
"timeout_seconds": 5,
"failure_threshold": 3,
}
Each deployment using the SMB PVC gets:
smb_healthcheck_volumeappended tovolumessmb_healthcheck_sidecarappended tocontainerssmb_liveness_probeasliveness_probeon the main container- The original HTTP probe moved to
readiness_probe(still useful for app health) termination_grace_period_seconds: 5on the pod spec
Detection latency: worst-case timeline
Worst case, the default setup — a 30-second sidecar interval, a 90-second heartbeat threshold,
periodSeconds: 30,failureThreshold: 3— takes about 3 minutes 5 seconds from NAS outage to pod termination. The sidecar needs up to 30 seconds to notice the TCP failure, then the liveness probe needs three consecutive failures spaced 30 seconds apart before Kubernetes kills the pod, plus a few seconds for termination itself. That bound comes entirely from two knobs — the sidecar’ssleepinterval and the heartbeat-age threshold in the probe — and lowering both together tightens the window. For a homelab that can tolerate an occasional multi-minute outage before recovery, 3 minutes is a reasonable tradeoff against fewer TCP checks hitting the NAS. Workloads with lower tolerance for stuck pods should reduce both knobs and keep the threshold aboveecho_interval × 3.
Worst case: SMB goes down immediately after a sidecar check.
t=0 NAS offline
t=30 sidecar detects failure, stops writing heartbeat
t=60 liveness probe: heartbeat 30s old → passes (< 90s threshold)
t=90 liveness probe: heartbeat 60s old → passes
t=120 liveness probe: heartbeat 90s old → FAIL (1/3)
t=150 liveness probe: FAIL (2/3)
t=180 liveness probe: FAIL (3/3) → pod killed
t=185 pod terminated
About 3 minutes 5 seconds. This is acceptable for a homelab. For tighter bounds, reduce periodSeconds on the sidecar sleep and the liveness probe, and lower the 90-second threshold accordingly.
The backstop: when a pod reaches D-state anyway
If a pod reaches D-state before the sidecar’s 90-second threshold trips — because I/O was already in flight when the NAS disappeared — no liveness probe can rescue it, since the probe process would itself block. The last line of defense is a kernel setting:
hung_task_panic=1withpanic=5on the boot command line. When the kernel’s hung-task watchdog detects a process stuck in D-state past its threshold (120 seconds of hung tasks by default in this setup), it panics the node, which then reboots automatically instead of sitting frozen until someone finds and power-cycles it. This does not prevent the freeze — it converts an indefinite hang requiring physical intervention into automatic recovery within a couple of minutes. Keep it enabled even after deploying the sidecar; the sidecar reduces how often the backstop fires, it does not make the backstop unnecessary.
If the pod is already in D-state before the liveness probe kills it, hung_task_panic=1 with panic=5 in the kernel command line is the last line of defense. The node will reboot automatically after 120 seconds of hung tasks. This doesn’t prevent the freeze — it just recovers from it without requiring physical intervention. The sidecar is meant to prevent pods from reaching D-state in the first place.
Why are network health checks the right abstraction for filesystem resources?
Any resource that can enter an uninterruptible kernel block — CIFS, NFS hard mounts, iSCSI — needs a health check operating at the network layer, never the filesystem layer, because a probe that touches the resource becomes part of the outage instead of detecting it. This sidecar-plus-heartbeat pattern does not fit every case: skip it for privileged containers that must stay alive through mount disruptions, or workloads that tolerate short I/O stalls within the
echo_intervalreconnection window. In those cases,hung_task_panicwith a shortpanictimeout works better as a standalone backstop, since there is no pod to surgically rescue before it reaches D-state. Keep an HTTP readiness probe running alongside this pattern regardless — it catches app-level failures like crashes or deadlocks that have nothing to do with the mount, so mount health and app health fail independently instead of masking each other.
The pattern this establishes: for any filesystem resource that can enter an uninterruptible block — CIFS, NFS hard mounts, iSCSI — the only probe that can’t itself get stuck is one that operates at the network layer. A probe that touches the filesystem becomes part of the problem the moment the resource is unavailable.
Where the sidecar pattern doesn’t apply
The sidecar approach doesn’t apply when the pod itself runs as a privileged container that needs to stay alive through mount disruptions, or when the workload tolerates short I/O stalls and the reconnection window from echo_interval is sufficient to cover them. In those cases, hung_task_panic with a short panic timeout is the more appropriate backstop — accept that the node reboots rather than try to surgically kill pods before they reach D-state.
Tuning the detection window
For tighter latency bounds than the 3-minute worst case above, the two knobs are the sidecar sleep interval and the heartbeat age threshold in the liveness probe. Dropping both to 10 s / 30 s gives a worst-case detection of about 70 seconds, at the cost of three times the TCP connection attempts against the NAS. Below echo_interval × 3 the sidecar can no longer reliably outrace the CIFS echo timeout, so the threshold and echo_interval should be tuned together.
The HTTP readiness probe is still worth keeping on the main container alongside this pattern: it catches app-level failures (crash, deadlock, misconfiguration) that have nothing to do with the mount. Separating the two concerns — mount health via network probe, app health via HTTP — means each probe does exactly one thing and can fail independently.
References
- mount.cifs(8) man page — mount options including
echo_interval,soft - Linux kernel CIFS usage documentation — module parameters vs mount options
- Samba Wiki: LinuxCIFS troubleshooting — D-state hang behavior and known limitations
- Red Hat: Is it possible to make CIFS timeout configurable? —
echo_retrieshistory as module parameter - Arch Linux Forums: Possible workaround for stuck cifs mounts — soft mount real-world behavior
- Kubernetes: Configure Liveness, Readiness and Startup Probes — liveness probe spec reference
- SMB CSI Driver for Kubernetes — CSI driver used in this setup