← All posts
Jun 26, 2026

SMB Mounts in Kubernetes Will Freeze Your Node — Here's What Actually Works

TL;DR: CIFS hard mounts (the default) put processes in TASK_UNINTERRUPTIBLE (D-state) when the server disappears — D-state ignores SIGKILL, so containers can’t be removed, pods get stuck Terminating, and if enough accumulate the hung-task watchdog panics the node; NFS-style options (timeo, retrans, nofail) are invalid arguments for the CIFS module, HTTP liveness probes don’t catch the hang because the app stays responsive until it physically touches the mount, and exec probes doing ls /mount/path block for the same reason. The solution is a sidecar that probes TCP 445 every 30 s via nc (no filesystem access), writes a Unix timestamp to a shared emptyDir on success, and a liveness probe on the main container that reads that timestamp and fails if it’s older than 90 s — so the probe can never block on a hung mount; hung_task_panic=1 with panic=5 remains as a backstop if a pod reaches D-state before the sidecar can intervene.

The right abstraction for detecting a hung CIFS mount is not filesystem access — it’s a network-level TCP check against port 445. Any probe that touches the mount itself will block in D-state alongside the workload it’s supposed to rescue. This post documents every standard approach that fails and why it fails at the kernel level, then builds a sidecar pattern that survives NAS outages without touching the filesystem. The echo_interval mount option does bound reconnection time, but it is not sufficient alone: pods that write continuously will enter D-state within that window, so the sidecar must kill the pod before the echo timeout expires.

kill -9 is supposed to kill anything. It won’t kill a process blocked on a hung CIFS mount. When a NAS goes offline and a pod’s containers are mid-I/O, those processes enter TASK_UNINTERRUPTIBLE — the kernel’s D-state — and become invisible to every signal, including SIGKILL. The container runtime can’t remove them. The pod stays Terminating indefinitely. Multiply that across eight pods in a media server namespace and the kernel’s hung-task watchdog fires, panicking the entire node.

This is the situation this cluster ended up in: a k3s homelab with ~8 pods mounting a 1 TB SMB share via the SMB CSI driver. When the NAS went offline, the node froze and required either a physical power-button hold or hung_task_panic=1 to trigger an automatic reboot. The search for a liveness probe that could catch this before it cascaded led through every obvious option — and most of them make the problem worse, not better.

This is a CIFS-specific hang, not to be confused with the unrelated memory-overcommit and NVMe-cascade freezes on the same cluster covered in k3s Node Freeze: Longhorn iSCSI Stalls and Memory Pressure — same symptom (unresponsive node), different mechanism. Why Your Kubernetes Node Freezes indexes this alongside the other storage failure modes found on the same cluster.

This post is for operators who already know their way around Kubernetes liveness probes and CIFS mounts and have hit (or want to avoid) the hung-mount failure mode. It assumes familiarity with pod specs, Kubernetes health checks, and basic Linux process states.

Table of Contents

Open Table of Contents

Why do D-state processes ignore SIGKILL — and why does that freeze entire nodes?

CIFS hard mounts put I/O-blocked processes into TASK_UNINTERRUPTIBLE (D-state), which ignores every signal, including SIGKILL, and is invisible to the OOM killer. The kernel retries I/O against the unreachable server indefinitely, so the process cannot be forced to exit until the mount recovers or the machine reboots. Kubernetes cannot reap a D-state process, so the pod hangs in Terminating, and if enough pods hit this at once, the kernel’s hung-task watchdog panics the whole node. The soft mount option is documented as the escape hatch, but on real kernels it does not reliably prevent D-state hangs under connection loss — the man page’s promise does not match observed behavior. Treat CIFS hard mounts as capable of freezing an entire node, not only the pod that owns the mount, whenever the backing server can disappear.

CIFS mounts are hard by default — the kernel retries I/O indefinitely until the server comes back. A process blocked on a hard mount sits in TASK_UNINTERRUPTIBLE, which means it ignores all signals, including SIGKILL. The OOM killer can’t touch it. kill -9 does nothing. The process will sit there until the mount comes back or the machine reboots.

This is the same behavior as hard NFS mounts. The CIFS module does have a soft option (it’s nominally the default per the man page), but in practice it doesn’t reliably prevent D-state hangs on most kernels — processes still block on I/O to an unreachable server. This is a longstanding kernel bug: the man page says soft mounts won’t hang, but real-world behavior contradicts that under connection-loss scenarios.

What was tried, and why did it fail?

Every standard Kubernetes health-check approach fails against a hung CIFS mount, for different reasons rooted in the kernel. NFS-style mount options (timeo, retrans, nofail) are not recognized by the CIFS module and cause the mount to fail outright with mount error(22). echo_interval is a real per-mount option that bounds reconnection time to roughly 3 × echo_interval, but echo_retries no longer exists as a tunable on modern kernels — it is absent from /sys/module/cifs/parameters/ on Debian 13 with kernel 6.12. HTTP liveness probes keep passing because the application stays responsive until it touches the mount, so the hang happens after the probe already reported healthy. Exec probes that run ls on the mount path block in D-state exactly like the workload they are meant to protect, adding stuck processes instead of catching the failure. None of these operate below the filesystem layer, so none of them can detect the hang before it happens.

timeo, retrans, nofail — NFS options the CIFS module rejects

These are NFS mount options. The CIFS module doesn’t recognize them. The mount fails immediately:

mount error(22): Invalid argument

echo_interval and echo_retries — what does the module actually support?

The CIFS kernel module uses server echo (keepalive) to detect dead connections. echo_interval is a valid per-mount option — it sets the interval in seconds between echo requests (default: 60s) and can be passed via mountOptions in a Kubernetes PV spec. The reconnection timeout is approximately 3 × echo_interval, so echo_interval=5 gives a ~15-second timeout before the kernel considers the server dead and starts returning errors.

The PV in this cluster uses echo_interval=30, giving a 90-second reconnection timeout — intentionally matched to the 90-second heartbeat threshold in the sidecar liveness probe. When the NAS goes offline, both mechanisms converge on the same window: the sidecar kills the pod before the CIFS echo timeout would kick in for idle connections, and the echo timeout bounds how long in-flight I/O can block before the kernel gives up on the connection.

echo_retries is a different story: it was a module-level parameter on older kernels but is absent from /sys/module/cifs/parameters/ on Debian 13 with kernel 6.12:

$ ls /sys/module/cifs/parameters/
CIFSMaxBufSize  cifs_max_pending  cifs_min_rcv  cifs_min_small
dir_cache_timeout  disable_legacy_dialects  enable_gcm_256
enable_negotiate_signing  enable_oplocks  require_gcm_256
slow_rsp_threshold

Even with echo_interval tuned down, any I/O already in-flight when the server disappears still blocks until the echo timeout elapses. For pods that continuously write (like a media scanner or a database), this window is enough to enter D-state and accumulate.

Why do HTTP liveness probes pass while the app sleeps into D-state?

The *arr stack (Sonarr, Radarr, Bazarr, etc.) exposes HTTP health endpoints. A liveness probe against /ping will restart the container if it stops responding. But if the SMB share goes down while the app is idle, the HTTP endpoint keeps returning 200. The liveness probe passes. The pod looks healthy. The next time the app tries to write a database record or scan a directory on the mount, it enters D-state. By then the probe can’t help — a D-state process won’t respond to the container restart signal either.

HTTP probes are useful for detecting app crashes. They don’t detect hung mounts.

Exec probes doing ls /mount/path block in D-state too

Same problem. ls on a hung CIFS mount enters D-state. The probe itself hangs. Kubernetes marks the probe as timed out after timeoutSeconds, but the blocked ls process stays alive in D-state, accumulating with each probe interval. The pod eventually gets restarted by the failure threshold, but the old D-state processes can prevent clean container removal.

What works: network-level sidecar with emptyDir heartbeat

The fix is a sidecar container that checks TCP port 445 with nc -z -w 5 <server> 445 every 30 seconds — no filesystem access, so it cannot enter D-state alongside the workload. On success it writes a Unix timestamp to a shared emptyDir volume; on failure it logs the event and stops writing. The main container’s liveness probe reads that timestamp file, not the CIFS mount, and fails once it is older than 90 seconds, giving a check that can never block. terminationGracePeriodSeconds: 5 matters too: without it, kubelet waits the default 30 seconds for graceful shutdown before sending SIGKILL, which a D-state process ignores regardless — a short grace period lets the container runtime force-destroy the cgroup sooner. The original HTTP probe moves to readinessProbe, keeping app-level health checks separate from mount-level health checks.

The key insight: any process that touches a hung CIFS mount will block. The solution is to never touch the mount to check its health. Instead, check at the network level — if the server is unreachable on TCP port 445, the mount will hang on the next I/O.

The sidecar’s TCP-only check loop

A sidecar container runs nc -z -w 5 <server> 445 every 30 seconds. No filesystem access. On success it writes a Unix timestamp to an emptyDir volume shared with the main container. On failure it logs the event and stops writing.

STATE=up
while true; do
  if nc -z -w 5 192.168.1.50 445 2>/dev/null; then
    date +%s > /healthcheck/alive
    if [ "$STATE" = "down" ]; then
      echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [RECOVERY] SMB reachable again"
      STATE=up
    fi
  else
    if [ "$STATE" = "up" ]; then
      echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [FAILURE] SMB unreachable — liveness probe will fail in ~90s"
      STATE=down
    else
      echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [FAILURE] SMB still unreachable"
    fi
  fi
  sleep 30
done

How does the liveness probe read the heartbeat without blocking?

The main container’s liveness probe reads the timestamp from emptyDir — not from the CIFS mount, so it can never block:

livenessProbe:
  exec:
    command:
      - /bin/sh
      - -c
      - "test $(( $(date +%s) - $(cat /healthcheck/alive 2>/dev/null || echo 0) )) -lt 90"
  initialDelaySeconds: 60
  periodSeconds: 30
  timeoutSeconds: 5
  failureThreshold: 3

When the NAS goes offline: the sidecar detects it within 30 seconds, stops writing the heartbeat, logs [FAILURE]. After 90 seconds of stale heartbeat the liveness probe fails. After three failures (another 90 seconds) the pod is killed.

The pod spec needs terminationGracePeriodSeconds: 5. Without it, if the app is already in D-state by the time the pod is killed, kubelet waits the full default 30 seconds for graceful shutdown before sending SIGKILL. With 5 seconds, the container runtime moves to force-kill faster. A D-state process still won’t exit on SIGKILL, but the container runtime can forcibly destroy the cgroup and move on.

Wiring the sidecar into the pod spec

The full pod spec addition, in Pulumi Python:

smb_healthcheck_volume = {
    "name": "smb-healthcheck",
    "emptyDir": {},
}

smb_healthcheck_sidecar = {
    "name": "smb-health-monitor",
    "image": "busybox:latest",
    "image_pull_policy": "IfNotPresent",
    "command": ["/bin/sh", "-c"],
    "args": [
        """
STATE=up
while true; do
  if nc -z -w 5 192.168.1.50 445 2>/dev/null; then
    date +%s > /healthcheck/alive
    if [ "$STATE" = "down" ]; then
      echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [RECOVERY] SMB reachable again"
      STATE=up
    fi
  else
    if [ "$STATE" = "up" ]; then
      echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [FAILURE] SMB unreachable — liveness probe will fail in ~90s"
      STATE=down
    else
      echo "$(date -u +%Y-%m-%dT%H:%M:%SZ) [FAILURE] SMB still unreachable"
    fi
  fi
  sleep 30
done
"""
    ],
    "volume_mounts": [{"name": "smb-healthcheck", "mount_path": "/healthcheck"}],
}

smb_liveness_probe = {
    "exec": {
        "command": [
            "/bin/sh",
            "-c",
            "test $(( $(date +%s) - $(cat /healthcheck/alive 2>/dev/null || echo 0) )) -lt 90",
        ]
    },
    "initial_delay_seconds": 60,
    "period_seconds": 30,
    "timeout_seconds": 5,
    "failure_threshold": 3,
}

Each deployment using the SMB PVC gets:

Detection latency: worst-case timeline

Worst case, the default setup — a 30-second sidecar interval, a 90-second heartbeat threshold, periodSeconds: 30, failureThreshold: 3 — takes about 3 minutes 5 seconds from NAS outage to pod termination. The sidecar needs up to 30 seconds to notice the TCP failure, then the liveness probe needs three consecutive failures spaced 30 seconds apart before Kubernetes kills the pod, plus a few seconds for termination itself. That bound comes entirely from two knobs — the sidecar’s sleep interval and the heartbeat-age threshold in the probe — and lowering both together tightens the window. For a homelab that can tolerate an occasional multi-minute outage before recovery, 3 minutes is a reasonable tradeoff against fewer TCP checks hitting the NAS. Workloads with lower tolerance for stuck pods should reduce both knobs and keep the threshold above echo_interval × 3.

Worst case: SMB goes down immediately after a sidecar check.

t=0    NAS offline
t=30   sidecar detects failure, stops writing heartbeat
t=60   liveness probe: heartbeat 30s old → passes (< 90s threshold)
t=90   liveness probe: heartbeat 60s old → passes
t=120  liveness probe: heartbeat 90s old → FAIL (1/3)
t=150  liveness probe: FAIL (2/3)
t=180  liveness probe: FAIL (3/3) → pod killed
t=185  pod terminated

About 3 minutes 5 seconds. This is acceptable for a homelab. For tighter bounds, reduce periodSeconds on the sidecar sleep and the liveness probe, and lower the 90-second threshold accordingly.

The backstop: when a pod reaches D-state anyway

If a pod reaches D-state before the sidecar’s 90-second threshold trips — because I/O was already in flight when the NAS disappeared — no liveness probe can rescue it, since the probe process would itself block. The last line of defense is a kernel setting: hung_task_panic=1 with panic=5 on the boot command line. When the kernel’s hung-task watchdog detects a process stuck in D-state past its threshold (120 seconds of hung tasks by default in this setup), it panics the node, which then reboots automatically instead of sitting frozen until someone finds and power-cycles it. This does not prevent the freeze — it converts an indefinite hang requiring physical intervention into automatic recovery within a couple of minutes. Keep it enabled even after deploying the sidecar; the sidecar reduces how often the backstop fires, it does not make the backstop unnecessary.

If the pod is already in D-state before the liveness probe kills it, hung_task_panic=1 with panic=5 in the kernel command line is the last line of defense. The node will reboot automatically after 120 seconds of hung tasks. This doesn’t prevent the freeze — it just recovers from it without requiring physical intervention. The sidecar is meant to prevent pods from reaching D-state in the first place.

Why are network health checks the right abstraction for filesystem resources?

Any resource that can enter an uninterruptible kernel block — CIFS, NFS hard mounts, iSCSI — needs a health check operating at the network layer, never the filesystem layer, because a probe that touches the resource becomes part of the outage instead of detecting it. This sidecar-plus-heartbeat pattern does not fit every case: skip it for privileged containers that must stay alive through mount disruptions, or workloads that tolerate short I/O stalls within the echo_interval reconnection window. In those cases, hung_task_panic with a short panic timeout works better as a standalone backstop, since there is no pod to surgically rescue before it reaches D-state. Keep an HTTP readiness probe running alongside this pattern regardless — it catches app-level failures like crashes or deadlocks that have nothing to do with the mount, so mount health and app health fail independently instead of masking each other.

The pattern this establishes: for any filesystem resource that can enter an uninterruptible block — CIFS, NFS hard mounts, iSCSI — the only probe that can’t itself get stuck is one that operates at the network layer. A probe that touches the filesystem becomes part of the problem the moment the resource is unavailable.

Where the sidecar pattern doesn’t apply

The sidecar approach doesn’t apply when the pod itself runs as a privileged container that needs to stay alive through mount disruptions, or when the workload tolerates short I/O stalls and the reconnection window from echo_interval is sufficient to cover them. In those cases, hung_task_panic with a short panic timeout is the more appropriate backstop — accept that the node reboots rather than try to surgically kill pods before they reach D-state.

Tuning the detection window

For tighter latency bounds than the 3-minute worst case above, the two knobs are the sidecar sleep interval and the heartbeat age threshold in the liveness probe. Dropping both to 10 s / 30 s gives a worst-case detection of about 70 seconds, at the cost of three times the TCP connection attempts against the NAS. Below echo_interval × 3 the sidecar can no longer reliably outrace the CIFS echo timeout, so the threshold and echo_interval should be tuned together.

The HTTP readiness probe is still worth keeping on the main container alongside this pattern: it catches app-level failures (crash, deadlock, misconfiguration) that have nothing to do with the mount. Separating the two concerns — mount health via network probe, app health via HTTP — means each probe does exactly one thing and can fail independently.

References


Breno Zanato Detomini
Breno Zanato Detomini

Embedded systems and network engineer based in Brazil.

← All posts