datadog2 min read

Curated summary

The trouble with mounting

Read original(opens in new tab)

Datadog found that some agents stopped reporting all metrics because they became stuck in an unkillable state during disk checks. The root cause was os.statvfs, whose glibc implementation can hang while inspecting NFS mounts configured with hard-mount behavior. Since agents run in unpredictable customer environments, Datadog isolated the call in a separate thread and allowed the main process to continue after a timeout.

Detecting the Hang

  • Customers reported gaps across every metric, indicating that the agent—not an individual check—had stopped functioning.
  • Logs showed the agent sometimes hung without producing an error.
  • A watchdog failed to terminate it because the process was stuck in an unkillable system call.
  • Developer-mode timing data identified os.statvfs as the consistently slow operation.

How NFS Causes Unkillable Processes

  • os.statvfs calls the Linux statvfs function through CPython and glibc.
  • statvfs can hang when examining a remote directory mounted through NFS.
  • NFS hard mounts retry indefinitely and do not time out system calls.
  • Soft mounts eventually return an error, while the intr option allows interruption of the calling process.
  • Hard mounts may be appropriate when reads and writes must eventually succeed, but they are risky with unreliable NFS connections because they are the default in many configurations.

The /proc/mounts Complication

  • Glibc’s statvfs implementation checks each directory listed in /proc/mounts until it finds the requested mount.
  • Consequently, a disconnected NFS mount can block statvfs even when the agent is checking a different filesystem.
  • This made changing NFS mount options impractical as a universal fix because Datadog cannot control customers’ system configurations.

Datadog’s Workaround

  • The agent now runs statvfs on a separate thread.
  • If the call exceeds a timeout, the main agent thread continues operating.
  • This approach avoids total metric loss across heterogeneous environments.
  • The trade-off is a modest increase in memory usage on systems with hard-mounted NFS volumes.

The practical lesson is to treat filesystem statistics as potentially blocking operations, especially in environments with NFS. Isolating such calls behind timeouts provides more reliable monitoring than assuming system calls will always return promptly.

Continue with another curated summary.