nfs-ha

This project formalizes two existing scripts for disk health monitoring and scheduled disk mirroring on Wintermute, a single NFS server with two local disks. Its purpose is to notify the operator when either disk starts failing, giving them time to act before a second disk failure causes data loss.

Background

Wintermute was rebuilt, but the automated disk checks and synchronization were not restored. The omission was discovered on October 8, 2026, and a manual sync was started. This project should make the setup straightforward to reinstall and verify after future rebuilds.

The existing scripts are in /opt/dev/utils/bin:

Disk layout

Role Mount point Filesystem UUID Existing sync path
Primary /exports/disk1 824971c7-d263-4b6a-a562-3452c3482ece /exports/old
Recovery copy /exports/disk2 b14c1ec0-df91-4f3d-a990-b0c820b81a9c /exports/new

Both filesystems are ext4 and are mounted by UUID through /etc/fstab. Synchronization runs from the primary to the recovery copy.

Agreed scope and behavior

The recovery disk is a mirror: deletions on the primary propagate to it. It does not retain earlier versions. Monitoring and mirroring reduce the risk of data loss but cannot guarantee protection against two sudden disk failures.

Manual response and boundaries

Disk replacement and recovery are manual. On receiving a disk failure alert, the operator will power down the server, disconnect the drives, obtain one or two replacements as needed, reconnect the appropriate drives, and synchronize the replacement storage.

NFS configuration, NFS failover, and automatic recovery are outside this project’s scope.

Decisions still to make

No implementation or scheduling changes have been made as part of this initial scope discussion.

Development

Coding guidelines