Both hosts were powered off and brought back from a cold start, which is the first real test of everything built on 2026-08-30. nas: both pools imported from /etc/zfs/zpool.cache, all instances autostarted, NFS exports republished, 0 failed units. nuc: /dev/dri present with 0 i915 warnings (booted with the display unplugged), NFS auto-remounted, VAAPI transcode at 9.5x realtime. Records the ordering that actually cleared tank's inherited <metadata>:<0x0> and <0x3d>: a scrub alone found 0 errors and repaired 0B but left them, and a plain `zpool clear` afterwards did not drop them — ZFS flushes the persistent error log on a scrub run *after* the clear. That matters beyond tidiness, because while those entries stand `zpool status -x` reports the pool unhealthy forever and zpool-health.sh cannot signal anything new. Two kiosk corrections, both from observed behaviour: - the Pioneer DAC being switched off is the most likely cause of "video, no sound" — asound.conf pins the ALSA default to it by card name, so `default` fails to open outright and mpv falls back to null silently. Adds the one-line aplay check. - hotplugging the display makes cage exit once and Restart=on-failure recovers it ~5s later. Do NOT restart it by hand; check ActiveEnterTimestamp against the hotplug time first. Also flags that the OS mirror is still untested with a disk physically unplugged — it is a guess until then. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
35 KiB
nas — build procedure
New storage host on the LAN (192.168.0.4), built 2026-08 from a
Supermicro A1SAi-2750F. It exists to solve one specific problem:
usb4t-dropouts.md ruled out the cable and
USB power management and left the JMicron 152d:0578 bridge as the
cause of 61 disconnects in 30 days. The durable fix named there is a
direct SATA connection — this box provides six of them.
Why it matters beyond the media library: usb4t is the intended home
of the ks4 off-site replicas, and "the pool holding the off-site copy
of ks4 must not be the least reliable device in the setup". That makes
this build a release gate for the ks2 decommission
(ks2/plan.md, deadline Sep 30, 2026).
It also takes over transmission-bt from nuc, so nuc keeps only what
needs its iGPU.
Hardware
- Supermicro A1SAi-2750F mini-ITX, Intel Atom C2750 (8 cores, 2.4 GHz Silvermont, 20 W SoC), AES-NI, no AVX
- RAM: 2× 4 GB DDR3-1600 ECC SO-DIMM fitted = 8 GB (
Single-bit ECCconfirmed), in DIMMA1/DIMMB1; 2 slots free. Board takes 32 GB officially, 64 GB with 16 GB modules. ⚠️ 8 GB is modest for a 3.6 TB pool — ARC lands around 4 GB. Fine for streaming and replication (neither benefits much from cache), but the first thing to raise if metadata-heavy operations feel slow. - SATA: 2× SATA3 + 4× SATA2 (6 total)
- BIOS 2.2 (2019-11-22) as shipped by the RMA
- NIC: 4× GbE (Intel i354) + dedicated IPMI LAN
- Video: ASPEED AST2400 BMC only — there is no iGPU. Avoton is a
headless server SoC;
/dev/driis empty. That is whyjellyfin-serverstays on nuc (§8).
⚠️ AVR54 — already handled
The C2750 is on the list of Atom C2000 parts affected by Intel's
AVR54 erratum: the SoC's LPC_CLKOUT0/1 signals degrade and stop,
after which the board never boots again — typically after ~18 months of
power-on, i.e. exactly an always-on duty cycle. This board was RMA'd
by Supermicro for that issue and replaced, so it carries the fix
(C0 stepping or the LPC pull-up rework). Recorded here so a dead C2000
board is not re-diagnosed from scratch later.
Disk plan
| Port | Device | Role |
|---|---|---|
| SATA3-0 | Intel SSDSC2CT120A3 (120 GB) |
md mirror + incus pool |
| SATA3-1 | Toshiba Q300. (120 GB) |
md mirror + incus pool |
| SATA2-0 | WD Red 4 TB (moved off the USB enclosure) | pool tank |
| SATA2-1/2/3 | free | a second 4 TB to mirror tank, later |
The SSDs get the SATA3 ports because they are the only devices that can use them: both negotiate 6 Gb/s and do ~450–500 MB/s, while the WD Red tops out near 180 MB/s and cannot saturate SATA2's ~270 MB/s. A future SLOG would also be fine on SATA2 — it is latency-bound on small sync writes, not bandwidth-bound.
⚠️ Disconnect the 4 TB before partitioning the SSDs. It carries the media library and nuc's replicas, nothing in the OS install needs it, and it keeps the two 120 GB disks unambiguous in the installer's list. Reconnect it before §5b.
Device letters shift with enumeration order and mean nothing — mdraid
assembles from superblock UUIDs, ZFS imports by GUID, fstab and GRUB use
UUIDs. Identify disks by model and serial
(lsblk -o NAME,SIZE,MODEL,SERIAL), never by letter: both SSDs are
120 GB, so size alone does not tell them apart.
SSD health measured 2026-08-30 (both read over the JMicron bridge on
nuc, smartctl -d sat), before deployment:
| Intel 330 (25 nm MLC) | Toshiba Q300 (15 nm TLC) | |
|---|---|---|
| host writes | 16.0 TiB | 7.5 TiB |
| endurance consumed | 0 % (Media_Wearout_Indicator 100) |
7 % (Percentage Used Endurance Indicator) |
| power-on hours | unreadable — attr 9 decodes to 914,563 h on this family | 3,877 |
| power cycles | 98 | 432 |
| defects | 0 reallocated / program-fail / erase-fail | 0 reported uncorrectable |
| SMART error log | not supported | supported, clean |
| device statistics log | absent | full ACS-2 set |
| short self-test | passed | passed |
| interface CRC errors | n/a | 16 (baseline — watch for growth) |
Both are healthy and far from wear-out; at OS-disk write rates endurance
is not the binding constraint for either. They were previously a
matched pair — identical layouts, bpool / rpool / ubuntu:0
labels from an Ubuntu ZFS-on-root mirror — which is why they go back
into a mirror here.
⚠️ The Intel is effectively unmonitorable: no error log, no device
statistics, no temperature, no usable hours counter. smartd can watch
the Toshiba properly and can only ask the Intel "are you still there".
Expect the Intel to be found dead rather than found degrading. That
asymmetry is the reason for the mirror.
tank is a single vdev — accepted
The 4 TB holds the media library, transmission's downloads and the ks4 replicas, with no redundancy. A single vdev can detect corruption but only self-heal metadata, not data — as seen in the usb4t incident.
Decision 2026-08-30: accepted, nothing on tank is irreplaceable.
Media and instances are re-fetchable from their sources; the ks4
replicas are leg 3 of ks4's 3-2-1 (local sdb5 + restic/S3 remain —
backup-strategy.md). Losing tank costs time,
not data.
Three SATA ports stay free, so zpool attach tank <existing> <new>
turns it into a mirror whenever a spare 4 TB turns up. Not a
prerequisite for anything. Scrub weekly regardless — on a single vdev
the scrub is the only thing that tells you a file has rotted.
1. IPMI and BIOS first
The AST2400 stack is old and has known vulnerabilities.
- Decision 2026-08-30: the BMC is never cabled. The dedicated IPMI port stays unplugged, so the AST2400's default credentials and its known vulnerabilities are not reachable from anything. This is the simplest correct answer for a box that sits on a flat home LAN — no management VLAN needed, nothing to harden, nothing to patch.
- ⚠️ Consequence: there is no out-of-band console. A boot that fails before sshd needs a physical monitor and keyboard. Worth knowing before changing anything that affects booting (GRUB, the md arrays, fstab).
- BIOS: enable restore-on-AC-loss, enable C-states, and disable the three unused i354 NICs (each costs about a watt).
Switching this board to UEFI
There is no "Boot Mode Select" entry in this BIOS — that option exists on later Supermicro generations, not here. What works (verified 2026-08-30):
- CSM → Disabled
- All OpROM policies → UEFI (storage, video and network)
Setting the video OpROM to UEFI is safe despite the console being the AST2400 BMC framebuffer — output survives, both on the BMC console and over IPMI KVM.
⚠️ Confirm the mode before you partition anything, from the
installer (Ctrl+Alt+F2):
ls /sys/firmware/efi # directory exists = UEFI. Missing = legacy
This single check is what makes the difference between a working install and an afternoon lost — see the trap in §3.
2. Prepare the SSDs
Both report ATA Security is: Disabled, NOT FROZEN, so a real secure
erase is available — do that rather than just repartitioning. It
restores the full spare-block pool on ten-year-old NAND.
# per disk, from a live system where the disk is NOT the running OS
hdparm --user-master u --security-set-pass Eins /dev/sdX
hdparm --user-master u --security-erase Eins /dev/sdX
hdparm -I /dev/sdX | grep -A2 Security # expect "not enabled"
Identical GPT layout on both, sized so ~27 GB (22 %) stays unallocated as over-provisioning:
| Part | Size | Type | Device | Use |
|---|---|---|---|---|
sdX1 |
1 GB | EFI System Partition | — | /boot/efi (two independent ESPs) |
sdX2 |
2 GB | Linux RAID | md0 |
/boot ext4 |
sdX3 |
24 GB | Linux RAID | md1 |
/ ext4 |
sdX4 |
2 GB | Linux RAID | md2 |
swap |
sdX5 |
64 GB | Solaris root (bf00) | — | zpool incus (mirror) |
| — | ~27 GB | unallocated | — | over-provisioning |
- ⚠️ The ESP is deliberately not a RAID1 array. ks4 has
md1 → /boot/efibecause OVH's installer builds it with mdadm metadata 1.0 (superblock at the end, so firmware still sees plain FAT).debian-installeronly creates metadata 1.2 arrays, whose superblock sits at the start and makes the ESP unreadable to firmware. So: one plain ESP per disk, only one mounted, the second filled by hand (§3). Deviation from ks4 is intentional. - A legacy-BIOS variant of this layout was tried first — 1 MB
bios_grubinstead of the ESP, which makes the mirror simpler (grub-installto both disks, nothing to keep in sync). It was abandoned because this board's firmware has no way to prefer legacy targets once CSM is off, and it kept falling through to the UEFI shell. Recorded so it is not retried: UEFI is the working path here. /at 24 GB matches the other hosts (nuc 46 GB, ks4 40 GB) — this box has no desktop and no container roots on/.- Mixing md partitions and a ZFS partition on the same disks is exactly
what ks4 does (
md1/2/3+ ZFS onsda5/sdb5).
Why not full root-on-ZFS
Decision 2026-08-30. These SSDs previously ran Ubuntu 20.04's
experimental ZFS-root installer (hence the leftover bpool / rpool
labels), so the option was on the table. Rejected because:
- Debian's installer cannot do it. ZFS is CDDL, shipped only in
contribaszfs-dkms;debian-installercan neither partition nor boot from ZFS. Root-on-ZFS means the manual OpenZFS Debian Trixie HOWTO — ~60 steps from a live ISO. Per the repo convention the doc is the rebuild procedure, and that is a bad thing to be executing during an actual failure. - DKMS failure mode.
zfs-dkmsrebuilds on every kernel upgrade. If that build fails, root-on-ZFS means the box does not boot; with an ext4 md root it boots normally and only the pools are missing — recoverable over IPMI with a shell. - GRUB's ZFS support lags OpenZFS, which is why every root-on-ZFS guide
needs a separate feature-limited
bpool;zpool upgrade bpoolis a known way to make a machine unbootable. - Ubuntu's version of this is a dead end anyway: the installer option
was nearly dropped in 22.04 and
zsys, which made boot environments useful, is abandoned.
What root-on-ZFS would buy — snapshot and roll back a bad upgrade — is
already covered where the state actually lives: container roots get
incus snapshots plus nightly replication (§9). The host is 24 GB of
packages reproducible from this file. Accepted cost: no pre-upgrade
rollback of the host itself (etckeeper covers /etc if wanted).
3. Install Debian 13 (trixie)
Netinst ISO (burned 2026-08-30, sha256 65273bee…664e7, verified
against cdimage.debian.org/debian-cd/13.6.0/amd64/iso-cd/SHA256SUMS).
- Manual partitioning per the table above:
sdX1as ESP,sdX2..4as RAID1 members (three arrays),sdX5left untouched. - Tasks: SSH server + standard utilities only. No desktop — there is no GPU and the console is a BMC framebuffer.
- Sources:
main contrib non-free-firmware(contribis required byzfs-dkms; the installer does not offer it — add it after first boot).
Confirm d-i picked the right bootloader once installed:
dpkg -l | grep grub-efi — grub-efi-amd64, not grub-pc.
⚠️ Trap: installer boot mode decides the bootloader (hit 2026-08-30)
debian-installer chooses grub-pc or grub-efi-amd64 from how the
installer itself booted, not from what the disks look like. Booting the
USB stick in legacy mode while the firmware prefers UEFI produces:
- d-i installs
grub-pc, targeting the MBR; - on a GPT disk that needs a 1 MB
bios_grubpartition — absent here, sogrub-installfails, easy to click past; - the firmware then tries the SSDs as UEFI targets, finds no
.efibinary, and drops to the UEFI shell.
Nothing is corrupt; the halves simply disagree. Symptoms and checks:
[ -d /sys/firmware/efi ] && echo UEFI || echo legacy # in the installer
lsblk -no PTTYPE,PARTTYPENAME /dev/sdX # gpt + "EFI System"?
dd if=/dev/sdX bs=440 count=1 2>/dev/null | od -c | head -3 # all \0 = no boot code
dpkg -l | grep -E '^ii.*grub-(pc|efi)' # which flavour got installed
The board's hybrid ISO offers both paths, so the F11 boot menu usually
lists the stick twice — picking the UEFI: entry avoids the whole
thing. Checking /sys/firmware/efi before partitioning is the one step
that prevents it.
Second ESP — do this before trusting the mirror
The installer populates only the ESP it mounted. Until the second one is written, losing that disk means the box does not boot, mirror or no mirror.
⚠️ Use /dev/disk/by-id/, never /dev/sdX. Reconnecting the 4 TB
after the install shifts every letter — observed 2026-08-30: the HDD on
SATA2-0 takes sda even with the SSDs on SATA3, because the SATA2
controller enumerates first on this SoC. A bare /dev/sdb1 written
during the install then points at a different disk, and here that
would mean reformatting the ESP the system actually boots from.
ESP2=/dev/disk/by-id/ata-TOSHIBA_Q300._36OB318OK1KU-part1 # the one NOT at /boot/efi
mkfs.vfat -F32 "$ESP2"
mkdir -p /boot/efi2 && mount "$ESP2" /boot/efi2
grub-install --target=x86_64-efi --efi-directory=/boot/efi2 \
--bootloader-id=debian-b --recheck
efibootmgr -v # expect: debian, debian-b
echo "UUID=$(blkid -s UUID -o value $ESP2) /boot/efi2 vfat umask=0077 0 1" >> /etc/fstab
mount -a && findmnt /boot/efi2
Re-run the grub-install after any GRUB or kernel change — the second
ESP is not kept in sync automatically. Verify by pulling one disk
and booting; the checklist item exists because an untested mirror is a
guess, and this is the component most likely to be silently wrong.
Do not fix the letter ordering by moving cables: the only arrangement
that makes the SSDs sda/sdb puts them on SATA2 and the HDD on SATA3,
which caps the only devices that can use 6 Gb/s and gives the bandwidth
to a disk that tops out near 180 MB/s.
4. Base system
apt update && apt full-upgrade -y
apt install -y \
linux-headers-amd64 zfs-dkms zfsutils-linux zfs-zed \
mdadm smartmontools nfs-kernel-server \
msmtp msmtp-mta bsd-mailx \
curl vim htop ripgrep sysstat dmidecode pciutils usbutils \
stress-ng fio
Static network — /etc/network/interfaces (ifupdown, matching nuc):
source /etc/network/interfaces.d/*
auto lo
iface lo inet loopback
allow-hotplug enp0s20f0
iface enp0s20f0 inet static
address 192.168.0.4
netmask 255.255.255.0
gateway 192.168.0.2
dns-nameservers 1.1.1.1 9.9.9.9
(Gateway is 192.168.0.2 — .1 is gone since 2026-08. The host
uses public resolvers, never blocky, to avoid a bootstrap loop.
Interface name is a guess until the board is up — check ip -br link.)
Restore /root/.ssh/authorized_keys (incl. id_rsa_claude.pub) and
timedatectl set-timezone Europe/Paris.
5. Pools
5a. incus — SSD mirror
zpool create -o ashift=12 \
-O compression=zstd -O atime=off -O xattr=sa -O acltype=posixacl \
incus mirror \
/dev/disk/by-id/ata-INTEL_SSDSC2CT120A3_CVMP250400ES120BGN-part5 \
/dev/disk/by-id/ata-TOSHIBA_Q300._36OB318OK1KU-part5
zpool set autotrim=on incus
Weekly scrubs come from the packaged systemd timers rather than cron:
systemctl enable --now zfs-scrub-weekly@tank.timer zfs-scrub-weekly@incus.timer
systemctl list-timers 'zfs-scrub*'
autotrim matters on ten-year-old NAND — it is what keeps the
unallocated 22 % actually available to the controller as spare.
5b. tank — move the 4 TB off USB onto SATA
The risky step. The pool holds usb4t/backup/nuc (nuc's replicas),
usb4t/backup/ks4 (empty, awaiting the FTTH seed) and usb4t/media.
⚠️ Between export and import, nuc has no replica target and Jellyfin has no media. Plan a maintenance window and disable nuc's 03:30 cron first, so it fails loudly rather than half-running.
# --- on nuc, first ---
zpool scrub usb4t # start clean; wait for it
zpool status usb4t
zpool export usb4t
Move the disk to SATA2-0 (reconnect it now if you unplugged it for the install), then:
# --- on nas ---
zpool import # confirm it is seen
zpool import usb4t tank # rename: it is not USB any more
zpool set cachefile=/etc/zfs/zpool.cache tank
zpool status tank
Properties survive from creation (ashift=12, compression=zstd,
atime=off, xattr=sa, acltype=posixacl). Re-point the mountpoints
and add the two new backup datasets:
zfs set mountpoint=/export/media tank/media # NFSv4 export root (§8)
zfs set mountpoint=none tank/backup
zfs list -o name,used,avail,mountpoint
Resulting layout:
incus mirror, 2× SSD — nas's own container roots
tank 4 TB, single vdev
├── tank/media → /export/media NFS ro → nuc; local device → transmission-bt
└── tank/backup
├── tank/backup/nuc → incus pool `nucbackup` (nuc pushes here)
└── tank/backup/ks4 → incus pool `ks4backup` (nas pulls from ks4)
nas's own replicas live on **nuc** (`data/backup/nas` → pool `nasbackup`),
not here — see §9.
Backup pools are named after the source host, matching ks4/incus-copy.md. nas's own instances are not backed up here — they cross-replicate to nuc (§9b), so neither host's instances depend on that host surviving.
Once tank has run a week on SATA with zero pool suspensions and no
CRC errors, the usb4t-dropouts gate is cleared — record that in
ks2/plan.md. Day 1 was clean (2026-08-31).
Clearing the inherited <metadata> errors — order matters
The pool imported carrying <metadata>:<0x0> and <metadata>:<0x3d> from
the 2026-08-29 USB dropout. A scrub found 0 errors and repaired 0B,
yet the entries stayed, and a plain zpool clear afterwards did not drop
them either. ZFS flushes its persistent error log on a scrub that runs
after the clear — so the working order is:
zpool clear tank
zpool scrub tank # this is the run that flushes the log
Result 2026-08-31: scrub repaired 0B in 02:34:24 with 0 errors,
errors: No known data errors, all pools are healthy. They were
artefacts of interrupted writes, not corruption — matching the
2026-08-28 incident.
⚠️ This matters for monitoring, not just tidiness: while those entries
stand, zpool status -x reports the pool unhealthy permanently, so
zpool-health.sh sits in the alarm state and cannot signal a new
problem. Clear them before trusting the watchdog.
6. Incus
Same Zabbly stable repo as nuc and ks4:
mkdir -p /etc/apt/keyrings
curl -fsSL https://pkgs.zabbly.com/key.asc -o /etc/apt/keyrings/zabbly.asc
cat > /etc/apt/sources.list.d/zabbly-incus-stable.sources <<EOF
Enabled: yes
Types: deb
URIs: https://pkgs.zabbly.com/incus/stable
Suites: trixie
Components: main
Architectures: amd64
Signed-By: /etc/apt/keyrings/zabbly.asc
EOF
apt update && apt install -y incus
macvlan like nuc, so instances get real LAN addresses — transmission-bt
keeps 192.168.0.7 when it moves:
cat <<EOF | incus admin init --preseed
config:
core.https_address: :8443
storage_pools:
- name: incus
driver: zfs
config:
source: incus
networks:
- name: macvlan
type: macvlan
config:
parent: enp0s20f0
profiles:
- name: default
devices:
eth0: {name: eth0, network: macvlan, type: nic}
root: {path: /, pool: incus, type: disk}
EOF
incus profile set default snapshots.schedule="0 3 * * *" snapshots.expiry=7d
Same macvlan quirk as nuc: the host cannot talk to its own instances,
and vice versa. Test container services from another LAN host or from
inside the container, never from nas.
Backup pools and the replica project:
incus storage create nucbackup zfs source=tank/backup/nuc
incus storage create ks4backup zfs source=tank/backup/ks4
incus project create backup -c features.images=false -c features.profiles=false
⚠️ tank/backup/nuc already contains nuc's replicas — they came
across with the pool. Re-register them so refreshes stay
ZFS-incremental instead of re-sending everything (the homeassistant VM
alone is a 50 GiB volume):
incus admin recover # point it at pool nucbackup; project backup
incus list --project backup
Same call nuc-install.md §5 uses after a rebuild. If recover is
skipped, the first push in §9a silently becomes a full re-send of every
nuc instance.
Order incus after the ZFS mounts, as on nuc —
/etc/systemd/system/incus.service.d/after-zfs.conf:
[Unit]
After=zfs-mount.service zfs.target
7. Move transmission-bt from nuc
Its WireGuard tunnel is entirely inside the container (wg0,
10.8.0.21, wg-quick@wg0, BindsTo= on the daemon), so the container
carries its own keys and ks4 needs no change at all — the peer stays
10.8.0.21/32. The kill-switch /32 route points at the gateway
192.168.0.2, which is the same from here.
# on nuc — remote already added in §9a
incus stop transmission-bt
incus move transmission-bt nas: --storage incus
Then on nas, re-point the media device at the local dataset — this is a
plain shift=true device again, because the data is local ZFS:
incus config device remove transmission-bt media
incus config device add transmission-bt media disk \
source=/export/media path=/media shift=true
incus config set transmission-bt boot.autostart=true
incus start transmission-bt
Verify the tunnel and the kill switch before trusting it:
incus exec transmission-bt -- wg show
incus exec transmission-bt -- curl -s ifconfig.me # must print 193.70.35.17
incus exec transmission-bt -- ip route # must have NO default route
⚠️ The watch-folder workflow moves with it —
transmission-bt.md says
scp some.torrent root@192.168.0.3:/srv/media/.watchdir/; it is now
root@192.168.0.4:/export/media/.watchdir/.
8. Media over NFS — jellyfin-server stays on nuc
jellyfin-server needs the Alder Lake-N iGPU for QSV/VAAPI; the C2750
has no render device at all, and software transcoding on Silvermont
manages 1–2 concurrent 1080p H.264 streams at best. So the container
stays on nuc and reaches the library over NFS, read-only —
transmission-bt is the only writer and it now lives here.
# --- nas: NFSv4 export, read-only, nuc only ---
cat >> /etc/exports <<'EOF'
/export 192.168.0.3(ro,fsid=0,crossmnt,no_subtree_check)
/export/media 192.168.0.3(ro,no_subtree_check,all_squash,anonuid=65534,anongid=65534)
EOF
exportfs -ra && exportfs -v
# --- nuc: mount at the SAME path, so jellyfin-server.md still applies ---
mkdir -p /srv/media
echo '192.168.0.4:/media /srv/media nfs4 ro,_netdev,soft,timeo=100,retrans=3 0 0' >> /etc/fstab
mount /srv/media && ls /srv/media
The container device changes only in losing the shift — per
nuc/jellyfin-server.md's own troubleshooting
note, "on CIFS files are world-readable synthetic ownership, enough for
a read-only library"; the same holds for NFS with all_squash:
incus stop jellyfin-server # shift cannot be hot-applied
incus config device set jellyfin-server media shift=false
incus config device set jellyfin-server media readonly=true
incus start jellyfin-server
incus exec jellyfin-server -- ls /media # must list the library
Requirements for that to work: media files must be world-readable
(find /export/media -type f ! -perm -o=r), and transmission must keep
creating them that way (it does — umask/0775 per its doc).
⚠️ Boot ordering on nuc: incus is already ordered after
zfs-mount.service; add remote-fs.target to that drop-in, or
jellyfin-server starts against an empty mountpoint and shows an empty
library.
⚠️ soft mount is deliberate: a hung NAS should fail Jellyfin's reads,
not wedge nuc's processes in uninterruptible sleep the way the suspended
usb4t pool did.
9. Backup legs
The driver is unchanged —
incus-copy.sh,
deployed to /root/scripts as everywhere else.
9a. nuc's instances -> nas (systemd timer, not cron)
⚠️ nuc is an on-demand media box (see nuc/README.md):
since 2026-08-30 it only runs when watching Jellyfin or using the Spotify
kiosk, so it is usually powered off at 03:30. Cron cannot catch up a
missed window; a systemd timer with Persistent=true runs the job
shortly after the next boot. That is the whole reason this is a timer.
/etc/systemd/system/incus-copy.service on nuc:
[Unit]
Description=Replicate all nuc instances to nas (project backup, pool nucbackup)
After=network-online.target incus.service
Wants=network-online.target
Requires=incus.service
[Service]
Type=oneshot
ExecStart=/bin/sh -c "/root/scripts/incus-copy.sh -d nas -m push -s nucbackup -p backup >> /var/log/incus-copy.log 2>&1"
/etc/systemd/system/incus-copy.timer:
[Timer]
OnCalendar=03:30
Persistent=true
RandomizedDelaySec=300
[Install]
WantedBy=timers.target
systemctl enable --now incus-copy.timer && systemctl list-timers incus-copy
⚠️ -p backup is not optional. Without it the replicas land in
default on nas and collide with nas's live instances — both hosts are
on the same macvlan LAN and the replicas carry the same static IPs.
9b. nas's own instances -> local pool nasbackup
Decision 2026-08-30: local, not cross-replicated. nuc is powered off
most of the time, so it is not a usable backup target — a nightly push to
it would fail noisily and, once mail works, alarm every morning. The
replicas instead go to tank/backup/nas, which is a different pool
from the instances themselves (incus, the SSD mirror), so it survives
losing that mirror. All three instances here are rebuildable from their
docs, so same-host is proportionate.
zfs create tank/backup/nas
zfs set mountpoint=legacy tank/backup/nas # see the trap below
incus storage create nasbackup zfs source=tank/backup/nas
# nas, /etc/cron.d/incus-copy
0 4 * * * root /root/scripts/incus-copy.sh -p backup -s nasbackup >> /var/log/incus-copy.log 2>&1
⚠️ Trap: incus storage create hangs forever on a mountpoint=none
dataset. tank/backup is set to mountpoint=none, so any child
created afterwards inherits it, and the pool create then blocks with no
error and no entry in incus operation list — it looks exactly like I/O
contention (a scrub was running, which sent me down that path for 20
minutes). Set the child to legacy to match its siblings first, and it
completes instantly.
9c. Cleanup on nuc after the pool move
Exporting usb4t leaves nuc with an incus storage pool whose backing
dataset is gone, plus replica records in the backup project pointing
at it. Remove them once §6's incus admin recover has re-registered the
same volumes on nas — verify there first, then delete here:
# on nas: confirm the replicas are registered
incus list --project backup -c ns -f csv
# on nuc: only then
incus delete --project backup --force <each-replica>
incus storage delete nucbackup
zpool status # only `data` should remain
9d. ks4 pull leg moves from nuc to nas
The leg ks2/nas-seed.md prepared. The target pool
moves, so the host WireGuard tunnel moves too — nuc drops wg-ks4
entirely once this works (its only other tunnel user, transmission-bt,
is now here and carries its own).
- ks4: add nas as a peer on the
wireguardcontainer —wg set wg0 peer <nas-pubkey> allowed-ips 10.8.0.22/32 && wg-quick save wg0. The existing ufw rule (allow in on incusbr0 from 192.168.1.18 to any port 8443 proto tcp) already covers it: traffic arrives masqueraded as the WG container whichever peer sent it. - nas:
/etc/wireguard/wg-ks4.confmodelled on nuc's —Address = 10.8.0.22/32, peer pubkeyTVs6d7…,Endpoint = 193.70.35.17:51845,AllowedIPs = 10.8.0.0/24, 192.168.1.1/32, keepalive 25 — thensystemctl enable --now wg-quick@wg-ks4. incus remote add ks4 https://192.168.1.1:8443 --accept-certificate --token '…'(cross-check the fingerprint against the token).- nuc, after verification:
systemctl disable --now wg-quick@wg-ks4, remove/etc/wireguard/wg-ks4.conf, and drop the10.8.0.20/32peer on ks4.
Seed manually in tmux — the first pass is ~1.75 TiB over the home downlink and only makes sense once FTTH is up:
/root/scripts/incus-copy.sh -r ks4 -s ks4backup 2>&1 | tee -a /var/log/incus-copy-ks4.log
Then test-restore one instance before ticking the gate in ks2/plan.md.
Resulting schedule
| When | Host | What |
|---|---|---|
| 03:00 | nuc, nas | instance snapshots (profile) |
| 03:30 | nuc | push all instances → nas:nucbackup (systemd timer, Persistent=true — catches up after boot) |
| 04:00 | nas | local copy of nas instances → nasbackup (tank/backup/nas) |
| 05:00 | nas | pull ks4:* → ks4backup (after FTTH) |
| 05:30 | nuc | apt upgrade all containers |
| Mon ~00:12 | nas | zfs-scrub-weekly@tank.timer / @incus.timer (systemd, not cron) |
Staggered around ks4's own 01:00 / 05:00 jobs.
⚠️ Keep snapshots.schedule on every source. A refresh with no common
snapshot degrades to a full re-send — the failure mode that cost
933 G on ks4 (ks4/local-backup-cron.md).
The homeassistant VM on nuc re-sends its whole volume without them.
Note incus-copy.sh takes a global flock: an overrunning ks4 pull
aborts that night's other run loudly rather than racing it.
10. Monitoring
Do not repeat nuc's 11-day blind spot (usb4t-dropouts.md — zed was running, but the host had no MTA).
Status: live since 2026-08-30, verified end to end (smtpstatus=250).
- msmtp, same shape as nuc: this box shares the dynamic home IP with
no PTR and no SPF alignment, so it must use submission (587) with
auth, not port 25 — rspamd rejects the port-25 path as spam.
/etc/msmtprcmode 600, hostmail.lutran.fr, STARTTLS. Give SMTP tests ≥ 30 s: the missing PTR delays the greeting. Thezed@lutran.fraccount is the same credential as nuc — it is SMTP-AUTH, not IP-bound, so/etc/msmtprccan simply be copied between hosts (mode 600, root:root). The password lives in the password manager; no host-specific setup is needed. - zed:
/etc/zfs/zed.d/zed.rcmode 600 withZED_EMAIL_ADDR="julien@lutran.fr",ZED_EMAIL_PROG="mail",ZED_NOTIFY_VERBOSE=1,ZED_NOTIFY_DATA=1,ZED_NOTIFY_INTERVAL_SECS=3600. zpool-health.shevery 15 min via/etc/cron.d/zpool-health— zed does not report a suspended pool (the vdev staysONLINE, sostatechange-notify.shnever fires). That watchdog is the only thing that catches the exact failure this box was built to prevent.- smartd —
/etc/smartd.conf. The Toshiba carries the useful attributes; the Intel is a liveness check only:
# Toshiba Q300 — endurance + temperature are real here
/dev/disk/by-id/ata-TOSHIBA_Q300._36OB318OK1KU -a -o on -S on -s (S/../.././02|L/../../6/03) -W 4,50,55 -m julien@lutran.fr
# Intel 330 — no error log, no devstat, no temperature: liveness only
/dev/disk/by-id/ata-INTEL_SSDSC2CT120A3_CVMP250400ES120BGN -H -s (S/../.././03|L/../../6/04) -m julien@lutran.fr
# WD Red
/dev/disk/by-id/ata-WDC_WD40EFRX-68WT0N0_WD-WCC4E6NLPJJE -a -o on -S on -s (S/../.././04|L/../../6/05) -W 4,45,50 -m julien@lutran.fr
Verify the whole chain the day you build it, not the day you need it:
/root/scripts/zpool-health.sh -m julien@lutran.fr -t
tail -2 /var/log/msmtp.log # expect smtpstatus=250
smartctl -d sat -l devstat /dev/sdX | grep -i endurance
11. Power baseline
Take the measurement before the box goes into service, so later readings mean something. Expect roughly 25–35 W idle with three disks: the 20 W SoC plus a BMC drawing several watts continuously, even at soft-off.
# baseline: 10 min idle, everything settled
# CPU in steps — the interesting curve for an always-on box
for l in 25 50 75 100; do echo "=== ${l}% $(date +%s)"; stress-ng --cpu 0 --cpu-load $l --timeout 120s; done
# disk: random I/O is what moves an HDD's power, not throughput
fio --name=rr --directory=/export/media --size=20G --rw=randread --bs=4k \
--iodepth=32 --numjobs=4 --ioengine=libaio --direct=1 --runtime=300 --time_based
Measure at the wall (PDU or smart plug) — RAPL is unreliable on Avoton and sees neither disks nor fans. Log epoch timestamps per step so the trace can be recut against the meter's series afterwards.
12. Post-install checklist
- IPMI left unplugged by decision (2026-08-30) — no BMC on the LAN, therefore no out-of-band console either
cat /proc/mdstat— all three arrays[UU]; GRUB written to both ESPs, both mounted,debian+debian-bboot entries present (verified across a cold boot 2026-08-31)- boot still untested with one disk physically unplugged — the mirror is a guess until that is done
zpool statushealthy forincusandtank; weekly scrubs scheduled; cold boot verified 2026-08-31 — both pools imported from/etc/zfs/zpool.cache, all instances autostarted, NFS exports republished, 0 failed units- 7 days with zero pool suspensions and zero CRC errors — the gate that closes usb4t-dropouts.md
smartdmonitoring all 3 disks;zpool-health.sh -tmail delivered (smtpstatus=250in/var/log/msmtp.log) — done 2026-08-30transmission-bton nas: egress is193.70.35.17, no default route, downloads land in/export/media/downloads, watch folder works from the new pathjellyfin-serveron nuc lists the library over NFS after a cold reboot of both hosts (the boot-ordering trap)incus admin recoverran on nas before the first push, so the inheritedtank/backup/nucreplicas refresh incrementally instead of re-sending (check the first run's duration, not justrc=0)- nuc's 03:30 leg →
nas:nucbackup,rc=0, all nuc instances present (incus list --project backup -c ns -f csvon nas) - nas's 04:00 local leg →
nasbackup,rc=0, and blocky, privoxy and transmission-bt all appear by name inincus list --project backup - stale
nucbackuppool removed from nuc (§9c),zpool statusshows onlydata - replicas are stopped with
boot.autostart=false— they hold the live containers' LAN addresses - ks4 pull leg seeded + one instance test-restored → tick the gate
in ks2/plan.md; then retire nuc's
wg-ks4 - power baseline recorded above, with the meter reading
- this file updated with what was built (RAM fitted, NIC name, WD Red serial)