docs: follow the storage move from nuc/USB to nas/SATA

backup-strategy (the entry point) still described nuc pulling the ks4
replicas: the leg, its WireGuard peer and the target pool now live on
nas with the 4 TB on direct SATA. Also: ks4 README flow chart redrawn
for the new topology, incus-copy leg 2 retargeted, usb4t-dropouts
marked RESOLVED (kept for the diagnosis method and the alerting gap),
and ks2/plan records that the interim push was deliberately not
re-enabled.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Julien Lutran
2026-08-31 09:32:29 +02:00
co-authored by Claude Fable 5
parent f46726b6ca
commit ef88ef9342
5 changed files with 59 additions and 37 deletions
+20 -14
View File
@@ -9,7 +9,7 @@ docs; this page is the map.
| leg | mechanism | protects | | leg | mechanism | protects |
|---|---|---| |---|---|---|
| **local replication** | `incus copy --refresh` → pool `backup` on ks4's second disk (sdb) | instances, against losing the `data` pool | | **local replication** | `incus copy --refresh` → pool `backup` on ks4's second disk (sdb) | instances, against losing the `data` pool |
| **remote replication** | nuc pulls the same replicas over WireGuard → pool `ks4backup` | instances, against losing ks4 or the site | | **remote replication** | **nas** pulls the same replicas over WireGuard → pool `ks4backup` (on `tank`, direct SATA) | instances, against losing ks4 or the site |
| **remote backup** | **restic** → S3 bucket `restic-data`: database dumps + selected filesystem trees | the data itself, versioned and encrypted, independent of every disk above (a nightly run walks 1.1 M files / 1.24 TiB in under two minutes) | | **remote backup** | **restic** → S3 bucket `restic-data`: database dumps + selected filesystem trees | the data itself, versioned and encrypted, independent of every disk above (a nightly run walks 1.1 M files / 1.24 TiB in under two minutes) |
Only two tools are involved: `incus copy` (ZFS-incremental, native to Only two tools are involved: `incus copy` (ZFS-incremental, native to
@@ -18,15 +18,22 @@ encryption, local metadata cache, so a night costs only the churn —
[ks4/restic-backup.md](ks4/restic-backup.md)). [ks4/restic-backup.md](ks4/restic-backup.md)).
The old backup server **ks2 is being retired** (decommission by The old backup server **ks2 is being retired** (decommission by
Sep 30, 2026 — [ks2/plan.md](ks2/plan.md)); until FTTH enables the nuc Sep 30, 2026 — [ks2/plan.md](ks2/plan.md)); nothing is written to it
leg it still receives an interim replica push. any more.
**Status 2026-08-28**: local replication and the S3 backup leg are **Status 2026-08-31**: local replication and the S3 backup leg are
live, and the S3 leg has been **restore-tested** (a file tree came live, and the S3 leg has been **restore-tested** (a file tree came
back identical to the live one; a database dump loaded into a scratch back identical to the live one; a database dump loaded into a scratch
server with all its tables). The **nuc pull leg waits for FTTH** server with all its tables). The **nas pull leg waits for FTTH**
(expected before end of September); until then the ks2 push stands (expected before end of September) — it moved off nuc on 2026-08-30,
in. Backing up whole onto a host where the 4 TB disk is on direct SATA rather than a USB
bridge that suspended the pool weekly
([nas/README.md](nas/README.md), [ks2/nas-seed.md](ks2/nas-seed.md)).
In the meantime instances have no *fresh* off-site copy: the ks2 push
was deliberately not re-enabled (a 3-week-old replica set on a
94 %-full pool that is about to be wiped), so off-site protection
rests on `restic-data`, which holds the data, the databases and the
incus configuration needed to rebuild. Backing up whole
instance *images* to S3 was considered and left out — S3 holds the instance *images* to S3 was considered and left out — S3 holds the
data, the databases and the incus configuration, which is what a data, the databases and the incus configuration, which is what a
rebuild needs. rebuild needs.
@@ -39,10 +46,10 @@ rebuild needs.
│ live instances pool `data` (sda) │ │ live instances pool `data` (sda) │
│ nextcloud, seafile, mail, git, ... │ │ nextcloud, seafile, mail, git, ... │
│ │ │ │ │ │
│ │ 01:00 incus copy --refresh │ nuc (home LAN, via WireGuard) │ │ 01:00 incus copy --refresh │ nas (home LAN, via WireGuard)
│ ▼ │ ┌───────────────────────────┐ │ ▼ │ ┌───────────────────────────┐
│ replicas (stopped) pool `backup` (sdb) │────▶│ 05:00 incus copy (pull) │ │ replicas (stopped) pool `backup` (sdb) │────▶│ 05:00 incus copy (pull) │
│ project `backup` ready to start │ │ pool `ks4backup` (USB) │ project `backup` ready to start │ │ ks4backup on `tank` (SATA)
│ │ └───────────────────────────┘ │ │ └───────────────────────────┘
│ ──────────────────────────────────────── │ │ ──────────────────────────────────────── │
│ │ OVH Object Storage (S3, sbg) │ │ OVH Object Storage (S3, sbg)
@@ -59,7 +66,7 @@ rebuild needs.
Two different kinds of protection, on purpose: Two different kinds of protection, on purpose:
- **instances** (the running systems) are protected by *replication* - **instances** (the running systems) are protected by *replication*
a ready-to-start copy on ks4's second disk and, after FTTH, on nuc. a ready-to-start copy on ks4's second disk and, after FTTH, on nas.
Restoring one is `incus copy` + `incus start`. Restoring one is `incus copy` + `incus start`.
- **the data inside them** (files, databases, incus configuration) is - **the data inside them** (files, databases, incus configuration) is
protected by *backup* — encrypted, deduplicated, versioned on S3, protected by *backup* — encrypted, deduplicated, versioned on S3,
@@ -77,10 +84,9 @@ instances two homes.
| when | what | log | | when | what | log |
|---|---|---| |---|---|---|
| 01:00 daily | `incus-copy.sh -p backup -s backup` — refresh all replicas onto sdb | `/var/log/incus-copy.log` | | 01:00 daily | `incus-copy.sh -p backup -s backup` — refresh all replicas onto sdb | `/var/log/incus-copy.log` |
| 02:00 daily | `incus-copy.sh -d ks2 -m push` — interim off-site replicas, until the nuc leg replaces it | `/var/log/incus-copy.log` |
| 03:00 daily | instance snapshots (incus profile, 7-day expiry) — what keeps the refreshes incremental | `incus info <inst>` | | 03:00 daily | instance snapshots (incus profile, 7-day expiry) — what keeps the refreshes incremental | `incus info <inst>` |
| 05:00 daily | `restic-backup.sh` — dumps + data trees → S3 | `/var/log/restic-backup.log` | | 05:00 daily | `restic-backup.sh` — dumps + data trees → S3 | `/var/log/restic-backup.log` |
| 05:00 daily (nuc) | nuc pulls all replicas over WireGuard | nuc: `/var/log/incus-copy-ks4.log` | | 05:00 daily (**nas**) | nas pulls all ks4 replicas over WireGuard`ks4backup` (after FTTH) | nas: `/var/log/incus-copy-ks4.log` |
| Sun 14:00 | `restic-maintenance.sh` — prune (capped), structure check, 1/52 data verification | `/var/log/restic-maintenance.log` | | Sun 14:00 | `restic-maintenance.sh` — prune (capped), structure check, 1/52 data verification | `/var/log/restic-maintenance.log` |
| 1st of month 06:00 / quarterly 15th | seafile GC dry-run report / real GC ([ks4/seafile-gc.md](ks4/seafile-gc.md)) | `/var/log/seafile-gc.log` | | 1st of month 06:00 / quarterly 15th | seafile GC dry-run report / real GC ([ks4/seafile-gc.md](ks4/seafile-gc.md)) | `/var/log/seafile-gc.log` |
@@ -101,8 +107,8 @@ never silently.
- **A whole instance, fast (same box)**: `incus copy backup:<inst>` - **A whole instance, fast (same box)**: `incus copy backup:<inst>`
style — copy the replica from project `backup` back into `default` style — copy the replica from project `backup` back into `default`
([ks4/incus-copy.md](ks4/incus-copy.md)); from nuc the same via the ([ks4/incus-copy.md](ks4/incus-copy.md)); from nas the same via the
remote. incus remote.
- **A file or directory** (any date within retention): - **A file or directory** (any date within retention):
`restic -r s3:s3.sbg.io.cloud.ovh.net/restic-data restore <snapshot> --target /backup/restore-x --include <path>` `restic -r s3:s3.sbg.io.cloud.ovh.net/restic-data restore <snapshot> --target /backup/restore-x --include <path>`
(never restore into `/tmp` — it is RAM). Env: `. /root/.restic-env`. (never restore into `/tmp` — it is RAM). Env: `. /root/.restic-env`.
+9 -5
View File
@@ -18,10 +18,14 @@ What remains on the box is **cold history**: instance replicas on pool
snapshotted daily there (`zfs-auto-snapshot.sh`, 2-month expiry). snapshotted daily there (`zfs-auto-snapshot.sh`, 2-month expiry).
⚠️ While the nas leg waits for FTTH, instances have no *fresh* ⚠️ While the nas leg waits for FTTH, instances have no *fresh*
off-site copy — the ks2 push is to be re-enabled as soon as the off-site copy. The ks2 push was **deliberately not re-enabled**
initial restic sync finishes (decided 2026-08-28), and retired again (2026-08-30): ks2's replicas are three weeks old, its `data` pool is
when nas takes over. (The leg moved from nuc to the new host `nas` 94 % full, no common snapshot survives ks4's 7-day expiry, and the box
on 2026-08-30 — see [nas-seed.md](nas-seed.md).) is wiped within the month — so a full ~1.5 T re-send buys four weeks
of freshness on hardware already scheduled for destruction. Off-site
protection meanwhile rests on `restic-data` (data, databases and the
incus configuration — enough to rebuild). (The leg moved from nuc to
the new host `nas` on 2026-08-30 — see [nas-seed.md](nas-seed.md).)
## Inventory findings (2026-08-22) ## Inventory findings (2026-08-22)
@@ -53,7 +57,7 @@ on 2026-08-30 — see [nas-seed.md](nas-seed.md).)
| off-site, S3 (data+DB) | **restic** → bucket `restic-data` ([restic-backup.md](../ks4/restic-backup.md)) | **live** (05:00) | | off-site, S3 (data+DB) | **restic** → bucket `restic-data` ([restic-backup.md](../ks4/restic-backup.md)) | **live** (05:00) |
| off-site, S3 (instances) | restic over `incus file mount` of the `backup`-project replicas | **shelved** 2026-08-28 (replication covers instances) | | off-site, S3 (instances) | restic over `incus file mount` of the `backup`-project replicas | **shelved** 2026-08-28 (replication covers instances) |
| off-site, nas | **nas** pulls `ks4:*` → pool `ks4backup` over WG ([nas-seed.md](nas-seed.md)) | waiting FTTH (< Sep 30); moved off nuc 2026-08-30 | | off-site, nas | **nas** pulls `ks4:*` → pool `ks4backup` over WG ([nas-seed.md](nas-seed.md)) | waiting FTTH (< Sep 30); moved off nuc 2026-08-30 |
| off-site, ks2 (interim) | `incus-copy.sh -d ks2 -m push` at 02:00 until the nas leg seeds | to re-enable once the restic seed finishes | | off-site, ks2 (interim) | `incus-copy.sh -d ks2 -m push` | **not re-enabled** 2026-08-30 — see above |
## Actions (backup work documented in [`../ks4/`](../ks4/), ks2-only tasks here) ## Actions (backup work documented in [`../ks4/`](../ks4/), ks2-only tasks here)
+17 -14
View File
@@ -21,40 +21,43 @@ Incus host at OVH — public-facing self-hosted services.
- ⚠️ The ZFS `data` pool is single-disk (not mirrored); durability - ⚠️ The ZFS `data` pool is single-disk (not mirrored); durability
rests on nightly cron jobs — 01:00 `incus copy --refresh` of all rests on nightly cron jobs — 01:00 `incus copy --refresh` of all
instances to the local `backup` pool on sdb5, then the instance leg instances to the local `backup` pool on sdb5, then the instance leg
to S3; 05:00 restic (DB dumps + data trees) to S3; nuc pulls the to S3; 05:00 restic (DB dumps + data trees) to S3; **nas** pulls the
replicas over WireGuard. Full picture and restore procedures: replicas over WireGuard (after FTTH). Full picture and restore procedures:
**[backup-strategy.md](../backup-strategy.md)** **[backup-strategy.md](../backup-strategy.md)**
([local-backup-cron.md](local-backup-cron.md), ([local-backup-cron.md](local-backup-cron.md),
[incus-copy.md](incus-copy.md), [incus-copy.md](incus-copy.md),
[restic-backup.md](restic-backup.md)). [restic-backup.md](restic-backup.md)).
## Network flows (nuc ↔ ks4) ## Network flows (home ↔ ks4)
``` ```
nuc — home LAN 192.168.0.0/24 ks4 — OVH 193.70.35.17 nas — home LAN 192.168.0.4 ks4 — OVH 193.70.35.17
+-----------------------------------+ +-------------------------------------+ +-----------------------------------+ +-------------------------------------+
| | | | | | | |
| host: wg-ks4 (10.8.0.20) | | [wireguard] 192.168.1.18 | | host: wg-ks4 (10.8.0.22) | | [wireguard] 192.168.1.18 |
| incus remote "ks4" ------+--WG-->| wg0 10.8.0.1/24, udp 51845 | | incus remote "ks4" ------+--WG-->| wg0 10.8.0.1/24, udp 51845 |
| pull ks4:* -> pool ks4backup | udp | | masquerade -> eth0 | | pull ks4:* -> pool ks4backup | udp | | masquerade -> eth0 |
| on usb4t [pending FTTH seed] | 51845 | | | | on tank (SATA) [pending FTTH] | 51845 | | |
| | | +-> incus API 192.168.1.1:8443 | | | | +-> incus API 192.168.1.1:8443 |
| [transmission-bt] wg0 (10.8.0.21) | | | (ufw: only from .18) | | [transmission-bt] wg0 (10.8.0.21) | | | (ufw: only from .18) |
| full tunnel 0.0.0.0/0 ------+--WG-->| | | | full tunnel 0.0.0.0/0 ------+--WG-->| | |
| kill switch: no default route | udp | +-> WAN egress: torrents + | | kill switch: no default route | udp | +-> WAN egress: torrents + |
| downloads -> /srv/media | 51845 | apt of transmission-bt | | downloads -> /export/media | 51845 | apt of transmission-bt |
| (usb4t/media, read by jellyfin) | | exit as 193.70.35.17 | | (NFS-exported to nuc) | | exit as 193.70.35.17 |
| | | | | | | |
| 03:00 instance snapshots | | 03:00 instance snapshots | | 03:30 nuc pushes its instances | | 03:00 instance snapshots |
| 03:30 incus-copy: all instances | | 01:00 incus-copy: all instances | | -> nucbackup on tank | | 01:00 incus-copy: all instances |
| -> project backup, pool | | -> project backup, zpool sdb5 | | 04:00 nas replicates its own | | -> project backup, zpool sdb5 |
| nucbackup (usb4t/backup/nuc) | | then restic instance leg -> S3 | | -> nasbackup on tank | | |
| 05:00 pull ks4:* -> ks4backup | | 05:00 restic: DB dumps + data | | 05:00 pull ks4:* -> ks4backup | | 05:00 restic: DB dumps + data |
| [pending FTTH] | | trees -> S3 (restic-data) | | [pending FTTH] | | trees -> S3 (restic-data) |
| 05:30 apt upgrade all containers | | Sun 14:00 restic maintenance | | | | Sun 14:00 restic maintenance |
+-----------------------------------+ +-------------------------------------+ +-----------------------------------+ +-------------------------------------+
phones/laptops: WG peers 10.8.0.2-3 reach 192.168.1.x through the same endpoint phones/laptops: WG peers 10.8.0.2-3 reach 192.168.1.x through the same endpoint
nuc (on-demand media box) mounts /export/media from nas over NFSv4
``` ```
Both tunnels initiate **from** nuc (home NAT, dynamic IP) toward ks4's Both tunnels initiate **from home** (NAT, dynamic IP) toward ks4's
fixed endpoint; ks4's incus API is never exposed to the internet. fixed endpoint; ks4's incus API is never exposed to the internet.
The pull leg and its tunnel moved from nuc to nas on 2026-08-30
([nas/README.md](../nas/README.md)).
+3 -3
View File
@@ -8,7 +8,7 @@ push to `ks2` (decommissioning):
1. **local** — replicas + dumps on a dedicated `backup` zpool on ks4's 1. **local** — replicas + dumps on a dedicated `backup` zpool on ks4's
second disk (`sdb5`), survives `sda` death second disk (`sdb5`), survives `sda` death
2. **off-site** — replicas pulled by **nuc** into pool `ks4backup` 2. **off-site** — replicas pulled by **nas** into pool `ks4backup`
(dataset `usb4t/backup/ks4`), survives losing ks4 entirely (dataset `usb4t/backup/ks4`), survives losing ks4 entirely
## The script ## The script
@@ -74,7 +74,7 @@ Cron (root on ks4) — replaces both ks2 jobs:
⚠️ Replicas in the `backup` project must stay **stopped** — they keep ⚠️ Replicas in the `backup` project must stay **stopped** — they keep
the live containers' static `192.168.1.x` addresses. the live containers' static `192.168.1.x` addresses.
## Leg 2 — off-site pull from nuc ## Leg 2 — off-site pull from nas (was nuc until 2026-08-30)
Storage pool on nuc (done 2026-08-09): `ks4backup`, backed by the Storage pool on nuc (done 2026-08-09): `ks4backup`, backed by the
dataset `usb4t/backup/ks4` on the USB 4 TB pool (quota on `usb4t/backup` dataset `usb4t/backup/ks4` on the USB 4 TB pool (quota on `usb4t/backup`
@@ -87,7 +87,7 @@ incus storage create ks4backup zfs source=usb4t/backup/ks4
Replicas live only on the USB drive — if it fails, only backups are Replicas live only on the USB drive — if it fails, only backups are
lost; nuc's own instances (pool `data` on the SSD) are unaffected. lost; nuc's own instances (pool `data` on the SSD) are unaffected.
**Direction: nuc pulls, through the WireGuard tunnel.** Verified **Direction: nas pulls, through the WireGuard tunnel.** Verified
2026-08-09: ks4's API listens on wildcard `:8443` (so it answers on 2026-08-09: ks4's API listens on wildcard `:8443` (so it answers on
`192.168.1.1`, the `incusbr0` host address) but is **firewalled from `192.168.1.1`, the `incusbr0` host address) but is **firewalled from
the internet** — the VPN path keeps it that way, needs no inbound port the internet** — the VPN path keeps it that way, needs no inbound port
+10 -1
View File
@@ -1,4 +1,13 @@
# usb4t: USB dropouts suspend the pool (2026-08) # usb4t: USB dropouts suspend the pool (2026-08) — RESOLVED
> **Outcome (2026-08-30): the disk moved off USB entirely.** It now
> runs on **direct SATA** in the new host `nas`
> ([nas/README.md](../nas/README.md)), as pool `tank`, and the ks4
> pull leg plus its WireGuard tunnel moved with it
> ([ks2/nas-seed.md](../ks2/nas-seed.md)). Everything below is the
> investigation that led there — worth keeping for the diagnosis
> method and for the alerting gap it exposed, which applies to any
> host.
Symptom seen first in the nightly backup log Symptom seen first in the nightly backup log
(`/var/log/incus-copy.log`, job in `/etc/cron.d/incus-copy`): every (`/var/log/incus-copy.log`, job in `/etc/cron.d/incus-copy`): every