Files
doc/backup-strategy.md
T
Julien LutranandClaude Fable 5 a0e3a7dd86 doc: split per-host READMEs, gitea cross-repo links, consistency pass
- nuc/README.md and ks4/README.md carry the host sections (+ network
  flows) that lived in the top-level README; links rebased
- top README: repo links (doc/scripts on git.lutran.fr), index points
  at the new per-host pages
- cross-repo references now use https://git.lutran.fr/julien/scripts
  instead of relative ../scripts paths that resolve nowhere
- plakar-s3-data.md and plakar-incus-integration.md marked SUPERSEDED
  / RETIRED with pointers to restic-backup.md; their measurements and
  rationale kept
- install.md, local-backup-cron.md, incus-copy.md: crontab sections
  updated to the live schedule (01:00 replicas, 05:00 restic, Sun
  maintenance); retired legs labelled as such
- restic-backup.md: status live, cutover recorded, post-GC memory
  estimate, seed plan dated
- seafile-gc.md: online GC noted, stale 'crons commented out' removed
- ks2/: what-ks2-does-today rewritten (nothing writes to it any more),
  legs table and gates reflect restic, decommission steps updated
- db-exclude replaces the plakar-era config name (script keeps a
  fallback)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 16:04:07 +02:00

6.2 KiB

ks4 backup strategy (2026-08)

Audience: anyone with root on ks4 who needs to know what is protected, where, by which tool, and how to check it. Details live in the linked docs; this page is the map.

Two tools, no more

tool job why this one
incus copy --refresh replicate whole instances (ready-to-start copies) ZFS-incremental, native to the platform, restores are incus start
restic off-site backups to S3 (files, database dumps, instance trees) dedup + encryption, local metadata cache → a night costs only the churn (restic-backup.md)

plakar was evaluated first and replaced (its nightly cost scaled with tree size — PlakarKorp/plakar#2338); the old backup server ks2 is being retired (decommission by Sep 30, 2026 — ks2/plan.md): it still holds an ageing copy of the instances until the nuc leg below takes over.

Status 2026-08-28: the local replica leg and the S3 data leg are live. The nuc pull leg is waiting for FTTH (expected before end of September) and the S3 instance leg is written but not yet seeded — until both land, instances have only the ks4-local replica off the live pool.

The map

 ks4 (OVH)                                            off-site
 ┌─────────────────────────────────────────────┐
 │  live instances        pool `data` (sda)    │
 │  nextcloud, seafile, mail, git, ...         │
 │        │                                    │
 │        │ 01:00  incus copy --refresh        │      nuc (home LAN, via WireGuard)
 │        ▼                                    │      ┌──────────────────────────┐
 │  replicas (stopped)    pool `backup` (sdb)  │─────▶│ 05:00 incus copy (pull)  │
 │  project `backup`                           │      │ pool `ks4backup` (USB)   │
 │        │                                    │      └──────────────────────────┘
 │        │ 01:00+ restic (via incus file mount)
 │        ▼                                    │      OVH Object Storage (S3, sbg)
 │  ─────────────────────────────────────────  │      ┌──────────────────────────┐
 │  05:00 restic-backup.sh                     │─────▶│ bucket restic-data       │
 │    • DB dumps (MariaDB/PostgreSQL, auto)    │      │  dumps + all data trees  │
 │    • data trees from /root/scripts/restic-paths     │                          │
 │  01:00+ restic-incus-backup.sh              │─────▶│ bucket restic-incus      │
 │    • every replica, per file + config yaml  │      │  instance filesystems    │
 │  Sun 14:00 restic-maintenance.sh            │      └──────────────────────────┘
 │    • prune + check (rotating full verify)   │
 └─────────────────────────────────────────────┘

Copies of any given byte: live → sdb replica (same box, other disk) → nuc replica (other site) → S3 (other site, other technology). Databases additionally get application-consistent dumps nightly.

Schedule (root crontab on ks4)

when what log
01:00 daily incus-copy.sh -p backup -s backup — refresh all replicas onto sdb, then restic-incus-backup.sh → S3 /var/log/incus-copy.log, /var/log/restic-incus.log
03:00 daily instance snapshots (incus profile, 7-day expiry) — what keeps the refreshes incremental incus info <inst>
05:00 daily restic-backup.sh — dumps + data trees → S3 /var/log/restic-backup.log
05:00 daily (nuc) nuc pulls all replicas over WireGuard nuc: /var/log/incus-copy-ks4.log
Sun 14:00 restic-maintenance.sh — prune (capped), structure check, 1/52 data verification /var/log/restic-maintenance.log
1st of month 06:00 / quarterly 15th seafile GC dry-run report / real GC (ks4/seafile-gc.md) /var/log/seafile-gc.log

Retention on S3: 14 daily, 8 weekly, 6 monthly snapshots.

"Did last night work?" — the 30-second check

grep -aE "done \(rc=|FAILED" /var/log/incus-copy.log | tail -2
grep -aE "snapshot .* saved|done \(rc=|failed" /var/log/restic-backup.log | tail -3
grep -aE "done \(rc=|failed" /var/log/restic-incus.log | tail -2

rc=0 everywhere = fine. Any failed/FAILED line names the culprit (instance, database or path). All scripts exit non-zero on any error, never silently.

Restoring

  • A whole instance, fast (same box): incus copy backup:<inst> style — copy the replica from project backup back into default (ks4/incus-copy.md); from nuc the same via the remote.
  • A file or directory (any date within retention): restic -r s3:s3.sbg.io.cloud.ovh.net/restic-data restore <snapshot> --target /backup/restore-x --include <path> (never restore into /tmp — it is RAM). Env: . /root/.restic-env.
  • A database: restore the .sql dump from restic-data (/backup/dumps/mariadb/<inst>/<db>.sql, grants.sql for users), load it with incus exec <inst> -- mariadb < dump.sql.
  • An instance from S3 only (ks4 and both replica sites gone): restore its tree + <inst>.yaml from restic-incus, recreate the instance from the yaml/profile, push the tree back.

Secrets you need for any of this: /root/.restic-passphrase and the S3 keys in /root/.restic-envboth are in the password manager; without the passphrase the S3 buckets are unreadable.

Adding something to the backups

New instances are picked up automatically by the replica and instance legs (all instances, opt-out only). Databases are auto-discovered. Only data trees need one line in restic-paths — see new-container.md.