Files
doc/backup-strategy.md
T
Julien LutranandClaude Fable 5 ef88ef9342 docs: follow the storage move from nuc/USB to nas/SATA
backup-strategy (the entry point) still described nuc pulling the ks4
replicas: the leg, its WireGuard peer and the target pool now live on
nas with the 4 TB on direct SATA. Also: ks4 README flow chart redrawn
for the new topology, incus-copy leg 2 retargeted, usb4t-dropouts
marked RESOLVED (kept for the diagnosis method and the alerting gap),
and ks2/plan records that the interim push was deliberately not
re-enabled.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 09:32:29 +02:00

151 lines
8.4 KiB
Markdown

# ks4 backup strategy (2026-08)
Audience: anyone with root on ks4 who needs to know **what is protected,
where, by which tool, and how to check it**. Details live in the linked
docs; this page is the map.
## Two tools, three legs
| leg | mechanism | protects |
|---|---|---|
| **local replication** | `incus copy --refresh` → pool `backup` on ks4's second disk (sdb) | instances, against losing the `data` pool |
| **remote replication** | **nas** pulls the same replicas over WireGuard → pool `ks4backup` (on `tank`, direct SATA) | instances, against losing ks4 or the site |
| **remote backup** | **restic** → S3 bucket `restic-data`: database dumps + selected filesystem trees | the data itself, versioned and encrypted, independent of every disk above (a nightly run walks 1.1 M files / 1.24 TiB in under two minutes) |
Only two tools are involved: `incus copy` (ZFS-incremental, native to
the platform — a restore is `incus start`) and restic (dedup +
encryption, local metadata cache, so a night costs only the churn —
[ks4/restic-backup.md](ks4/restic-backup.md)).
The old backup server **ks2 is being retired** (decommission by
Sep 30, 2026 — [ks2/plan.md](ks2/plan.md)); nothing is written to it
any more.
**Status 2026-08-31**: local replication and the S3 backup leg are
live, and the S3 leg has been **restore-tested** (a file tree came
back identical to the live one; a database dump loaded into a scratch
server with all its tables). The **nas pull leg waits for FTTH**
(expected before end of September) — it moved off nuc on 2026-08-30,
onto a host where the 4 TB disk is on direct SATA rather than a USB
bridge that suspended the pool weekly
([nas/README.md](nas/README.md), [ks2/nas-seed.md](ks2/nas-seed.md)).
In the meantime instances have no *fresh* off-site copy: the ks2 push
was deliberately not re-enabled (a 3-week-old replica set on a
94 %-full pool that is about to be wiped), so off-site protection
rests on `restic-data`, which holds the data, the databases and the
incus configuration needed to rebuild. Backing up whole
instance *images* to S3 was considered and left out — S3 holds the
data, the databases and the incus configuration, which is what a
rebuild needs.
## The map
```
ks4 (OVH) off-site
┌────────────────────────────────────────────┐
│ live instances pool `data` (sda) │
│ nextcloud, seafile, mail, git, ... │
│ │ │
│ │ 01:00 incus copy --refresh │ nas (home LAN, via WireGuard)
│ ▼ │ ┌───────────────────────────┐
│ replicas (stopped) pool `backup` (sdb) │────▶│ 05:00 incus copy (pull) │
│ project `backup` ready to start │ │ ks4backup on `tank` (SATA)│
│ │ └───────────────────────────┘
│ ──────────────────────────────────────── │
│ │ OVH Object Storage (S3, sbg)
│ 05:00 restic-backup.sh │ ┌───────────────────────────┐
│ • DB dumps (MariaDB/PostgreSQL, auto) │────▶│ bucket restic-data │
│ • incus config + global DB dump │ │ everything needed to │
│ • data trees (restic-paths) │ │ rebuild: dumps + data │
│ │ └───────────────────────────┘
│ Sun 14:00 restic-maintenance.sh │
│ • prune + check (rotating full verify) │
└────────────────────────────────────────────┘
```
Two different kinds of protection, on purpose:
- **instances** (the running systems) are protected by *replication*
a ready-to-start copy on ks4's second disk and, after FTTH, on nas.
Restoring one is `incus copy` + `incus start`.
- **the data inside them** (files, databases, incus configuration) is
protected by *backup* — encrypted, deduplicated, versioned on S3,
independent of ks4 and of the disks. Restoring means recreating the
container (its config is in the S3 dumps, its install steps are in
this repo) and pouring the data back.
Backing up whole instance images to S3 as well was designed
([ks4/restic-backup.md](ks4/restic-backup.md) §6) and shelved: it
duplicates data already covered, and the replication legs already give
instances two homes.
## Schedule (root crontab on ks4)
| when | what | log |
|---|---|---|
| 01:00 daily | `incus-copy.sh -p backup -s backup` — refresh all replicas onto sdb | `/var/log/incus-copy.log` |
| 03:00 daily | instance snapshots (incus profile, 7-day expiry) — what keeps the refreshes incremental | `incus info <inst>` |
| 05:00 daily | `restic-backup.sh` — dumps + data trees → S3 | `/var/log/restic-backup.log` |
| 05:00 daily (**nas**) | nas pulls all ks4 replicas over WireGuard → `ks4backup` (after FTTH) | nas: `/var/log/incus-copy-ks4.log` |
| Sun 14:00 | `restic-maintenance.sh` — prune (capped), structure check, 1/52 data verification | `/var/log/restic-maintenance.log` |
| 1st of month 06:00 / quarterly 15th | seafile GC dry-run report / real GC ([ks4/seafile-gc.md](ks4/seafile-gc.md)) | `/var/log/seafile-gc.log` |
Retention on S3: 14 daily, 8 weekly, 6 monthly snapshots.
## "Did last night work?" — the 30-second check
```sh
grep -aE "done \(rc=|FAILED" /var/log/incus-copy.log | tail -2
grep -aE "snapshot .* saved|done \(rc=|failed" /var/log/restic-backup.log | tail -3
```
`rc=0` everywhere = fine. Any `failed`/`FAILED` line names the culprit
(instance, database or path). All scripts exit non-zero on any error,
never silently.
## Restoring
- **A whole instance, fast (same box)**: `incus copy backup:<inst>`
style — copy the replica from project `backup` back into `default`
([ks4/incus-copy.md](ks4/incus-copy.md)); from nas the same via the
incus remote.
- **A file or directory** (any date within retention):
`restic -r s3:s3.sbg.io.cloud.ovh.net/restic-data restore <snapshot> --target /backup/restore-x --include <path>`
(never restore into `/tmp` — it is RAM). Env: `. /root/.restic-env`.
- **A database**: restore the `.sql` dump from `restic-data`
(`/backup/dumps/mariadb/<inst>/<db>.sql`, `grants.sql` for users), load
it with `incus exec <inst> -- mariadb < dump.sql`.
- **An instance when no replica survives** (worst case: ks4 and the
replica sites are gone): recreate the container from its page in
[ks4/](ks4/) (the doc *is* the install script), restore its data
trees and database dumps from `restic-data`, and use the incus
configuration captured nightly in the dump tree
(`/backup/dumps/incus/incus-global-db.sql`) to check devices,
profiles and addresses.
Secrets you need for any of this: `/root/.restic-passphrase` and the S3
keys in `/root/.restic-env`**both are in the password manager**;
without the passphrase the S3 buckets are unreadable.
## Adding something to the backups
New instances are picked up automatically by the replica and
instance legs (all instances, opt-out only). Databases are
auto-discovered. Only **data trees** need one line in
[`restic-paths`](https://git.lutran.fr/julien/scripts/src/branch/main/restic-paths) — see [new-container.md](new-container.md).
## History (reference only)
The road to the setup above, kept for the measurements rather than for
operations: an rsync-based leg to ks2
([ks2/plan.md](ks2/plan.md)) and a first S3 implementation with
**plakar** ([ks4/plakar-s3-data.md](ks4/plakar-s3-data.md), plus a
self-written incus connector,
[ks4/plakar-incus-integration.md](ks4/plakar-incus-integration.md)).
plakar was dropped because its nightly cost scaled with the size of
the tree rather than with the churn
([PlakarKorp/plakar#2338](https://github.com/PlakarKorp/plakar/issues/2338)).
It is still installed on ks4 and its `plakar-data` bucket still exists
— kept only so the upstream issue can be reproduced; nothing schedules
it any more.