backup-strategy (the entry point) still described nuc pulling the ks4 replicas: the leg, its WireGuard peer and the target pool now live on nas with the 4 TB on direct SATA. Also: ks4 README flow chart redrawn for the new topology, incus-copy leg 2 retargeted, usb4t-dropouts marked RESOLVED (kept for the diagnosis method and the alerting gap), and ks2/plan records that the interim push was deliberately not re-enabled. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
151 lines
8.4 KiB
Markdown
151 lines
8.4 KiB
Markdown
# ks4 backup strategy (2026-08)
|
|
|
|
Audience: anyone with root on ks4 who needs to know **what is protected,
|
|
where, by which tool, and how to check it**. Details live in the linked
|
|
docs; this page is the map.
|
|
|
|
## Two tools, three legs
|
|
|
|
| leg | mechanism | protects |
|
|
|---|---|---|
|
|
| **local replication** | `incus copy --refresh` → pool `backup` on ks4's second disk (sdb) | instances, against losing the `data` pool |
|
|
| **remote replication** | **nas** pulls the same replicas over WireGuard → pool `ks4backup` (on `tank`, direct SATA) | instances, against losing ks4 or the site |
|
|
| **remote backup** | **restic** → S3 bucket `restic-data`: database dumps + selected filesystem trees | the data itself, versioned and encrypted, independent of every disk above (a nightly run walks 1.1 M files / 1.24 TiB in under two minutes) |
|
|
|
|
Only two tools are involved: `incus copy` (ZFS-incremental, native to
|
|
the platform — a restore is `incus start`) and restic (dedup +
|
|
encryption, local metadata cache, so a night costs only the churn —
|
|
[ks4/restic-backup.md](ks4/restic-backup.md)).
|
|
|
|
The old backup server **ks2 is being retired** (decommission by
|
|
Sep 30, 2026 — [ks2/plan.md](ks2/plan.md)); nothing is written to it
|
|
any more.
|
|
|
|
**Status 2026-08-31**: local replication and the S3 backup leg are
|
|
live, and the S3 leg has been **restore-tested** (a file tree came
|
|
back identical to the live one; a database dump loaded into a scratch
|
|
server with all its tables). The **nas pull leg waits for FTTH**
|
|
(expected before end of September) — it moved off nuc on 2026-08-30,
|
|
onto a host where the 4 TB disk is on direct SATA rather than a USB
|
|
bridge that suspended the pool weekly
|
|
([nas/README.md](nas/README.md), [ks2/nas-seed.md](ks2/nas-seed.md)).
|
|
In the meantime instances have no *fresh* off-site copy: the ks2 push
|
|
was deliberately not re-enabled (a 3-week-old replica set on a
|
|
94 %-full pool that is about to be wiped), so off-site protection
|
|
rests on `restic-data`, which holds the data, the databases and the
|
|
incus configuration needed to rebuild. Backing up whole
|
|
instance *images* to S3 was considered and left out — S3 holds the
|
|
data, the databases and the incus configuration, which is what a
|
|
rebuild needs.
|
|
|
|
## The map
|
|
|
|
```
|
|
ks4 (OVH) off-site
|
|
┌────────────────────────────────────────────┐
|
|
│ live instances pool `data` (sda) │
|
|
│ nextcloud, seafile, mail, git, ... │
|
|
│ │ │
|
|
│ │ 01:00 incus copy --refresh │ nas (home LAN, via WireGuard)
|
|
│ ▼ │ ┌───────────────────────────┐
|
|
│ replicas (stopped) pool `backup` (sdb) │────▶│ 05:00 incus copy (pull) │
|
|
│ project `backup` ready to start │ │ ks4backup on `tank` (SATA)│
|
|
│ │ └───────────────────────────┘
|
|
│ ──────────────────────────────────────── │
|
|
│ │ OVH Object Storage (S3, sbg)
|
|
│ 05:00 restic-backup.sh │ ┌───────────────────────────┐
|
|
│ • DB dumps (MariaDB/PostgreSQL, auto) │────▶│ bucket restic-data │
|
|
│ • incus config + global DB dump │ │ everything needed to │
|
|
│ • data trees (restic-paths) │ │ rebuild: dumps + data │
|
|
│ │ └───────────────────────────┘
|
|
│ Sun 14:00 restic-maintenance.sh │
|
|
│ • prune + check (rotating full verify) │
|
|
└────────────────────────────────────────────┘
|
|
```
|
|
|
|
Two different kinds of protection, on purpose:
|
|
|
|
- **instances** (the running systems) are protected by *replication* —
|
|
a ready-to-start copy on ks4's second disk and, after FTTH, on nas.
|
|
Restoring one is `incus copy` + `incus start`.
|
|
- **the data inside them** (files, databases, incus configuration) is
|
|
protected by *backup* — encrypted, deduplicated, versioned on S3,
|
|
independent of ks4 and of the disks. Restoring means recreating the
|
|
container (its config is in the S3 dumps, its install steps are in
|
|
this repo) and pouring the data back.
|
|
|
|
Backing up whole instance images to S3 as well was designed
|
|
([ks4/restic-backup.md](ks4/restic-backup.md) §6) and shelved: it
|
|
duplicates data already covered, and the replication legs already give
|
|
instances two homes.
|
|
|
|
## Schedule (root crontab on ks4)
|
|
|
|
| when | what | log |
|
|
|---|---|---|
|
|
| 01:00 daily | `incus-copy.sh -p backup -s backup` — refresh all replicas onto sdb | `/var/log/incus-copy.log` |
|
|
| 03:00 daily | instance snapshots (incus profile, 7-day expiry) — what keeps the refreshes incremental | `incus info <inst>` |
|
|
| 05:00 daily | `restic-backup.sh` — dumps + data trees → S3 | `/var/log/restic-backup.log` |
|
|
| 05:00 daily (**nas**) | nas pulls all ks4 replicas over WireGuard → `ks4backup` (after FTTH) | nas: `/var/log/incus-copy-ks4.log` |
|
|
| Sun 14:00 | `restic-maintenance.sh` — prune (capped), structure check, 1/52 data verification | `/var/log/restic-maintenance.log` |
|
|
| 1st of month 06:00 / quarterly 15th | seafile GC dry-run report / real GC ([ks4/seafile-gc.md](ks4/seafile-gc.md)) | `/var/log/seafile-gc.log` |
|
|
|
|
Retention on S3: 14 daily, 8 weekly, 6 monthly snapshots.
|
|
|
|
## "Did last night work?" — the 30-second check
|
|
|
|
```sh
|
|
grep -aE "done \(rc=|FAILED" /var/log/incus-copy.log | tail -2
|
|
grep -aE "snapshot .* saved|done \(rc=|failed" /var/log/restic-backup.log | tail -3
|
|
```
|
|
|
|
`rc=0` everywhere = fine. Any `failed`/`FAILED` line names the culprit
|
|
(instance, database or path). All scripts exit non-zero on any error,
|
|
never silently.
|
|
|
|
## Restoring
|
|
|
|
- **A whole instance, fast (same box)**: `incus copy backup:<inst>`
|
|
style — copy the replica from project `backup` back into `default`
|
|
([ks4/incus-copy.md](ks4/incus-copy.md)); from nas the same via the
|
|
incus remote.
|
|
- **A file or directory** (any date within retention):
|
|
`restic -r s3:s3.sbg.io.cloud.ovh.net/restic-data restore <snapshot> --target /backup/restore-x --include <path>`
|
|
(never restore into `/tmp` — it is RAM). Env: `. /root/.restic-env`.
|
|
- **A database**: restore the `.sql` dump from `restic-data`
|
|
(`/backup/dumps/mariadb/<inst>/<db>.sql`, `grants.sql` for users), load
|
|
it with `incus exec <inst> -- mariadb < dump.sql`.
|
|
- **An instance when no replica survives** (worst case: ks4 and the
|
|
replica sites are gone): recreate the container from its page in
|
|
[ks4/](ks4/) (the doc *is* the install script), restore its data
|
|
trees and database dumps from `restic-data`, and use the incus
|
|
configuration captured nightly in the dump tree
|
|
(`/backup/dumps/incus/incus-global-db.sql`) to check devices,
|
|
profiles and addresses.
|
|
|
|
Secrets you need for any of this: `/root/.restic-passphrase` and the S3
|
|
keys in `/root/.restic-env` — **both are in the password manager**;
|
|
without the passphrase the S3 buckets are unreadable.
|
|
|
|
## Adding something to the backups
|
|
|
|
New instances are picked up automatically by the replica and
|
|
instance legs (all instances, opt-out only). Databases are
|
|
auto-discovered. Only **data trees** need one line in
|
|
[`restic-paths`](https://git.lutran.fr/julien/scripts/src/branch/main/restic-paths) — see [new-container.md](new-container.md).
|
|
|
|
## History (reference only)
|
|
|
|
The road to the setup above, kept for the measurements rather than for
|
|
operations: an rsync-based leg to ks2
|
|
([ks2/plan.md](ks2/plan.md)) and a first S3 implementation with
|
|
**plakar** ([ks4/plakar-s3-data.md](ks4/plakar-s3-data.md), plus a
|
|
self-written incus connector,
|
|
[ks4/plakar-incus-integration.md](ks4/plakar-incus-integration.md)).
|
|
plakar was dropped because its nightly cost scaled with the size of
|
|
the tree rather than with the churn
|
|
([PlakarKorp/plakar#2338](https://github.com/PlakarKorp/plakar/issues/2338)).
|
|
It is still installed on ks4 and its `plakar-data` bucket still exists
|
|
— kept only so the upstream issue can be reproduced; nothing schedules
|
|
it any more.
|