docs: schedules live in root crontab, not /etc/cron.d or timers
Standardised 2026-08-31 on regular crontabs: nuc's push moved off its systemd timer (units kept disabled on disk), and the cron.d files on nuc and nas were folded into root's crontab. Notes the consequence accepted for nuc: a night with the box powered off is skipped rather than caught up after boot. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
4c6396a8fa
commit
1721d7b05d
+19
-36
@@ -615,43 +615,26 @@ The driver is unchanged —
|
|||||||
[`incus-copy.sh`](https://git.lutran.fr/julien/scripts/src/branch/main/incus-copy.sh),
|
[`incus-copy.sh`](https://git.lutran.fr/julien/scripts/src/branch/main/incus-copy.sh),
|
||||||
deployed to `/root/scripts` as everywhere else.
|
deployed to `/root/scripts` as everywhere else.
|
||||||
|
|
||||||
### 9a. nuc's instances -> nas (systemd timer, not cron)
|
### 9a. nuc's instances -> nas
|
||||||
|
|
||||||
⚠️ **nuc is an on-demand media box** (see [nuc/README.md](../nuc/README.md)):
|
Root crontab on nuc:
|
||||||
since 2026-08-30 it only runs when watching Jellyfin or using the Spotify
|
|
||||||
kiosk, so it is usually powered off at 03:30. Cron cannot catch up a
|
|
||||||
missed window; a **systemd timer with `Persistent=true`** runs the job
|
|
||||||
shortly after the next boot. That is the whole reason this is a timer.
|
|
||||||
|
|
||||||
`/etc/systemd/system/incus-copy.service` on nuc:
|
```cron
|
||||||
|
30 3 * * * /root/scripts/incus-copy.sh -d nas -m push -s nucbackup -p backup >> /var/log/incus-copy.log 2>&1
|
||||||
```ini
|
|
||||||
[Unit]
|
|
||||||
Description=Replicate all nuc instances to nas (project backup, pool nucbackup)
|
|
||||||
After=network-online.target incus.service
|
|
||||||
Wants=network-online.target
|
|
||||||
Requires=incus.service
|
|
||||||
|
|
||||||
[Service]
|
|
||||||
Type=oneshot
|
|
||||||
ExecStart=/bin/sh -c "/root/scripts/incus-copy.sh -d nas -m push -s nucbackup -p backup >> /var/log/incus-copy.log 2>&1"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
`/etc/systemd/system/incus-copy.timer`:
|
⚠️ **nuc is an on-demand media box** (see
|
||||||
|
[nuc/README.md](../nuc/README.md)): since 2026-08-30 it only runs when
|
||||||
```ini
|
watching Jellyfin or using the Spotify kiosk, so it is often powered
|
||||||
[Timer]
|
off at 03:30 and **that night's push is simply skipped** — cron does
|
||||||
OnCalendar=03:30
|
not catch up missed windows. Accepted deliberately (2026-08-31):
|
||||||
Persistent=true
|
nuc's instances change rarely and the next time it is up the refresh
|
||||||
RandomizedDelaySec=300
|
is incremental anyway.
|
||||||
|
This ran briefly as a `systemd` timer with `Persistent=true` (which
|
||||||
[Install]
|
*does* catch up after boot); the units are still on disk, disabled, at
|
||||||
WantedBy=timers.target
|
`/etc/systemd/system/incus-copy.{service,timer}` if that behaviour is
|
||||||
```
|
ever wanted back:
|
||||||
|
`systemctl enable --now incus-copy.timer` (and remove the cron line).
|
||||||
```sh
|
|
||||||
systemctl enable --now incus-copy.timer && systemctl list-timers incus-copy
|
|
||||||
```
|
|
||||||
|
|
||||||
⚠️ **`-p backup` is not optional.** Without it the replicas land in
|
⚠️ **`-p backup` is not optional.** Without it the replicas land in
|
||||||
`default` on nas and collide with nas's *live* instances — both hosts are
|
`default` on nas and collide with nas's *live* instances — both hosts are
|
||||||
@@ -674,7 +657,7 @@ incus storage create nasbackup zfs source=tank/backup/nas
|
|||||||
```
|
```
|
||||||
|
|
||||||
```cron
|
```cron
|
||||||
# nas, /etc/cron.d/incus-copy
|
# nas, root crontab
|
||||||
0 4 * * * root /root/scripts/incus-copy.sh -p backup -s nasbackup >> /var/log/incus-copy.log 2>&1
|
0 4 * * * root /root/scripts/incus-copy.sh -p backup -s nasbackup >> /var/log/incus-copy.log 2>&1
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -742,7 +725,7 @@ Then test-restore one instance before ticking the gate in
|
|||||||
| When | Host | What |
|
| When | Host | What |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| 03:00 | nuc, nas | instance snapshots (profile) |
|
| 03:00 | nuc, nas | instance snapshots (profile) |
|
||||||
| 03:30 | nuc | push all instances → `nas:nucbackup` (**systemd timer**, `Persistent=true` — catches up after boot) |
|
| 03:30 | nuc | push all instances → `nas:nucbackup` (root crontab; skipped when nuc is off) |
|
||||||
| 04:00 | nas | local copy of nas instances → `nasbackup` (`tank/backup/nas`) |
|
| 04:00 | nas | local copy of nas instances → `nasbackup` (`tank/backup/nas`) |
|
||||||
| 05:00 | nas | pull `ks4:*` → `ks4backup` (after FTTH) |
|
| 05:00 | nas | pull `ks4:*` → `ks4backup` (after FTTH) |
|
||||||
| 05:30 | nuc | apt upgrade all containers |
|
| 05:30 | nuc | apt upgrade all containers |
|
||||||
@@ -779,7 +762,7 @@ host had no MTA).
|
|||||||
`ZED_NOTIFY_VERBOSE=1`, **`ZED_NOTIFY_DATA=1`**,
|
`ZED_NOTIFY_VERBOSE=1`, **`ZED_NOTIFY_DATA=1`**,
|
||||||
`ZED_NOTIFY_INTERVAL_SECS=3600`.
|
`ZED_NOTIFY_INTERVAL_SECS=3600`.
|
||||||
- **[`zpool-health.sh`](https://git.lutran.fr/julien/scripts/src/branch/main/zpool-health.sh)**
|
- **[`zpool-health.sh`](https://git.lutran.fr/julien/scripts/src/branch/main/zpool-health.sh)**
|
||||||
every 15 min via `/etc/cron.d/zpool-health` — zed does *not* report a
|
every 15 min from root's crontab — zed does *not* report a
|
||||||
suspended pool (the vdev stays `ONLINE`, so `statechange-notify.sh`
|
suspended pool (the vdev stays `ONLINE`, so `statechange-notify.sh`
|
||||||
never fires). That watchdog is the only thing that catches the exact
|
never fires). That watchdog is the only thing that catches the exact
|
||||||
failure this box was built to prevent.
|
failure this box was built to prevent.
|
||||||
|
|||||||
+3
-3
@@ -31,7 +31,7 @@ incus project create backup -c features.images=false -c features.profiles=false
|
|||||||
/root/scripts/incus-copy.sh -p backup -s nucbackup
|
/root/scripts/incus-copy.sh -p backup -s nucbackup
|
||||||
```
|
```
|
||||||
|
|
||||||
Runs nightly via `/etc/cron.d/incus-copy` at **03:30** (30 min after
|
Runs nightly from root's crontab at **03:30** (30 min after
|
||||||
the profile-scheduled 03:00 instance snapshots, so VM refreshes stay
|
the profile-scheduled 03:00 instance snapshots, so VM refreshes stay
|
||||||
incremental), logging to `/var/log/incus-copy.log` (logrotate:
|
incremental), logging to `/var/log/incus-copy.log` (logrotate:
|
||||||
`/etc/logrotate.d/incus-copy`). Note the script's `flock` is global:
|
`/etc/logrotate.d/incus-copy`). Note the script's `flock` is global:
|
||||||
@@ -245,8 +245,8 @@ once `usb4t` is imported.)
|
|||||||
| When | What | Where |
|
| When | What | Where |
|
||||||
|-------|------|-------|
|
|-------|------|-------|
|
||||||
| 03:00 | instance snapshots (`snapshots.schedule` on the default profile, expiry 7d) | incus |
|
| 03:00 | instance snapshots (`snapshots.schedule` on the default profile, expiry 7d) | incus |
|
||||||
| 03:30 | replicate all instances to the USB pool (`incus-copy.sh -p backup -s nucbackup`) | `/etc/cron.d/incus-copy` → `/var/log/incus-copy.log` |
|
| 03:30 | replicate all instances to the USB pool (`incus-copy.sh -p backup -s nucbackup`) | root crontab → `/var/log/incus-copy.log` |
|
||||||
| 05:00 | apt dist-upgrade all running containers (`incus-container-upgrade.sh`; VMs and non-apt containers skipped; jellyfin pinned to the 10.11 series in-container) | `/etc/cron.d/incus-container-upgrade` → `/var/log/incus-container-upgrade.log` |
|
| 05:00 | apt dist-upgrade all running containers (`incus-container-upgrade.sh`; VMs and non-apt containers skipped; jellyfin pinned to the 10.11 series in-container) | root crontab → `/var/log/incus-container-upgrade.log` |
|
||||||
|
|
||||||
The ordering is deliberate: snapshot → backup → upgrade, so a broken
|
The ordering is deliberate: snapshot → backup → upgrade, so a broken
|
||||||
upgrade is always one snapshot-restore away and the replicas predate it.
|
upgrade is always one snapshot-restore away and the replicas predate it.
|
||||||
|
|||||||
@@ -10,7 +10,7 @@
|
|||||||
> host.
|
> host.
|
||||||
|
|
||||||
Symptom seen first in the nightly backup log
|
Symptom seen first in the nightly backup log
|
||||||
(`/var/log/incus-copy.log`, job in `/etc/cron.d/incus-copy`): every
|
(`/var/log/incus-copy.log`, job in root's crontab): every
|
||||||
instance fails with
|
instance fails with
|
||||||
|
|
||||||
```
|
```
|
||||||
@@ -140,7 +140,7 @@ before). Two reasons:
|
|||||||
|
|
||||||
Fixed 2026-08-30 by adding
|
Fixed 2026-08-30 by adding
|
||||||
[`zpool-health.sh`](https://git.lutran.fr/julien/scripts/src/branch/main/zpool-health.sh),
|
[`zpool-health.sh`](https://git.lutran.fr/julien/scripts/src/branch/main/zpool-health.sh),
|
||||||
run every 15 min from `/etc/cron.d/zpool-health`. It mails only on
|
run every 15 min from root's crontab. It mails only on
|
||||||
`healthy <-> problem` **transitions**, so it is silent in normal
|
`healthy <-> problem` **transitions**, so it is silent in normal
|
||||||
operation and cannot spam; `-t` sends a test. Worth deploying on ks4
|
operation and cannot spam; `-t` sends a test. Worth deploying on ks4
|
||||||
too — its `data` pool is single-disk and has the same blind spot.
|
too — its `data` pool is single-disk and has the same blind spot.
|
||||||
|
|||||||
Reference in New Issue
Block a user