Add Atlas media and storage services [Phase 1] (#9)

* Add Atlas media and storage services

* Document Atlas backend phase one and WireGuard deployment

* Enable Atlas NAS management and document bootstrap workflow

* Harden Atlas network, SSH, firewall, and sharing

* Rotate Ansible Vault secrets

* Allow configurable Aegis SSH users and authorized keys

* Manage Aegis SSH authorized key fragments

* Manage SSH authorized key fragments for infrastructure hosts

* Harden Rocky storage and sharing configuration

* Verify WireGuard handshakes and restore Podman networking
This commit is contained in:
Fabio Scotto di Santolo
2026-09-15 22:39:28 +02:00
committed by GitHub
parent 73bf2cd62a
commit 160d63c02d
50 changed files with 2071 additions and 460 deletions

149
README.md
View File

@@ -108,16 +108,25 @@ That gives it Fedora packages through DNF, Docker from the official repository,
dotfiles and templates. The profile provisions configuration only: it does not transfer data, start
the Compose stack, update DNS, or perform a cutover.
The server profile installs platform-specific packages, Podman and podman-compose,
declared systemd services, the server Compose stack behind the `podman-compose-server` systemd unit, and firewalld. The Rocky server excludes
Syncthing. Rocky bind mounts use private SELinux relabeling for application data while host system
files remain unchanged.
The server profile installs platform-specific packages, Podman and podman-compose, declared systemd
services, and firewalld. The manually activated `podman-compose-server` unit contains the existing
Nginx Proxy Manager and Gitea services. The desired Compose file no longer includes Navidrome,
Syncthing, or the obsolete Navidrome PostgreSQL database. Navidrome and Syncthing belong to Atlas;
official Navidrome uses SQLite instead. Applying the profile does not stop or remove legacy
containers and does not delete `/opt/postgres/data`.
Firewalld enables SSH, Cockpit (`9090/tcp`), HTTP and HTTPS. Nginx Proxy Manager publishes only
`80/tcp` and `443/tcp`; its administration interface is bound to `127.0.0.1:81` and can be reached
from Ikaros or Nymph with the `npm-tunnel` Bash alias. Nextcloud remains disabled and the profile
does not provision any `/srv/nextcloud` directories.
The Atlas phase-one work does not change this NPM deployment or its persistent data. Once WireGuard
and the Atlas services are active, configure the current NPM proxy hosts with Navidrome upstream
`http://10.0.0.2:4533` and Syncthing GUI upstream `http://10.0.0.2:8384`. Only the Syncthing web GUI
uses NPM; synchronization traffic remains on explicitly published native ports bound only to the Atlas
WireGuard address. Configure both Syncthing authentication and an appropriate NPM access policy before
publishing its GUI.
Server identity comes from `server_username`, `server_user_group`, and `server_user_home` in `ansible/inventory/group_vars/server.yml`. `server_username` defaults to `username`, but it can be overridden, for example:
```bash
@@ -128,6 +137,8 @@ ansible-playbook ansible/site.yml --limit prometheus \
```
The target must already provide `server_username` with local sudo access.
Prometheus authorizes its declared SSH public keys through separate files below
`~/.ssh/authorized_keys.d/`, while `sshd` is configured to read those files directly.
### DuckDNS
@@ -150,7 +161,7 @@ back in; preserve any uncommitted work separately without copying secrets.
### Data migration
Provision Rocky first, then run the migration script **on the retired Ubuntu source host**. It is
dry-run by default and requires an explicit source-stack stop before it can copy PostgreSQL data:
dry-run by default and requires an explicit source-stack stop before it can copy application data:
```bash
sudo ./scripts/migrate_prometheus_data.sh \
@@ -163,11 +174,11 @@ sudo ./scripts/migrate_prometheus_data.sh \
--quiesce-source --execute
```
The script copies Navidrome, music, Nginx Proxy Manager, PostgreSQL and Gitea data. It does not
delete data, move Syncthing, copy `/home/git/.ssh`, start containers, update DNS, or perform a
cutover. The destination SSH host key must already be trusted and the destination account needs
passwordless sudo for `rsync`. It preserves ACLs but not extended attributes, so source SELinux labels
are not transferred; the Rocky Compose bind mounts apply their own `:Z` labels when containers start.
The script copies only Nginx Proxy Manager and Gitea data. It does not delete data, move
Navidrome/Syncthing, copy `/home/git/.ssh`, start containers, update DNS, or perform a cutover. The
destination SSH host key must already be trusted and the destination account needs passwordless sudo
for `rsync`. It preserves ACLs but not extended attributes, so source SELinux labels are not
transferred; the Rocky Compose bind mounts apply their own `:Z` labels when containers start.
## DNS Filter
@@ -191,6 +202,10 @@ for AdGuard while retaining DNS learned from the router. Define
`vault_aegis_icloudpd_apple_id` in Vault before applying it. iCloudPD still requires interactive MFA
initialization after its first deployment.
New Aegis images create the `admin` account in Butane. Before configuring a newly imaged node, run its
first playbook execution with `-e ansible_user=admin`; the SSH hardening role then permits that same
account. Keep the inventory on `pi` until the existing node has been replaced.
Validate the profile before deployment:
```bash
@@ -200,28 +215,97 @@ ansible-playbook ansible/site.yml --limit aegis --check --diff --ask-become-pass
## NAS
`atlas` is a Rocky Linux 9 NAS reached through SSH. Its pool already exists: the profile only
manages child datasets and must never create, partition, destroy, roll back, or otherwise alter the
pool itself. Linux clients use NFSv4 and Windows/WSL clients use SMB; both are restricted to the
configured LAN.
`atlas` is a Rocky Linux 9 NAS reached through SSH. Normally its pool already exists and the profile
only manages child datasets. A one-time RAIDZ2 bootstrap is available only with explicit confirmation
(`atlas_create_pool=true`) and exactly four verified `/dev/disk/by-id/...` paths in `atlas_zpool_disks`.
It never partitions, forces, destroys, rolls back, or changes the vdev layout of an existing pool. Linux
clients use NFSv4 and Windows/WSL clients use SMB; both are restricted to the configured LAN.
For the first run, replace the Atlas placeholders and provide
`vault_atlas_authorized_ssh_keys`, `vault_atlas_admin_password_hash`, and
`vault_atlas_samba_password`. Bootstrap the host through its existing administrator:
For the first run, provide `vault_atlas_admin_password_hash`, `vault_atlas_samba_password`, and
`vault_atlas_immich_db_password`. Bootstrap the host through its
existing administrator. Open `51820/udp` towards Prometheus in the provider firewall first, then
include both WireGuard peers in the same idempotent playbook run:
```bash
ansible-playbook ansible/site.yml --limit atlas \
-e atlas_connection_username=<existing-admin>
ansible-playbook ansible/site.yml --limit prometheus,atlas \
-e atlas_connection_username=<existing-admin> \
-e atlas_create_pool=true
```
`vault_atlas_admin_password_hash` must be an `/etc/shadow`-compatible hash, not a clear-text
Cockpit password. Subsequent runs use `atlas_admin_username`. Enable
`atlas_manage_storage` only after checking the existing pool and mountpoints; enable
`atlas_manage_firewall` only after checking the LAN subnet and active firewalld zone.
The explicit pool gate is safe to repeat: the role creates the RAIDZ2 pool only when it is absent.
WireGuard waits for a real peer handshake before the play continues.
Snapshot retention, Syncthing topology, VPN access, Prometheus pulls, encrypted Borg backups to a
Hetzner Storage Box, USB backup, monitoring, and disaster-recovery tests remain follow-up work. The
detailed operational backlog is kept in `AGENTS.md`.
`vault_atlas_admin_password_hash` must be an `/etc/shadow`-compatible hash, not a clear-text
Cockpit password. Subsequent runs use `atlas_admin_username`. Atlas declares storage, sharing, and its
LAN firewall rules enabled. Before the first apply, check the existing pool and mountpoints, LAN subnet,
and active firewalld zone. `atlas_manage_media_stack` remains disabled until `/dev/dri`, the container
paths, and the Immich database secret are validated. Atlas reads its declared SSH public keys from
separate files below `~/.ssh/authorized_keys.d/`.
With storage management enabled, Atlas creates the complete dataset hierarchy below the existing or
explicitly bootstrapped `zpool`: `work`, `archive`, `archive/app_data`, the separate `archive/app_data/navidrome` and
`archive/app_data/syncthing` application datasets, `media`, `media/music`, `media/photobook`,
`backups`, `backups/services`, and `backup_prometheus`. Application/archive datasets use `zstd`,
while media, Syncthing and service-backup datasets use `lz4`; `backups/services` also has a `500G`
refreservation. Atlas enforces targeted SELinux persistently and reports, without initiating, any reboot required to activate it. It assigns its primary LAN interface explicitly to the managed firewalld zone and applies persistent kernel network hardening: redirects and source routes are rejected, martians logged, reverse-path filtering remains loose for WireGuard, and IPv4 forwarding is disabled. SSH permits only the declared administrator using public-key authentication; root login, passwords,
agent and remote forwarding are disabled, while local forwarding remains available for private administrative tunnels. SMB3 exposes `Archive` only to the configured Vault-backed
Samba accounts on encrypted, signed SMB3 over TCP/445 only and admits the configured LAN without host-specific
exclusions. NFSv4 exports only `media/photobook` to the configured Aegis IP over TCP/2049, using
`all_squash` with anonymous UID/GID `1100`.
The `immich` system account is fixed to UID/GID `1100`, has no login shell or `wheel` membership, and
receives `video` and `render` access. The rootful Immich Server, ML, Redis-compatible cache, PostgreSQL,
and NPM Quadlets share one Podman network. Immich runs as `1100:1100`; Server and ML receive `/dev/dri`,
and Photobook is mounted read-only at `/external/photobook`. NPM publishes ports `80` and `443`; its
administration interface remains restricted to `127.0.0.1:81` for SSH-tunnel access.
Phase 1 is limited to rootless Navidrome and Syncthing user Quadlets on Atlas. It is enabled in the
Atlas host configuration and can be set to `false` only for a deliberate suspension. Official Navidrome `0.63.2` uses its SQLite database below `/data`; it does
not support `ND_DATABASE_URL` or an external PostgreSQL backend. The obsolete `navidromedb` service
was therefore removed from Prometheus instead of being reproduced on Atlas. The role derives all
storage paths from the `zpool` mounted at `/zpool`: music is read-only at
`/zpool/media/music`, Navidrome application state and `navidrome.db` are stored at
`/zpool/archive/app_data/navidrome`, and Syncthing persists at
`/zpool/archive/app_data/syncthing`. `profile_atlas` creates these datasets when
`atlas_manage_storage` is enabled; the backend role verifies their exact mountpoints before starting
containers. The backend role never creates the pool. The separate `wireguard_overlay` role manages `wg0`
between Prometheus (`10.0.0.1`) and Atlas (`10.0.0.2`), generating private keys once
on their respective hosts and exchanging only public keys through Ansible. Prometheus alone opens
`51820/udp` publicly. When the WireGuard zone is created, Ansible reloads firewalld and immediately
reloads Prometheus' rootful Podman networks so the existing proxy stack retains container DNS and
connectivity. Backend ports are admitted only in the WireGuard firewalld zone.
`backend_phase1_start_services` stays false during the application-state transfer, so the first real
backend run renders the Quadlets without creating an empty Atlas database. After stopping Navidrome
on Prometheus, copy the complete `/opt/navidrome/data/` directory into
`/zpool/archive/app_data/navidrome/`, preserving `navidrome.db` and any SQLite sidecar files. Then set
this variable to true and rerun the role to enable and start Navidrome and Syncthing. The playbook
never copies or deletes application data.
Validate and render the Atlas services with:
```bash
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --limit atlas --tags storage
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --limit prometheus,atlas --tags wireguard
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --limit atlas --tags backend_phase1 --check --diff
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
ansible-playbook ansible/site.yml --limit atlas --tags backend_phase1
```
For the cutover, stop the old Navidrome writer before copying its data directory, verify ownership by
the Atlas `admin` account and confirm that the copied SQLite database is present before changing
`backend_phase1_start_services: true` in `host_vars/atlas.yml`. Keep the source data and the stopped
legacy `navidromedb` container until Navidrome on Atlas and a restore test have been validated.
Snapshot retention, Syncthing topology, WireGuard/firewall validation, Prometheus backup pulls,
encrypted Borg backups to a Hetzner Storage Box, USB backup, monitoring, and disaster-recovery tests
remain follow-up work. The detailed operational backlog is kept in `AGENTS.md`.
## How layering works
@@ -308,6 +392,8 @@ ansible-playbook ansible/site.yml --limit deadalus --tags ai_agents --check --di
| `profile_workstation_dev_wsl` | WSL development setup. |
| `profile_server` | Server setup. |
| `profile_atlas` | Rocky Linux 9 NAS setup. |
| `profile_backend_phase1` | Rootless Navidrome and Syncthing on Atlas. |
| `wireguard_overlay` | Prometheus/Atlas WireGuard overlay. |
| `profile_aegis` | Fedora IoT always-on LAN node. |
| `dotfiles_common` | Shared user dotfiles. |
@@ -319,8 +405,10 @@ platform_void -> packages_void + services_runit
platform_void & graphical_desktop -> profile_desktop_common + profile_desktop_sway + profile_desktop_niri + profile_desktop_host
platform_fedora -> packages_fedora + services_systemd
platform_rocky -> packages_rocky + services_systemd
wireguard_overlay -> wireguard_overlay (after platform_rocky)
role_aegis -> profile_aegis
atlas -> profile_atlas
role_backend_phase1 -> profile_backend_phase1 (after atlas)
rocky_server -> dotfiles_common + profile_server (after platform_rocky)
platform_fedora & role_personal_workstation -> profile_personal_workstation
platform_fedora & desktop_gnome -> profile_desktop_gnome
@@ -389,6 +477,7 @@ ansible-playbook ansible/site.yml --limit <host> --start-at-task "<task name>" -
ansible-lint ansible/roles/<role>
yamllint ansible/path/to/file.yml
podman-compose -f /opt/docker/server/docker-compose.yml config
ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff
```
## Tags
@@ -403,6 +492,9 @@ ansible-playbook ansible/site.yml --list-tags
| --- | --- |
| `always` | Common pre-tasks, including optional vault loading. |
| `ai_agents` | AI coding-agent install, configuration deployment, and managed-binary removal. |
| `atlas` | Atlas NAS account, storage, sharing, and container configuration. |
| `backend_phase1` | Rootless Atlas Navidrome and Syncthing Quadlets. |
| `containers` | Rootful Atlas Quadlets. |
| `dotfiles` | User configuration across all profiles. |
| `dotfiles:common` | Shared dotfiles. |
| `dotfiles:desktop` | Void and Fedora/GNOME desktop dotfiles. |
@@ -411,10 +503,15 @@ ansible-playbook ansible/site.yml --list-tags
| `dotfiles:workstation` | Personal workstation and WSL dotfiles. |
| `emacs` | Shared Emacs setup and authoring dependencies. |
| `gnome` | Fedora/GNOME desktop configuration. |
| `immich` | Atlas Immich account and Quadlets. |
| `npm` | Global npm packages. |
| `packages` | Package installation and updates. |
| `podman` | Podman Compose and rootless Quadlet integration. |
| `services` | runit and systemd services. |
| `sharing` | Atlas NFSv4 and SMB3 configuration. |
| `storage` | Atlas child ZFS datasets. |
| `tmux` | tmux configuration and plugins. |
| `wireguard` | Prometheus/Atlas WireGuard overlay. |
| `wsl` | WSL bootstrap and configuration. |
## Bootstrapping a new machine