mirror of
https://github.com/fscotto/infra.git
synced 2026-09-27 19:03:47 +00:00
Add Atlas media and storage services [Phase 1] (#9)
* Add Atlas media and storage services * Document Atlas backend phase one and WireGuard deployment * Enable Atlas NAS management and document bootstrap workflow * Harden Atlas network, SSH, firewall, and sharing * Rotate Ansible Vault secrets * Allow configurable Aegis SSH users and authorized keys * Manage Aegis SSH authorized key fragments * Manage SSH authorized key fragments for infrastructure hosts * Harden Rocky storage and sharing configuration * Verify WireGuard handshakes and restore Podman networking
This commit is contained in:
committed by
GitHub
parent
73bf2cd62a
commit
160d63c02d
149
README.md
149
README.md
@@ -108,16 +108,25 @@ That gives it Fedora packages through DNF, Docker from the official repository,
|
||||
dotfiles and templates. The profile provisions configuration only: it does not transfer data, start
|
||||
the Compose stack, update DNS, or perform a cutover.
|
||||
|
||||
The server profile installs platform-specific packages, Podman and podman-compose,
|
||||
declared systemd services, the server Compose stack behind the `podman-compose-server` systemd unit, and firewalld. The Rocky server excludes
|
||||
Syncthing. Rocky bind mounts use private SELinux relabeling for application data while host system
|
||||
files remain unchanged.
|
||||
The server profile installs platform-specific packages, Podman and podman-compose, declared systemd
|
||||
services, and firewalld. The manually activated `podman-compose-server` unit contains the existing
|
||||
Nginx Proxy Manager and Gitea services. The desired Compose file no longer includes Navidrome,
|
||||
Syncthing, or the obsolete Navidrome PostgreSQL database. Navidrome and Syncthing belong to Atlas;
|
||||
official Navidrome uses SQLite instead. Applying the profile does not stop or remove legacy
|
||||
containers and does not delete `/opt/postgres/data`.
|
||||
|
||||
Firewalld enables SSH, Cockpit (`9090/tcp`), HTTP and HTTPS. Nginx Proxy Manager publishes only
|
||||
`80/tcp` and `443/tcp`; its administration interface is bound to `127.0.0.1:81` and can be reached
|
||||
from Ikaros or Nymph with the `npm-tunnel` Bash alias. Nextcloud remains disabled and the profile
|
||||
does not provision any `/srv/nextcloud` directories.
|
||||
|
||||
The Atlas phase-one work does not change this NPM deployment or its persistent data. Once WireGuard
|
||||
and the Atlas services are active, configure the current NPM proxy hosts with Navidrome upstream
|
||||
`http://10.0.0.2:4533` and Syncthing GUI upstream `http://10.0.0.2:8384`. Only the Syncthing web GUI
|
||||
uses NPM; synchronization traffic remains on explicitly published native ports bound only to the Atlas
|
||||
WireGuard address. Configure both Syncthing authentication and an appropriate NPM access policy before
|
||||
publishing its GUI.
|
||||
|
||||
Server identity comes from `server_username`, `server_user_group`, and `server_user_home` in `ansible/inventory/group_vars/server.yml`. `server_username` defaults to `username`, but it can be overridden, for example:
|
||||
|
||||
```bash
|
||||
@@ -128,6 +137,8 @@ ansible-playbook ansible/site.yml --limit prometheus \
|
||||
```
|
||||
|
||||
The target must already provide `server_username` with local sudo access.
|
||||
Prometheus authorizes its declared SSH public keys through separate files below
|
||||
`~/.ssh/authorized_keys.d/`, while `sshd` is configured to read those files directly.
|
||||
|
||||
### DuckDNS
|
||||
|
||||
@@ -150,7 +161,7 @@ back in; preserve any uncommitted work separately without copying secrets.
|
||||
### Data migration
|
||||
|
||||
Provision Rocky first, then run the migration script **on the retired Ubuntu source host**. It is
|
||||
dry-run by default and requires an explicit source-stack stop before it can copy PostgreSQL data:
|
||||
dry-run by default and requires an explicit source-stack stop before it can copy application data:
|
||||
|
||||
```bash
|
||||
sudo ./scripts/migrate_prometheus_data.sh \
|
||||
@@ -163,11 +174,11 @@ sudo ./scripts/migrate_prometheus_data.sh \
|
||||
--quiesce-source --execute
|
||||
```
|
||||
|
||||
The script copies Navidrome, music, Nginx Proxy Manager, PostgreSQL and Gitea data. It does not
|
||||
delete data, move Syncthing, copy `/home/git/.ssh`, start containers, update DNS, or perform a
|
||||
cutover. The destination SSH host key must already be trusted and the destination account needs
|
||||
passwordless sudo for `rsync`. It preserves ACLs but not extended attributes, so source SELinux labels
|
||||
are not transferred; the Rocky Compose bind mounts apply their own `:Z` labels when containers start.
|
||||
The script copies only Nginx Proxy Manager and Gitea data. It does not delete data, move
|
||||
Navidrome/Syncthing, copy `/home/git/.ssh`, start containers, update DNS, or perform a cutover. The
|
||||
destination SSH host key must already be trusted and the destination account needs passwordless sudo
|
||||
for `rsync`. It preserves ACLs but not extended attributes, so source SELinux labels are not
|
||||
transferred; the Rocky Compose bind mounts apply their own `:Z` labels when containers start.
|
||||
|
||||
## DNS Filter
|
||||
|
||||
@@ -191,6 +202,10 @@ for AdGuard while retaining DNS learned from the router. Define
|
||||
`vault_aegis_icloudpd_apple_id` in Vault before applying it. iCloudPD still requires interactive MFA
|
||||
initialization after its first deployment.
|
||||
|
||||
New Aegis images create the `admin` account in Butane. Before configuring a newly imaged node, run its
|
||||
first playbook execution with `-e ansible_user=admin`; the SSH hardening role then permits that same
|
||||
account. Keep the inventory on `pi` until the existing node has been replaced.
|
||||
|
||||
Validate the profile before deployment:
|
||||
|
||||
```bash
|
||||
@@ -200,28 +215,97 @@ ansible-playbook ansible/site.yml --limit aegis --check --diff --ask-become-pass
|
||||
|
||||
## NAS
|
||||
|
||||
`atlas` is a Rocky Linux 9 NAS reached through SSH. Its pool already exists: the profile only
|
||||
manages child datasets and must never create, partition, destroy, roll back, or otherwise alter the
|
||||
pool itself. Linux clients use NFSv4 and Windows/WSL clients use SMB; both are restricted to the
|
||||
configured LAN.
|
||||
`atlas` is a Rocky Linux 9 NAS reached through SSH. Normally its pool already exists and the profile
|
||||
only manages child datasets. A one-time RAIDZ2 bootstrap is available only with explicit confirmation
|
||||
(`atlas_create_pool=true`) and exactly four verified `/dev/disk/by-id/...` paths in `atlas_zpool_disks`.
|
||||
It never partitions, forces, destroys, rolls back, or changes the vdev layout of an existing pool. Linux
|
||||
clients use NFSv4 and Windows/WSL clients use SMB; both are restricted to the configured LAN.
|
||||
|
||||
For the first run, replace the Atlas placeholders and provide
|
||||
`vault_atlas_authorized_ssh_keys`, `vault_atlas_admin_password_hash`, and
|
||||
`vault_atlas_samba_password`. Bootstrap the host through its existing administrator:
|
||||
For the first run, provide `vault_atlas_admin_password_hash`, `vault_atlas_samba_password`, and
|
||||
`vault_atlas_immich_db_password`. Bootstrap the host through its
|
||||
existing administrator. Open `51820/udp` towards Prometheus in the provider firewall first, then
|
||||
include both WireGuard peers in the same idempotent playbook run:
|
||||
|
||||
```bash
|
||||
ansible-playbook ansible/site.yml --limit atlas \
|
||||
-e atlas_connection_username=<existing-admin>
|
||||
ansible-playbook ansible/site.yml --limit prometheus,atlas \
|
||||
-e atlas_connection_username=<existing-admin> \
|
||||
-e atlas_create_pool=true
|
||||
```
|
||||
|
||||
`vault_atlas_admin_password_hash` must be an `/etc/shadow`-compatible hash, not a clear-text
|
||||
Cockpit password. Subsequent runs use `atlas_admin_username`. Enable
|
||||
`atlas_manage_storage` only after checking the existing pool and mountpoints; enable
|
||||
`atlas_manage_firewall` only after checking the LAN subnet and active firewalld zone.
|
||||
The explicit pool gate is safe to repeat: the role creates the RAIDZ2 pool only when it is absent.
|
||||
WireGuard waits for a real peer handshake before the play continues.
|
||||
|
||||
Snapshot retention, Syncthing topology, VPN access, Prometheus pulls, encrypted Borg backups to a
|
||||
Hetzner Storage Box, USB backup, monitoring, and disaster-recovery tests remain follow-up work. The
|
||||
detailed operational backlog is kept in `AGENTS.md`.
|
||||
`vault_atlas_admin_password_hash` must be an `/etc/shadow`-compatible hash, not a clear-text
|
||||
Cockpit password. Subsequent runs use `atlas_admin_username`. Atlas declares storage, sharing, and its
|
||||
LAN firewall rules enabled. Before the first apply, check the existing pool and mountpoints, LAN subnet,
|
||||
and active firewalld zone. `atlas_manage_media_stack` remains disabled until `/dev/dri`, the container
|
||||
paths, and the Immich database secret are validated. Atlas reads its declared SSH public keys from
|
||||
separate files below `~/.ssh/authorized_keys.d/`.
|
||||
|
||||
With storage management enabled, Atlas creates the complete dataset hierarchy below the existing or
|
||||
explicitly bootstrapped `zpool`: `work`, `archive`, `archive/app_data`, the separate `archive/app_data/navidrome` and
|
||||
`archive/app_data/syncthing` application datasets, `media`, `media/music`, `media/photobook`,
|
||||
`backups`, `backups/services`, and `backup_prometheus`. Application/archive datasets use `zstd`,
|
||||
while media, Syncthing and service-backup datasets use `lz4`; `backups/services` also has a `500G`
|
||||
refreservation. Atlas enforces targeted SELinux persistently and reports, without initiating, any reboot required to activate it. It assigns its primary LAN interface explicitly to the managed firewalld zone and applies persistent kernel network hardening: redirects and source routes are rejected, martians logged, reverse-path filtering remains loose for WireGuard, and IPv4 forwarding is disabled. SSH permits only the declared administrator using public-key authentication; root login, passwords,
|
||||
agent and remote forwarding are disabled, while local forwarding remains available for private administrative tunnels. SMB3 exposes `Archive` only to the configured Vault-backed
|
||||
Samba accounts on encrypted, signed SMB3 over TCP/445 only and admits the configured LAN without host-specific
|
||||
exclusions. NFSv4 exports only `media/photobook` to the configured Aegis IP over TCP/2049, using
|
||||
`all_squash` with anonymous UID/GID `1100`.
|
||||
|
||||
The `immich` system account is fixed to UID/GID `1100`, has no login shell or `wheel` membership, and
|
||||
receives `video` and `render` access. The rootful Immich Server, ML, Redis-compatible cache, PostgreSQL,
|
||||
and NPM Quadlets share one Podman network. Immich runs as `1100:1100`; Server and ML receive `/dev/dri`,
|
||||
and Photobook is mounted read-only at `/external/photobook`. NPM publishes ports `80` and `443`; its
|
||||
administration interface remains restricted to `127.0.0.1:81` for SSH-tunnel access.
|
||||
|
||||
Phase 1 is limited to rootless Navidrome and Syncthing user Quadlets on Atlas. It is enabled in the
|
||||
Atlas host configuration and can be set to `false` only for a deliberate suspension. Official Navidrome `0.63.2` uses its SQLite database below `/data`; it does
|
||||
not support `ND_DATABASE_URL` or an external PostgreSQL backend. The obsolete `navidromedb` service
|
||||
was therefore removed from Prometheus instead of being reproduced on Atlas. The role derives all
|
||||
storage paths from the `zpool` mounted at `/zpool`: music is read-only at
|
||||
`/zpool/media/music`, Navidrome application state and `navidrome.db` are stored at
|
||||
`/zpool/archive/app_data/navidrome`, and Syncthing persists at
|
||||
`/zpool/archive/app_data/syncthing`. `profile_atlas` creates these datasets when
|
||||
`atlas_manage_storage` is enabled; the backend role verifies their exact mountpoints before starting
|
||||
containers. The backend role never creates the pool. The separate `wireguard_overlay` role manages `wg0`
|
||||
between Prometheus (`10.0.0.1`) and Atlas (`10.0.0.2`), generating private keys once
|
||||
on their respective hosts and exchanging only public keys through Ansible. Prometheus alone opens
|
||||
`51820/udp` publicly. When the WireGuard zone is created, Ansible reloads firewalld and immediately
|
||||
reloads Prometheus' rootful Podman networks so the existing proxy stack retains container DNS and
|
||||
connectivity. Backend ports are admitted only in the WireGuard firewalld zone.
|
||||
|
||||
`backend_phase1_start_services` stays false during the application-state transfer, so the first real
|
||||
backend run renders the Quadlets without creating an empty Atlas database. After stopping Navidrome
|
||||
on Prometheus, copy the complete `/opt/navidrome/data/` directory into
|
||||
`/zpool/archive/app_data/navidrome/`, preserving `navidrome.db` and any SQLite sidecar files. Then set
|
||||
this variable to true and rerun the role to enable and start Navidrome and Syncthing. The playbook
|
||||
never copies or deletes application data.
|
||||
|
||||
Validate and render the Atlas services with:
|
||||
|
||||
```bash
|
||||
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags storage
|
||||
|
||||
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
||||
ansible-playbook ansible/site.yml --limit prometheus,atlas --tags wireguard
|
||||
|
||||
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags backend_phase1 --check --diff
|
||||
|
||||
ANSIBLE_LOCAL_TEMP=/tmp/ansible-local \
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags backend_phase1
|
||||
```
|
||||
|
||||
For the cutover, stop the old Navidrome writer before copying its data directory, verify ownership by
|
||||
the Atlas `admin` account and confirm that the copied SQLite database is present before changing
|
||||
`backend_phase1_start_services: true` in `host_vars/atlas.yml`. Keep the source data and the stopped
|
||||
legacy `navidromedb` container until Navidrome on Atlas and a restore test have been validated.
|
||||
|
||||
Snapshot retention, Syncthing topology, WireGuard/firewall validation, Prometheus backup pulls,
|
||||
encrypted Borg backups to a Hetzner Storage Box, USB backup, monitoring, and disaster-recovery tests
|
||||
remain follow-up work. The detailed operational backlog is kept in `AGENTS.md`.
|
||||
|
||||
## How layering works
|
||||
|
||||
@@ -308,6 +392,8 @@ ansible-playbook ansible/site.yml --limit deadalus --tags ai_agents --check --di
|
||||
| `profile_workstation_dev_wsl` | WSL development setup. |
|
||||
| `profile_server` | Server setup. |
|
||||
| `profile_atlas` | Rocky Linux 9 NAS setup. |
|
||||
| `profile_backend_phase1` | Rootless Navidrome and Syncthing on Atlas. |
|
||||
| `wireguard_overlay` | Prometheus/Atlas WireGuard overlay. |
|
||||
| `profile_aegis` | Fedora IoT always-on LAN node. |
|
||||
| `dotfiles_common` | Shared user dotfiles. |
|
||||
|
||||
@@ -319,8 +405,10 @@ platform_void -> packages_void + services_runit
|
||||
platform_void & graphical_desktop -> profile_desktop_common + profile_desktop_sway + profile_desktop_niri + profile_desktop_host
|
||||
platform_fedora -> packages_fedora + services_systemd
|
||||
platform_rocky -> packages_rocky + services_systemd
|
||||
wireguard_overlay -> wireguard_overlay (after platform_rocky)
|
||||
role_aegis -> profile_aegis
|
||||
atlas -> profile_atlas
|
||||
role_backend_phase1 -> profile_backend_phase1 (after atlas)
|
||||
rocky_server -> dotfiles_common + profile_server (after platform_rocky)
|
||||
platform_fedora & role_personal_workstation -> profile_personal_workstation
|
||||
platform_fedora & desktop_gnome -> profile_desktop_gnome
|
||||
@@ -389,6 +477,7 @@ ansible-playbook ansible/site.yml --limit <host> --start-at-task "<task name>" -
|
||||
ansible-lint ansible/roles/<role>
|
||||
yamllint ansible/path/to/file.yml
|
||||
podman-compose -f /opt/docker/server/docker-compose.yml config
|
||||
ansible-playbook ansible/site.yml --limit atlas --tags storage,sharing,containers --check --diff
|
||||
```
|
||||
|
||||
## Tags
|
||||
@@ -403,6 +492,9 @@ ansible-playbook ansible/site.yml --list-tags
|
||||
| --- | --- |
|
||||
| `always` | Common pre-tasks, including optional vault loading. |
|
||||
| `ai_agents` | AI coding-agent install, configuration deployment, and managed-binary removal. |
|
||||
| `atlas` | Atlas NAS account, storage, sharing, and container configuration. |
|
||||
| `backend_phase1` | Rootless Atlas Navidrome and Syncthing Quadlets. |
|
||||
| `containers` | Rootful Atlas Quadlets. |
|
||||
| `dotfiles` | User configuration across all profiles. |
|
||||
| `dotfiles:common` | Shared dotfiles. |
|
||||
| `dotfiles:desktop` | Void and Fedora/GNOME desktop dotfiles. |
|
||||
@@ -411,10 +503,15 @@ ansible-playbook ansible/site.yml --list-tags
|
||||
| `dotfiles:workstation` | Personal workstation and WSL dotfiles. |
|
||||
| `emacs` | Shared Emacs setup and authoring dependencies. |
|
||||
| `gnome` | Fedora/GNOME desktop configuration. |
|
||||
| `immich` | Atlas Immich account and Quadlets. |
|
||||
| `npm` | Global npm packages. |
|
||||
| `packages` | Package installation and updates. |
|
||||
| `podman` | Podman Compose and rootless Quadlet integration. |
|
||||
| `services` | runit and systemd services. |
|
||||
| `sharing` | Atlas NFSv4 and SMB3 configuration. |
|
||||
| `storage` | Atlas child ZFS datasets. |
|
||||
| `tmux` | tmux configuration and plugins. |
|
||||
| `wireguard` | Prometheus/Atlas WireGuard overlay. |
|
||||
| `wsl` | WSL bootstrap and configuration. |
|
||||
|
||||
## Bootstrapping a new machine
|
||||
|
||||
Reference in New Issue
Block a user