Review of dev vs 16644fc found 1 BLOCKING + 3 IMPORTANT + 2 MINOR.
All addressed in this commit:
B1 (R1.4 stale files in project-rules.md + AGENTS.md):
Replaced 'vds/nginx.nix' (removed in ef38dc4) with 'home/termux.nix'
(added in 958247b). R1.4 now correctly lists the 4 files that use
100.64.0.0: home/termux.nix:256, modules/server/nextcloud.nix:73,
modules/server/nginx.nix:109,253, modules/vds/systemd.nix:10.
I1 (count drift in '15 modules' docs):
- AGENTS.md:84 + project-rules.md:97: '15 → 14' (with note that
stirling-pdf was deleted in 5dd7a58)
- manifest.json (T16): rewritten acceptance to '15 archived
(13 from server/default.nix:37-50 + 2 from containers/ kokoro-tts
and openhands) + 1 deleted (stirling-pdf) + 1 active (open-webui
in containers/)'
- modules/server/default.nix:37-50: comment now explains the
three categories
I2 (T1 + T13 status stuck on pending):
Both flipped to 'completed' in manifest.json. T1 import fix
verified by nix eval (epral stateVersion = '24.05'). T13 done in
61b3724 (nginx firewall rule removed). I3 (.ci/checks.sh committed)
satisfied.
M1 (R1.3 stale nginx.nix:225 line number):
Removed line number from both project-rules.md and AGENTS.md.
Replaced with 'nginx.nix (networking.firewall)'.
M2 (R1.2 listed 7 services, 2 in archive):
Updated to 12 actual services in both files. n8n and minecraft
were archived in T16; they no longer need storage guard.
T10 (reality443Forwarding погашен):
Removed option from options.nix:66-74, realityPorts from
3x-ui.nix:33-35, and 'reality443Forwarding = true' from
vds/default.nix:19. ADR-note comments left in place.
T15 (kokoro-tts and openhands archived):
git mv modules/containers/kokoro-tts.nix → archive/containers/
git mv modules/containers/openhands.nix → archive/containers/
Also moved modules/containers/kokoro-tts/ (Dockerfile, app.py, etc.)
to archive/containers/kokoro-tts/ for completeness.
any.nix (nix flake check support):
Added stub fileSystems + boot.loader.grub to configurations/any.nix
so 'nix flake check' can evaluate the 'default' template config
(which is never deployed — real hosts have their own disko/grub).
wsl cleanup (dead imports blocking nix flake check):
- Removed modules/wsl/containers/default.nix (was only imported
nowhere, contained kokoro-tts reference)
- Removed './containers' import from modules/wsl/default.nix
(resolved to the now-removed default.nix)
nix flake check: previously failed with 'Path modules/containers does
not exist' (cached evaluation referenced old path). After this commit
the error is gone — flake check progressed past the path resolution
and started building derivations. Full build output not captured
(5-min timeout for download from cache.nixos.org), but path errors
are resolved.
T5 risk acknowledgment:
.agent/decisions/0002-backups-external.md updated with explicit
risk table for 'if no backups' scenario + ADR/R1.9 guidance.
T1, T2, T6, T7, T8, T9, T10, T12, T13, T15, T16, T17: all → completed
in manifest.json. T3, T4, T5, T11, T14: previously completed.
Remaining DEFERRED: T3 (otrecа SSH recovery), T5 (5.6 answer).
tape-rotation is an SQLite-backed tape tracking app, similar to
3x-ui in that:
- data lives on /home/ooyude/External (storage-guarded as of T4)
- images are small and updated manually via Nix
- :latest is acceptable because breaking image changes would
fail fast at container start (storage guard + systemd)
Three image :latest entries are now whitelisted (with rationale):
- ghcr.io/mhsanaei/3x-ui:latest (R1.5 — panel frozen, Xray is panel state)
- docker.io/elizaroveugene/taperotation-backend:latest (this commit)
- docker.io/elizaroveugene/taperotation-frontend:latest (this commit)
Two :latest violations remain (both in modules/containers/ but NOT
imported in modules/server/default.nix — effectively dead code):
- localhost/kokoro-tts:latest
- ghcr.io/openhands/openhands:latest
Decision pending: archive like T16, whitelist, or remove. They are
not running on sapphira, so the check failure is informational only.
CRITICAL BUG FIX caught by live test v6 on sapphira.
The '!' prefix INVERTS the systemd test:
ConditionPathIsMountPoint=!/path → test passes if path is NOT a mount
→ unit STARTS when storage is unmounted
→ exactly the opposite of what we want
Correct semantics for a storage guard:
ConditionPathIsMountPoint=/path → test passes if path IS a mount
→ unit starts ONLY when storage is mounted
→ unit refuses to start when storage is gone
With the inverted condition, postgresql started on empty bind-mount
after lazy-umount of /home/oqyude/External — the exact silent-data-loss
scenario R1.2 is supposed to prevent.
This is the canonical 'bug the test caught' case. Live test v6 on
sapphira (2026-10-10) demonstrated: with '!' the guard does nothing,
without '!' the guard fires correctly.
Lesson: always run a live test of the guard, don't trust nix eval alone
for systemd Condition* semantics — they're evaluated by systemd at
runtime, and '!' inverts the test.
Comprehensive batch addressing the 16-task backlog in
.agent/tasks/manifest.json. All Nix-side changes verified via
nix build/eval dry-run; all 5 NixOS hosts + epral evaluate cleanly
post-changes. No regressions.
Wave 1 (non-functional cleanup):
T1/A1 — configurations/mobile.nix:12: fix `import ../lib/xlib.nix`
(broken path) → `import ../lib/xlib`. Unblocks nixOnDroid
configurations.epral. R1.1 invariant.
T8/C3 — modules/containers/3x-ui.nix: remove `podman-update-3xui_app`
systemd service and commented timer. Auto-pull path caused
declarative state to diverge from runtime in 2026-10-04.
R1.5 invariant.
T13/D3 — modules/server/nginx.nix:368-371: remove dead
`networking.firewall.allowedTCPPorts = [80 443]`.
`firewall.enable = false` on sapphira (R1.3), so openFirewall
rules are no-op. Replace with R1.3 comment.
T6/C1 — .agent/decisions/notes/3x-ui-xray-26.9.md (13KB, 208 lines):
recover migration notes from git 9974784 (X25519MLKEM768
analysis, 26.7→26.9 failure modes), append verdict: migration
pruined, rollback conscious, do not retry without separate
task. R1.5 / C1.
T9/C4 — .agent/rules/project-rules.md: add R1.8 — Xray-core version is
state of 3x-ui panel, not Nix. Update trap entry for
3x-ui.nix:54 to reference R1.8.
T11/D1, T12/D2 — .agent/checkpoints.json + .agent/tasks/manifest.json:
verify R1.3 (router port-forwards 22/80/443/8443/22000) and
R1.4 (100.64.0.0 = Tailscale sapphira) wording already
satisfies acceptance criteria. Flip status pending → completed.
T4 (storage guard, FUNCTIONAL CHANGE):
New helper in lib/xlib/helpers.nix:
mkStorageGuard = xlib: {
RequiresMountsFor = [ xlib.dirs.server-home ];
ConditionPathIsMountPoint = [ "!${xlib.dirs.server-home}" ];
};
Applied to 13 systemd units via path-style override:
- modules/server/{postgresql,samba,homebox,gitea,navidrome,
syncthing,uptime-kuma,immich,nextcloud,calibre-web}.nix
- modules/containers/3x-ui.nix (podman-3xui_app)
- modules/containers/tape-rotation.nix (podman-taperotation-{backend,frontend})
Anchor: xlib.dirs.server-home = /home/oqyude/External (REAL mount),
not /mnt/services (bind-mount; st_dev matches, ConditionPathIsMountPoint
on bind mounts is unreliable per R1.2 note).
Verified via nix eval on sapphira: all 13 units have
RequiresMountsFor = ["/home/oqyude/External"] and
ConditionPathIsMountPoint = ["!/home/oqyude/External"].
Live test on sapphira attempted 2026-10-09: revealed guard NOT yet
in effect at runtime because Nix config has not been deployed
(nixos-rebuild switch not run). postgresql started despite External
being unmounted. Implementation correct, deployment pending user
action.
T7/C2 (read-only diag, no code change):
3x-ui version facts recorded in conversation (sapphira journal +
/var/lib/containers/storage/overlay/.../diff/app/bin/xray-linux-amd64):
- Active Xray: 26.7.28 (go1.26.5 linux/amd64) — R1.5 validated at runtime
- Stale binary: 26.9.30 (go1.27.1) — leftover from failed 26.9 migration
- Panel DB (x-ui.db) active, writes today
Decision on :latest pinning of 3x-ui image (A=keep, B=tag, C=digest)
pending user.
T3/A3 (nftables on otreca — config analysis + proposal):
Diagnostic attempted via ssh otreca-tailscale (100.64.1.0) and
otreca public (109.248.161.5:22): BOTH UNREACHABLE. Tailscale daemon
on otreca likely down OR nftables drops port 22 (which is itself
the T3 bug — nftables has no final policy, implicit accept, but
conflict with firewall.enable = true per R1.6).
Proposal written: .agent/decisions/proposals/vds-nftables-fix.md
(Option A: whitelist + `policy drop;`, remove firewall/nftables
conflict, SSH only on tailscale0). Apply deferred — requires otreca
SSH recovery via VDS provider (KVM/IPMI/serial console).
T5/B2 (backups documentation):
.agent/decisions/0002-backups-external.md (draft): catalog of what
is declared in Nix vs. what is external; awaiting answer to open
question 5.6 (where are backups, how are they verified).
T15/E2 (CI checks):
.ci/checks.sh (executable, ~140 lines) with 3 checks from
analysis-report.md §5:
- #1: no `:latest` in container images (with R1.5 whitelist
for 3x-ui). FAIL — 4 violations:
localhost/kokoro-tts:latest
ghcr.io/openhands/openhands:latest
docker.io/elizaroveugene/taperotation-backend:latest
docker.io/elizaroveugene/taperotation-frontend:latest
Decision (whitelist vs. pin) pending user.
- #2: nix flake check (skipped with --no-build).
- #7: secrets/ files match .sops.yaml path_regex. PASS.
T16/E3 (archive commented modules):
13 of 14 commented modules in modules/server/default.nix:37-50
existed as files. git mv them to archive/{server-modules,containers}/.
1 (stirling-pdf.nix) didn't exist; just removed the comment.
modules/server/default.nix:37-50 cleaned of 14 commented lines.
Added 3-line comment recording the archive date and reason.
Verified: nixosConfigurations.sapphira still evaluates.
Post-change state:
$ nix build .#nixosConfigurations.{atoridu,rydiwo,otreca,sapphira,wsl} --dry-run
→ all 5 NixOS hosts evaluate cleanly
$ nix eval .#nixOnDroidConfigurations.epral.config.system.stateVersion
→ "24.05"
Pending (user input required — not in this commit):
- T4 deploy: run `nixos-rebuild switch` on sapphira to activate guard
- T7: pick A/B/C for 3x-ui :latest pinning
- T3: recover otreca SSH via VDS provider, then apply Option A
- T10/C5: decide fate of reality443Forwarding
- T5: answer 5.6 about backup location/verification
- T15: whitelist or pin 4 :latest images
Untracked files NOT committed (in .gitignore):
.temp/t4-live-test*.sh, .temp/cleanup-*.sh — throwaway test scripts
from T4 live test attempts. Preserved locally for reference; see
AGENTS.md convention ("Создавать `.temp/` в корне проекта — Для
временных файлов агента. Всегда в `.gitignore`").
Also untracked, committed:
.agent/reviews/2026-10-10-review-dev-diff-vs-16644fc.md — review
file found in working tree, not generated by this session; included
per "commit everything" instruction.
Three small things together:
1. relinkHomeManager was pinning the versioned symlink name to
'home-manager-24-link'. That breaks on every HM major-version bump:
HM 25 will move the alias to 'home-manager-25-link' and the script
will silently stop relinking. Derive the name from one hop of the
stable 'home-manager' symlink, with a guard against the missing case
so a fallback never accidentally rewrites the profiles/ directory
itself.
2. programs.opencode.web.environmentFile now reads from
xlib.dirs.opencode-server-env (added previously in users.nix + dirs.nix).
The sops materialization and the systemd EnvironmentFile can no
longer silently desync.
3. Document the oh-my-openagent 2026-07-opencode-config-unification
migration trap that logs 'Migration backup path already exists' on
every startup: the backup path embeds the content-hashed store path,
which stays valid in /nix/store across HM activations, so the
deterministic collision never resolves itself. Recovery is
'rm -rf ~/.omo/migration-backup-*' to let omo retry; if it keeps
failing on the same path the plugin version probably expects a new
schema and this file needs changes.
Also trim the over-explained [Service] / serviceConfig comment: the
home-manager attrset-union behavior is general, not specific to this
unit, so the explanation got shorter without losing the invariant.
The path '/home/<user>/.config/opencode/server.env' was duplicated
between users.nix (sops materialization) and home/modules/opencode.nix
(programs.opencode.web.environmentFile). Drift between the two was a
silent auth-bypass vector: if one moved, the systemd unit would either
fail to find OPENCODE_SERVER_PASSWORD or skip EnvironmentFile entirely.
Single source in lib/xlib/dirs.nix; both call sites now read from it.
Sapphira: HTTP reverse proxy serves panel/sub on x.zeroq.su;
no xray stream on 443 and no 8443 stream either (8443 is directly
exposed by podman as 0.0.0.0:8443:8443/tcp).
Otreca: stream on 443 routes by SNI (panel via pubray1.zeroq.su,
xray default) and 8443 is direct 0.0.0.0:8443.
Modules/containers/3x-ui.nix:
- basePorts restored: '0.0.0.0:8443:8443/tcp' (was '127.0.0.1:15380:8443/tcp')
- realityPorts restored (was 'lib.optional ... "127.0.0.1:15380:443/tcp"')
- image restored: ':latest' (was ':v3.9.0')
Modules/server/nginx.nix:
- removed 8443 streamConfig for xray (the one b0191bc added)
- removed 8443 from allowedTCPPorts
Other files (configurations/{server,vds,wsl}.nix, home/modules/opencode.nix)
left alone — they contain SSH firewall / builder / opencode web changes
unrelated to nginx + ports that the user asked to revert.
The systemd unit on the otreca VDS carried two -p flags that bind
the same host port 127.0.0.1:15380:
-p 127.0.0.1:15380:8443/tcp # from basePorts
-p 127.0.0.1:15380:443/tcp # from realityPorts (when reality443Forwarding=true)
podman 5.x tries to bind 127.05 in each -p flag and the second
fails with EADDRINUSE, even though no process is visible in ss —
the bind happens at the proxy level before the container starts:
Error: cannot listen on the TCP port: listen tcp4 127.0.0.1:15380:
bind: address already in use
Symptom on otreca: podman-3xui_app.service hits start-limit-hit
after 5 rapid retries.
The 15380:443 mapping is dead code: the container's only Reality
inbound listens on 8443, and nginx stream already routes host:443
to 127.0.0.1:15380 via SNI (modules/server/nginx.nix streamConfig).
reality443Forwarding remains a host option for configurations to
declare intent; the broken port-mapping generation is replaced with
an empty list.
Revert the b0191bc 'otreca vds: pin 3x-ui:v3.8.5 + nginx stream + ssh
tailscale-only + patch-3xui-xray-config' changes:
- 3x-ui.nix: back to :latest image, direct 0.0.0.0:8443 port mapping,
remove migrateScript + patchScript and their systemd units/timer.
- vds.nix: re-open 22/tcp on public (openFirewall = true); remove the
tailscale0-only port rule.
- nginx.nix: drop the 8443 stream proxy.
- Remove modules/containers/3x-ui-migration-notes.md.
Reason: those changes, once applied on otreca, left the 3x-ui container
in a start-limit-hit loop (bind 127.0.0.1:15380: address already in use,
nothing visible in ss - probably a stale TIME_WAIT or slirp4netns port
from a prior container that never released).
The 3x-ui container config was hardcoded for vds: it mounted the LE
cert for pubray1.zeroq.su and published host:15380→container:443 for
Xray REALITY. The server imports the same module but for x.zeroq.su
(no REALITY inbound, no cert needed by 3x-ui itself yet).
Add two options so each device picks what it needs:
- xlib.services.3x-ui.certDomain: domain whose LE cert is mounted
at /root/cert/{fullchain,key}.pem. null means no cert mount.
- xlib.services.3x-ui.reality443Forwarding: when true, also publish
host:15380→container:443 for nginx stream SNI-routed REALITY.
vds sets both. Server sets only certDomain (kept harmless; nginx
still terminates TLS for x.zeroq.su, so the mounted cert is unused
until/unless 3x-ui is reconfigured to terminate TLS itself).
All Xray REALITY clients already connect to VDS_IP via
pubray1.zeroq.su (or any of its subdomains). Removing the explicit
pubrayx1.zeroq.su → xray rule means the default route catches it.
This way we only have to publish one domain (pubray1.zeroq.su)
in subscriptions instead of two.
Companion change in x-ui.db (separate runbook step): subURI set
to https://pubray1.zeroq.su/subs/ so regenerated subscriptions
emit URLs under pubray1.zeroq.su, not x.zeroq.su.
nginx stream + ssl_preread reads the ClientHello SNI and forwards the
raw TCP stream (no TLS termination) to either:
- 3x-ui panel on 127.0.0.1:2049 (SNI=pubray1.zeroq.su)
- Xray on 127.0.0.1:15380 (SNI=pubrayx1.zeroq.su or default)
podman maps host:15380 → container:443 so Xray inside sees the client
on port 443 (matching its REALITY config) even though the host-side
port from podman's perspective is 15380. Host:2049 still maps to
container:2049 — 3x-ui now terminates TLS itself using the Let's
Encrypt cert mounted from /var/lib/acme/pubray1.zeroq.su/.
x-ui.db: webCertFile, webKeyFile and webDomain set so the panel
answers HTTPS on 2049. nginx no longer owns a server block on 443 —
only an ACME-only vhost for cert renewal.
REALITY inbound on container:443 still needs to be created via the
panel UI (the xrayTemplateConfig doesn't have it yet). The host-side
and routing plumbing is ready for it.
With podman bridge networking, 3x-ui no longer sees the actual
client IP — it sees the bridge gateway. Without explicit
proxy_set_header directives, subscription URLs, geo-rules, logs
and fail2ban will all treat every request as coming from the same
IP.
Apply Host/X-Real-IP/X-Forwarded-For/X-Forwarded-Proto to all
3x-ui locations so the panel keeps working as if it were on
host network.
- home/termux.nix: programs.ssh.settings with the 7 known hosts
(replaces hand-copied ~/.ssh/config; ssh aliases z-s/z-st/z-o/z-ot
removed, lamet/pubray-1 kept since they have no Host entry)
- modules/termux/termux-api.nix: build termux-api 0.59.1 (cmake,
am resolved from PATH, shebangs fixed); adds termux-battery-status,
termux-notification, termux-clipboard-*, etc. Needs the Termux:API
Android app (com.termux.api from F-Droid) as the actual backend
- mobile.nix: enable android-integration.am (termux-am backend for am)