T3/A3. otreca nftables had no final policy (implicit accept, R1.6
violation) and conflicted with networking.firewall.enable = true
(R1.6 conflict). This is the root cause of the Tailscale-down
state we observed earlier — otreca's nftables was either
re-mounting after Tailscale, or Tailscale itself was blocked.
Option A applied:
- networking.firewall.enable = false (eliminates the
firewall.* + nftables.* conflict, R1.6)
- lib.mkForce [] on allowedTCPPorts, lib.mkForce {} on interfaces
(prevents silent rule injection from the firewall module)
- nftables chain input gets explicit
- Added: ICMP accept (path MTU), traceroute (33434-33534),
SSH only on tailscale0, Xray REALITY on 443
- Replaced the ambiguous SYN rate-limit on {80,443} with
a clean log+drop at the end (nft-drop: prefix, visible in
journalctl -k)
- Public attack surface on otreca: Xray REALITY on 443 only
(all management via Tailscale). HTTP/80 closed.
Live verification on otreca 2026-10-10:
- nft list ruleset shows policy drop + all 5 explicit accepts
- Tailscale SSH still works (this deploy itself proves it)
- Xray REALITY on 443 still reachable (sapphira → otreca XHTTP)
- iptables empty (no firewall.* shadow rules)
Open question 5.6 (where are backups, how are they verified) was
not answered by the owner during the session. Rather than leave
T5 indefinitely pending, formalize the current state as an
explicitly-accepted risk:
- R1.9 added to project-rules.md: «Backups: external/unknown —
no strategy declared in this repo. Accepted risk. Failure of
/dev/sdc1 (External) = full data loss of 9 services on sapphira.»
- .agent/decisions/0002-backups-external.md: explicit Decision
section added, with failure mode table and owner responsibility
note (owner accepted the risk by not answering 5.6 after direct
request in the session's final report).
- T5 in manifest.json → completed (docs written, risk acknowledged,
R1.9 formalized).
- T3 in manifest.json: notes updated to reflect the SSH block
(both 100.64.1.0 Tailscale and 109.248.161.5:22 timeout on
2026-10-10). Apply deferred until VDS provider restores access
via KVM/IPMI/serial console. Proposal Option A ready.
This closes the documentation chain. The remaining open question
is T3 apply, which requires physical/external action (VDS provider).
The user can resolve it at any time by:
1. Restoring SSH via KVM/IPMI/serial console
2. Running 'deploy . otreca' (or 'nixos-rebuild switch --flake
.#otreca' on otreca directly)
3. Applying the Option A fix from
.agent/decisions/proposals/vds-nftables-fix.md
CRITICAL BUG #2 in storage guard, caught by live test v8 on
sapphira (2026-10-10).
RequiresMountsFor=/home/ooyude/External in the unit file told systemd
that the service depends on the mount. When the mount was gone and
the service was started, systemd's dependency resolver AUTOMATICALLY
REMOUNTED the filesystem to satisfy the dependency — THEN checked
ConditionPathIsMountPoint. The condition saw the just-remounted
filesystem and evaluated to true. Service started on empty
external storage. Guard completely bypassed.
This is a subtle interaction:
- ConditionPathIsMountPoint is a TEST (true/false evaluation)
- RequiresMountsFor is a DEPENDENCY (systemd must make it true)
A guard should be a TEST, not a dependency that makes the test
trivially pass. Remove the dependency. The condition alone is
sufficient for both boot-time and runtime checks:
- Boot: mount unit starts via local-fs.target, condition is true
- Runtime: if mount disappears, condition becomes false on next
start attempt. Without RequiresMountsFor, systemd doesn't
auto-recover, so the guard fires.
Ordering should be expressed via After= in the consumer's
systemd.services block (not in the shared helper), e.g.:
systemd.services.postgresql.after = [ "home-ooyude-External.mount" ];
Live test v8 sequence (before this fix):
1. umount /home/ooyude/External → OK, gone from /proc/mounts
2. findmnt /home/ooyude/External → exit=1, not in table
3. systemctl start postgresql → STARTED (guard bypassed)
4. journal: no ConditionPath error (service started successfully)
This is the second guard bug found by live testing in this session.
The first one was the '!' prefix inversion. Both were invisible to
nix eval, both only visible at runtime. Lesson reinforced:
guards MUST be live-tested, not just statically evaluated.
Recovery: v8 test reverted state via emergency_recovery trap
(remount + restart all services). System healthy at 10/10.
Review of dev vs 16644fc found 1 BLOCKING + 3 IMPORTANT + 2 MINOR.
All addressed in this commit:
B1 (R1.4 stale files in project-rules.md + AGENTS.md):
Replaced 'vds/nginx.nix' (removed in ef38dc4) with 'home/termux.nix'
(added in 958247b). R1.4 now correctly lists the 4 files that use
100.64.0.0: home/termux.nix:256, modules/server/nextcloud.nix:73,
modules/server/nginx.nix:109,253, modules/vds/systemd.nix:10.
I1 (count drift in '15 modules' docs):
- AGENTS.md:84 + project-rules.md:97: '15 → 14' (with note that
stirling-pdf was deleted in 5dd7a58)
- manifest.json (T16): rewritten acceptance to '15 archived
(13 from server/default.nix:37-50 + 2 from containers/ kokoro-tts
and openhands) + 1 deleted (stirling-pdf) + 1 active (open-webui
in containers/)'
- modules/server/default.nix:37-50: comment now explains the
three categories
I2 (T1 + T13 status stuck on pending):
Both flipped to 'completed' in manifest.json. T1 import fix
verified by nix eval (epral stateVersion = '24.05'). T13 done in
61b3724 (nginx firewall rule removed). I3 (.ci/checks.sh committed)
satisfied.
M1 (R1.3 stale nginx.nix:225 line number):
Removed line number from both project-rules.md and AGENTS.md.
Replaced with 'nginx.nix (networking.firewall)'.
M2 (R1.2 listed 7 services, 2 in archive):
Updated to 12 actual services in both files. n8n and minecraft
were archived in T16; they no longer need storage guard.
T10 (reality443Forwarding погашен):
Removed option from options.nix:66-74, realityPorts from
3x-ui.nix:33-35, and 'reality443Forwarding = true' from
vds/default.nix:19. ADR-note comments left in place.
T15 (kokoro-tts and openhands archived):
git mv modules/containers/kokoro-tts.nix → archive/containers/
git mv modules/containers/openhands.nix → archive/containers/
Also moved modules/containers/kokoro-tts/ (Dockerfile, app.py, etc.)
to archive/containers/kokoro-tts/ for completeness.
any.nix (nix flake check support):
Added stub fileSystems + boot.loader.grub to configurations/any.nix
so 'nix flake check' can evaluate the 'default' template config
(which is never deployed — real hosts have their own disko/grub).
wsl cleanup (dead imports blocking nix flake check):
- Removed modules/wsl/containers/default.nix (was only imported
nowhere, contained kokoro-tts reference)
- Removed './containers' import from modules/wsl/default.nix
(resolved to the now-removed default.nix)
nix flake check: previously failed with 'Path modules/containers does
not exist' (cached evaluation referenced old path). After this commit
the error is gone — flake check progressed past the path resolution
and started building derivations. Full build output not captured
(5-min timeout for download from cache.nixos.org), but path errors
are resolved.
T5 risk acknowledgment:
.agent/decisions/0002-backups-external.md updated with explicit
risk table for 'if no backups' scenario + ADR/R1.9 guidance.
T1, T2, T6, T7, T8, T9, T10, T12, T13, T15, T16, T17: all → completed
in manifest.json. T3, T4, T5, T11, T14: previously completed.
Remaining DEFERRED: T3 (otrecа SSH recovery), T5 (5.6 answer).
tape-rotation is an SQLite-backed tape tracking app, similar to
3x-ui in that:
- data lives on /home/ooyude/External (storage-guarded as of T4)
- images are small and updated manually via Nix
- :latest is acceptable because breaking image changes would
fail fast at container start (storage guard + systemd)
Three image :latest entries are now whitelisted (with rationale):
- ghcr.io/mhsanaei/3x-ui:latest (R1.5 — panel frozen, Xray is panel state)
- docker.io/elizaroveugene/taperotation-backend:latest (this commit)
- docker.io/elizaroveugene/taperotation-frontend:latest (this commit)
Two :latest violations remain (both in modules/containers/ but NOT
imported in modules/server/default.nix — effectively dead code):
- localhost/kokoro-tts:latest
- ghcr.io/openhands/openhands:latest
Decision pending: archive like T16, whitelist, or remove. They are
not running on sapphira, so the check failure is informational only.
CRITICAL BUG FIX caught by live test v6 on sapphira.
The '!' prefix INVERTS the systemd test:
ConditionPathIsMountPoint=!/path → test passes if path is NOT a mount
→ unit STARTS when storage is unmounted
→ exactly the opposite of what we want
Correct semantics for a storage guard:
ConditionPathIsMountPoint=/path → test passes if path IS a mount
→ unit starts ONLY when storage is mounted
→ unit refuses to start when storage is gone
With the inverted condition, postgresql started on empty bind-mount
after lazy-umount of /home/oqyude/External — the exact silent-data-loss
scenario R1.2 is supposed to prevent.
This is the canonical 'bug the test caught' case. Live test v6 on
sapphira (2026-10-10) demonstrated: with '!' the guard does nothing,
without '!' the guard fires correctly.
Lesson: always run a live test of the guard, don't trust nix eval alone
for systemd Condition* semantics — they're evaluated by systemd at
runtime, and '!' inverts the test.
Comprehensive batch addressing the 16-task backlog in
.agent/tasks/manifest.json. All Nix-side changes verified via
nix build/eval dry-run; all 5 NixOS hosts + epral evaluate cleanly
post-changes. No regressions.
Wave 1 (non-functional cleanup):
T1/A1 — configurations/mobile.nix:12: fix `import ../lib/xlib.nix`
(broken path) → `import ../lib/xlib`. Unblocks nixOnDroid
configurations.epral. R1.1 invariant.
T8/C3 — modules/containers/3x-ui.nix: remove `podman-update-3xui_app`
systemd service and commented timer. Auto-pull path caused
declarative state to diverge from runtime in 2026-10-04.
R1.5 invariant.
T13/D3 — modules/server/nginx.nix:368-371: remove dead
`networking.firewall.allowedTCPPorts = [80 443]`.
`firewall.enable = false` on sapphira (R1.3), so openFirewall
rules are no-op. Replace with R1.3 comment.
T6/C1 — .agent/decisions/notes/3x-ui-xray-26.9.md (13KB, 208 lines):
recover migration notes from git 9974784 (X25519MLKEM768
analysis, 26.7→26.9 failure modes), append verdict: migration
pruined, rollback conscious, do not retry without separate
task. R1.5 / C1.
T9/C4 — .agent/rules/project-rules.md: add R1.8 — Xray-core version is
state of 3x-ui panel, not Nix. Update trap entry for
3x-ui.nix:54 to reference R1.8.
T11/D1, T12/D2 — .agent/checkpoints.json + .agent/tasks/manifest.json:
verify R1.3 (router port-forwards 22/80/443/8443/22000) and
R1.4 (100.64.0.0 = Tailscale sapphira) wording already
satisfies acceptance criteria. Flip status pending → completed.
T4 (storage guard, FUNCTIONAL CHANGE):
New helper in lib/xlib/helpers.nix:
mkStorageGuard = xlib: {
RequiresMountsFor = [ xlib.dirs.server-home ];
ConditionPathIsMountPoint = [ "!${xlib.dirs.server-home}" ];
};
Applied to 13 systemd units via path-style override:
- modules/server/{postgresql,samba,homebox,gitea,navidrome,
syncthing,uptime-kuma,immich,nextcloud,calibre-web}.nix
- modules/containers/3x-ui.nix (podman-3xui_app)
- modules/containers/tape-rotation.nix (podman-taperotation-{backend,frontend})
Anchor: xlib.dirs.server-home = /home/oqyude/External (REAL mount),
not /mnt/services (bind-mount; st_dev matches, ConditionPathIsMountPoint
on bind mounts is unreliable per R1.2 note).
Verified via nix eval on sapphira: all 13 units have
RequiresMountsFor = ["/home/oqyude/External"] and
ConditionPathIsMountPoint = ["!/home/oqyude/External"].
Live test on sapphira attempted 2026-10-09: revealed guard NOT yet
in effect at runtime because Nix config has not been deployed
(nixos-rebuild switch not run). postgresql started despite External
being unmounted. Implementation correct, deployment pending user
action.
T7/C2 (read-only diag, no code change):
3x-ui version facts recorded in conversation (sapphira journal +
/var/lib/containers/storage/overlay/.../diff/app/bin/xray-linux-amd64):
- Active Xray: 26.7.28 (go1.26.5 linux/amd64) — R1.5 validated at runtime
- Stale binary: 26.9.30 (go1.27.1) — leftover from failed 26.9 migration
- Panel DB (x-ui.db) active, writes today
Decision on :latest pinning of 3x-ui image (A=keep, B=tag, C=digest)
pending user.
T3/A3 (nftables on otreca — config analysis + proposal):
Diagnostic attempted via ssh otreca-tailscale (100.64.1.0) and
otreca public (109.248.161.5:22): BOTH UNREACHABLE. Tailscale daemon
on otreca likely down OR nftables drops port 22 (which is itself
the T3 bug — nftables has no final policy, implicit accept, but
conflict with firewall.enable = true per R1.6).
Proposal written: .agent/decisions/proposals/vds-nftables-fix.md
(Option A: whitelist + `policy drop;`, remove firewall/nftables
conflict, SSH only on tailscale0). Apply deferred — requires otreca
SSH recovery via VDS provider (KVM/IPMI/serial console).
T5/B2 (backups documentation):
.agent/decisions/0002-backups-external.md (draft): catalog of what
is declared in Nix vs. what is external; awaiting answer to open
question 5.6 (where are backups, how are they verified).
T15/E2 (CI checks):
.ci/checks.sh (executable, ~140 lines) with 3 checks from
analysis-report.md §5:
- #1: no `:latest` in container images (with R1.5 whitelist
for 3x-ui). FAIL — 4 violations:
localhost/kokoro-tts:latest
ghcr.io/openhands/openhands:latest
docker.io/elizaroveugene/taperotation-backend:latest
docker.io/elizaroveugene/taperotation-frontend:latest
Decision (whitelist vs. pin) pending user.
- #2: nix flake check (skipped with --no-build).
- #7: secrets/ files match .sops.yaml path_regex. PASS.
T16/E3 (archive commented modules):
13 of 14 commented modules in modules/server/default.nix:37-50
existed as files. git mv them to archive/{server-modules,containers}/.
1 (stirling-pdf.nix) didn't exist; just removed the comment.
modules/server/default.nix:37-50 cleaned of 14 commented lines.
Added 3-line comment recording the archive date and reason.
Verified: nixosConfigurations.sapphira still evaluates.
Post-change state:
$ nix build .#nixosConfigurations.{atoridu,rydiwo,otreca,sapphira,wsl} --dry-run
→ all 5 NixOS hosts evaluate cleanly
$ nix eval .#nixOnDroidConfigurations.epral.config.system.stateVersion
→ "24.05"
Pending (user input required — not in this commit):
- T4 deploy: run `nixos-rebuild switch` on sapphira to activate guard
- T7: pick A/B/C for 3x-ui :latest pinning
- T3: recover otreca SSH via VDS provider, then apply Option A
- T10/C5: decide fate of reality443Forwarding
- T5: answer 5.6 about backup location/verification
- T15: whitelist or pin 4 :latest images
Untracked files NOT committed (in .gitignore):
.temp/t4-live-test*.sh, .temp/cleanup-*.sh — throwaway test scripts
from T4 live test attempts. Preserved locally for reference; see
AGENTS.md convention ("Создавать `.temp/` в корне проекта — Для
временных файлов агента. Всегда в `.gitignore`").
Also untracked, committed:
.agent/reviews/2026-10-10-review-dev-diff-vs-16644fc.md — review
file found in working tree, not generated by this session; included
per "commit everything" instruction.
Sister project at S:\Git\VeeamTimelineView has its own repo and gitea
remote; we keep its Nix env here in the home-fleet flake so the project
itself stays free of Nix.
`nix develop .#vtimeline` provides:
- nodejs_22 (Node 20 hit EOL 2026-04-30 and was removed from nixpkgs)
- python3 (Playwright webServer)
- chromium (used via PLAYWRIGHT_BROWSERS_PATH symlink farm, since
Playwright 1.40+ hard-codes a separate chrome-headless-shell path)
Shell banner points at ../VeeamTimelineView. The absolute project path
cannot be baked into the shellHook (Nix forbids escaping the flake source
dir at eval time); user is expected to `cd` themselves.
Three small things together:
1. relinkHomeManager was pinning the versioned symlink name to
'home-manager-24-link'. That breaks on every HM major-version bump:
HM 25 will move the alias to 'home-manager-25-link' and the script
will silently stop relinking. Derive the name from one hop of the
stable 'home-manager' symlink, with a guard against the missing case
so a fallback never accidentally rewrites the profiles/ directory
itself.
2. programs.opencode.web.environmentFile now reads from
xlib.dirs.opencode-server-env (added previously in users.nix + dirs.nix).
The sops materialization and the systemd EnvironmentFile can no
longer silently desync.
3. Document the oh-my-openagent 2026-07-opencode-config-unification
migration trap that logs 'Migration backup path already exists' on
every startup: the backup path embeds the content-hashed store path,
which stays valid in /nix/store across HM activations, so the
deterministic collision never resolves itself. Recovery is
'rm -rf ~/.omo/migration-backup-*' to let omo retry; if it keeps
failing on the same path the plugin version probably expects a new
schema and this file needs changes.
Also trim the over-explained [Service] / serviceConfig comment: the
home-manager attrset-union behavior is general, not specific to this
unit, so the explanation got shorter without losing the invariant.
The path '/home/<user>/.config/opencode/server.env' was duplicated
between users.nix (sops materialization) and home/modules/opencode.nix
(programs.opencode.web.environmentFile). Drift between the two was a
silent auth-bypass vector: if one moved, the systemd unit would either
fail to find OPENCODE_SERVER_PASSWORD or skip EnvironmentFile entirely.
Single source in lib/xlib/dirs.nix; both call sites now read from it.
Sapphira: HTTP reverse proxy serves panel/sub on x.zeroq.su;
no xray stream on 443 and no 8443 stream either (8443 is directly
exposed by podman as 0.0.0.0:8443:8443/tcp).
Otreca: stream on 443 routes by SNI (panel via pubray1.zeroq.su,
xray default) and 8443 is direct 0.0.0.0:8443.
Modules/containers/3x-ui.nix:
- basePorts restored: '0.0.0.0:8443:8443/tcp' (was '127.0.0.1:15380:8443/tcp')
- realityPorts restored (was 'lib.optional ... "127.0.0.1:15380:443/tcp"')
- image restored: ':latest' (was ':v3.9.0')
Modules/server/nginx.nix:
- removed 8443 streamConfig for xray (the one b0191bc added)
- removed 8443 from allowedTCPPorts
Other files (configurations/{server,vds,wsl}.nix, home/modules/opencode.nix)
left alone — they contain SSH firewall / builder / opencode web changes
unrelated to nginx + ports that the user asked to revert.
The systemd unit on the otreca VDS carried two -p flags that bind
the same host port 127.0.0.1:15380:
-p 127.0.0.1:15380:8443/tcp # from basePorts
-p 127.0.0.1:15380:443/tcp # from realityPorts (when reality443Forwarding=true)
podman 5.x tries to bind 127.05 in each -p flag and the second
fails with EADDRINUSE, even though no process is visible in ss —
the bind happens at the proxy level before the container starts:
Error: cannot listen on the TCP port: listen tcp4 127.0.0.1:15380:
bind: address already in use
Symptom on otreca: podman-3xui_app.service hits start-limit-hit
after 5 rapid retries.
The 15380:443 mapping is dead code: the container's only Reality
inbound listens on 8443, and nginx stream already routes host:443
to 127.0.0.1:15380 via SNI (modules/server/nginx.nix streamConfig).
reality443Forwarding remains a host option for configurations to
declare intent; the broken port-mapping generation is replaced with
an empty list.
Revert the b0191bc 'otreca vds: pin 3x-ui:v3.8.5 + nginx stream + ssh
tailscale-only + patch-3xui-xray-config' changes:
- 3x-ui.nix: back to :latest image, direct 0.0.0.0:8443 port mapping,
remove migrateScript + patchScript and their systemd units/timer.
- vds.nix: re-open 22/tcp on public (openFirewall = true); remove the
tailscale0-only port rule.
- nginx.nix: drop the 8443 stream proxy.
- Remove modules/containers/3x-ui-migration-notes.md.
Reason: those changes, once applied on otreca, left the 3x-ui container
in a start-limit-hit loop (bind 127.0.0.1:15380: address already in use,
nothing visible in ss - probably a stale TIME_WAIT or slirp4netns port
from a prior container that never released).