Multi-site active failover pilot: home etcd member + Postgres NodePort exposure

Part of the Authentik HA pilot (see plan doc). Home's etcd quorum member
(StatefulSet, pinned to talos-cp-01 for a stable advertised address) plus
a NodePort exposing pg-authentik's current primary - both reachable from
the VPS/witness over the already-authenticated Netbird mesh (confirmed
live: home nodes are directly reachable from Netbird peers on their real
node IP, via the netbird-egress DaemonSet's route). Deliberately NOT
going through UniFi/public-internet exposure - this stays entirely
inside the private mesh, a materially safer path than the WAN port-
forward originally considered.

Needs a scoped Kyverno PolicyException (ha-failover-nodeport-exception.yaml)
since disallow-nodeport-services is enforced cluster-wide - narrowly
scoped to Services named ha-*, matching the existing netbird-egress-
exemption.yaml precedent for exceptions.
This commit is contained in:
Scooby Husky
2026-08-18 21:12:45 -05:00
parent 9b45cfd542
commit 5f52f2b36b
4 changed files with 183 additions and 0 deletions
@@ -0,0 +1,106 @@
---
# Home's member of the 3-way etcd quorum used by the failover-watcher
# (see /home/scooby/.claude/plans/jiggly-snacking-iverson.md, "Multi-Site
# Active Failover" section) to safely decide when to trigger a CNPG
# distributed-topology promotion. VPS and witness run this same etcd
# version as plain systemd services (simpler - no k3s/Kyverno involved
# there); home has to run it as a real K8s workload since Talos doesn't
# allow bare-metal process installs.
#
# Pinned to a specific node (talos-cp-01) via nodeSelector so its
# advertised peer/client URLs (which must be stable, not "whichever node
# it landed on today") stay correct. Reachable from the VPS/witness via
# that node's real IP - confirmed live 2026-08-19 that home nodes are
# directly reachable from Netbird mesh peers on their node IP (the
# netbird-egress DaemonSet's route) - and exposed via a NodePort Service,
# which needs infrastructure/kyverno/policies/ha-failover-nodeport-exception.yaml
# (disallow-nodeport-services is enforced cluster-wide otherwise).
apiVersion: v1
kind: Namespace
metadata:
name: ha-failover
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: ha-etcd
namespace: ha-failover
spec:
serviceName: ha-etcd
replicas: 1
selector:
matchLabels:
app: ha-etcd
template:
metadata:
labels:
app: ha-etcd
spec:
nodeSelector:
kubernetes.io/hostname: talos-cp-01
securityContext:
runAsNonRoot: true
runAsUser: 1000
fsGroup: 1000
seccompProfile:
type: RuntimeDefault
containers:
- name: etcd
image: gcr.io/etcd-development/etcd:v3.5.17
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
command:
- etcd
- --name=home
- --data-dir=/var/lib/etcd/data
- --listen-client-urls=http://0.0.0.0:2379
- --advertise-client-urls=http://100.108.42.109:32379
- --listen-peer-urls=http://0.0.0.0:2380
- --initial-advertise-peer-urls=http://100.108.42.109:32380
- --initial-cluster=home=http://100.108.42.109:32380,vps=http://100.108.113.41:2380,witness=http://100.108.130.74:2380
- --initial-cluster-state=new
- --initial-cluster-token=ha-failover-quorum
ports:
- containerPort: 2379
name: client
- containerPort: 2380
name: peer
volumeMounts:
- name: data
mountPath: /var/lib/etcd
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
memory: 256Mi
volumeClaimTemplates:
- metadata:
name: data
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: rook-ceph-block
resources:
requests:
storage: 1Gi
---
apiVersion: v1
kind: Service
metadata:
name: ha-etcd
namespace: ha-failover
spec:
type: NodePort
selector:
app: ha-etcd
ports:
- name: client
port: 2379
targetPort: 2379
nodePort: 32379
- name: peer
port: 2380
targetPort: 2380
nodePort: 32380