Files
Homelabv4/infrastructure/authentik/values.yaml
T
Scooby HuskyandClaude Sonnet 5 f7a9a03816 authentik HA: floating DNS hostname + failover-watcher
Builds the one genuinely new component from the plan - everything else
reuses CNPG/etcd as-is.

- infrastructure/ha-failover/manifests/heartbeat-writer.yaml: home writes
  a fresh timestamp to etcd (/ha-failover/home-heartbeat) every 10s via
  its local etcd member. No custom quorum/voting logic - etcd's own Raft
  consensus (majority write/linearizable read) already gives the 'a
  majority agrees' guarantee.

- scripts/ha-failover-watcher.sh + .service: runs on the VPS as a
  systemd service (not k3s - must survive the VPS's own cluster being
  unhealthy). Reads the heartbeat via its own local etcd member; if the
  read itself fails, we can't tell if home is down or if this VPS is the
  one partitioned, so it does NOT act (etcd's consensus requirement
  provides the split-brain safety here, not custom code). If the read
  succeeds and the heartbeat is stale (>45s, ~4-5 missed beats), home is
  confirmed down by majority - promotes pg-authentik's Cluster CR
  (spec.replica self/primary/source -> vps, no promotionToken - confirmed
  live via dry-run that CNPG's webhook doesn't require one, it's only for
  planned/graceful switchover LSN cross-checks that aren't available
  during a genuine unplanned outage) and flips the
  pg-authentik.ha.huskypup.net Cloudflare A record to the VPS's IP.

- scripts/ha-failback-authentik.sh: the deliberately manual reverse -
  human confirms home is healthy and caught up before running this.

- Floating hostname plumbing: both sites' Authentik values now read
  AUTHENTIK_POSTGRESQL__HOST/PORT as pg-authentik.ha.huskypup.net:61432
  instead of the local pg-authentik-app secret's host/port, so app pods
  on either site always reach whichever site is actually primary. Both
  sites listen on the same external port (VPS gets a new
  pg-authentik-forward.service socat forward, mirroring the existing
  minio-forward.service pattern, so its NodePort 32433 is externally
  reachable on 61432 same as home's UniFi-forwarded port).

- Found and worked around two real bugs surfaced while wiring this up
  (fixed manually via scripts handed to the user - both blocked by the
  Claude Code auto-mode classifier as credential-transmission /
  cluster-DNS-edit actions):
  - VPS's pg-authentik-app secret had a STALE password from before
    streaming replication existed - the live Postgres role password now
    replicates from home via WAL, but VPS's local K8s secret never got
    updated to match. Needs a one-time sync (and again on any future
    rotation).
  - This UniFi does not support NAT hairpin/loopback for its own WAN
    port-forwards - home's own pods resolving the floating hostname need
    a local CoreDNS rewrite straight to pg-authentik-rw, confirmed live by
    a DNS-resolves-but-TCP-connect-fails test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-20 17:43:08 -05:00

198 lines
5.8 KiB
YAML

# Authentik configuration
authentik:
secret_key: env://AUTHENTIK_SECRET_KEY
postgresql:
host: env://AUTHENTIK_POSTGRESQL__HOST
port: 5432
name: env://AUTHENTIK_POSTGRESQL__NAME
user: env://AUTHENTIK_POSTGRESQL__USER
password: env://AUTHENTIK_POSTGRESQL__PASSWORD
global:
envFrom:
# First secret: Authentik-specific settings (AUTHENTIK_SECRET_KEY)
- secretRef:
name: authentik
# Second secret: Database credentials from CNPG
- secretRef:
name: pg-authentik-app
prefix: AUTHENTIK_POSTGRESQL__
env:
# Configure external URL for proper OIDC discovery responses
- name: AUTHENTIK_URL
value: "https://auth.kube.huskypup.net"
# Multi-site active failover pilot (see
# /home/scooby/.claude/plans/jiggly-snacking-iverson.md): HOST/PORT
# point at the floating pg-authentik.ha.huskypup.net Cloudflare A
# record instead of the local pg-authentik-app secret's host/port -
# the failover-watcher flips that record between home's and the VPS's
# public IP, both listening on the SAME external port 61432 (home via
# UniFi WAN forward, VPS via pg-authentik-forward.service's local
# socat forward - see infrastructure/ha-failover/manifests/). This way
# app pods on EITHER site always reach whichever site is currently
# primary, without needing per-site Helm value differences that would
# go stale on failover.
#
# NAME/USER/PASSWORD still come from the local pg-authentik-app secret
# (CNPG-generated, doesn't change on failover) - but the two sites'
# copies of this secret must hold the SAME password, since either site
# may end up dialing the other. CNPG only sets it once at cluster
# creation from an independently-generated value per cluster, so this
# needed a one-time manual sync (VPS's copy was stale from before
# streaming replication existed - confirmed live 2026-08-20, VPS's
# local secret still had its original bootstrap-time password even
# though the live Postgres role itself now replicates from home via
# WAL). If the app user's password is ever rotated, it must be synced
# to both sites' secrets the same way.
- name: AUTHENTIK_POSTGRESQL__HOST
value: "pg-authentik.ha.huskypup.net"
- name: AUTHENTIK_POSTGRESQL__NAME
valueFrom:
secretKeyRef:
name: pg-authentik-app
key: dbname
- name: AUTHENTIK_POSTGRESQL__USER
valueFrom:
secretKeyRef:
name: pg-authentik-app
key: username
- name: AUTHENTIK_POSTGRESQL__PASSWORD
valueFrom:
secretKeyRef:
name: pg-authentik-app
key: password
- name: AUTHENTIK_POSTGRESQL__PORT
value: "61432"
# Blueprints - Mount ConfigMaps for auto-discovery
blueprints:
configMaps:
- authentik-blueprints-gitlab
- authentik-blueprints-nextcloud
- authentik-blueprints-argocd
- authentik-blueprints-grafana
- authentik-blueprints-forward-auth
- authentik-blueprints-prometheus
- authentik-blueprints-n8n
- authentik-blueprints-ceph
- authentik-blueprints-home-assistant
- authentik-blueprints-guacamole
- authentik-blueprints-percona-everest
- authentik-blueprints-rancher
- authentik-blueprints-netbird
- authentik-blueprints-vault
# Enable Prometheus metrics
server:
# Single replica required: the embedded outpost stores forward-auth sessions
# in-memory (no Redis). Multiple replicas cause callback failures because
# the session created during ext_authz check is only on one pod.
replicas: 1
containerSecurityContext:
runAsNonRoot: true
allowPrivilegeEscalation: false
capabilities:
drop:
- ALL
seccompProfile:
type: RuntimeDefault
metrics:
enabled: true
serviceMonitor:
enabled: true
labels:
release: prometheus
# Resource limits to prevent OOM
resources:
requests:
cpu: 50m
memory: 512Mi
limits:
memory: 1Gi
# Health probes - removed invalid 'enabled' fields
startupProbe:
initialDelaySeconds: 10
periodSeconds: 5
timeoutSeconds: 3
failureThreshold: 24 # 120 seconds for Authentik startup
readinessProbe:
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 3
failureThreshold: 2
livenessProbe:
initialDelaySeconds: 10
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3
# Spread across nodes
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
app.kubernetes.io/name: authentik
app.kubernetes.io/component: server
topologyKey: kubernetes.io/hostname
worker:
containerSecurityContext:
runAsNonRoot: true
allowPrivilegeEscalation: false
capabilities:
drop:
- ALL
seccompProfile:
type: RuntimeDefault
metrics:
enabled: true
serviceMonitor:
enabled: true
labels:
release: prometheus
# Resource limits to prevent OOM
resources:
requests:
cpu: 50m
memory: 512Mi
limits:
memory: 1Gi
# Spread across nodes
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
app.kubernetes.io/name: authentik
app.kubernetes.io/component: worker
topologyKey: kubernetes.io/hostname
# Mount cert-manager secret for auto-import of Let's Encrypt cert
volumeMounts:
- name: certs
mountPath: /certs
readOnly: true
volumes:
- name: certs
secret:
secretName: authentik-tls
optional: true
# Disable bundled Postgres
postgresql:
enabled: false