Root cause of the 'PostgreSQL connection failed, retrying... (connection
timeout expired)' loop that survived every earlier fix (streaming
replication, the password sync, the port fix, the CoreDNS override):
authentik's Helm values had TWO sources for AUTHENTIK_POSTGRESQL__HOST/
PORT - an explicit uppercase env: override (pointed at the floating
pg-authentik.ha.huskypup.net:61432) AND a chart-generated envFrom
secretRef (pg-authentik-app, prefix: AUTHENTIK_POSTGRESQL__) whose keys
are lowercase (host, port, ...), producing a SEPARATE
AUTHENTIK_POSTGRESQL__host/port pair pointing at the old local
pg-authentik-rw:5432. Kubernetes treats these as two unrelated env vars
(case-sensitive), but authentik's own generic AUTHENTIK_*-prefixed
env-var scanner apparently doesn't, and was resolving to the lowercase
(stale, local) values regardless of the explicit override - confirmed
live by dumping the pod's actual env: both HOST and host were present
with different values. Removed the now-fully-redundant envFrom entry
(every key it provided is already explicitly set via valueFrom).
Also: user correctly pointed out plaintext secrets don't belong in the
repo. Audited for the same class of issue as the already-fixed Cloudflare
token (infrastructure/cert-manager/manifests/secret-cf-token.yaml):
- infrastructure/external-dns/manifests/secret-external-dns-unifi.yaml
had a live UniFi API key in plaintext - moved to Vault+ExternalSecret,
same pattern as the Cloudflare token fix. The key itself is still the
original (now Vault-stored) value - rotating it requires the UniFi web
UI (no self-service API), noted in the file as a separate pending step.
- infrastructure/vps-eso/manifests/clustersecretstore.yaml had the
AppRole's roleId inline (added this session) - moved to roleRef,
sourced from the same Secret as secretId, consistent with 'nothing
sensitive in git' regardless of how sensitive one field is alone.
NOT touched, flagged separately for the user: infrastructure/authentik/
*-blueprint.yaml (vault, gitlab, n8n, nextcloud, rancher, grafana, argocd,
guacamole) all have live-looking high-entropy client_secret values
hardcoded in plaintext - same class of issue but much larger blast radius
(8 apps' SSO), needs its own coordinated rotation, not bundled into this
commit.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
User correctly flagged this: manually copying secrets between home and
the VPS defeats the entire point of having Vault. Every VPS secret this
session (MinIO creds, Cloudflare token, CNPG certs, the pg-authentik-app
password) was a one-off kubectl create secret copy-paste, because the
VPS's k3s cluster had no Vault/ESO pipeline at all - this builds one.
- infrastructure/authentik/manifests/ha-postgres-app-pushsecret.yaml:
home pushes pg-authentik-app's password into Vault at
secret/vps/pg-authentik-app (PushSecret, not just Get - this password is
CNPG-generated, not Vault-native, so it has to originate from a push).
- infrastructure/vps-eso/manifests/clustersecretstore.yaml: ESO on the
VPS (installed via helm, out-of-band like k3s/cert-manager - see the
file's own header) authenticates to home's Vault via AppRole (not
Kubernetes auth - the VPS is a separate cluster with no federation to
home's API server). Reachable via a new public
https://vault.kube.huskypup.net record - Vault was deliberately kept
off the public internet before this, explicitly confirmed with the user
before opening it. Traffic goes through the same Istio ingress gateway
already serving other public hosts, so unlike the ha-authentik-postgres
NodePort case, no PeerAuthentication/AuthorizationPolicy change was
needed - it arrives as a normal in-mesh call from the gateway's own
identity, not raw external TCP to a pod.
Found and fixed a real Vault gotcha while wiring this up: tried to
remove the AppRole's token_bound_cidrs restriction (added first, before
discovering Vault can't see the VPS's real source IP through the
gateway - it only ever sees the gateway's own pod IP) by omitting the
field from a follow-up - that does NOT clear it, the
AppRole role endpoint preserves omitted fields rather than resetting
them to default. Had to explicitly write token_bound_cidrs=. Spent a
while chasing a misleading 403 permission denied on auth/token/
lookup-self before finding this - vault token capabilities said read
was allowed (policy was fine), the actual rejection was IP-bound token
use from an unbound context.
- infrastructure/vps-standby/authentik/manifests/
pg-authentik-app-externalsecret.yaml: pulls it back down, Merge policy
(only overwrites the password key - host/dbname/username stay local,
CNPG still needs its own local -rw hostname for internal use).
- argocd-apps/vps-standby/vps-eso.yaml: new Application for the
ClusterSecretStore.
Verified end-to-end: password now matches between home and VPS's
pg-authentik-app secrets via this pipeline (not the earlier manual patch),
confirmed by comparing both live secret values after ESO's sync.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>