Fix real cause of VPS authentik's DB timeout loop + repo secrets cleanup

Root cause of the 'PostgreSQL connection failed, retrying... (connection
timeout expired)' loop that survived every earlier fix (streaming
replication, the password sync, the port fix, the CoreDNS override):
authentik's Helm values had TWO sources for AUTHENTIK_POSTGRESQL__HOST/
PORT - an explicit uppercase env: override (pointed at the floating
pg-authentik.ha.huskypup.net:61432) AND a chart-generated envFrom
secretRef (pg-authentik-app, prefix: AUTHENTIK_POSTGRESQL__) whose keys
are lowercase (host, port, ...), producing a SEPARATE
AUTHENTIK_POSTGRESQL__host/port pair pointing at the old local
pg-authentik-rw:5432. Kubernetes treats these as two unrelated env vars
(case-sensitive), but authentik's own generic AUTHENTIK_*-prefixed
env-var scanner apparently doesn't, and was resolving to the lowercase
(stale, local) values regardless of the explicit override - confirmed
live by dumping the pod's actual env: both HOST and host were present
with different values. Removed the now-fully-redundant envFrom entry
(every key it provided is already explicitly set via valueFrom).

Also: user correctly pointed out plaintext secrets don't belong in the
repo. Audited for the same class of issue as the already-fixed Cloudflare
token (infrastructure/cert-manager/manifests/secret-cf-token.yaml):
- infrastructure/external-dns/manifests/secret-external-dns-unifi.yaml
  had a live UniFi API key in plaintext - moved to Vault+ExternalSecret,
  same pattern as the Cloudflare token fix. The key itself is still the
  original (now Vault-stored) value - rotating it requires the UniFi web
  UI (no self-service API), noted in the file as a separate pending step.
- infrastructure/vps-eso/manifests/clustersecretstore.yaml had the
  AppRole's roleId inline (added this session) - moved to roleRef,
  sourced from the same Secret as secretId, consistent with 'nothing
  sensitive in git' regardless of how sensitive one field is alone.

NOT touched, flagged separately for the user: infrastructure/authentik/
*-blueprint.yaml (vault, gitlab, n8n, nextcloud, rancher, grafana, argocd,
guacamole) all have live-looking high-entropy client_secret values
hardcoded in plaintext - same class of issue but much larger blast radius
(8 apps' SSO), needs its own coordinated rotation, not bundled into this
commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Scooby Husky
2026-08-20 18:47:50 -05:00
co-authored by Claude Sonnet 5
parent cfbe844ecd
commit 185e9c292e
4 changed files with 60 additions and 15 deletions
@@ -57,13 +57,14 @@
# vault write -f auth/approle/role/vps-eso/secret-id # -> secret_id
#
# vault-approle-creds is a plain Secret created manually on the VPS
# (kubectl, not git - same reasoning as every other VPS secret):
# (kubectl, not git - same reasoning as every other VPS secret, and
# consistent regardless of "how sensitive" a given field is on its own -
# role_id and secret_id both live in the same Secret, neither in git):
# kubectl -n external-secrets create secret generic vault-approle-creds \
# --from-literal=role_id=<from role-id above> \
# --from-literal=secret_id=<from secret-id above>
# Rotate the secret_id periodically by writing a new one and patching this
# Secret - role_id is stable and not sensitive on its own (useless without
# a valid secret_id).
# Secret.
apiVersion: external-secrets.io/v1
kind: ClusterSecretStore
metadata:
@@ -77,7 +78,10 @@ spec:
auth:
appRole:
path: approle
roleId: ecaf2eda-d922-f7cf-1143-690bbbb4d8ea # not sensitive alone, see note above
roleRef:
name: vault-approle-creds
namespace: external-secrets
key: role_id
secretRef:
name: vault-approle-creds
namespace: external-secrets