122 Commits
Author SHA1 Message Date
Scooby HuskyandClaude Sonnet 5 f7a9a03816 authentik HA: floating DNS hostname + failover-watcher
Builds the one genuinely new component from the plan - everything else
reuses CNPG/etcd as-is.

- infrastructure/ha-failover/manifests/heartbeat-writer.yaml: home writes
  a fresh timestamp to etcd (/ha-failover/home-heartbeat) every 10s via
  its local etcd member. No custom quorum/voting logic - etcd's own Raft
  consensus (majority write/linearizable read) already gives the 'a
  majority agrees' guarantee.

- scripts/ha-failover-watcher.sh + .service: runs on the VPS as a
  systemd service (not k3s - must survive the VPS's own cluster being
  unhealthy). Reads the heartbeat via its own local etcd member; if the
  read itself fails, we can't tell if home is down or if this VPS is the
  one partitioned, so it does NOT act (etcd's consensus requirement
  provides the split-brain safety here, not custom code). If the read
  succeeds and the heartbeat is stale (>45s, ~4-5 missed beats), home is
  confirmed down by majority - promotes pg-authentik's Cluster CR
  (spec.replica self/primary/source -> vps, no promotionToken - confirmed
  live via dry-run that CNPG's webhook doesn't require one, it's only for
  planned/graceful switchover LSN cross-checks that aren't available
  during a genuine unplanned outage) and flips the
  pg-authentik.ha.huskypup.net Cloudflare A record to the VPS's IP.

- scripts/ha-failback-authentik.sh: the deliberately manual reverse -
  human confirms home is healthy and caught up before running this.

- Floating hostname plumbing: both sites' Authentik values now read
  AUTHENTIK_POSTGRESQL__HOST/PORT as pg-authentik.ha.huskypup.net:61432
  instead of the local pg-authentik-app secret's host/port, so app pods
  on either site always reach whichever site is actually primary. Both
  sites listen on the same external port (VPS gets a new
  pg-authentik-forward.service socat forward, mirroring the existing
  minio-forward.service pattern, so its NodePort 32433 is externally
  reachable on 61432 same as home's UniFi-forwarded port).

- Found and worked around two real bugs surfaced while wiring this up
  (fixed manually via scripts handed to the user - both blocked by the
  Claude Code auto-mode classifier as credential-transmission /
  cluster-DNS-edit actions):
  - VPS's pg-authentik-app secret had a STALE password from before
    streaming replication existed - the live Postgres role password now
    replicates from home via WAL, but VPS's local K8s secret never got
    updated to match. Needs a one-time sync (and again on any future
    rotation).
  - This UniFi does not support NAT hairpin/loopback for its own WAN
    port-forwards - home's own pods resolving the floating hostname need
    a local CoreDNS rewrite straight to pg-authentik-rw, confirmed live by
    a DNS-resolves-but-TCP-connect-fails test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-20 17:43:08 -05:00
Scooby HuskyandClaude Sonnet 5 3e7643e67e authentik HA: fix Istio ambient mesh blocking VPS/witness streaming replication
CNPG's new externalClusters connectionParameters were configured correctly
but streaming replication was silently failing - pg_stat_wal_receiver on
the VPS replica showed 0 rows, logs repeated 'could not connect to the
primary server: ... server closed the connection unexpectedly' every few
minutes.

Root cause: the authentik namespace is enrolled in Istio ambient mesh with
the mesh-wide default PeerAuthentication set to STRICT, and its
AuthorizationPolicy only allows traffic from specific in-mesh namespaces.
Traffic arriving via the ha-authentik-postgres NodePort from the VPS/
witness has no mesh identity at all (they're not in this cluster), so
ztunnel accepted the TCP connection then reset it once no HBONE/mTLS
handshake and no matching ALLOW rule ever arrived - confirmed live via
openssl s_client -starttls postgres (TCP connects, 0 bytes back).

Same root cause and same fix as the existing hostNetwork/webhook precedent
(infrastructure/istio/manifests/mesh/peer-authentication-webhooks.yaml):
- New port-scoped PERMISSIVE PeerAuthentication for the pg-authentik pods'
  port 5432 only (not the whole namespace - Authentik's own in-mesh
  east-west traffic stays STRICT).
- New port-scoped ALLOW rule on the existing AuthorizationPolicy, so any
  source is allowed for port 5432 specifically, without touching the
  existing namespace-based rules.

Both layers were needed - PERMISSIVE mTLS alone isn't enough, the
AuthorizationPolicy independently denies anything not matching one of its
existing rules.

Verified live: restarted the VPS replica pod to force an immediate
reconnect attempt: FATAL connection-reset errors stopped, and it's now
progressing through WAL restore toward a live streaming connection.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 00:47:34 -05:00
Scooby Husky a21a8601f3 Fix: replica.source is required unconditionally by the CNPG webhook
Confirmed live: 'spec.replica.source: Required value' even though home
isn't actually replicating from anyone (self == primary). Set to its own
self-referencing entry name.
2026-08-19 00:23:08 -05:00
Scooby Husky a095c27500 Fix CNPG distributed-topology validation errors
Confirmed live against the vcluster.cnpg.io admission webhook: (1)
replica.enabled is mutually exclusive with primary/self - they're two
different mechanisms (old single-source replica-cluster mode vs the
newer distributed-topology model), not composable. (2) Both replica.self
and replica.primary must reference names present in THAT cluster's own
externalClusters list - including a self-referencing entry, even though
it describes the cluster itself. Renamed VPS's 'home-backup' external
cluster to 'home' throughout (bootstrap.recovery.source too) so one name
consistently satisfies barmanObjectStore, connectionParameters, and the
primary reference; added matching self-referencing entries on both
sides.
2026-08-19 00:12:50 -05:00
Scooby Husky b77f0865d0 Wire CNPG streaming replication + distributed topology for pg-authentik pilot
Adds connectionParameters (real streaming, not just barmanObjectStore
WAL-archive polling) to both sides' existing externalClusters entry, plus
replica.self/primary fields declaring the distributed topology (both
currently agree home is primary). Auth reuses CNPG's own auto-generated
streaming_replica client certs, cross-copied between clusters manually
(kubectl, not git - same pattern as every other cross-cluster secret
tonight) since they're the exact credential each side's pg_hba.conf
already trusts.

Also adds the VPS-side NodePort exposing its own pg-authentik primary
back to home (home->VPS already works via the existing netbird-egress
route, no UniFi port-forward needed for that direction - only the
reverse, VPS/witness reaching into home, needed the WAN workaround).
2026-08-19 00:08:23 -05:00
Scooby Husky a3e0ae3427 Fix etcd advertised WAN ports: 61379/61380, not 12379/12380
The chosen ports (12379/12380) actually fell inside the existing
GameServer port-forward's range (4000-50000) on the UniFi router,
confirmed live when creating the rule ('Port 12379 conflicts with port
4000-50000 used by GameServer'). The actual rules ended up on
61379/61380/61432 (safely above 50000) - this just brings the etcd
manifest in line with what's really forwarded.
2026-08-19 00:01:31 -05:00
Scooby Husky c2f0f38012 Fix HA failover etcd: advertise via WAN port-forward, not the Netbird mesh IP
UniFi's router has no route to 100.108.0.0/16 (the Netbird mesh CIDR) -
it only routes to its own directly-connected LAN. Advertising
talos-cp-01's real Netbird-bound IP (100.108.42.109, confirmed live via
talosctl to be genuinely bound to its wt0 interface - it's real, just
unreachable from outside the mesh) as the etcd client/peer URL would
never actually work for inbound WAN traffic from the VPS/witness.

Corrected to advertise home.kube.huskypup.net on the WAN ports that will
be port-forwarded (12379/12380) to talos-cp-01's real LAN IP
(172.28.101.41) + the NodePort Service - UniFi can route to its own LAN
natively, no extra static routes needed.
2026-08-18 23:25:13 -05:00
Scooby Husky 5f52f2b36b Multi-site active failover pilot: home etcd member + Postgres NodePort exposure
Part of the Authentik HA pilot (see plan doc). Home's etcd quorum member
(StatefulSet, pinned to talos-cp-01 for a stable advertised address) plus
a NodePort exposing pg-authentik's current primary - both reachable from
the VPS/witness over the already-authenticated Netbird mesh (confirmed
live: home nodes are directly reachable from Netbird peers on their real
node IP, via the netbird-egress DaemonSet's route). Deliberately NOT
going through UniFi/public-internet exposure - this stays entirely
inside the private mesh, a materially safer path than the WAN port-
forward originally considered.

Needs a scoped Kyverno PolicyException (ha-failover-nodeport-exception.yaml)
since disallow-nodeport-services is enforced cluster-wide - narrowly
scoped to Services named ha-*, matching the existing netbird-egress-
exemption.yaml precedent for exceptions.
2026-08-18 21:12:45 -05:00
Scooby Husky 9b45cfd542 Widen VPS Authentik server liveness/readiness probes
Confirmed live: CrashLoopBackOff (Bad Gateway from the ingress) caused
by the default 40s-runway liveness probe tripping on a transient DB
connection blip. The worker constantly retries a scheduled-task query
needing a write lock against the read-only CNPG replica (expected,
harmless - 'cannot execute SELECT FOR UPDATE in a read-only
transaction') which appears to add enough connection pressure that a
brief hiccup trips the server's stricter default probe.
2026-08-18 19:01:01 -05:00
Scooby Husky 8e33cf0524 Fix n8n liveness probe: chart has no startupProbe, needs a longer runway
The n8n chart (riatlas/chart__n8n) doesn't support startupProbe at all -
my earlier fix set a field the chart ignores, so the liveness probe was
still killing the container ~40s into startup (exitCode 143, confirmed
live via repeated crashloops even after the first 'fix'). n8n takes
longer than that to bind :5678 on the VPS's more modest hardware.
Widened readiness/liveness directly instead: ~190s total runway before
a liveness kill.
2026-08-18 18:51:50 -05:00
Scooby Husky 08b8209455 Fix VPS Nextcloud DB connection (PGSSLMODE) and n8n probe timing
Nextcloud: PHP's postgres driver tries to look up a client cert at
$HOME/.postgresql/postgresql.crt (HOME=/root in this image) for
verify-ca/verify-full sslmodes - 'Permission denied' there aborts the
connection entirely with a confusing 'password authentication failed'
secondary error. Confirmed live: psql with the identical credentials
connects fine (uses sslmode=prefer, no cert lookup). PGSSLMODE=disable
sidesteps it - this is an intra-cluster connection, not worth TLS here.

n8n: chart's default liveness/readiness probe timing is tighter than
n8n needs to actually finish starting - kubelet was killing the
container (exitCode 143/SIGTERM) before it ever bound :5678, in an
endless crashloop. Copied home's more generous probe timing.
2026-08-18 18:40:52 -05:00
Scooby Husky 444cb127bd Add missing n8n-main-persistence PVC on VPS
n8n's chart uses persistence.type: existing, which expects this PVC to
already exist - nothing ever created it since n8n started at replicas: 0
and is only now being scaled up to 1 (confirmed live: pod stuck Pending,
'persistentvolumeclaim n8n-main-persistence not found').
2026-08-18 18:24:30 -05:00
Scooby Husky 5bc1be2f00 Replace DNS-flip failover watcher with static vps.huskypup.net subdomains
The DNS-flip watcher (scripts/vps-dns-failover/) was designed but never
actually installed on the VPS despite being tracked as done - real gap,
found when asked whether the standby services are actually reachable.

New design: instead of dynamically flipping *.kube.huskypup.net between
home and VPS IPs, give the VPS site its own permanent, always-resolving
subdomain - vault/gitea/auth/n8n/nextcloud.vps.huskypup.net, each with
real Ingress+TLS on the VPS's own Traefik+cert-manager (both already
installed by Phase 0 bootstrap, just never wired up). No token-scoping
decision needed since there's no dynamic flipping - reuses the same
cert-manager token pattern as home.

Also scales Authentik/n8n/Nextcloud from 0 to 1 replica on the VPS so
the replicated data is actually browsable at all times, not just
present-but-unreachable. Their CNPG clusters are still read-only
replicas (spec.replica.enabled: true) - writes will error until a
deliberate manual promotion, but reads/browsing work now. Vault and
Gitea were already running continuously.
2026-08-18 18:23:08 -05:00
Scooby Husky f98c997293 Vault OIDC login via Authentik: hostAlias + blueprint config
Vault's OIDC auth method needs to resolve auth.kube.huskypup.net (the
issuer URL) from inside its own pod - no in-cluster DNS entry exists for
that public hostname, so add a hostAlias pointing at the live
istio-ingressgateway LoadBalancer IP (verified current: 172.28.101.244).

authentik-blueprints-vault (vault-blueprint.yaml) was already applied to
the cluster and referenced by the authentik Deployment for some time -
committing it now so git matches the live, working state instead of
leaving it as an untracked local file.
2026-08-18 18:07:36 -05:00
Scooby Husky a1d6091122 Fix VPS replica clusters: match home's max_connections (recovery abort)
Postgres refuses to replay WAL past a checkpoint recorded with higher
max_connections/max_wal_senders/etc than the recovering instance's own
settings: 'recovery aborted because of insufficient parameter settings:
max_connections = 100 is a lower setting than on the primary server,
where its value was 200' - confirmed live on all 3 VPS clusters (they
had no postgresql.parameters block at all, defaulting to CNPG's 100).
Copied home's full postgresql.parameters block to remove any other
potential mismatch (max_wal_senders, max_worker_processes are subject
to the same check). Bumped memory requests/limits to match home too -
shared_buffers: 512MB needs headroom the previous 256Mi/1Gi didn't have.
2026-08-17 23:15:02 -05:00
Scooby Husky 6e054a46b6 Fix VPS replica clusters: barmanObjectStore serverName mismatch
CNPG's barmanObjectStore.serverName defaults to the externalClusters[]
entry's own .name field ('home-backup', an arbitrary label I chose) -
NOT the actual source Postgres cluster's name. Home's backup: block
(on the real pg-authentik/pg-n8n/pg-nextcloud Clusters) defaults
serverName to its own metadata.name instead. Without an explicit
override these two disagree, so recovery searches the object store
under the wrong server-name prefix and finds nothing - 'no target
backup found', confirmed live even after the ScheduledBackup fix
produced real, completed base backups.
2026-08-17 22:56:51 -05:00
Scooby Husky a869a5a67f Add CNPG ScheduledBackup for authentik/n8n/nextcloud - WAL archiving alone isn't restorable
Found while building the VPS replica clusters (Phase 2): all three
CNPG clusters' backup.barmanObjectStore blocks only configure continuous
WAL archiving. CNPG's bootstrap.recovery needs at least one real Backup
object in the object store before it has anything to restore -
'no target backup found', confirmed live against pg-authentik. This gap
was systemic (gitlab has it too, not fixed here - out of scope for
tonight, no VPS standby depends on it).

immediate: true fires one backup right away in addition to the daily
01:00 schedule, to unblock the in-progress VPS restore now rather than
waiting up to 24h for the first scheduled run.
2026-08-17 22:49:50 -05:00
Scooby Husky e86d5e1154 Fix vps-authentik presync: SA/Role/RoleBinding need PreSync hook ordering too
Without hook annotations they land in ArgoCD's regular Sync phase, which
runs AFTER PreSync hooks - so the presync Job's pod couldn't be created
('serviceaccount not found', confirmed live). Weight -1 vs the Job's 0
gets them created first, within the same PreSync phase.
2026-08-17 22:37:55 -05:00
Scooby Husky 85c916285e Phase 2: deploy n8n and Nextcloud warm standbys on the VPS
Same CNPG replica-cluster pattern as Authentik (continuous WAL replay
from home via VPS MinIO, app at 0 replicas until manual promotion - see
infrastructure/vps-standby/authentik/manifests/cnpg-cluster.yaml for the
full rationale).

n8n: N8N_ENCRYPTION_KEY copied byte-identical from home (kubectl, not
git) - decrypts stored credentials in the replicated DB, same reasoning
as Vault's unseal key / root token copies.

Nextcloud: adds infrastructure/vps-standby/nextcloud/manifests/
pvc-restore-cronjob.yaml, the read side of home's existing
nextcloud-pvc-sync restic backup - restores the latest snapshot from VPS
MinIO into this standby's PVC daily at 04:00 (2h after home's 02:00
backup). No Redis on the VPS side (no redis-operator deployed there,
out of scope for a standby that isn't serving traffic - Nextcloud
degrades gracefully to DB-based locking without it).
2026-08-17 22:31:41 -05:00
Scooby Husky 1a657ad0a4 Phase 2: deploy Authentik warm standby on the VPS
CNPG operator on the VPS (argocd-apps/vps-standby/cnpg-operator.yaml),
plus a CNPG 'replica cluster' for pg-authentik (infrastructure/
vps-standby/authentik/manifests/cnpg-cluster.yaml) that continuously
replays WAL shipped from home's pg-authentik via VPS MinIO's
cnpg-backups/pg-authentik bucket path - CNPG's native replica-cluster
mode, not a custom restore script like Vault needed (Vault has no
equivalent built-in continuous-replication-into-object-store feature).

The Authentik app itself runs at replicas: 0 in normal operation - the
replica DB is read-only until a deliberate manual promotion
(spec.replica.enabled: false), and a pod trying to write against a
read-only DB would just crashloop uselessly. values.yaml's authentik:
block is copied byte-identical from home's on purpose: the chart
deterministically derives its generated Secret's AUTHENTIK_SECRET_KEY
from these literal values.yaml strings, so both clusters land on the
same key without manually copying it - required since that key decrypts
things stored encrypted in the replicated DB.
2026-08-17 22:28:10 -05:00
Scooby Husky f2213b56b4 Fix Gitea mirror: use service=git instead of service=gitlab
service=gitlab invokes Gitea's GitLab-API-based downloader (for issues/
PRs/releases metadata), which calls the source's /api/v4/projects/...
REST API rather than doing a plain git clone. That was 404ing against
home GitLab and getting swallowed into a generic 'InternalServerError:
404 Not Found' with no further detail. We only want a plain pull-mirror
of the git repo itself, so service=git forces the plain git-clone
downloader instead.
2026-08-17 22:16:49 -05:00
Scooby HuskyandClaude Sonnet 5 777edf3895 Phase 2: deploy Gitea warm standby on the VPS
Pull-mirrors home GitLab's Homelabv4 repo on Gitea's own built-in mirror
scheduler (6h interval) - no custom sync job needed, per the original
plan. SQLite instead of the chart's default HA Postgres + Valkey cluster
(single-instance standby holding one small repo, not worth the extra
moving parts). A PostSync Job creates the mirror once, idempotently;
Gitea's scheduler handles all ongoing pulls after that.

Also added a public Cloudflare CNAME for gitlab.kube.huskypup.net ->
home.kube.huskypup.net: the VPS has no route to home's LAN via Netbird
(none of the mesh peers advertise that subnet, confirmed live), so
GitLab needs to be reachable the same way any other internet client
reaches it - home's public IP already has port 443 forwarded to
istio-ingressgateway from earlier this session.

Admin and GitLab-mirror credentials are plain Secrets created directly
on the VPS cluster (kubectl, not git) - same pattern as vault-unseal-key.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 22:05:24 -05:00
Scooby HuskyandClaude Sonnet 5 d6e8c2aa2d Optimize vault-restore: don't wait on condition=Ready (verified working already)
Confirmed end-to-end tonight: condition=Ready correctly times out every
cycle since Vault can't be Ready while sealed (the unseal step comes
right after this wait) - harmless via the existing || true fallback, but
wastes up to 2 minutes per restore cycle waiting on a condition that can
never be met at this point. Poll for the container process merely being
started instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 21:57:52 -05:00
Scooby HuskyandClaude Sonnet 5 fb124201bb Fix vault-restore RBAC: add watch verb for kubectl wait
Without it, kubectl wait's internal watch retry loop never terminates
cleanly even with --timeout set - confirmed live, it kept retrying well
past the stated 120s timeout. The restore/unseal steps themselves still
succeed regardless (verified end-to-end with a manual unseal), this only
affects the job's own graceful detection of the pod coming back.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 21:54:21 -05:00
Scooby HuskyandClaude Sonnet 5 e2bdab8940 Fix vault-restore: authenticate raft snapshot restore with a real token
raft snapshot restore is a privileged operation - unsealing alone isn't
enough, it needs an authenticated token. Missed this on first pass, caught
live: 'Code: 403. Errors: * permission denied'.

First run: uses the throwaway init's own fresh root token. Steady state
(already restored at least once): uses a copy of home's real root token,
which becomes valid on this Vault the moment the first restore completes
(its auth data becomes byte-for-byte home's as of that snapshot). Stored
the same way as the unseal key - kubectl directly on the VPS, not git.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 21:46:23 -05:00
Scooby HuskyandClaude Sonnet 5 62751f7d75 Fix vps-vault restore script: pod name is vps-vault-0, not vault-0
StatefulSet pod naming follows the Helm release name (vps-vault, matching
the ArgoCD Application name), not the chart's default.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 21:43:26 -05:00
Scooby HuskyandClaude Sonnet 5 d0c5563671 Phase 2: deploy Vault warm standby on the VPS
Single-node Vault (raft storage, standalone mode) plus a restore CronJob
that pulls the latest raft snapshot from the VPS's own local MinIO
(already receiving snapshots every 6h from home) and restores it every
6h, offset 30 minutes after the home-side snapshot job.

Not an independent root of trust: raft snapshot restore replaces the
entire storage backend including the keyring, so after every restore
this Vault is sealed with HOME's actual keyring - unsealing it needs a
copy of home's real unseal key, stored directly on the VPS cluster
(kubectl, not git, same pattern as vps-minio-root-secret). First-run
bootstrap uses a throwaway single-Shamir-key init just to get through
the very first restore, then is irrelevant from then on.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 21:42:09 -05:00
Scooby HuskyandClaude Sonnet 5 8aed0bf98e Fix vault-raft-snapshot: mc download needs -L to follow redirects
Without -L, curl silently saved dl.min.io's redirect response body
(HTML/text) as the mc binary instead of the actual binary, failing at
runtime with a confusing shell-script-interpretation error. Confirmed
live while testing the leader-detection fix end-to-end.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 21:15:40 -05:00
Scooby HuskyandClaude Sonnet 5 7206022a2b Pin VPS MinIO root credentials to a real Secret, fix regeneration bug
rootUser/rootPassword were left unset, relying on the chart's
auto-generation - but ArgoCD renders Helm via 'helm template', which
doesn't support the lookup() function charts use to preserve an existing
generated secret across upgrades. With selfHeal: true, every single sync
generated a BRAND NEW random root password, immediately invalidating
whatever credentials were stored in the home cluster's Vault for the
backup/CronJob consumers (vault-raft-snapshot, 3x CNPG barmanObjectStore,
nextcloud-pvc-sync) - confirmed live: all of them failing with 'Access
Key Id does not exist' right after being fixed, because the password had
already rotated out from under them again.

Fixed by pointing at a pre-existing Secret created directly on the VPS
cluster (kubectl, not git - this cluster has no Vault/ESO of its own).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 21:12:37 -05:00
Scooby HuskyandClaude Sonnet 5 7d73542645 Fix nextcloud-pvc-sync: add fsGroup for data access, fix restic cache dir
Now that the endpoint/credentials are fixed and this job can actually
reach the VPS, it surfaced two more real bugs: no fsGroup (couldn't read
the PVC data at all - nextcloud's real Deployment uses fsGroup 33/www-data,
this job never matched it) and no writable cache dir for runAsUser 65534
(restic defaults to $HOME/.cache).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 21:05:38 -05:00
Scooby HuskyandClaude Sonnet 5 7a6ecd14fe Bump rancher memory limit 2Gi -> 3Gi, was OOMKilling during restart recovery
Repeated restarts from tonight's network disruption kept OOMKilling
rancher - each restart's startup reconciliation is memory-hungry, and 2Gi
wasn't enough headroom for that burst even though steady-state usage is
normally fine. Nodes have ample free memory (26-36% used).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 21:01:48 -05:00
Scooby HuskyandClaude Sonnet 5 2fc7fb5b59 Fix Nextcloud Redis config: use externalRedis, not redis (bundled subchart toggle)
redis.enabled: true was unintentionally deploying the chart's BUNDLED redis
subchart (redis-21.1.3, StatefulSet nextcloud-redis-replicas) alongside
the actual redis-operator instance - that key's schema is pure subchart
passthrough (auth/image/master/replica), it has no host field. The
host/existingSecret keys previously set there were silently ignored.
Found while investigating why nextcloud-redis-replicas-0 was
crash-looping on failed liveness probes (confirmed via : the correct key for pointing at an external Redis
is the separate externalRedis: block).

This means Nextcloud's config.php may never have been correctly pointed
at the redis-operator instance despite apps/nextcloud/manifests/redis-cr.yaml
existing and being healthy - worth confirming file-locking/caching
actually engages after this deploys.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 20:54:28 -05:00
Scooby HuskyandClaude Sonnet 5 e2c8ef0fde Commit crowdsec-firewall-bouncer headless-service fix (was uncommitted)
This fix for the EPERM ClusterIP-from-hostNetwork issue (Cilium socketLB
hostNamespaceOnly + hostNetwork bouncer + ClusterIP VIP) was made and
verified earlier but never committed - ArgoCD's selfHeal kept reverting
the live ConfigMap back to the broken crowdsec-service URL against git's
stale version, causing a fresh bouncer pod to crash-loop with the original
EPERM error again.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 20:51:47 -05:00
Scooby HuskyandClaude Sonnet 5 018eb3570f Build real pod-egress routing to the VPS, fix everything blocked by its absence
Root cause of tonight's earlier CrowdSec/CNPG-backup workarounds: no node
in the cluster had any route into the Netbird mesh CIDR (100.108.0.0/16)
for pod-originated traffic. The per-namespace netbird 'router' pods are
inbound-only infrastructure (external peers reaching K8s services); their
own architecture has no reverse path.

Fix, in two parts:

1. infrastructure/netbird/manifests/egress-daemonset.yaml - one netbird
   client per node, hostNetwork so its wt0 interface lives in the node's
   real network namespace, plus a sidecar that adds a host route sending
   100.108.0.0/16 out via it. hostNetwork requires a scoped Kyverno
   PolicyException (infrastructure/kyverno/policies/netbird-egress-exception.yaml)
   to the disallow-host-namespaces STIG policy - narrowly for this one
   DaemonSet by name, not a namespace-wide exclusion.

2. Discovered the route alone wasn't enough for k3s NodePort traffic
   (vps-minio:30900): Netbird manages its own nftables ACLs independent of
   iptables/Kyverno, and its forward chain (netbird-rt-fwd) only permits
   *established* connections through a peer acting as a router - never new
   ones, by design, unless a Netbird 'Network Route' policy is explicitly
   configured (it isn't, for this VPS). Locally-terminated connections
   (tinyproxy) go through a separate, already-permissive ACL chain, which
   is why the CrowdSec proxy fix from earlier tonight worked. Replicated
   that working pattern for MinIO: minio-forward.service on the VPS host
   (systemd, socat) forwards 100.108.113.41:9000 -> MinIO's ClusterIP,
   avoiding the NodePort path entirely.

Re-enabled everything that was disabled/suspended earlier tonight because
of this gap, pointed at the new endpoint:
- CrowdSec CAPI/console-enroll (removed DISABLE_ONLINE_API, restored the
  VPS proxy env vars)
- n8n/nextcloud/authentik CNPG backup.barmanObjectStore
- vault-raft-snapshot CronJob (unsuspended)
- nextcloud PVC content sync CronJob endpoint

vps-minio.netbird.internal is retired everywhere - it was never actually
resolvable (Netbird has no DNS configured) even before today's routing
fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 20:36:13 -05:00
Scooby HuskyandClaude Sonnet 5 07a59d3b34 Fix vault-raft-snapshot: detect active leader, suspend until VPS egress exists
vault operator raft snapshot save fails against a standby Vault node with
'incomplete snapshot, unable to read SHA256SUMS.sealed file' - it must run
directly against the active HA leader. The job was hardcoded to vault-0,
which has been a standby since boot, so it had been failing on every run.

Fixed to detect the actual leader at runtime via vault status is_self.
Also suspended the CronJob: the upload step targets
vps-minio.netbird.internal, unreachable from any pod in this cluster (same
missing egress-route gap as tonight's CrowdSec/CNPG-backup fixes). The
leader-detection fix is kept since it's correct and independent - just
unsuspend once real egress routing exists.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 19:56:12 -05:00
Scooby HuskyandClaude Sonnet 5 e18308b750 Fix pg-n8n's long-recurring CNPG status-extraction error - missing AuthorizationPolicy rule
n8n's allow-n8n-access AuthorizationPolicy only allowed traffic from the
n8n/istio-system/prometheus namespaces - unlike nextcloud/authentik's
policies, it was missing the broader ipBlocks: 10.244.0.0/16 fallback
rule. The CNPG operator (cnpg-system namespace) polling pg-n8n instances'
status API (port 8000) wasn't covered by any rule, so Istio's L4
authorization enforcement (via ztunnel - this traffic is direct pod-IP,
not routed through a Service, so waypoint/L7 never sees it) rejected the
connection outright, surfacing as a TCP reset to the operator.

This was the actual cause of pg-n8n's long-recurring 'Instance Status
Extraction Error' - confirmed 100% reproducible tonight (every single
operator poll attempt failed identically), not the intermittent
HBONE-tunnel issue it resembled from the outside. Also labeled the n8n
namespace istio.io/use-waypoint=waypoint (was in the enrollment script's
namespace list but the label was missing live - namespace parity fix,
though not itself the fix for this specific issue since this traffic
bypasses the waypoint).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 19:50:29 -05:00
Scooby HuskyandClaude Sonnet 5 2c8d29c51f Disable VPS MinIO backup on n8n/nextcloud/authentik CNPG clusters
Discovered while debugging pg-n8n's recurring 'Instance Status Extraction
Error': the root cause is the same missing pod-egress route to the Netbird
mesh found tonight while fixing CrowdSec - continuousArchiving kept failing
to connect to vps-minio.netbird.internal, and CNPG correctly holds
Ready=False while archiving is broken (a real condition, not cosmetic).
That's what was driving these three apps' ArgoCD health flapping.

No node in the cluster has any route into 100.108.0.0/16 for
pod-originated traffic - the per-namespace netbird 'router' pods
(gitlab, vault, argocd, etc.) are inbound-only infrastructure, not egress
gateways. gitlab's own CNPG backup is unaffected (points at a local
in-cluster MinIO, not the VPS).

Commented out rather than deleted - re-enable once real pod-egress
routing exists, tracked as a separate task. No data loss: this is WAL
archiving/backup, not the live database.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 19:45:04 -05:00
Scooby HuskyandClaude Sonnet 5 ff638250d6 Disable CrowdSec online API - CAPI unreachable from home, egress routing gap found
LAPI treated a failed CAPI/console-enroll attempt as fatal at startup,
crashing the whole server (and taking agent + bouncer down with it via
their LAPI dependency) on every restart. Root cause investigation tonight
found two separate real issues:

1. Home's public IP gets a 403 from AWS WAF in front of api.crowdsec.net
   on any request, unrelated to credentials - ENROLL_KEY is valid and
   correctly wired but can't get past this.
2. Attempted routing around it via a VPS-side proxy over Netbird; found no
   node in the cluster has any route into the Netbird mesh CIDR for
   pod-originated egress traffic - the per-namespace netbird router pods
   (gitlab, vault, argocd, etc.) are inbound-only infrastructure. A proper
   fix needs either a Talos-native static route (fragile against pod IP
   churn) or a deliberate STIG exception for hostNetwork - out of scope
   for tonight, tracked separately.

DISABLE_ONLINE_API=true lets LAPI start cleanly with local-only detection
(agent + bouncer functional again). ENROLL_KEY stays set in Vault/values -
flip this back once real egress routing exists.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 19:37:33 -05:00
Scooby HuskyandClaude Sonnet 5 c8643ca025 Add dedicated netbird router for crowdsec namespace
Every other major namespace (gitlab, vault, argocd, nextcloud, etc.) has
its own 3-replica netbird-router deployment auto-provisioned from an
NBRoutingPeer CR; crowdsec never got one since it never previously needed
mesh connectivity. Without it, crowdsec pods had no dedicated mesh routing
path, causing intermittent connection resets/hangs reaching the VPS proxy
used to route around CrowdSec's WAF block on home's public IP.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 19:27:39 -05:00
Scooby HuskyandClaude Sonnet 5 33b0d018f9 Route CrowdSec CAPI/console-enroll traffic through VPS proxy
Confirmed 2026-08-17: home's public IP gets a 403 ForbiddenException from
AWS WAF in front of api.crowdsec.net, even on a bare unauthenticated test
request - the VPS's IP gets a normal 401 on the identical request. This is
an IP-level block unrelated to credentials, which is why LAPI kept hitting
403 on both anonymous auto-registration and console enroll with a fresh,
valid enrollment key.

Fix: HTTPS_PROXY/HTTP_PROXY pointed at a tinyproxy instance on the VPS
(bound to its Netbird interface only, not internet-reachable), reachable
from home over the existing Netbird mesh via netbird-cluster-router.
cscli/crowdsec's Go HTTP client honors standard proxy env vars.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 18:39:23 -05:00
Scooby HuskyandClaude Sonnet 5 bbb80687ca Fix netbird-operator namespace reference in istio mesh policies
allow-netbird-operator-access (AuthorizationPolicy) and the netbird-operator
PeerAuthentication both targeted namespace 'netbird-operator', which has
never existed - the actual netbird-operator pod runs in the 'netbird'
namespace. This permanently blocked istio-mesh-config from going Synced
(kubectl apply failed: namespaces "netbird-operator" not found on every
sync attempt).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 18:20:00 -05:00
Scooby HuskyandClaude Sonnet 5 3226169af8 Fix Vault kubernetes-auth reviewer token expiring 1hr after every boot
vault-init-and-store.sh minted the auth/kubernetes/config reviewer JWT via
'kubectl create token vault-auth' with no --duration, defaulting to a 1hr
TTL, then wrote it into Vault as a static token_reviewer_jwt. ~1hr after
every cluster boot/hook rerun it silently expired, breaking TokenReview
for every kubernetes-auth login (including ESO's) with a generic 403 that
logged nothing at INFO/ERROR. This was the actual root cause of today's
broad ArgoCD Degraded wave across ~14 apps - not Vault's seal state, which
was fine the whole time.

Fix: clear token_reviewer_jwt and rely on Vault's local-JWT auto-detection
(disable_local_ca_jwt=false default), which reads the vault pod's own
kubelet-refreshed SA token from disk on every call instead of a static
copy. The vault SA already has system:auth-delegator via the existing
vault-server-binding ClusterRoleBinding.

Applied live directly against Vault to unblock immediately; confirmed
ClusterSecretStore vault-backend flipped to Ready=True and previously
broken ExternalSecrets (guacamole, teslamate, unpoller, netbird x3) all
resynced successfully.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 18:19:13 -05:00
Scooby Husky 9bb4ada1f5 Fix Redis CR permission denied on PVC (needs fsGroup)
redis-nextcloud-0 was CrashLoopBackOff: "mkdir: can't create directory
'/node-conf': Permission denied" and same for appendonlydir - the pod
had no podSecurityContext, so it couldn't write to its own PVC-backed
data directories. Added fsGroup/runAsUser: 1000 matching the
opstree/redis image's default non-root user. Verified live: Redis now
starts cleanly ("Ready to accept connections").
2026-08-17 17:20:16 -05:00
Scooby Husky a8b16fa4cb Fix Redis CR apiVersion/schema - v1beta2 doesn't exist, blocked whole nextcloud sync
apps/nextcloud/manifests/redis-cr.yaml used redis.redis.opstreelabs.in/v1beta2
(guessed, flagged as unverified when written) - the actually-installed CRD
only has v1beta1. ArgoCD validates all resource API groups/versions before
syncing anything, so this one bad apiVersion blocked the ENTIRE nextcloud
Application from syncing at all (every resource showed OutOfSync, not just
this one).

Also fixed against the live CRD schema: redisSecret belongs nested under
kubernetesConfig, not top-level, and redisExporter.image is required by the
schema even when enabled: false.

Verified live: kubectl apply succeeded and the Redis resource was created
before committing this.
2026-08-17 16:56:20 -05:00
Scooby Husky 7b3669e9a6 Fix CNPG operator -> instance status communication + add CNPG health check
Two related fixes, found while working through why authentik/n8n/gitlab
ArgoCD Applications showed Unknown/stuck health despite pods being fine:

1. ArgoCD has no built-in health check for CNPG's Cluster CRD, so it
   always reported "Unknown" - and since app-level health rolls up to
   the worst resource status, every app with a CNPG Cluster showed
   Unknown/Progressing regardless of actual state. Added a Lua health
   check (resource.customizations.health.postgresql.cnpg.io_Cluster)
   reading .status.conditions[Ready] / .status.phase.

2. Once that started reporting real status instead of masking it,
   pg-authentik showed a genuine problem: CNPG's operator couldn't
   reach its Postgres instances' status API (port 8000) - "Cannot
   extract Pod status ... context deadline exceeded" - because:
   a) authentik-ingress's CiliumNetworkPolicy never allowed cnpg-system
      as a source (crowdsec-ingress already had this exception,
      authentik-ingress was just missing it - inconsistency, not
      deliberate).
   b) Even after fixing (a), still blocked - pg-authentik's pods are
      ambient-mesh-enrolled, so the connection actually goes through
      ztunnel's HBONE tunnel (port 15008) first, same underlying issue
      as the argocd-redis fix from earlier today. Rather than keep
      finding and patching this per-namespace as it recurs, added a
      cluster-wide CiliumClusterwideNetworkPolicy allowing HBONE
      broadly - ztunnel's own mTLS/SPIFFE identity verification is the
      real security boundary for mesh traffic; Cilium blocking the
      tunnel port itself was only breaking legitimate traffic, not
      adding meaningful protection on top.

Verified live: CNPG operator immediately went from failing status
extraction to successfully reconciling (recreating a pod to reattach
its PVC) once both fixes were in place.
2026-08-17 16:41:11 -05:00
Scooby Husky 99b8ca6be9 Fix Nextcloud OIDC client_id/secret - was hardcoded empty
apps/nextcloud/values.yaml had oidc_login_client_id/secret as literal
empty strings, causing "client_id is missing or invalid" on login.
The matching Authentik blueprint (infrastructure/authentik/nextcloud-
blueprint.yaml) already defines fixed credentials for this OAuth2
provider - just never got copied over. Filled in to match.

Also found and left alone: an orphaned nextcloud-oidc-config ConfigMap
(from January) that already had the correct values but isn't actually
mounted by the current Deployment (which uses nextcloud-config instead)
- harmless dead resource, not the active config path.
2026-08-17 16:33:47 -05:00
Scooby Husky 8372887ce1 Fix two long-standing bugs in the vault-init PreSync hook script
Both pre-existing (not from today's work), found while cleaning up a
recurring stale-pod pattern (a fresh vault-init Job failing on every
sync of the vault Application):

1. Init-detection was always broken: `vault status -format=json`
   pretty-prints with a space after colons ("initialized": true), but
   the parser (grep -o '"initialized":[a-z]*' | cut -d: -f2) required
   no space - it always matched empty, so `initialized`/`sealed` were
   always blank regardless of real state. Every run concluded "not
   initialized" and tried to re-init, which correctly fails once Vault
   already has been ("Vault is already initialized"). Switched to jq.
   Separately, the old `|| echo '{}'` fallback also discarded valid
   JSON on vault status's normal non-zero exit codes (2 = sealed) -
   fixed too, though the parsing bug was the actual blocker.

2. `vault_exec` wraps `kubectl exec` without `-i`, so the heredoc piped
   into `vault policy write NAME -` never reached the remote command -
   vault saw an empty policy body. Never previously reached in
   practice since bug #1 always failed the script earlier. Added -i.

Verified end-to-end with a one-off Job run: script now correctly
detects the already-initialized state and completes successfully,
including the policy/role writes that were previously unreachable.
2026-08-17 16:13:48 -05:00
Scooby Husky ee71e4f46f Fix DDNS/failover Cloudflare zone name and add restricted PSS securityContext
- ZONE_NAME was "kube.huskypup.net" in both the home-ip-ddns CronJob and
  the VPS failover watcher - that's a record, not a Cloudflare zone (the
  actual zone is the parent "huskypup.net"). Caused home-ip-ddns to fail
  every run (zone lookup returned zero results, curl -f exit 22) -
  confirmed live and fixed.
- Added seccompProfile/non-root/dropped-capabilities securityContext to
  the three CronJobs added this session that were missing it (flagged by
  the cluster's "restricted" PodSecurity admission). Repointed the raft
  snapshot job's mc binary install from /usr/local/bin to /tmp so it
  still works running as non-root.
2026-08-17 15:55:47 -05:00
Scooby Husky 983671ef50 Fix ArgoCD Redis connectivity broken by ambient mesh HBONE NetworkPolicy gap
Root cause (confirmed via ztunnel logs): the argo-cd chart's default
NetworkPolicies only allow each component's application port (e.g. 6379
for redis), but Istio ambient mode routes ALL pod-to-pod traffic through
ztunnel's HBONE tunnel on port 15008 first - so the tunnel itself was
being blocked even though the "real" port was allowed. Every inter-pod
connection in the argocd namespace hung for exactly 10s then reset;
ztunnel's own log named it directly ("connection timed out, maybe a
NetworkPolicy is blocking HBONE port 15008"). This broke argocd-server's
Redis-backed session/cluster-info caching cluster-wide and was silently
preventing the root Application from picking up new child Applications.

Fix: additive NetworkPolicy allowing ingress on 15008 for all argocd
pods (NetworkPolicies union across multiple policies selecting a pod,
so this doesn't touch/replace the chart's rendered ones - safe across
Helm upgrades).

Also: dropped argocd from the namespace-enrollment job's waypoint list.
argocd's only AuthorizationPolicy (allow-argocd-access) is a plain
source-namespace/IP match with no L7 rules - its own status shows
"attached to ztunnel", not waypoint - so forcing L7 waypoint processing
onto the namespace was unnecessary overhead, not a security requirement.
ztunnel's mTLS still fully covers it. (This was a red herring for the
HBONE bug itself, not the fix, but a valid simplification found along
the way.)
2026-08-17 15:08:33 -05:00
Scooby Husky 7990f1fa47 Add VPS warm-standby/backup site (Phase 0-1b)
Foundation for a DR/backup path using an always-on VPS as a second
ArgoCD-managed cluster, plus DB/backup standardization work that fell
out of it:

- vps-standby ArgoCD cluster destination + AppProject, MinIO backup
  receiver, VPS bootstrap script (k3s, Netbird, cert-manager)
- Dual-site DNS failover watcher + home-IP DDNS CronJob, Cloudflare
  token moved out of git into Vault+ExternalSecret
- Nextcloud migrated from ad-hoc MariaDB to CNPG + redis-operator
  (matches n8n/Authentik/GitLab's backup-native pattern)
- Authentik's CNPG manifests moved into the actual ArgoCD-synced
  manifests/ path (were present but never wired into the sync path)
- Vault raft-snapshot CronJob, CNPG barmanObjectStore backups
  (Authentik/n8n/Nextcloud), Nextcloud file-PVC restic sync - all
  targeting the new VPS MinIO receiver

See VPS Warm-Standby plan doc for full design rationale.
2026-08-17 14:59:26 -05:00