13 Commits
Author SHA1 Message Date
Scooby HuskyandClaude Sonnet 5 8aed0bf98e Fix vault-raft-snapshot: mc download needs -L to follow redirects
Without -L, curl silently saved dl.min.io's redirect response body
(HTML/text) as the mc binary instead of the actual binary, failing at
runtime with a confusing shell-script-interpretation error. Confirmed
live while testing the leader-detection fix end-to-end.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 21:15:40 -05:00
Scooby HuskyandClaude Sonnet 5 018eb3570f Build real pod-egress routing to the VPS, fix everything blocked by its absence
Root cause of tonight's earlier CrowdSec/CNPG-backup workarounds: no node
in the cluster had any route into the Netbird mesh CIDR (100.108.0.0/16)
for pod-originated traffic. The per-namespace netbird 'router' pods are
inbound-only infrastructure (external peers reaching K8s services); their
own architecture has no reverse path.

Fix, in two parts:

1. infrastructure/netbird/manifests/egress-daemonset.yaml - one netbird
   client per node, hostNetwork so its wt0 interface lives in the node's
   real network namespace, plus a sidecar that adds a host route sending
   100.108.0.0/16 out via it. hostNetwork requires a scoped Kyverno
   PolicyException (infrastructure/kyverno/policies/netbird-egress-exception.yaml)
   to the disallow-host-namespaces STIG policy - narrowly for this one
   DaemonSet by name, not a namespace-wide exclusion.

2. Discovered the route alone wasn't enough for k3s NodePort traffic
   (vps-minio:30900): Netbird manages its own nftables ACLs independent of
   iptables/Kyverno, and its forward chain (netbird-rt-fwd) only permits
   *established* connections through a peer acting as a router - never new
   ones, by design, unless a Netbird 'Network Route' policy is explicitly
   configured (it isn't, for this VPS). Locally-terminated connections
   (tinyproxy) go through a separate, already-permissive ACL chain, which
   is why the CrowdSec proxy fix from earlier tonight worked. Replicated
   that working pattern for MinIO: minio-forward.service on the VPS host
   (systemd, socat) forwards 100.108.113.41:9000 -> MinIO's ClusterIP,
   avoiding the NodePort path entirely.

Re-enabled everything that was disabled/suspended earlier tonight because
of this gap, pointed at the new endpoint:
- CrowdSec CAPI/console-enroll (removed DISABLE_ONLINE_API, restored the
  VPS proxy env vars)
- n8n/nextcloud/authentik CNPG backup.barmanObjectStore
- vault-raft-snapshot CronJob (unsuspended)
- nextcloud PVC content sync CronJob endpoint

vps-minio.netbird.internal is retired everywhere - it was never actually
resolvable (Netbird has no DNS configured) even before today's routing
fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 20:36:13 -05:00
Scooby HuskyandClaude Sonnet 5 07a59d3b34 Fix vault-raft-snapshot: detect active leader, suspend until VPS egress exists
vault operator raft snapshot save fails against a standby Vault node with
'incomplete snapshot, unable to read SHA256SUMS.sealed file' - it must run
directly against the active HA leader. The job was hardcoded to vault-0,
which has been a standby since boot, so it had been failing on every run.

Fixed to detect the actual leader at runtime via vault status is_self.
Also suspended the CronJob: the upload step targets
vps-minio.netbird.internal, unreachable from any pod in this cluster (same
missing egress-route gap as tonight's CrowdSec/CNPG-backup fixes). The
leader-detection fix is kept since it's correct and independent - just
unsuspend once real egress routing exists.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 19:56:12 -05:00
Scooby HuskyandClaude Sonnet 5 3226169af8 Fix Vault kubernetes-auth reviewer token expiring 1hr after every boot
vault-init-and-store.sh minted the auth/kubernetes/config reviewer JWT via
'kubectl create token vault-auth' with no --duration, defaulting to a 1hr
TTL, then wrote it into Vault as a static token_reviewer_jwt. ~1hr after
every cluster boot/hook rerun it silently expired, breaking TokenReview
for every kubernetes-auth login (including ESO's) with a generic 403 that
logged nothing at INFO/ERROR. This was the actual root cause of today's
broad ArgoCD Degraded wave across ~14 apps - not Vault's seal state, which
was fine the whole time.

Fix: clear token_reviewer_jwt and rely on Vault's local-JWT auto-detection
(disable_local_ca_jwt=false default), which reads the vault pod's own
kubelet-refreshed SA token from disk on every call instead of a static
copy. The vault SA already has system:auth-delegator via the existing
vault-server-binding ClusterRoleBinding.

Applied live directly against Vault to unblock immediately; confirmed
ClusterSecretStore vault-backend flipped to Ready=True and previously
broken ExternalSecrets (guacamole, teslamate, unpoller, netbird x3) all
resynced successfully.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 18:19:13 -05:00
Scooby Husky 8372887ce1 Fix two long-standing bugs in the vault-init PreSync hook script
Both pre-existing (not from today's work), found while cleaning up a
recurring stale-pod pattern (a fresh vault-init Job failing on every
sync of the vault Application):

1. Init-detection was always broken: `vault status -format=json`
   pretty-prints with a space after colons ("initialized": true), but
   the parser (grep -o '"initialized":[a-z]*' | cut -d: -f2) required
   no space - it always matched empty, so `initialized`/`sealed` were
   always blank regardless of real state. Every run concluded "not
   initialized" and tried to re-init, which correctly fails once Vault
   already has been ("Vault is already initialized"). Switched to jq.
   Separately, the old `|| echo '{}'` fallback also discarded valid
   JSON on vault status's normal non-zero exit codes (2 = sealed) -
   fixed too, though the parsing bug was the actual blocker.

2. `vault_exec` wraps `kubectl exec` without `-i`, so the heredoc piped
   into `vault policy write NAME -` never reached the remote command -
   vault saw an empty policy body. Never previously reached in
   practice since bug #1 always failed the script earlier. Added -i.

Verified end-to-end with a one-off Job run: script now correctly
detects the already-initialized state and completes successfully,
including the policy/role writes that were previously unreachable.
2026-08-17 16:13:48 -05:00
Scooby Husky ee71e4f46f Fix DDNS/failover Cloudflare zone name and add restricted PSS securityContext
- ZONE_NAME was "kube.huskypup.net" in both the home-ip-ddns CronJob and
  the VPS failover watcher - that's a record, not a Cloudflare zone (the
  actual zone is the parent "huskypup.net"). Caused home-ip-ddns to fail
  every run (zone lookup returned zero results, curl -f exit 22) -
  confirmed live and fixed.
- Added seccompProfile/non-root/dropped-capabilities securityContext to
  the three CronJobs added this session that were missing it (flagged by
  the cluster's "restricted" PodSecurity admission). Repointed the raft
  snapshot job's mc binary install from /usr/local/bin to /tmp so it
  still works running as non-root.
2026-08-17 15:55:47 -05:00
Scooby Husky 7990f1fa47 Add VPS warm-standby/backup site (Phase 0-1b)
Foundation for a DR/backup path using an always-on VPS as a second
ArgoCD-managed cluster, plus DB/backup standardization work that fell
out of it:

- vps-standby ArgoCD cluster destination + AppProject, MinIO backup
  receiver, VPS bootstrap script (k3s, Netbird, cert-manager)
- Dual-site DNS failover watcher + home-IP DDNS CronJob, Cloudflare
  token moved out of git into Vault+ExternalSecret
- Nextcloud migrated from ad-hoc MariaDB to CNPG + redis-operator
  (matches n8n/Authentik/GitLab's backup-native pattern)
- Authentik's CNPG manifests moved into the actual ArgoCD-synced
  manifests/ path (were present but never wired into the sync path)
- Vault raft-snapshot CronJob, CNPG barmanObjectStore backups
  (Authentik/n8n/Nextcloud), Nextcloud file-PVC restic sync - all
  targeting the new VPS MinIO receiver

See VPS Warm-Standby plan doc for full design rationale.
2026-08-17 14:59:26 -05:00
Scooby HuskyandClaude Opus 4.6 af573c647f Switch hook jobs from bitnami/kubectl to alpine/k8s:1.32.13
bitnami/kubectl has removed all version tags from Docker Hub,
leaving only :latest which is blocked by Kyverno disallow-latest-tag
policy. Switch to alpine/k8s:1.32.13 which has kubectl, bash, jq,
and uses proper version tags.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 18:08:22 -05:00
Scooby HuskyandClaude Opus 4.6 263dd3f6cc Use bitnami/kubectl:1.29 for hook jobs
The :1.31 tag is gone from Docker Hub and :latest is blocked by Kyverno
disallow-latest-tag policy. The 1.29 image is cached on cluster nodes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 18:00:16 -05:00
Scooby HuskyandClaude Opus 4.6 306d374ede Fix bitnami/kubectl image tag in hook jobs
The bitnami/kubectl:1.31 tag doesn't exist on Docker Hub. Updated all
hook jobs to use bitnami/kubectl:latest.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 17:57:08 -05:00
Scooby HuskyandClaude Opus 4.6 e296400053 Add argocd-hook-sa ServiceAccount to app namespaces
PreSync/PostSync hook Jobs reference argocd-hook-sa but it only existed
in the argocd namespace. Jobs running in app namespaces (crowdsec,
authentik, netbird, vault, istio-system) need the SA in their own
namespace.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 17:51:04 -05:00
Scooby HuskyandClaude Opus 4.6 131cbca4a6 Fix ArgoCD ServerSideDiff permanent OutOfSync diffs
Add explicit CRD/API defaults to manifests that were causing ArgoCD's
SSA dry-run to produce results different from live state:

- HTTPRoutes: add group, kind, weight defaults to parentRefs/backendRefs
- Kyverno ClusterPolicies: add skipBackgroundRequests, allowExistingViolations
- Tetragon TracingPolicies: add return, maxData, resolve, returnCopy defaults
- Gateway certificateRefs: add group="" default
- Guacamole Gateway: add group="" to certificateRefs

Add ignoreDifferences for resources that legitimately differ:
- Cilium cert Secrets (auto-generated, data always differs)
- Istio ValidatingWebhookConfiguration failurePolicy (istiod mutates)
- Crowdsec LAPI Secrets (randomly generated)
- ServiceMonitor/PodMonitor relabeling action defaults
- StatefulSet volumeClaimTemplates apiVersion/kind defaults

Persist argocd-cm ignoreDifferences config in ArgoCD Helm values.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 17:32:59 -05:00
Scooby Husky aacb8eebbe Initial commit 2026-03-09 20:21:35 -05:00