Commit Graph
47 Commits
Author SHA1 Message Date
Scooby Husky 983671ef50 Fix ArgoCD Redis connectivity broken by ambient mesh HBONE NetworkPolicy gap
Root cause (confirmed via ztunnel logs): the argo-cd chart's default
NetworkPolicies only allow each component's application port (e.g. 6379
for redis), but Istio ambient mode routes ALL pod-to-pod traffic through
ztunnel's HBONE tunnel on port 15008 first - so the tunnel itself was
being blocked even though the "real" port was allowed. Every inter-pod
connection in the argocd namespace hung for exactly 10s then reset;
ztunnel's own log named it directly ("connection timed out, maybe a
NetworkPolicy is blocking HBONE port 15008"). This broke argocd-server's
Redis-backed session/cluster-info caching cluster-wide and was silently
preventing the root Application from picking up new child Applications.

Fix: additive NetworkPolicy allowing ingress on 15008 for all argocd
pods (NetworkPolicies union across multiple policies selecting a pod,
so this doesn't touch/replace the chart's rendered ones - safe across
Helm upgrades).

Also: dropped argocd from the namespace-enrollment job's waypoint list.
argocd's only AuthorizationPolicy (allow-argocd-access) is a plain
source-namespace/IP match with no L7 rules - its own status shows
"attached to ztunnel", not waypoint - so forcing L7 waypoint processing
onto the namespace was unnecessary overhead, not a security requirement.
ztunnel's mTLS still fully covers it. (This was a red herring for the
HBONE bug itself, not the fix, but a valid simplification found along
the way.)
2026-08-17 15:08:33 -05:00
Scooby Husky 7990f1fa47 Add VPS warm-standby/backup site (Phase 0-1b)
Foundation for a DR/backup path using an always-on VPS as a second
ArgoCD-managed cluster, plus DB/backup standardization work that fell
out of it:

- vps-standby ArgoCD cluster destination + AppProject, MinIO backup
  receiver, VPS bootstrap script (k3s, Netbird, cert-manager)
- Dual-site DNS failover watcher + home-IP DDNS CronJob, Cloudflare
  token moved out of git into Vault+ExternalSecret
- Nextcloud migrated from ad-hoc MariaDB to CNPG + redis-operator
  (matches n8n/Authentik/GitLab's backup-native pattern)
- Authentik's CNPG manifests moved into the actual ArgoCD-synced
  manifests/ path (were present but never wired into the sync path)
- Vault raft-snapshot CronJob, CNPG barmanObjectStore backups
  (Authentik/n8n/Nextcloud), Nextcloud file-PVC restic sync - all
  targeting the new VPS MinIO receiver

See VPS Warm-Standby plan doc for full design rationale.
2026-08-17 14:59:26 -05:00
Scooby HuskyandClaude Opus 4.6 5163403e24 Scale down non-critical replicas for homelab resource savings
Reduce replicas to 1 for workloads that don't need HA in a homelab:
- Prometheus 2→1, Alertmanager 2→1 (~4.4GB RAM saved)
- cert-manager 3→1
- GitLab: registry 2→1, kas 2→1, sidekiq 2→1, praefect 2→1,
  pgbouncer-ro 3→1
- Guacamole + guacd 2→1
- Kiali 2→1, ArgoCD server 2→1
- Kyverno background-controller 2→1
- Scylla operator 2→1
- ext-authz-proxy 2→1, netbird-cluster-router 2→1

Kept multi-replica: coredns, envoy-gateway, kyverno admission,
vault, argocd-repo-server, gitlab-webservice, istiod, rook-ceph CSI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 23:00:48 -05:00
Scooby HuskyandClaude Opus 4.6 a8b9064248 Set Grafana deployment strategy to Recreate for RWO PVC compatibility
RollingUpdate with a ReadWriteOnce PVC causes a Multi-Attach deadlock
where the new pod can't start because the old pod still holds the
volume. Recreate strategy ensures the old pod is terminated before
the new one is created.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 22:33:16 -05:00
Scooby HuskyandClaude Opus 4.6 02632d06a4 Fix NetBird service annotations: use netbird.io/expose for operator
The operator watches for netbird.io/expose (presence-based), not
netbird.io/resource which was from the old Helmfile bootstrap script.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 16:25:35 -05:00
Scooby HuskyandClaude Opus 4.6 2edc8a156e Fix netbird-operator: add self-hosted management URL
The operator was missing managementURL, causing it to default to the
NetBird SaaS cloud (api.netbird.io) instead of the self-hosted instance.
This resulted in an infinite reconcile loop and no networks/resources
being created.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 16:08:51 -05:00
Scooby HuskyandClaude Opus 4.6 987f91c877 Enable CephFS CSI driver for SMB gateway and CephFS volumes
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 12:27:00 -05:00
Scooby HuskyandClaude Opus 4.6 7eae427d1a Scale services to 2+ replicas for HA
- ArgoCD: server and repo-server → 2 replicas
- GitLab: webservice, sidekiq, registry, KAS, praefect → min 2 replicas
- Guacamole: client and guacd → 2 replicas
- Kiali: 1 → 2 replicas
- Alertmanager: add 2 replicas
- TeslaMate CNPG: 1 → 2 instances

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 03:08:49 -05:00
Scooby HuskyandClaude Opus 4.6 c0fd93c73b Upgrade Ceph to v19.2.3 (Squid)
Includes security fixes (CVE-2023-43040) and RBD/RADOS improvements.
Rook operator upgraded to v1.19.2 via Helm.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:38:10 -05:00
Scooby HuskyandClaude Opus 4.6 a93df9375a Add talos-cp-02 back to Rook-Ceph nodes list
LVM filter now applied via Talos machine config on all nodes to prevent
lvs hang on BIOS partitions. talos-cp-02 can safely run OSD prepare jobs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 16:31:13 -05:00
Scooby HuskyandClaude Opus 4.6 2643d5e0bf Fix TeslaMate DB corruption and Rook-Ceph reconcile hang
- Scale CNPG pg-teslamate to 1 instance: on-disk storage corruption
  (corrupted sequences, system catalog indexes) means replicas cloned
  via pg_basebackup inherit corruption and can't open any database.
  Primary works from cached catalogs. Sequences already rebuilt live.

- Remove talos-cp-02 from Rook-Ceph nodes list: ceph-volume inventory
  hangs on BIOS partition via nsenter lvs in host namespace. Talos
  read-only /etc prevents LVM filter fix. OSD 1 keeps running from
  its existing deployment.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 14:53:31 -05:00
Scooby HuskyandClaude Opus 4.6 8df30c11ee Re-enable Ceph dashboard SSL to fix SAML http/https URL mismatch
X-Forwarded-Proto approach didn't work - Ceph dashboard SAML doesn't
read it. Re-enabling native SSL so the dashboard sees https:// URLs
directly. Updated both Istio and Envoy Gateway backends for TLS with
insecureSkipVerify for the self-signed cert.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 13:56:33 -05:00
Scooby HuskyandClaude Opus 4.6 fb0ed0071f Set X-Forwarded-Proto header for Ceph dashboard SAML authentication
Dashboard runs with ssl: false behind TLS-terminating proxy, so SAML
library sees http:// URLs and rejects responses expecting https://.
Explicitly set X-Forwarded-Proto: https on both Istio VirtualService
and Envoy Gateway HTTPRoute so the SAML ACS URL matches.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 13:48:35 -05:00
Scooby HuskyandClaude Opus 4.6 7c184f9d14 Fix Ceph dashboard DestinationRule: disable TLS to match dashboard config
Dashboard has ssl: false (plain HTTP on port 7000), but the DestinationRule
had tls.mode: SIMPLE which caused Istio to attempt TLS to the backend,
resulting in WRONG_VERSION_NUMBER errors.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 13:41:57 -05:00
Scooby HuskyandClaude Opus 4.6 bc37f9d1a4 Complete OSD migration: remove temp Kyverno policy, fix talos-cp-02 device name
All 4 OSDs successfully recreated in clean raw mode (NVMe only, no DB devices).
Fixed talos-cp-02 NVMe device path from nvme0n1 to nvme1n1 to match actual hardware.
Removed temporary fix-rook-osd-metadata-device Kyverno policy as it's no longer needed.
Operator runs normally without manual patches.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 03:47:14 -05:00
Scooby HuskyandClaude Opus 4.6 1a8ef4347f Restore Ceph dashboard Istio VirtualService on port 7000
The dashboard has SSL disabled and listens on port 7000. Restoring the
VirtualService with the correct port for the Istio gateway path.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 00:19:11 -05:00
Scooby HuskyandClaude Opus 4.6 d27b90ba19 Revert external-dns gateway-httproute source
Services behind the Envoy gateway cannot reach Istio ambient mesh
backends due to STRICT mTLS. Reverting DNS to point back to the Istio
gateway until the migration path is resolved.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 00:11:23 -05:00
Scooby HuskyandClaude Opus 4.6 3244fdf611 Allow envoy-gateway-system ingress to vault and authentik
All namespaces with CiliumNetworkPolicies need envoy-gateway-system
access now that services are migrating from Istio to Envoy gateway.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 23:59:53 -05:00
Scooby HuskyandClaude Opus 4.6 7a9703f3bc Add gateway-httproute source to external-dns
Allows external-dns to create DNS records from Gateway API HTTPRoute
resources, needed for services using the Envoy edge gateway.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 23:50:58 -05:00
Scooby HuskyandClaude Opus 4.6 13ff2dafd6 Remove old Istio VirtualService for Ceph dashboard
The dashboard now uses the Envoy edge gateway via HTTPRoute. The old
Istio VirtualService was causing DNS to point at the wrong gateway IP.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 23:49:32 -05:00
Scooby HuskyandClaude Opus 4.6 c199715276 Allow envoy-gateway-system ingress to rook-ceph for dashboard
The Ceph dashboard moved from Istio to the Envoy edge gateway but the
CiliumNetworkPolicy was still only allowing istio-system, causing 503s.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 23:24:26 -05:00
Scooby HuskyandClaude Opus 4.6 31113ecf16 Fix Ceph dashboard Backend to use HTTP port 7000
The dashboard has SSL disabled and listens on port 7000, but the
Backend was configured for TLS on port 8443.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 22:32:09 -05:00
Scooby HuskyandClaude Opus 4.6 b795f5aa8f Add Kyverno PolicyExceptions for third-party charts
ArgoCD, NetBird operator, and Scylla Manager are upstream Helm charts
that don't set resource limits, probes, drop-all capabilities, or
disable SA token automount. These audit-mode policy violations inflated
the cumulative Prometheus failure counter to ~3800.

- argocd: exempt from require-resource-limits, require-probes
- netbird-operator: exempt from require-resource-limits,
  disallow-automount-sa-token, require-probes
- scylla-manager: exempt from require-drop-all-capabilities,
  disallow-automount-sa-token

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:43:22 -05:00
Scooby HuskyandClaude Opus 4.6 0d39bcddc9 Add ServiceEntry for wildcard host to fix Kiali KIA0104
The global-authentik-authz AuthorizationPolicy uses *.kube.huskypup.net
as a host match, but Kiali flagged KIA0104 because the wildcard had no
matching entry in Istio's service registry. Adding a MESH_EXTERNAL
ServiceEntry with resolution NONE registers the hostname without
affecting routing.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:35:20 -05:00
Scooby HuskyandClaude Opus 4.6 e1960d7071 Fix create-block-db-symlink to use activate-osd volume
Use the OSD-specific activate-osd volume mount instead of rook-data
parent directory. Multiple old OSD dirs exist on each node and the
script was creating the symlink in the wrong (stale LVM) directory.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:25:31 -05:00
Scooby HuskyandClaude Opus 4.6 550ee8fd13 Fix Kyverno OSD policy init container ordering
Use patchesJson6902 with append (/-) for create-block-db-symlink to
ensure it runs after activate, not before. The patchStrategicMerge was
inserting it before activate, causing prime-osd-dir to clear the symlink.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:15:20 -05:00
Scooby HuskyandClaude Opus 4.6 64723eb0ba Allow kube-apiserver webhook calls to netbird namespace
The netbird-operator webhook service runs in the netbird namespace and
receives admission webhook calls from kube-apiserver using node IPs
(172.28.101.0/24). Without this, ztunnel rejects the connections with
"allow policies exist, but none allowed".

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:09:55 -05:00
Scooby HuskyandClaude Opus 4.6 511ba55b30 Remove dedicated DB devices from Rook-Ceph OSD configuration
Remove metadataDevice (sdb) from all 4 nodes in cluster-values.yaml
since the SATA SSDs are slower than the NVMe data drives. Add temporary
Kyverno mutate policy to protect existing OSD deployments when the
operator is re-enabled - replaces expand-bluefs with a no-op and adds
conditional block.db symlink creation based on BlueStore label detection.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:06:34 -05:00
Scooby HuskyandClaude Opus 4.6 f8eea2ed45 Remove nessus from cluster configuration
Remove all nessus references: ArgoCD project destination, Istio ambient
enrollment, waypoint gateway, authorization policies, ext-authz policy,
TLS certificate, ingress gateway host, Kyverno exclusion, and Authentik
forward-auth blueprint.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 21:02:45 -05:00
Scooby HuskyandClaude Opus 4.6 c5e13377bd Fix talos client key PEM header for nginx compatibility
ED25519 PRIVATE KEY header is not recognized by nginx's OpenSSL.
Convert to standard PKCS#8 PRIVATE KEY header (same DER content).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 20:31:31 -05:00
Scooby HuskyandClaude Opus 4.6 4ed06c1aa1 Add envoy-gateway to Kiali gateway_api_classes
Kiali was reporting "Gateway API Class not found" errors for the
envoy-gateway Gateways (gateway/edge and guacamole/guacamole).
Explicitly register istio, istio-waypoint, and envoy-gateway classes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 20:28:15 -05:00
Scooby HuskyandClaude Opus 4.6 712307da14 Fix ztunnel PodMonitor port name to match actual container port
The ztunnel container exposes metrics on port named 'ztunnel-stats' (15020),
not 'http-monitoring'. This was preventing Prometheus from scraping ztunnel
L4 metrics, causing missing traffic in Kiali.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 20:03:03 -05:00
Scooby HuskyandClaude Opus 4.6 6646943a98 Fix talos-client-cert secret namespace
Secret was in prometheus namespace but the talos-metrics-proxy
deployment referencing it runs in kube-system namespace.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 20:00:36 -05:00
Scooby HuskyandClaude Opus 4.6 7a8eb2046f Fix n8n, gitlab, and netbird-operator degraded states
- n8n: Set replicaCount=1 (RWO PVC incompatible with multiple replicas),
  add resource limits to satisfy Kyverno policy
- gitlab: Add ignoreDifferences for redis-gitlab-additional service
  (port names and selectors managed by Redis operator)
- netbird-operator: Add ExternalSecret for netbird-mgmt-api-key in
  netbird namespace and add manifests source to Application

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 19:53:10 -05:00
Scooby HuskyandClaude Opus 4.6 af573c647f Switch hook jobs from bitnami/kubectl to alpine/k8s:1.32.13
bitnami/kubectl has removed all version tags from Docker Hub,
leaving only :latest which is blocked by Kyverno disallow-latest-tag
policy. Switch to alpine/k8s:1.32.13 which has kubectl, bash, jq,
and uses proper version tags.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 18:08:22 -05:00
Scooby HuskyandClaude Opus 4.6 263dd3f6cc Use bitnami/kubectl:1.29 for hook jobs
The :1.31 tag is gone from Docker Hub and :latest is blocked by Kyverno
disallow-latest-tag policy. The 1.29 image is cached on cluster nodes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 18:00:16 -05:00
Scooby HuskyandClaude Opus 4.6 306d374ede Fix bitnami/kubectl image tag in hook jobs
The bitnami/kubectl:1.31 tag doesn't exist on Docker Hub. Updated all
hook jobs to use bitnami/kubectl:latest.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 17:57:08 -05:00
Scooby HuskyandClaude Opus 4.6 e296400053 Add argocd-hook-sa ServiceAccount to app namespaces
PreSync/PostSync hook Jobs reference argocd-hook-sa but it only existed
in the argocd namespace. Jobs running in app namespaces (crowdsec,
authentik, netbird, vault, istio-system) need the SA in their own
namespace.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 17:51:04 -05:00
Scooby HuskyandClaude Opus 4.6 308145cc90 Add followChildren default to tetragon shell-execution TracingPolicy
The Tetragon CRD defaults followChildren to false in matchBinaries,
causing ArgoCD ServerSideDiff to detect a permanent diff.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 17:39:38 -05:00
Scooby HuskyandClaude Opus 4.6 131cbca4a6 Fix ArgoCD ServerSideDiff permanent OutOfSync diffs
Add explicit CRD/API defaults to manifests that were causing ArgoCD's
SSA dry-run to produce results different from live state:

- HTTPRoutes: add group, kind, weight defaults to parentRefs/backendRefs
- Kyverno ClusterPolicies: add skipBackgroundRequests, allowExistingViolations
- Tetragon TracingPolicies: add return, maxData, resolve, returnCopy defaults
- Gateway certificateRefs: add group="" default
- Guacamole Gateway: add group="" to certificateRefs

Add ignoreDifferences for resources that legitimately differ:
- Cilium cert Secrets (auto-generated, data always differs)
- Istio ValidatingWebhookConfiguration failurePolicy (istiod mutates)
- Crowdsec LAPI Secrets (randomly generated)
- ServiceMonitor/PodMonitor relabeling action defaults
- StatefulSet volumeClaimTemplates apiVersion/kind defaults

Persist argocd-cm ignoreDifferences config in ArgoCD Helm values.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 17:32:59 -05:00
Scooby HuskyandClaude Opus 4.6 782a271e17 Revert Ceph to v19.2.0 - v19.2.3 has same expand-bluefs issue
BlueStore has UUID inconsistency at block position 0x0 that causes
expand_devices assertion in both v19.2.0 and v19.2.3. The OSD
daemon itself runs fine; only expand-bluefs init container crashes.
Will use deployment patches with do-not-reconcile annotation instead.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 06:00:45 -05:00
Scooby HuskyandClaude Opus 4.6 ce23f9282c Fix prometheus secret namespace and frigate capabilities
- Revert talos-client-cert secret to prometheus namespace (prometheus
  pods mount this secret, not kube-system pods)
- Add CHOWN/DAC_OVERRIDE/SETUID/SETGID capabilities to frigate
  (s6-overlay needs chown for log directories)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 05:59:23 -05:00
Scooby HuskyandClaude Opus 4.6 fefa183db7 Upgrade Ceph to v19.2.3, fix netbird webhook failurePolicy
- Ceph v19.2.0 has BlueStore::expand_devices assertion bug that
  crashes OSD expand-bluefs init containers. v19.2.3 includes fixes.
- NetBird operator webhook failurePolicy: Fail blocks all pod
  creation cluster-wide when webhook service is unreachable.
  Changed to Ignore.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 05:29:39 -05:00
Scooby HuskyandClaude Opus 4.6 aff4ce7441 Fix talos-client-cert secret namespace to match deployment (kube-system)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 04:17:03 -05:00
Scooby HuskyandClaude Opus 4.6 1cdc6ac68c Fix crowdsec duplicate REGISTRATION_TOKEN env var
The chart auto-generates this env var; remove the manual one from values.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 04:10:01 -05:00
Scooby HuskyandClaude Opus 4.6 7640fb895b Fix degraded ArgoCD applications
- frigate: Remove duplicate PVCs (Helm manages them), update config size to 15Gi
- gitlab: Fix sync-admin-job secretKeyRef schema error, fetch password via kubectl
- nextcloud: Remove duplicate PVC, remove invalid pod-level securityContext
- rook-ceph: Update cephfs-smb-pvc to 50Gi, remove unsupported dashboard.config
- nessus: Add privileged PSS label for NET_ADMIN/NET_RAW capabilities
- scylla-manager: Add privileged PSS label for SYS_NICE capability
- n8n: Create missing n8n-main-persistence PVC
- projects: Add authentik namespace to applications project destinations

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 00:05:29 -05:00
Scooby Husky aacb8eebbe Initial commit 2026-03-09 20:21:35 -05:00