Commit Graph
83 Commits
Author SHA1 Message Date
Scooby HuskyandClaude Sonnet 5 33b0d018f9 Route CrowdSec CAPI/console-enroll traffic through VPS proxy
Confirmed 2026-08-17: home's public IP gets a 403 ForbiddenException from
AWS WAF in front of api.crowdsec.net, even on a bare unauthenticated test
request - the VPS's IP gets a normal 401 on the identical request. This is
an IP-level block unrelated to credentials, which is why LAPI kept hitting
403 on both anonymous auto-registration and console enroll with a fresh,
valid enrollment key.

Fix: HTTPS_PROXY/HTTP_PROXY pointed at a tinyproxy instance on the VPS
(bound to its Netbird interface only, not internet-reachable), reachable
from home over the existing Netbird mesh via netbird-cluster-router.
cscli/crowdsec's Go HTTP client honors standard proxy env vars.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 18:39:23 -05:00
Scooby HuskyandClaude Sonnet 5 bbb80687ca Fix netbird-operator namespace reference in istio mesh policies
allow-netbird-operator-access (AuthorizationPolicy) and the netbird-operator
PeerAuthentication both targeted namespace 'netbird-operator', which has
never existed - the actual netbird-operator pod runs in the 'netbird'
namespace. This permanently blocked istio-mesh-config from going Synced
(kubectl apply failed: namespaces "netbird-operator" not found on every
sync attempt).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 18:20:00 -05:00
Scooby HuskyandClaude Sonnet 5 3226169af8 Fix Vault kubernetes-auth reviewer token expiring 1hr after every boot
vault-init-and-store.sh minted the auth/kubernetes/config reviewer JWT via
'kubectl create token vault-auth' with no --duration, defaulting to a 1hr
TTL, then wrote it into Vault as a static token_reviewer_jwt. ~1hr after
every cluster boot/hook rerun it silently expired, breaking TokenReview
for every kubernetes-auth login (including ESO's) with a generic 403 that
logged nothing at INFO/ERROR. This was the actual root cause of today's
broad ArgoCD Degraded wave across ~14 apps - not Vault's seal state, which
was fine the whole time.

Fix: clear token_reviewer_jwt and rely on Vault's local-JWT auto-detection
(disable_local_ca_jwt=false default), which reads the vault pod's own
kubelet-refreshed SA token from disk on every call instead of a static
copy. The vault SA already has system:auth-delegator via the existing
vault-server-binding ClusterRoleBinding.

Applied live directly against Vault to unblock immediately; confirmed
ClusterSecretStore vault-backend flipped to Ready=True and previously
broken ExternalSecrets (guacamole, teslamate, unpoller, netbird x3) all
resynced successfully.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 18:19:13 -05:00
Scooby Husky 9bb4ada1f5 Fix Redis CR permission denied on PVC (needs fsGroup)
redis-nextcloud-0 was CrashLoopBackOff: "mkdir: can't create directory
'/node-conf': Permission denied" and same for appendonlydir - the pod
had no podSecurityContext, so it couldn't write to its own PVC-backed
data directories. Added fsGroup/runAsUser: 1000 matching the
opstree/redis image's default non-root user. Verified live: Redis now
starts cleanly ("Ready to accept connections").
2026-08-17 17:20:16 -05:00
Scooby Husky a8b16fa4cb Fix Redis CR apiVersion/schema - v1beta2 doesn't exist, blocked whole nextcloud sync
apps/nextcloud/manifests/redis-cr.yaml used redis.redis.opstreelabs.in/v1beta2
(guessed, flagged as unverified when written) - the actually-installed CRD
only has v1beta1. ArgoCD validates all resource API groups/versions before
syncing anything, so this one bad apiVersion blocked the ENTIRE nextcloud
Application from syncing at all (every resource showed OutOfSync, not just
this one).

Also fixed against the live CRD schema: redisSecret belongs nested under
kubernetesConfig, not top-level, and redisExporter.image is required by the
schema even when enabled: false.

Verified live: kubectl apply succeeded and the Redis resource was created
before committing this.
2026-08-17 16:56:20 -05:00
Scooby Husky 7b3669e9a6 Fix CNPG operator -> instance status communication + add CNPG health check
Two related fixes, found while working through why authentik/n8n/gitlab
ArgoCD Applications showed Unknown/stuck health despite pods being fine:

1. ArgoCD has no built-in health check for CNPG's Cluster CRD, so it
   always reported "Unknown" - and since app-level health rolls up to
   the worst resource status, every app with a CNPG Cluster showed
   Unknown/Progressing regardless of actual state. Added a Lua health
   check (resource.customizations.health.postgresql.cnpg.io_Cluster)
   reading .status.conditions[Ready] / .status.phase.

2. Once that started reporting real status instead of masking it,
   pg-authentik showed a genuine problem: CNPG's operator couldn't
   reach its Postgres instances' status API (port 8000) - "Cannot
   extract Pod status ... context deadline exceeded" - because:
   a) authentik-ingress's CiliumNetworkPolicy never allowed cnpg-system
      as a source (crowdsec-ingress already had this exception,
      authentik-ingress was just missing it - inconsistency, not
      deliberate).
   b) Even after fixing (a), still blocked - pg-authentik's pods are
      ambient-mesh-enrolled, so the connection actually goes through
      ztunnel's HBONE tunnel (port 15008) first, same underlying issue
      as the argocd-redis fix from earlier today. Rather than keep
      finding and patching this per-namespace as it recurs, added a
      cluster-wide CiliumClusterwideNetworkPolicy allowing HBONE
      broadly - ztunnel's own mTLS/SPIFFE identity verification is the
      real security boundary for mesh traffic; Cilium blocking the
      tunnel port itself was only breaking legitimate traffic, not
      adding meaningful protection on top.

Verified live: CNPG operator immediately went from failing status
extraction to successfully reconciling (recreating a pod to reattach
its PVC) once both fixes were in place.
2026-08-17 16:41:11 -05:00
Scooby Husky 99b8ca6be9 Fix Nextcloud OIDC client_id/secret - was hardcoded empty
apps/nextcloud/values.yaml had oidc_login_client_id/secret as literal
empty strings, causing "client_id is missing or invalid" on login.
The matching Authentik blueprint (infrastructure/authentik/nextcloud-
blueprint.yaml) already defines fixed credentials for this OAuth2
provider - just never got copied over. Filled in to match.

Also found and left alone: an orphaned nextcloud-oidc-config ConfigMap
(from January) that already had the correct values but isn't actually
mounted by the current Deployment (which uses nextcloud-config instead)
- harmless dead resource, not the active config path.
2026-08-17 16:33:47 -05:00
Scooby Husky 8372887ce1 Fix two long-standing bugs in the vault-init PreSync hook script
Both pre-existing (not from today's work), found while cleaning up a
recurring stale-pod pattern (a fresh vault-init Job failing on every
sync of the vault Application):

1. Init-detection was always broken: `vault status -format=json`
   pretty-prints with a space after colons ("initialized": true), but
   the parser (grep -o '"initialized":[a-z]*' | cut -d: -f2) required
   no space - it always matched empty, so `initialized`/`sealed` were
   always blank regardless of real state. Every run concluded "not
   initialized" and tried to re-init, which correctly fails once Vault
   already has been ("Vault is already initialized"). Switched to jq.
   Separately, the old `|| echo '{}'` fallback also discarded valid
   JSON on vault status's normal non-zero exit codes (2 = sealed) -
   fixed too, though the parsing bug was the actual blocker.

2. `vault_exec` wraps `kubectl exec` without `-i`, so the heredoc piped
   into `vault policy write NAME -` never reached the remote command -
   vault saw an empty policy body. Never previously reached in
   practice since bug #1 always failed the script earlier. Added -i.

Verified end-to-end with a one-off Job run: script now correctly
detects the already-initialized state and completes successfully,
including the policy/role writes that were previously unreachable.
2026-08-17 16:13:48 -05:00
Scooby Husky ee71e4f46f Fix DDNS/failover Cloudflare zone name and add restricted PSS securityContext
- ZONE_NAME was "kube.huskypup.net" in both the home-ip-ddns CronJob and
  the VPS failover watcher - that's a record, not a Cloudflare zone (the
  actual zone is the parent "huskypup.net"). Caused home-ip-ddns to fail
  every run (zone lookup returned zero results, curl -f exit 22) -
  confirmed live and fixed.
- Added seccompProfile/non-root/dropped-capabilities securityContext to
  the three CronJobs added this session that were missing it (flagged by
  the cluster's "restricted" PodSecurity admission). Repointed the raft
  snapshot job's mc binary install from /usr/local/bin to /tmp so it
  still works running as non-root.
2026-08-17 15:55:47 -05:00
Scooby Husky 983671ef50 Fix ArgoCD Redis connectivity broken by ambient mesh HBONE NetworkPolicy gap
Root cause (confirmed via ztunnel logs): the argo-cd chart's default
NetworkPolicies only allow each component's application port (e.g. 6379
for redis), but Istio ambient mode routes ALL pod-to-pod traffic through
ztunnel's HBONE tunnel on port 15008 first - so the tunnel itself was
being blocked even though the "real" port was allowed. Every inter-pod
connection in the argocd namespace hung for exactly 10s then reset;
ztunnel's own log named it directly ("connection timed out, maybe a
NetworkPolicy is blocking HBONE port 15008"). This broke argocd-server's
Redis-backed session/cluster-info caching cluster-wide and was silently
preventing the root Application from picking up new child Applications.

Fix: additive NetworkPolicy allowing ingress on 15008 for all argocd
pods (NetworkPolicies union across multiple policies selecting a pod,
so this doesn't touch/replace the chart's rendered ones - safe across
Helm upgrades).

Also: dropped argocd from the namespace-enrollment job's waypoint list.
argocd's only AuthorizationPolicy (allow-argocd-access) is a plain
source-namespace/IP match with no L7 rules - its own status shows
"attached to ztunnel", not waypoint - so forcing L7 waypoint processing
onto the namespace was unnecessary overhead, not a security requirement.
ztunnel's mTLS still fully covers it. (This was a red herring for the
HBONE bug itself, not the fix, but a valid simplification found along
the way.)
2026-08-17 15:08:33 -05:00
Scooby Husky 7990f1fa47 Add VPS warm-standby/backup site (Phase 0-1b)
Foundation for a DR/backup path using an always-on VPS as a second
ArgoCD-managed cluster, plus DB/backup standardization work that fell
out of it:

- vps-standby ArgoCD cluster destination + AppProject, MinIO backup
  receiver, VPS bootstrap script (k3s, Netbird, cert-manager)
- Dual-site DNS failover watcher + home-IP DDNS CronJob, Cloudflare
  token moved out of git into Vault+ExternalSecret
- Nextcloud migrated from ad-hoc MariaDB to CNPG + redis-operator
  (matches n8n/Authentik/GitLab's backup-native pattern)
- Authentik's CNPG manifests moved into the actual ArgoCD-synced
  manifests/ path (were present but never wired into the sync path)
- Vault raft-snapshot CronJob, CNPG barmanObjectStore backups
  (Authentik/n8n/Nextcloud), Nextcloud file-PVC restic sync - all
  targeting the new VPS MinIO receiver

See VPS Warm-Standby plan doc for full design rationale.
2026-08-17 14:59:26 -05:00
Scooby HuskyandClaude Opus 4.6 5163403e24 Scale down non-critical replicas for homelab resource savings
Reduce replicas to 1 for workloads that don't need HA in a homelab:
- Prometheus 2→1, Alertmanager 2→1 (~4.4GB RAM saved)
- cert-manager 3→1
- GitLab: registry 2→1, kas 2→1, sidekiq 2→1, praefect 2→1,
  pgbouncer-ro 3→1
- Guacamole + guacd 2→1
- Kiali 2→1, ArgoCD server 2→1
- Kyverno background-controller 2→1
- Scylla operator 2→1
- ext-authz-proxy 2→1, netbird-cluster-router 2→1

Kept multi-replica: coredns, envoy-gateway, kyverno admission,
vault, argocd-repo-server, gitlab-webservice, istiod, rook-ceph CSI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 23:00:48 -05:00
Scooby HuskyandClaude Opus 4.6 fa23e5084c Reduce GitLab latency: increase Gitaly CPU limit and Puma workers
Gitaly CPU limit was 250m causing throttling on git operations (485ms
p99 to webservice). Increased to 1 core with 100m request. Also added
workerProcesses: 3 to webservice for better request concurrency
(was default 2).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 22:47:39 -05:00
Scooby HuskyandClaude Opus 4.6 a8b9064248 Set Grafana deployment strategy to Recreate for RWO PVC compatibility
RollingUpdate with a ReadWriteOnce PVC causes a Multi-Attach deadlock
where the new pod can't start because the old pod still holds the
volume. Recreate strategy ensures the old pod is terminated before
the new one is created.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 22:33:16 -05:00
Scooby HuskyandClaude Opus 4.6 ff00f666a0 Add privileged PodSecurity label to 8 namespaces for NetBird router
NetBird router deployments require NET_ADMIN capability, which is
blocked by the default baseline PodSecurity standard. Add
managedNamespaceMetadata with pod-security.kubernetes.io/enforce:
privileged to argocd, frigate, gitlab, grafana, guacamole,
home-assistant, teslamate, and vault ArgoCD Applications.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 22:27:39 -05:00
Scooby HuskyandClaude Opus 4.6 02632d06a4 Fix NetBird service annotations: use netbird.io/expose for operator
The operator watches for netbird.io/expose (presence-based), not
netbird.io/resource which was from the old Helmfile bootstrap script.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 16:25:35 -05:00
Scooby HuskyandClaude Opus 4.6 2edc8a156e Fix netbird-operator: add self-hosted management URL
The operator was missing managementURL, causing it to default to the
NetBird SaaS cloud (api.netbird.io) instead of the self-hosted instance.
This resulted in an infinite reconcile loop and no networks/resources
being created.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 16:08:51 -05:00
Scooby HuskyandClaude Opus 4.6 0c16785116 Fix Guacamole OIDC session persistence via Istio DestinationRule
Traffic goes through Istio (not Envoy Gateway), so cookie-based
consistent hashing must be on the Istio DestinationRule. Removes
the unused Envoy Gateway BackendTrafficPolicy.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 15:11:15 -05:00
Scooby HuskyandClaude Opus 4.6 f71dbdb689 Add Envoy Gateway cookie-based session persistence for Guacamole OIDC
Fixes OIDC nonce validation failures with 2 replicas by using a
BackendTrafficPolicy with cookie-based consistent hashing instead of
Service-level ClientIP affinity (which doesn't work behind Envoy proxy).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 14:55:53 -05:00
Scooby HuskyandClaude Opus 4.6 b9fd5a9461 Add session affinity to Guacamole service for OIDC flow
With 2 replicas, OIDC callbacks can hit a different pod than the
one that generated the state/nonce, causing auth failures. Session
affinity ensures the same client always reaches the same pod.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 14:47:25 -05:00
Scooby HuskyandClaude Opus 4.6 e29eb9611b Make Authentik the primary login method for Guacamole
Change EXTENSION_PRIORITY from "*,openid" to "openid,*" so
unauthenticated users are redirected to Authentik OIDC instead
of seeing the database login form.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 14:24:53 -05:00
Scooby HuskyandClaude Opus 4.6 987f91c877 Enable CephFS CSI driver for SMB gateway and CephFS volumes
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 12:27:00 -05:00
Scooby HuskyandClaude Opus 4.6 7eae427d1a Scale services to 2+ replicas for HA
- ArgoCD: server and repo-server → 2 replicas
- GitLab: webservice, sidekiq, registry, KAS, praefect → min 2 replicas
- Guacamole: client and guacd → 2 replicas
- Kiali: 1 → 2 replicas
- Alertmanager: add 2 replicas
- TeslaMate CNPG: 1 → 2 instances

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 03:08:49 -05:00
Scooby HuskyandClaude Opus 4.6 38b1c16c1e Update rook-ceph-operator chart to v1.19.2 and remove stale comment
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 21:25:22 -05:00
Scooby HuskyandClaude Opus 4.6 5d09fb86e7 Add GitLab Redis initializer ordering fix
Fixes NameError: uninitialized constant Gitlab::Redis::ALL_CLASSES
caused by 7_prometheus_metrics.rb loading before 7_redis.rb
(alphabetically 'p' < 'r'). The ConfigMap provides a 0_redis_early.rb
initializer that requires gitlab/redis before other initializers run.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 20:26:24 -05:00
Scooby HuskyandClaude Opus 4.6 c0fd93c73b Upgrade Ceph to v19.2.3 (Squid)
Includes security fixes (CVE-2023-43040) and RBD/RADOS improvements.
Rook operator upgraded to v1.19.2 via Helm.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:38:10 -05:00
Scooby HuskyandClaude Opus 4.6 a93df9375a Add talos-cp-02 back to Rook-Ceph nodes list
LVM filter now applied via Talos machine config on all nodes to prevent
lvs hang on BIOS partitions. talos-cp-02 can safely run OSD prepare jobs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 16:31:13 -05:00
Scooby HuskyandClaude Opus 4.6 2643d5e0bf Fix TeslaMate DB corruption and Rook-Ceph reconcile hang
- Scale CNPG pg-teslamate to 1 instance: on-disk storage corruption
  (corrupted sequences, system catalog indexes) means replicas cloned
  via pg_basebackup inherit corruption and can't open any database.
  Primary works from cached catalogs. Sequences already rebuilt live.

- Remove talos-cp-02 from Rook-Ceph nodes list: ceph-volume inventory
  hangs on BIOS partition via nsenter lvs in host namespace. Talos
  read-only /etc prevents LVM filter fix. OSD 1 keeps running from
  its existing deployment.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 14:53:31 -05:00
Scooby HuskyandClaude Opus 4.6 8df30c11ee Re-enable Ceph dashboard SSL to fix SAML http/https URL mismatch
X-Forwarded-Proto approach didn't work - Ceph dashboard SAML doesn't
read it. Re-enabling native SSL so the dashboard sees https:// URLs
directly. Updated both Istio and Envoy Gateway backends for TLS with
insecureSkipVerify for the self-signed cert.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 13:56:33 -05:00
Scooby HuskyandClaude Opus 4.6 fb0ed0071f Set X-Forwarded-Proto header for Ceph dashboard SAML authentication
Dashboard runs with ssl: false behind TLS-terminating proxy, so SAML
library sees http:// URLs and rejects responses expecting https://.
Explicitly set X-Forwarded-Proto: https on both Istio VirtualService
and Envoy Gateway HTTPRoute so the SAML ACS URL matches.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 13:48:35 -05:00
Scooby HuskyandClaude Opus 4.6 7c184f9d14 Fix Ceph dashboard DestinationRule: disable TLS to match dashboard config
Dashboard has ssl: false (plain HTTP on port 7000), but the DestinationRule
had tls.mode: SIMPLE which caused Istio to attempt TLS to the backend,
resulting in WRONG_VERSION_NUMBER errors.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 13:41:57 -05:00
Scooby HuskyandClaude Opus 4.6 bc37f9d1a4 Complete OSD migration: remove temp Kyverno policy, fix talos-cp-02 device name
All 4 OSDs successfully recreated in clean raw mode (NVMe only, no DB devices).
Fixed talos-cp-02 NVMe device path from nvme0n1 to nvme1n1 to match actual hardware.
Removed temporary fix-rook-osd-metadata-device Kyverno policy as it's no longer needed.
Operator runs normally without manual patches.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 03:47:14 -05:00
Scooby HuskyandClaude Opus 4.6 1a8ef4347f Restore Ceph dashboard Istio VirtualService on port 7000
The dashboard has SSL disabled and listens on port 7000. Restoring the
VirtualService with the correct port for the Istio gateway path.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 00:19:11 -05:00
Scooby HuskyandClaude Opus 4.6 d27b90ba19 Revert external-dns gateway-httproute source
Services behind the Envoy gateway cannot reach Istio ambient mesh
backends due to STRICT mTLS. Reverting DNS to point back to the Istio
gateway until the migration path is resolved.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 00:11:23 -05:00
Scooby HuskyandClaude Opus 4.6 3244fdf611 Allow envoy-gateway-system ingress to vault and authentik
All namespaces with CiliumNetworkPolicies need envoy-gateway-system
access now that services are migrating from Istio to Envoy gateway.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 23:59:53 -05:00
Scooby HuskyandClaude Opus 4.6 7a9703f3bc Add gateway-httproute source to external-dns
Allows external-dns to create DNS records from Gateway API HTTPRoute
resources, needed for services using the Envoy edge gateway.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 23:50:58 -05:00
Scooby HuskyandClaude Opus 4.6 13ff2dafd6 Remove old Istio VirtualService for Ceph dashboard
The dashboard now uses the Envoy edge gateway via HTTPRoute. The old
Istio VirtualService was causing DNS to point at the wrong gateway IP.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 23:49:32 -05:00
Scooby HuskyandClaude Opus 4.6 59dd4dcda2 Enable directory recursion for Cilium manifests source
Network policies in the network-policies/ subdirectory were not being
synced because ArgoCD was not recursing into subdirectories.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 23:35:09 -05:00
Scooby HuskyandClaude Opus 4.6 c199715276 Allow envoy-gateway-system ingress to rook-ceph for dashboard
The Ceph dashboard moved from Istio to the Envoy edge gateway but the
CiliumNetworkPolicy was still only allowing istio-system, causing 503s.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 23:24:26 -05:00
Scooby HuskyandClaude Opus 4.6 31113ecf16 Fix Ceph dashboard Backend to use HTTP port 7000
The dashboard has SSL disabled and listens on port 7000, but the
Backend was configured for TLS on port 8443.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 22:32:09 -05:00
Scooby HuskyandClaude Opus 4.6 b795f5aa8f Add Kyverno PolicyExceptions for third-party charts
ArgoCD, NetBird operator, and Scylla Manager are upstream Helm charts
that don't set resource limits, probes, drop-all capabilities, or
disable SA token automount. These audit-mode policy violations inflated
the cumulative Prometheus failure counter to ~3800.

- argocd: exempt from require-resource-limits, require-probes
- netbird-operator: exempt from require-resource-limits,
  disallow-automount-sa-token, require-probes
- scylla-manager: exempt from require-drop-all-capabilities,
  disallow-automount-sa-token

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:43:22 -05:00
Scooby HuskyandClaude Opus 4.6 0d39bcddc9 Add ServiceEntry for wildcard host to fix Kiali KIA0104
The global-authentik-authz AuthorizationPolicy uses *.kube.huskypup.net
as a host match, but Kiali flagged KIA0104 because the wildcard had no
matching entry in Istio's service registry. Adding a MESH_EXTERNAL
ServiceEntry with resolution NONE registers the hostname without
affecting routing.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:35:20 -05:00
Scooby HuskyandClaude Opus 4.6 9385a13d5f Add manifests source to scylla-manager ArgoCD Application
The scylla-manager Application was missing the third source for raw
manifests, so the VirtualService in infrastructure/scylla/manifests/
was never deployed. This caused Kiali to flag the ext-authz policy
host as having no matching service entry.

Also removed stale PERMISSIVE PeerAuthentication from authentik
namespace (manually applied, not in git, unnecessary with mesh-wide
STRICT mTLS).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:28:10 -05:00
Scooby HuskyandClaude Opus 4.6 e1960d7071 Fix create-block-db-symlink to use activate-osd volume
Use the OSD-specific activate-osd volume mount instead of rook-data
parent directory. Multiple old OSD dirs exist on each node and the
script was creating the symlink in the wrong (stale LVM) directory.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:25:31 -05:00
Scooby HuskyandClaude Opus 4.6 550ee8fd13 Fix Kyverno OSD policy init container ordering
Use patchesJson6902 with append (/-) for create-block-db-symlink to
ensure it runs after activate, not before. The patchStrategicMerge was
inserting it before activate, causing prime-osd-dir to clear the symlink.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:15:20 -05:00
Scooby HuskyandClaude Opus 4.6 64723eb0ba Allow kube-apiserver webhook calls to netbird namespace
The netbird-operator webhook service runs in the netbird namespace and
receives admission webhook calls from kube-apiserver using node IPs
(172.28.101.0/24). Without this, ztunnel rejects the connections with
"allow policies exist, but none allowed".

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:09:55 -05:00
Scooby HuskyandClaude Opus 4.6 511ba55b30 Remove dedicated DB devices from Rook-Ceph OSD configuration
Remove metadataDevice (sdb) from all 4 nodes in cluster-values.yaml
since the SATA SSDs are slower than the NVMe data drives. Add temporary
Kyverno mutate policy to protect existing OSD deployments when the
operator is re-enabled - replaces expand-bluefs with a no-op and adds
conditional block.db symlink creation based on BlueStore label detection.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 22:06:34 -05:00
Scooby HuskyandClaude Opus 4.6 34a7cf95d0 Add HTTP proxy trust config to Home Assistant init container
Home Assistant was returning 400 errors because it received
X-Forwarded-For headers from the Istio ingress gateway but wasn't
configured to trust reverse proxies. Add use_x_forwarded_for and
trusted_proxies to the init container's configuration.yaml template.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 21:51:24 -05:00
Scooby HuskyandClaude Opus 4.6 f8eea2ed45 Remove nessus from cluster configuration
Remove all nessus references: ArgoCD project destination, Istio ambient
enrollment, waypoint gateway, authorization policies, ext-authz policy,
TLS certificate, ingress gateway host, Kyverno exclusion, and Authentik
forward-auth blueprint.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 21:02:45 -05:00
Scooby HuskyandClaude Opus 4.6 c5e13377bd Fix talos client key PEM header for nginx compatibility
ED25519 PRIVATE KEY header is not recognized by nginx's OpenSSL.
Convert to standard PKCS#8 PRIVATE KEY header (same DER content).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 20:31:31 -05:00