Root cause of tonight's earlier CrowdSec/CNPG-backup workarounds: no node
in the cluster had any route into the Netbird mesh CIDR (100.108.0.0/16)
for pod-originated traffic. The per-namespace netbird 'router' pods are
inbound-only infrastructure (external peers reaching K8s services); their
own architecture has no reverse path.
Fix, in two parts:
1. infrastructure/netbird/manifests/egress-daemonset.yaml - one netbird
client per node, hostNetwork so its wt0 interface lives in the node's
real network namespace, plus a sidecar that adds a host route sending
100.108.0.0/16 out via it. hostNetwork requires a scoped Kyverno
PolicyException (infrastructure/kyverno/policies/netbird-egress-exception.yaml)
to the disallow-host-namespaces STIG policy - narrowly for this one
DaemonSet by name, not a namespace-wide exclusion.
2. Discovered the route alone wasn't enough for k3s NodePort traffic
(vps-minio:30900): Netbird manages its own nftables ACLs independent of
iptables/Kyverno, and its forward chain (netbird-rt-fwd) only permits
*established* connections through a peer acting as a router - never new
ones, by design, unless a Netbird 'Network Route' policy is explicitly
configured (it isn't, for this VPS). Locally-terminated connections
(tinyproxy) go through a separate, already-permissive ACL chain, which
is why the CrowdSec proxy fix from earlier tonight worked. Replicated
that working pattern for MinIO: minio-forward.service on the VPS host
(systemd, socat) forwards 100.108.113.41:9000 -> MinIO's ClusterIP,
avoiding the NodePort path entirely.
Re-enabled everything that was disabled/suspended earlier tonight because
of this gap, pointed at the new endpoint:
- CrowdSec CAPI/console-enroll (removed DISABLE_ONLINE_API, restored the
VPS proxy env vars)
- n8n/nextcloud/authentik CNPG backup.barmanObjectStore
- vault-raft-snapshot CronJob (unsuspended)
- nextcloud PVC content sync CronJob endpoint
vps-minio.netbird.internal is retired everywhere - it was never actually
resolvable (Netbird has no DNS configured) even before today's routing
fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
LAPI treated a failed CAPI/console-enroll attempt as fatal at startup,
crashing the whole server (and taking agent + bouncer down with it via
their LAPI dependency) on every restart. Root cause investigation tonight
found two separate real issues:
1. Home's public IP gets a 403 from AWS WAF in front of api.crowdsec.net
on any request, unrelated to credentials - ENROLL_KEY is valid and
correctly wired but can't get past this.
2. Attempted routing around it via a VPS-side proxy over Netbird; found no
node in the cluster has any route into the Netbird mesh CIDR for
pod-originated egress traffic - the per-namespace netbird router pods
(gitlab, vault, argocd, etc.) are inbound-only infrastructure. A proper
fix needs either a Talos-native static route (fragile against pod IP
churn) or a deliberate STIG exception for hostNetwork - out of scope
for tonight, tracked separately.
DISABLE_ONLINE_API=true lets LAPI start cleanly with local-only detection
(agent + bouncer functional again). ENROLL_KEY stays set in Vault/values -
flip this back once real egress routing exists.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Confirmed 2026-08-17: home's public IP gets a 403 ForbiddenException from
AWS WAF in front of api.crowdsec.net, even on a bare unauthenticated test
request - the VPS's IP gets a normal 401 on the identical request. This is
an IP-level block unrelated to credentials, which is why LAPI kept hitting
403 on both anonymous auto-registration and console enroll with a fresh,
valid enrollment key.
Fix: HTTPS_PROXY/HTTP_PROXY pointed at a tinyproxy instance on the VPS
(bound to its Netbird interface only, not internet-reachable), reachable
from home over the existing Netbird mesh via netbird-cluster-router.
cscli/crowdsec's Go HTTP client honors standard proxy env vars.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>