Files
Homelabv4/apps/gitlab/manifests/praefect-cnpg-cluster.yaml
T
Scooby HuskyandClaude Sonnet 5 04fa80ece3 GitLab cross-site replication Phase 1a: Postgres network plumbing
Per the approved plan (jiggly-snacking-iverson.md, 'GitLab True
Cross-Site Replication' section) - replacing the git-mirror CronJob
with real CNPG streaming replication for pg-gitlab and pg-praefect,
reusing the exact pattern already proven for pg-authentik.

This commit is network plumbing only, no DB replication yet:
- Repoint both CNPG clusters' barmanObjectStore backup target from
  in-cluster gitlab-minio-svc to VPS MinIO (also becomes the initial
  seed source for the VPS's replicas in Phase 1c).
- New ha-gitlab-postgres/ha-praefect-postgres NodePort Services
  (bypass CNPG's own -rw Services) on both home and VPS.
- New PeerAuthentication with portLevelMtls PERMISSIVE on the replication
  ports - ambient mesh STRICT default would otherwise reset the
  VPS/witness's connection before the Postgres TLS handshake starts.
- Extended allow-gitlab-access AuthorizationPolicy with a port-scoped
  rule for the same traffic (no source.namespaces rule can match
  external, non-mesh peers).
- CoreDNS rewrites for pg-gitlab.ha.huskypup.net / pg-praefect.ha.huskypup.net
  so home's own pods reach the floating hostname locally instead of
  round-tripping through the WAN forward (no NAT hairpin support).

Still needed before Phase 1b/1c: UniFi WAN port-forwards for external
ports 61442/61443 -> talos-cp-01:32442/32443, source-restricted to the
VPS + witness public IPs (manual, same as the existing 61432 rule).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-22 12:47:03 -05:00

79 lines
2.2 KiB
YAML

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: pg-praefect
namespace: gitlab
spec:
imageName: ghcr.io/cloudnative-pg/postgresql:16
instances: 2
# Resources: keep memory, keep CPU requests low for scheduling.
# CPU limits omitted so Postgres can burst when available.
resources:
requests:
memory: "512Mi"
cpu: "100m"
limits:
memory: "2Gi"
# Spread replicas across different nodes
affinity:
topologyKey: kubernetes.io/hostname
storage:
size: 10Gi
storageClass: rook-ceph-block
primaryUpdateStrategy: unsupervised
# Explicit params added 2026-08-22 (GitLab cross-site replication plan)
# - this cluster ran on bare Postgres defaults before. A VPS replica
# can't replay WAL past a checkpoint recorded with higher settings than
# its own ("recovery aborted because of insufficient parameter
# settings", confirmed live for pg-authentik) - giving both sides an
# identical explicit block up front means nobody can tune one side
# later and silently break replication.
postgresql:
parameters:
max_connections: "200"
shared_buffers: "512MB"
effective_cache_size: "1536MB"
maintenance_work_mem: "128MB"
checkpoint_completion_target: "0.9"
wal_buffers: "16MB"
default_statistics_target: "100"
random_page_cost: "1.1"
effective_io_concurrency: "200"
work_mem: "4MB"
min_wal_size: "1GB"
max_wal_size: "4GB"
max_worker_processes: "4"
bootstrap:
initdb:
database: praefect_production
owner: app
# Backup to VPS MinIO, also the initial-seed source for the VPS's own
# pg-praefect replica (bootstrap.recovery, Phase 1c) - this cluster had
# no backup block at all before. See the matching comment on
# cnpg-cluster.yaml (pg-gitlab) in this same directory.
backup:
barmanObjectStore:
destinationPath: s3://cnpg-backups/pg-praefect
endpointURL: http://100.108.113.41:9000
s3Credentials:
accessKeyId:
name: vps-minio-secret
key: accesskey
secretAccessKey:
name: vps-minio-secret
key: secretkey
wal:
compression: gzip
maxParallel: 2
retentionPolicy: "30d"
monitoring:
enablePodMonitor: true