Prereqs
Beforehelm install, the following must be in place on the cluster side.
Cluster
- GKE 1.30+ (Autopilot or Standard mode both work; Standard gives more control over node pools)
- Workload Identity enabled on the cluster (
--workload-pool=<project>.svc.id.goog) — required for keyless SA binding to GCP IAM - OIDC provider is implicit on GKE when Workload Identity is enabled; no separate step needed
Addons + controllers
GKE Standard clusters create node pools manually; size them to match the workload sizing table below. GKE Autopilot provisions nodes on-demand from pod resource requests — set
resources.requests precisely so Autopilot selects the right machine family.
GCP-managed resources (recommended)
- Cloud SQL for PostgreSQL 16 in the same region as the cluster, with Private IP enabled. Nebula requires
vector,pg_partman, andpg_cron; confirm all three are available in your Cloud SQL version, enable the required database flags forvector/pg_cron, then runnebula-enterprise postgres provisionto create the Nebula database, user, extensions, and chart credential Secret. - GCS bucket in the same region. Grant the Nebula service account
roles/storage.objectAdminon the bucket. - DynamoDB-compatible service for orchestration state, reachable from Nebula pods. Before Helm install, run
nebula-enterprise orchestration dynamodb ensureto create or verify the fourpk/skorchestration tables and writer-authority records; use--endpoint-urlfor non-AWS endpoints and set the matchingNEBULA_ORCHESTRATION_DYNAMODB_*values in the chart.
objectStorage block uses S3-protocol env vars. GCS exposes an S3-compatible XML API at https://storage.googleapis.com. Use HMAC keys (Service Accounts → HMAC keys in the Cloud Console) as the credentialsSecret, and set objectStorage.forcePathStyle: false for the GCS XML API. Alternatively, run a MinIO gateway in front of GCS.
Workload Identity setup
-
Create a GCP service account for Nebula:
-
Bind it to the Kubernetes service account the chart creates:
Replace
<release>with yourhelm installrelease name. -
Grant the GCP service account access to GCS:
-
If ESO uses the same GCP service account for Secret Manager access, also grant
roles/secretmanager.secretAccessoron the secrets. -
Annotate the Kubernetes service account in your values file:
Install
1. Push images to Artifact Registry
2. Seed secrets in GCP Secret Manager
Generate the JWT RSA private key with the commands in Service authentication before creatingNEBULA_JWT_PRIVATE_KEY_PEM.
NEBULA_JWT_RETIRED_PUBLIC_KEYS_JSON can stay [] on a fresh install. Populate it only during JWT signing-key rotation; see Service authentication.
For an empty Cloud SQL instance, use the bundle helper as the canonical logical bootstrap:
username and password keys. Run the read-only verifier before Helm install:
3. Copy + fill the reference values file
The bundle shipshelm/examples/gke/values.yaml with GKE-specific knobs pre-wired (Workload Identity annotation, GCS endpoint, nginx ingress, Secret Manager ESO). Copy it, fill in the <placeholder> markers, and save as your-values.yaml.
4. Install
_common/production-sizing.yaml is the shared production-shape sizing block (replicas, CPU/memory requests + limits, persistence) used by all three cloud-managed K8s examples (EKS/AKS/GKE). Omit it to keep the chart’s minimal-dev defaults; override per-workload in your-values.yaml to fit your GKE node SKUs.
The chart runs schema migrations and catalog-apply automatically via a per-revision Job (<release>-nebula-migrations-<revision>); API and worker pods gate startup on an init container that polls public.nebula_release_contract for the install’s release row. releaseContract.releaseId and releaseContract.gitSha are stamped by bundle.sh and consumed automatically.
5. Verify
Upgrade
Pull the new bundle, push new images to Artifact Registry, then:Sizing reference
Recommended GKE machine types:
n2-standard-4 (4 vCPU / 16 GB) for API, worker, Orchestration; n2-highmem-4 (4 vCPU / 32 GB) for graph-engine and compactor.
Troubleshooting
Workload Identity not bound — pods receive permission denied from GCS
Workload Identity not bound — pods receive permission denied from GCS
Confirm the Kubernetes SA annotation is set:
kubectl -n nebula describe sa <release>-nebula-sa should show iam.gke.io/gcp-service-account. Also verify the IAM binding: gcloud iam service-accounts get-iam-policy nebula-sa@<project>.iam.gserviceaccount.com should list the workloadIdentityUser binding for the K8s SA. Ensure the cluster’s Workload Identity pool (<project>.svc.id.goog) is enabled.GCE Ingress (not nginx) provisioning slow
GCE Ingress (not nginx) provisioning slow
The GCE Ingress controller provisions a Google Cloud Load Balancer which can take 5-10 minutes. Check
kubectl -n nebula describe ingress nebula for events. If you need faster provisioning, switch ingress.className: nginx and install the nginx Ingress controller instead.pgvector missing on Cloud SQL — 'extension vector does not exist'
pgvector missing on Cloud SQL — 'extension vector does not exist'
Cloud SQL for PostgreSQL 16.3+ supports pgvector via the
vector extension. Enable the Cloud SQL flag (cloudsql.enable_pgvector=on), then run nebula-enterprise postgres provision or have your platform workflow satisfy nebula-enterprise postgres verify. Cloud SQL docs: Use pgvector.GCS HMAC credentials rejected by graph-engine
GCS HMAC credentials rejected by graph-engine
Verify the HMAC key is created for a service account (not a user account). HMAC keys for service accounts are under IAM & Admin → Service Accounts → select the account → Keys tab → HMAC keys. Store the Access ID and Secret in the Kubernetes Secret referenced by
objectStorage.credentialsSecret. The Secret must have AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY keys — those exact uppercase names — the chart’s nebula.objectStorageEnv helper reads them via secretKeyRef.key.