Skip to content

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

☁️ Google Cloud GKE Cluster Terraform Module

Provisions a single Google Kubernetes Engine cluster (google_container_cluster) with remove_default_node_pool = true, plus one or more separately managed node pools (google_container_node_pool, for_each-managed). Targets hashicorp/google ~> 7.0, Terraform >= 1.12.0.

Terraform Provider Version Type Resources Posture


🧩 Overview

  • ☸️ Provisions one GKE cluster (google_container_cluster.this) with remove_default_node_pool = true + initial_node_count = 1 hardcoded — every GKE cluster manages its node pools separately, never via the cluster's own inline node_pool argument.
  • 🧵 Manages any number of node pools (google_container_node_pool.node_pools) via for_each over var.node_pools, keyed by node pool name.
  • 🔒 Private-by-default networking: enable_private_nodes = true, and master_authorized_networks_config is always rendered with an empty cidr_blocks list — the control-plane API is locked down to automatically whitelisted node IPs until the caller explicitly opens a CIDR.
  • 🪪 Workload Identity, GKE Network Policy (Calico), Shielded GKE Nodes, and a broad logging/monitoring component set are all enabled by default.
  • 🔑 Optional CMEK for both database_encryption (etcd secrets) and each node pool's boot disk — accepted, never defaulted to a specific key.
  • 🌐 VPC-native (alias IP) networking only — ip_allocation_policy requires named pod/service secondary ranges provisioned by terraform-google-vpc-network; this module does not support routes-based clusters.

💡 Why it matters: GKE is our primary Kubernetes compute platform. This module encodes the house-mandated separately-managed-node-pool pattern once, so every consuming composition gets a private, Workload-Identity-enabled, Shielded-Nodes cluster by default — without every caller having to re-derive the same ~15 secure-default decisions from scratch.


❤️ Support this project

If these Terraform modules have been helpful to you or your organization, I'd appreciate your support in any of the following ways:

Whether it's a star, a professional connection, or a coffee, every gesture helps keep these modules actively maintained and continually improving. Thank you for being part of the community!


🗺️ Where this fits

This module sits downstream of our networking, IAM, and security foundation modules. It is expected to be a leaf/terminal module in the initial catalog — no sibling module currently consumes its outputs, and the diagram says so explicitly rather than inventing a downstream consumer.

flowchart LR
 vpcnet["terraform-google-vpc-network"]:::neutral
 svcacct["terraform-google-service-account"]:::neutral
 kms["terraform-google-kms-keyring<br/>(optional)"]:::neutral
 gke["terraform-google-gke-cluster"]:::thismodule
 keystone["google_container_cluster.this"]:::keystone
 leaf["No confirmed downstream consumer yet<br/>(leaf module in current catalog)"]:::neutral

 vpcnet -- "network + subnetwork self_link,<br/>pod/services secondary range names" --> gke
 svcacct -- "node pool service account email" --> gke
 kms -. "CMEK crypto key id (optional)".-> gke
 gke --> keystone
 gke -.-> leaf

 classDef thismodule fill:#4285F4,color:#ffffff,stroke:#174EA6;
 classDef keystone fill:#174EA6,color:#ffffff,stroke:#174EA6;
 classDef neutral fill:#E8EAED,color:#202124,stroke:#9AA0A6;
Loading

Validated via the Mermaid Chart MCP (validate_and_render_mermaid_diagram) before embedding.


🧬 What this builds

flowchart TB
 subgraph module["terraform-google-gke-cluster"]
 direction TB
 cluster["google_container_cluster.this<br/>(keystone)"]:::keystone
 np["google_container_node_pool.node_pools[key]<br/>(for_each over var.node_pools)"]:::child
 cluster --> np
 end

 classDef keystone fill:#174EA6,color:#ffffff,stroke:#174EA6;
 classDef child fill:#4285F4,color:#ffffff,stroke:#174EA6;
Loading

Validated via the Mermaid Chart MCP before embedding.

Resource inventory (2 resource types):

Resource Count Role
google_container_cluster.this 1 Keystone GKE cluster — control plane, cluster-wide networking/security config
google_container_node_pool.node_pools for_each over var.node_pools Separately managed node pools, each with its own sizing/autoscaling/node_config

✅ Provider / Versions

Requirement Value
Terraform >= 1.12.0
hashicorp/google ~> 7.0
Provider block None — the caller's root module configures google (ADC, WIF, or a service-account key per our authentication model)

Schema notes that bite:

  • Inline node_pool {} replacement warning. Node pools defined inside a cluster resource can't be changed, added, or removed after cluster creation without deleting and recreating the entire cluster — confirmed in the live docs. This module never uses that argument; every node pool is a separate google_container_node_pool.
  • enable_private_nodes + master_authorized_networks_config lockout trap. With private nodes enabled (the default) and an empty cidr_blocks list (also this module's default), an operator without a bastion/VPN/Cloud Interconnect path into the VPC has no way to reach kubectl against the control plane. terraform plan/apply succeed regardless.
  • enable_private_endpoint compounds the lockout risk. It disables the public endpoint entirely and only applies when enable_private_nodes = true. Never defaulted to true — always an explicit caller opt-in.
  • Workload Identity's Kubernetes-side annotation gap. workload_identity_config.workload_pool only attaches the pool at the cluster level. Binding a specific Kubernetes ServiceAccount to a google_service_account additionally requires an iam.gke.io/gcp-service-account annotation set directly on the Kubernetes object — entirely outside this module's (and Terraform's GKE provider's) reach.
  • Closed logging_config/monitoring_config enum lists — confirmed against the live v7.39.0 docs. monitoring_config does not accept WORKLOADS in the GA provider (beta-only, deprecated/removed as of GKE 1.24) even though logging_config does.
  • node_config.image_type recreates nodes, not the cluster or the pool resource. A rolling node replacement — a materially smaller blast radius than the inline-node_pool warning above; do not conflate the two.
  • Asymmetric timeouts. Cluster: create/read/update/delete, 90 min default each (unusually includes read). Node pool: create/update/delete only, 60 min default each, no read.
  • Two-step deletion_protection apply. The GCP API itself rejects a destroy while deletion_protection = true (the default) — the caller must apply with it explicitly flipped to false first, distinct from a Terraform-only lifecycle.prevent_destroy.
  • initial_node_count is force-new on google_container_node_pool. Confirmed via the live docs. This module only renders it when autoscaling is set; changing it after the autoscaler has taken over destroys and recreates the whole pool.

🔑 Required IAM Roles

  • roles/container.admin — full lifecycle management (create/update/delete) of GKE clusters and node pools on the target project. This is the only role confirmed for this module; do not grant a broader project-level role than this without a documented reason.

(Sourced directly from SCOPE.md — not re-derived.)


☁️ GCP Prerequisites

  • container.googleapis.com API enabled on the target project (via terraform-google-project-services — this module does not enable it itself).
  • Quota: regional in-use IP address quota (each node consumes a private RFC 1918 address) and Compute Engine CPU quota sized to node_count × machine_type across all node pools in the target region — failures surface at apply time, not plan time.
  • Org policy: this module's secure default is enable_private_nodes = true. If a caller opts out (enable_private_nodes = false), confirm constraints/compute.vmExternalIpAccess is not set to deny at the org/folder level — otherwise node creation fails at apply with an org-policy-violation error terraform validate/plan cannot catch in advance.

(Sourced directly from SCOPE.md — not re-derived.)


📁 Module Structure

terraform-google-gke-cluster/
├── providers.tf # required_providers (hashicorp/google ~> 7.0) + required_version — no provider {} block
├── variables.tf # cluster identity/networking/security nested blocks + node_pools map(object(...))
├── main.tf # google_container_cluster.this + google_container_node_pool.node_pools (for_each)
├── outputs.tf # id, self_link, name, endpoint, node_pool_ids, node_pool_instance_group_urls
├── README.md # this file
├── SCOPE.md # cross-module contract
└── examples/ # runnable example(s) matching the Quick Start below

⚙️ Quick Start

# Caller's root module configures the google provider (ADC, WIF, or a service-account key) —
# this module never declares project/region/zone/credentials variables.

module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-prod-cluster"
  location   = "us-east1"
  network    = "https://www.googleapis.com/compute/v1/projects/casey-prod/global/networks/casey-prod-networking"
  subnetwork = "https://www.googleapis.com/compute/v1/projects/casey-prod/regions/us-east1/subnetworks/app-subnet-use1"

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  private_cluster_config = {
    master_ipv4_cidr_block = "172.16.0.0/28"
  }

  master_authorized_networks_config = {
    cidr_blocks = [
      { cidr_block = "10.50.0.0/24", display_name = "corp VPN range" }
    ]
  }

  workload_identity_config = {
    workload_pool = "casey-prod.svc.id.goog"
  }

  node_pools = {
    "default-pool" = {
      node_count = 3
      node_config = {
        service_account = "gke-nodes@casey-prod.iam.gserviceaccount.com"
      }
    }
  }
}

🔌 Cross-Module Contract

Consumes

Input Type Source module
Network identity (network) string (name or self_link) terraform-google-vpc-network (self_link output)
Subnetwork identity (subnetwork) string (name or self_link) terraform-google-vpc-network (subnetwork_self_links map output)
Pod/services secondary range names (ip_allocation_policy.*) string terraform-google-vpc-network (subnetwork secondary_ranges keys)
Node pool service account email (node_config.service_account, per pool) string terraform-google-service-account (email output)
CMEK crypto key id (database_encryption.key_name, node_config.boot_disk_kms_key) — optional string terraform-google-kms-keyring (crypto key id output)
Workload Identity pool string (workload_identity_config.workload_pool) string (PROJECT_ID.svc.id.goog) Caller's own composition (e.g. data "google_project") — never a sibling module output

Emits

Output Description Consumed by
id Cluster resource id (projects/{{project}}/locations/{{location}}/clusters/{{name}} form) None identified yet — leaf/terminal module in the initial catalog
self_link Cluster self_link (full GCP API URL form) None identified yet
name Cluster name None identified yet
endpoint Control-plane IP address (kubeconfig construction) CI/CD pipelines; no sibling module yet
node_pool_ids Map of node pool key → id None identified yet
node_pool_instance_group_urls Map of node pool key → list of instance_group_urls Potential future firewall-target/load-balancer-backend module

📚 Example Library

1 · Minimal cluster, single fixed-size node pool
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-dev-cluster"
  location   = "us-east1-b"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-dev.svc.id.goog"
  }

  node_pools = {
    "default-pool" = {
      node_count = 3
      node_config = {
        service_account = module.node_service_account.email
      }
    }
  }
}

ℹ️ No autoscaling block on the pool, so node_count = 3 is required and rendered directly — a fixed-size pool, resizable in place.

2 · Regional cluster with a per-zone autoscaling node pool
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-prod-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-prod.svc.id.goog"
  }

  node_pools = {
    "general-pool" = {
      initial_node_count = 1
      autoscaling = {
        min_node_count = 1
        max_node_count = 5
      }
      node_config = {
        service_account = module.node_service_account.email
      }
    }
  }
}

⚠️ initial_node_count is force-new — set it once and never change it again after the autoscaler has taken over, or the entire pool is destroyed and recreated.

3 · Autoscaling node pool with total (cluster-wide) limits
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-burst-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-prod.svc.id.goog"
  }

  node_pools = {
    "burst-pool" = {
      autoscaling = {
        total_min_node_count = 0
        total_max_node_count = 20
        location_policy      = "ANY"
      }
      node_config = {
        service_account = module.node_service_account.email
      }
    }
  }
}

ℹ️ The total pair and the per-zone pair are mutually exclusive — this module's variables.tf rejects a pool that sets both or neither. location_policy = "ANY" prioritizes unused reservations/Spot capacity over balanced zone spread.

4 · Multiple node pools — general-purpose + GPU
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-mixed-fleet-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-prod.svc.id.goog"
  }

  node_pools = {
    "general-pool" = {
      node_count = 3
      node_config = {
        machine_type    = "e2-standard-4"
        service_account = module.node_service_account.email
      }
    }
    "gpu-pool" = {
      node_count = 2
      node_config = {
        machine_type    = "n1-standard-8"
        service_account = module.node_service_account.email
        guest_accelerators = [
          { type = "nvidia-tesla-t4", count = 1 }
        ]
      }
    }
  }
}

🔒 guest_accelerators requires on_host_maintenance = "TERMINATE" at the GCP API level for GPU-attached nodes — this is a live API constraint invisible to terraform validate; an incompatible combination fails only at apply.

5 · Preemptible node pool for cost-optimized batch workloads
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-batch-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-prod.svc.id.goog"
  }

  node_pools = {
    "batch-preemptible" = {
      autoscaling = {
        min_node_count = 0
        max_node_count = 10
      }
      node_config = {
        preemptible     = true
        machine_type    = "e2-standard-8"
        service_account = module.node_service_account.email
      }
    }
  }
}

⚠️ Preemptible/Spot nodes can be reclaimed by GCP at any time with short notice — only schedule fault-tolerant, restartable workloads here (use node taints + tolerations to keep other workloads off this pool).

6 · Spot VM node pool (the modern preemptible alternative)
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-batch-spot-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-prod.svc.id.goog"
  }

  node_pools = {
    "batch-spot" = {
      autoscaling = {
        min_node_count = 0
        max_node_count = 10
      }
      node_config = {
        spot            = true
        machine_type    = "e2-standard-8"
        service_account = module.node_service_account.email
      }
    }
  }
}

ℹ️ spot is the newer, more flexible successor to preemptible (no fixed 24-hour lifetime cap). Do not set both preemptible and spot on the same node pool.

7 · CMEK for etcd secrets (database_encryption)
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-regulated-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-regulated.svc.id.goog"
  }

  database_encryption = {
    key_name = module.kms_keyring.crypto_key_ids["gke-etcd-key"]
  }

  node_pools = {
    "default-pool" = {
      node_count = 3
      node_config = {
        service_account = module.node_service_account.email
      }
    }
  }
}

🔒 The GKE service agent needs roles/cloudkms.cryptoKeyEncrypterDecrypter on the target key — grant this via terraform-google-kms-keyring's own IAM surface or terraform-google-project-iam-bindings before this module's apply.

8 · CMEK for a node pool's boot disk
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-encrypted-node-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-regulated.svc.id.goog"
  }

  node_pools = {
    "encrypted-pool" = {
      node_count = 3
      node_config = {
        service_account   = module.node_service_account.email
        boot_disk_kms_key = module.kms_keyring.crypto_key_ids["gke-node-boot-key"]
      }
    }
  }
}

ℹ️ boot_disk_kms_key is independent of database_encryption — the former encrypts each node's boot disk; the latter encrypts etcd secrets at the control-plane level. Set either, both, or neither.

9 · Opting out of private nodes (public cluster)
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-public-demo-cluster"
  location   = "us-east1-b"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  private_cluster_config = {
    enable_private_nodes = false
  }

  workload_identity_config = {
    workload_pool = "casey-sandbox.svc.id.goog"
  }

  node_pools = {
    "default-pool" = {
      node_count = 2
      node_config = {
        service_account = module.node_service_account.email
      }
    }
  }
}

⚠️ When enable_private_nodes = false, main.tf omits the entire private_cluster_config block (per the live provider docs' own recommendation) — nodes receive public IPs. Confirm constraints/compute.vmExternalIpAccess is not set to deny at the org/folder level before relying on this.

10 · Opening operator access through master_authorized_networks_config
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-prod-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-prod.svc.id.goog"
  }

  master_authorized_networks_config = {
    cidr_blocks = [
      { cidr_block = "10.50.0.0/24", display_name = "corp VPN range" },
      { cidr_block = "203.0.113.4/32", display_name = "bastion host" }
    ]
  }

  node_pools = {
    "default-pool" = {
      node_count = 3
      node_config = {
        service_account = module.node_service_account.email
      }
    }
  }
}

🔒 Without at least one cidr_blocks entry, only automatically whitelisted node IPs can reach the control-plane API — an operator with no VPN/bastion/Cloud Interconnect path into the VPC is locked out of kubectl, even though plan/apply succeed.

11 · Disabling Workload Identity explicitly
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-no-workload-identity-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    enabled = false
  }

  node_pools = {
    "default-pool" = {
      node_count = 3
      node_config = {
        service_account = module.node_service_account.email
      }
    }
  }
}

⚠️ An explicit, documented opt-out — this module suite defaults Workload Identity to enabled. Disabling it means Kubernetes workloads cannot federate to Google service accounts without static key material, which our key-material policy otherwise avoids.

12 · Blue/green node pool upgrades
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-blue-green-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-prod.svc.id.goog"
  }

  node_pools = {
    "default-pool" = {
      node_count = 6
      upgrade_settings = {
        strategy = "BLUE_GREEN"
        blue_green_settings = {
          node_pool_soak_duration = "3600s"
          standard_rollout_policy = {
            batch_percentage    = 0.5
            batch_soak_duration = "600s"
          }
        }
      }
      node_config = {
        service_account = module.node_service_account.email
      }
    }
  }
}

ℹ️ blue_green_settings is required whenever strategy = "BLUE_GREEN" — this module's variables.tf validates that combination.

13 · Custom logging/monitoring component sets
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-trimmed-observability-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-prod.svc.id.goog"
  }

  logging_config = {
    enable_components = ["SYSTEM_COMPONENTS"]
  }

  monitoring_config = {
    enable_components = ["SYSTEM_COMPONENTS", "APISERVER", "SCHEDULER", "CONTROLLER_MANAGER"]
  }

  node_pools = {
    "default-pool" = {
      node_count = 3
      node_config = {
        service_account = module.node_service_account.email
      }
    }
  }
}

⚠️ Trims the module's broader default component set — document the observability trade-off in the calling composition (fewer signals shipped to Cloud Logging/Monitoring, lower cost).

14 · Tainted node pool for a dedicated workload class
module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name       = "casey-gpu-dedicated-cluster"
  location   = "us-east1"
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["app-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  workload_identity_config = {
    workload_pool = "casey-prod.svc.id.goog"
  }

  node_pools = {
    "gpu-dedicated" = {
      node_count = 2
      node_config = {
        machine_type    = "n1-standard-8"
        service_account = module.node_service_account.email
        guest_accelerators = [
          { type = "nvidia-tesla-t4", count = 1 }
        ]
        taints = [
          { key = "nvidia.com/gpu", value = "present", effect = "NO_SCHEDULE" }
        ]
      }
    }
  }
}

💡 Pair this taint with a matching Kubernetes toleration on GPU-workload pod specs so non-GPU workloads are never scheduled onto (and don't waste) this pool's capacity.

15 · 🏗️ End-to-end composition

Wires terraform-google-vpc-network's network/subnetwork/secondary-range outputs, terraform-google-service-account's node pool identity, and terraform-google-kms-keyring's CMEK key (optional) into this module — reflecting the Consumes relationships documented in SCOPE.md. This module's own outputs (id/self_link/endpoint/node_pool_ids) end the composition here, since SCOPE.md documents no confirmed downstream consumer yet.

module "vpc_network" {
  source = "git::https://github.com/microsoftexpert/terraform-google-vpc-network.git?ref=v1.0.0"

  network_name = "casey-prod-networking"

  subnetworks = {
    "gke-subnet-use1" = {
      ip_cidr_range = "10.30.0.0/20"
      region        = "us-east1"
      secondary_ranges = {
        "pods"     = "10.31.0.0/16"
        "services" = "10.32.0.0/20"
      }
    }
  }
}

module "node_service_account" {
  source = "git::https://github.com/microsoftexpert/terraform-google-service-account.git?ref=v1.0.0"

  account_id   = "gke-nodes"
  display_name = "GKE node pool identity"
}

module "kms_keyring" {
  source = "git::https://github.com/microsoftexpert/terraform-google-kms-keyring.git?ref=v1.0.0"

  key_ring_name = "casey-prod-gke-keyring"
  location      = "us-east1"

  crypto_keys = {
    "gke-etcd-key" = {}
  }
}

module "gke_cluster" {
  source = "git::https://github.com/microsoftexpert/terraform-google-gke-cluster.git?ref=v1.0.0"

  name     = "casey-prod-cluster"
  location = "us-east1"

  # Consumes terraform-google-vpc-network's Emits: self_link, subnetwork_self_links.
  network    = module.vpc_network.self_link
  subnetwork = module.vpc_network.subnetwork_self_links["gke-subnet-use1"]

  ip_allocation_policy = {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  private_cluster_config = {
    master_ipv4_cidr_block = "172.16.0.0/28"
  }

  master_authorized_networks_config = {
    cidr_blocks = [
      { cidr_block = "10.50.0.0/24", display_name = "corp VPN range" }
    ]
  }

  workload_identity_config = {
    workload_pool = "casey-prod.svc.id.goog"
  }

  # Optional CMEK — consumes terraform-google-kms-keyring's crypto key id.
  database_encryption = {
    key_name = module.kms_keyring.crypto_key_ids["gke-etcd-key"]
  }

  node_pools = {
    "general-pool" = {
      autoscaling = {
        min_node_count = 1
        max_node_count = 5
      }
      node_config = {
        machine_type = "e2-standard-4"
        # Consumes terraform-google-service-account's Emits: email.
        service_account = module.node_service_account.email
      }
    }
  }
}

💡 module.gke_cluster.endpoint is the value a CI/CD pipeline would use to build a kubeconfig; no sibling module currently consumes any of this module's outputs (leaf/terminal in the current catalog).


📥 Inputs

Variable Type Default Notes
name string — (required) Force-new; 1-40 chars, RFC1035-style
location string — (required) Force-new; zone or region
description string null
network string — (required) Name or self_link — source from terraform-google-vpc-network
subnetwork string — (required) Name or self_link — source from terraform-google-vpc-network
ip_allocation_policy object({...}) — (required) VPC-native secondary range names
private_cluster_config object({...}) {} (enable_private_nodes = true) Omitted entirely when enable_private_nodes = false
master_authorized_networks_config object({...}) {} (cidr_blocks = []) Always rendered — empty list is the secure default
workload_identity_config object({...}) {} (enabled = true) workload_pool required when enabled
network_policy object({...}) {} (enabled = true, provider = "CALICO")
logging_config object({...}) {} (SYSTEM_COMPONENTS + WORKLOADS) Closed enum
monitoring_config object({...}) {} (broad multi-signal set) Closed enum; no WORKLOADS
database_encryption object({...}) null CMEK — never defaulted to a key
enable_shielded_nodes bool true
enable_legacy_abac bool false
enable_intranode_visibility bool true
release_channel object({channel = string}) {} ("REGULAR") Closed enum
node_locations list(string) []
deletion_protection bool true Two-step apply to destroy
node_pools map(object({...})) — (required, ≥1 entry) Keyed by node pool name — see full schema below
labels map(string) {} GCP label-format validated
timeouts object({...}) null Cluster-only (4 keys, includes read)
Full ip_allocation_policy object schema
variable "ip_allocation_policy" {
  type = object({
    cluster_secondary_range_name  = string
    services_secondary_range_name = string
    stack_type                    = optional(string, "IPV4")
  })
}
Full private_cluster_config object schema
variable "private_cluster_config" {
  type = object({
    enable_private_nodes        = optional(bool, true)
    enable_private_endpoint     = optional(bool, false)
    master_ipv4_cidr_block      = optional(string)
    private_endpoint_subnetwork = optional(string)
    master_global_access_config = optional(object({
      enabled = optional(bool, false)
    }))
  })
  default = {}
}
Full master_authorized_networks_config object schema
variable "master_authorized_networks_config" {
  type = object({
    gcp_public_cidrs_access_enabled      = optional(bool)
    private_endpoint_enforcement_enabled = optional(bool)
    cidr_blocks = optional(list(object({
      cidr_block   = string
      display_name = optional(string)
    })), [])
  })
  default = {}
}
Full workload_identity_config / network_policy object schemas
variable "workload_identity_config" {
  type = object({
    enabled       = optional(bool, true)
    workload_pool = optional(string)
  })
  default = {}
}

variable "network_policy" {
  type = object({
    enabled  = optional(bool, true)
    provider = optional(string, "CALICO")
  })
  default = {}
}
Full logging_config / monitoring_config / database_encryption object schemas
variable "logging_config" {
  type = object({
    enable_components = optional(list(string), ["SYSTEM_COMPONENTS", "WORKLOADS"])
  })
  default = {}
}

variable "monitoring_config" {
  type = object({
    enable_components = optional(list(string), [
      "SYSTEM_COMPONENTS", "APISERVER", "SCHEDULER", "CONTROLLER_MANAGER",
      "STORAGE", "HPA", "POD", "DAEMONSET", "DEPLOYMENT", "STATEFULSET"
    ])
  })
  default = {}
}

variable "database_encryption" {
  type = object({
    state    = optional(string, "ENCRYPTED")
    key_name = string
  })
  default = null
}
Full node_pools object schema
variable "node_pools" {
  type = map(object({
    node_locations     = optional(list(string), [])
    initial_node_count = optional(number, 1)
    node_count         = optional(number)
    max_pods_per_node  = optional(number)

    timeouts = optional(object({
      create = optional(string)
      update = optional(string)
      delete = optional(string)
    }))

    autoscaling = optional(object({
      min_node_count       = optional(number)
      max_node_count       = optional(number)
      total_min_node_count = optional(number)
      total_max_node_count = optional(number)
      location_policy      = optional(string, "BALANCED")
    }))

    management = optional(object({
      auto_repair  = optional(bool, true)
      auto_upgrade = optional(bool, true)
    }), {})

    network_config = optional(object({
      create_pod_range    = optional(bool, false)
      pod_range           = optional(string)
      pod_ipv4_cidr_block = optional(string)
    }))

    upgrade_settings = optional(object({
      strategy        = optional(string, "SURGE")
      max_surge       = optional(number, 1)
      max_unavailable = optional(number, 0)
      blue_green_settings = optional(object({
        node_pool_soak_duration = optional(string)
        standard_rollout_policy = optional(object({
          batch_percentage    = optional(number)
          batch_node_count    = optional(number)
          batch_soak_duration = optional(string)
        }))
      }))
    }), {})

    node_config = object({
      machine_type      = optional(string, "e2-medium")
      disk_size_gb      = optional(number, 100)
      disk_type         = optional(string, "pd-balanced")
      image_type        = optional(string, "COS_CONTAINERD")
      labels            = optional(map(string), {})
      metadata          = optional(map(string), {})
      tags              = optional(list(string), [])
      oauth_scopes      = optional(list(string), ["https://www.googleapis.com/auth/cloud-platform"])
      service_account   = string
      preemptible       = optional(bool, false)
      spot              = optional(bool, false)
      min_cpu_platform  = optional(string)
      local_ssd_count   = optional(number, 0)
      boot_disk_kms_key = optional(string)

      shielded_instance_config = optional(object({
        enable_secure_boot          = optional(bool, true)
        enable_integrity_monitoring = optional(bool, true)
      }), {})

      workload_metadata_config = optional(object({
        mode = optional(string, "GKE_METADATA")
      }), {})

      taints = optional(list(object({
        key    = string
        value  = string
        effect = string
      })), [])

      guest_accelerators = optional(list(object({
        type               = string
        count              = number
        gpu_partition_size = optional(string)
      })), [])
    })
  }))
}

🧾 Outputs

Output Description Sensitive
id Cluster Terraform-internal resource id No
self_link Cluster self_link (full GCP API URL form) No
name Cluster name No
endpoint Control-plane IP address No — see Architecture Notes
node_pool_ids Map of node pool key → id No
node_pool_instance_group_urls Map of node pool key → list of instance_group_urls No

🧠 Architecture Notes

  • endpoint is deliberately non-sensitive. It is a network endpoint, not a secret; marking it sensitive would suppress it from terraform output and CI logs where operators legitimately need it to build a kubeconfig. This is a documented judgment call (see SCOPE.md's Emits table), not an oversight.
  • remove_default_node_pool and initial_node_count = 1 are hardcoded, not variables. This is a structural invariant of the composite design (every GKE cluster manages node pools separately) — there is no legitimate reason for a caller to want a different value.
  • master_authorized_networks_config is always rendered. An empty cidr_blocks list IS the secure default; this module never omits the block to "simplify" the render. Combined with enable_private_nodes = true (also default), an operator has no path to kubectl without an explicit cidr_blocks entry or a bastion/VPN/Cloud Interconnect route into the VPC — see Troubleshooting.
  • private_cluster_config is omitted entirely when enable_private_nodes = false, per the live provider docs' own recommendation that its validation is unreliable otherwise.
  • workload_metadata_config.mode defaults to GKE_METADATA on every node pool. This is required for the cluster-level workload_identity_config to be usable by workloads scheduled on that specific pool — a node pool still running GCE_METADATA/MODE_UNSPECIFIED cannot bind a Kubernetes ServiceAccount to a Google service account, even with Workload Identity enabled cluster-wide.
  • Node pool sizing is driven by presence/absence of autoscaling. When set, initial_node_count is rendered (force-new on subsequent change); when absent, node_count is rendered instead (resizable in place). The two are never rendered together, matching the live provider docs' own guidance.
  • for_each key stability. Node pool map keys become the GCP resource name. Renaming a key is a destroy/recreate of that specific pool (Terraform sees a removed + an added resource), not an in-place rename.
  • IAM propagation delay. If terraform-google-service-account/terraform-google-project-iam-bindings grant a role to the node pool's service account in the same composition run immediately before this module creates the node pool, the grant can take up to ~60 seconds to propagate — a transient permission-denied error at apply is an operational timing issue, not a bug in this module's graph ordering.
  • Scope for v1.0.0 is deliberately narrower than the full combined schema of both resources (one of the largest in the provider). See variables.tf's header comment for the complete list of modeled vs. deliberately excluded nested blocks — every exclusion is additive.

🧱 Design Principles

Concern Secure default Opt-out (explicit)
Node pool management remove_default_node_pool = true + separately managed google_container_node_pool for_each, hardcoded N/A — never the cluster's inline node_pool argument
Private cluster networking enable_private_nodes = true Caller sets false — omits the block entirely
Private endpoint enable_private_endpoint = false Caller sets true explicitly (compounds lockout risk)
Control-plane API access master_authorized_networks_config always rendered, empty cidr_blocks by default Caller adds explicit cidr_blocks entries
Workload Identity enabled = true; workload_pool required when enabled Caller sets enabled = false
Network Policy enabled = true, provider = "CALICO" Caller sets enabled = false
Node pool metadata mode workload_metadata_config.mode = "GKE_METADATA" per pool Caller overrides per pool (breaks Workload Identity on that pool)
Shielded GKE Nodes enable_shielded_nodes = true (cluster); shielded_instance_config secure boot + integrity monitoring true (per pool) Caller disables per-flag
Legacy ABAC enable_legacy_abac = false Caller sets true (not recommended)
Logging/monitoring components Broad recommended sets, not the provider's bare-minimum default Caller trims the component list explicitly
Deletion protection deletion_protection = true (cluster) Caller sets false, two-step apply required
Release channel channel = "REGULAR" Caller sets RAPID/STABLE/UNSPECIFIED
CMEK (database_encryption, boot_disk_kms_key) Accepted, never defaulted to a specific key Caller supplies a key from terraform-google-kms-keyring

🚀 Runbook

cd C:\GitHubCode\newgooglecloudmodules\terraform-google-gke-cluster
terraform init -backend=false
terraform validate
terraform fmt -check

Pin ?ref=v1.0.0 when consuming this module — never a branch. This library is plan-only; a human applies from CI with valid Workload Identity Federation or ADC credentials.


🧪 Testing

  • terraform init -backend=false, terraform validate, and terraform fmt -check are the entire offline proof gate for this module — all three pass cleanly as of this authoring session.
  • validate/fmt confirm internal type/reference consistency and canonical formatting only. Neither can catch GCP API-level rejections (quota, org policy, IAM propagation, an incompatible GPU/scheduling combination, an org policy blocking public node IPs) — those surface only at apply time, against a real project, from a human-run CI pipeline with valid credentials.
  • The examples/ directory exists so a consuming GitHub Actions workflow can run a real terraform plan against a real project as part of that pipeline's own review gate; this library only guarantees the example is syntactically and structurally sound in isolation.

💬 Example Output

$ terraform output

id = "projects/casey-prod/locations/us-east1/clusters/casey-prod-cluster"
name = "casey-prod-cluster"
self_link = "https://container.googleapis.com/v1/projects/casey-prod/locations/us-east1/clusters/casey-prod-cluster"
endpoint = "34.71.128.12"
node_pool_ids = {
 "general-pool" = "casey-prod/us-east1/casey-prod-cluster/general-pool"
}
node_pool_instance_group_urls = {
 "general-pool" = [
 "https://www.googleapis.com/compute/v1/projects/casey-prod/zones/us-east1-b/instanceGroupManagers/gke-casey-prod-cluster-general-pool-abcd1234-grp",
 ]
}

🔍 Troubleshooting

Symptom Cause Fix
kubectl times out reaching the cluster after a clean apply Private-cluster lockout: enable_private_nodes = true (default) with an empty master_authorized_networks_config.cidr_blocks (also default) and no VPN/bastion/Cloud Interconnect path into the VPC Add an explicit cidr_blocks entry for your operator range, or apply from a host inside the VPC (bastion/Cloud Shell with IAP tunnel)
terraform plan shows a full node pool destroy/recreate you didn't expect initial_node_count was changed after the autoscaler took over (force-new), or the node_pools map key was renamed Never touch initial_node_count post-creation; treat map keys as a stable long-term contract
Workload Identity federation silently fails for pods on one specific node pool That pool's node_config.workload_metadata_config.mode is not GKE_METADATA (overridden away from the default) Restore the default, or explicitly set mode = "GKE_METADATA" on that pool
Error: Invalid value for "database_encryption.state" (or similar) at plan time Unsupported enum value supplied for a closed-list field Check the field's description in variables.tf for the exact confirmed value set
Node pool creation fails at apply with a quota or org-policy error despite a clean plan This library is plan-only — quota/org-policy constraints are invisible to validate/plan Confirm CPU/IP quota and any org policy (e.g. constraints/compute.vmExternalIpAccess when opting into public nodes) before apply
Node pool apply fails with a transient permission-denied error for the node service account IAM propagation lag — the service account's role grant is only seconds old Retry after ~60 seconds, or add an explicit wait/dependency in the composing pipeline
GPU node pool apply fails with a scheduling-compatibility error guest_accelerators set without a compatible on_host_maintenance = "TERMINATE" scheduling policy at the underlying instance level This is a live GCP API constraint outside this module's node_config scope for v1.0.0 — confirm the machine type/accelerator combination is supported before applying

🔗 Related Docs


💙 "Infrastructure as Code should be standardized, consistent, and secure."

Releases

Packages

Contributors

Languages