Provisioning a GKE GPU Node Pool With Terraform
10 min read · updated August 11, 2026
A GKE node pool with guest_accelerator { type, count } and nothing else will create nodes that have GPUs attached and no driver installed. Pods requesting nvidia.com/gpu stay pending, and nothing in the Terraform output suggests why.
The driver block is the whole page
Inside node_config.guest_accelerator, the Google provider documents an optional gpu_driver_installation_config block whose required gpu_driver_version field accepts GPU_DRIVER_VERSION_UNSPECIFIED, INSTALLATION_DISABLED, DEFAULT and LATEST. Those four values are the decision.
DEFAULT— GKE installs the default driver for the node’s GPU type and GKE version. This is the right answer for almost everybody.LATEST— the newest driver GKE offers for that combination. Choose it when a CUDA version in your image needs a newer driver than the default supplies; be aware that “latest” is a moving target, so two node pools created a month apart may not match.INSTALLATION_DISABLED— you are installing drivers yourself, typically through the NVIDIA GPU operator or a DaemonSet. A deliberate choice, and a wrong one if made by accident.GPU_DRIVER_VERSION_UNSPECIFIED— do not select this expecting it to mean “default”.
Historically this block was documented as optional while behaving as required, which is recorded in the provider’s issue tracker. Treat it as required regardless of what the schema says: omitting it is how you get the pending-pod symptom above.
The node pool
resource "google_container_node_pool" "gpu" {
name = "gpu-l4"
cluster = google_container_cluster.main.id
location = "us-central1"
node_locations = ["us-central1-a", "us-central1-b"]
autoscaling {
min_node_count = 0
max_node_count = 4
}
management {
auto_repair = true
auto_upgrade = true
}
node_config {
machine_type = "g2-standard-8"
disk_size_gb = 200
disk_type = "pd-balanced"
guest_accelerator {
type = "nvidia-l4"
count = 1
gpu_driver_installation_config {
gpu_driver_version = "DEFAULT"
}
}
oauth_scopes = [
"https://www.googleapis.com/auth/cloud-platform",
]
labels = {
workload = "inference"
}
}
}Two arguments here are chosen rather than copied. node_locations is set explicitly because GPU availability is per-zone, not per-region, and a regional pool that spans a zone with no capacity for your accelerator will fail to scale into it. Pinning the zones you have verified capacity in turns a mysterious autoscaler failure into a predictable one.
disk_size_gb is generous on purpose. Model weights, container layers and CUDA images are large, and the default boot disk is a frequent cause of image pull failures on a node pool that was otherwise correct. It is also worth remembering that disk throughput on smaller persistent disks scales with size — a small disk makes the first pull slow as well as tight.
As with ECS, resolve the accelerator name against current documentation rather than an example. Accelerator type strings are product names — older material still shows nvidia-tesla-k80, which is not hardware you want.
Time-sharing, MPS and partitions
If your inference workload does not saturate a GPU, the provider exposes gpu_sharing_config inside the same guest_accelerator block, with a required gpu_sharing_strategy of either TIME_SHARING or MPS, plus max_shared_clients_per_gpu.
guest_accelerator {
type = "nvidia-l4"
count = 1
gpu_driver_installation_config {
gpu_driver_version = "DEFAULT"
}
gpu_sharing_config {
gpu_sharing_strategy = "TIME_SHARING"
max_shared_clients_per_gpu = 4
}
}The two strategies are not interchangeable. Google describes time sharing as allowing multiple containers time-shared access to a single GPU device, and MPS as letting co-operative multi-process CUDA workloads run concurrently on one device. Time sharing context-switches and gives no memory isolation, so one greedy process can still exhaust GPU memory and take the others down with it. MPS runs them concurrently and is better suited to small, well-behaved, cooperating workloads.
There is also gpu_partition_size, which is documented as the size of partitions to create on the GPU. That is MIG, available only on hardware that supports it, and unlike the sharing strategies it gives genuine memory isolation between partitions. If you need isolation rather than density, that is the one to reach for.
The taint you did not add
GKE applies a taint to GPU nodes so that ordinary workloads do not land on expensive hardware. That is helpful, and it interacts badly with Terraform: the taint exists in the API but not in your configuration, so a subsequent plan can show a perpetual diff trying to remove it.
Declare it in the configuration instead of fighting it, so state and reality agree, and add the matching toleration to your workload:
# In node_config
taint {
key = "nvidia.com/gpu"
value = "present"
effect = "NO_SCHEDULE"
}On the workload side, request the resource and tolerate the taint — requesting nvidia.com/gpu alone is not enough if the taint is present, and this is the second most common reason for a pending pod after the missing driver.
resources:
limits:
nvidia.com/gpu: 1
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoScheduleNote that nvidia.com/gpu is only ever a limit, never a fractional request. Kubernetes extended resources are integers, which is why the sharing strategies above exist at all.
Almost every change replaces the pool
This is the property that makes GPU node pools different to operate from ordinary ones. Most of node_config is immutable in the GKE API — machine type, accelerator configuration, disk size, OAuth scopes and taints among them — so editing any of them destroys the node pool and creates a new one. Every node drains, every pod reschedules, and on scarce accelerator capacity you may not get the hardware back at the moment you need it.
The plan tells you this, and it is worth reading rather than skimming. Terraform prints # google_container_node_pool.gpu must be replaced and annotates the specific attribute with # forces replacement. Find that annotation before applying, because the attribute that triggered it is frequently one you did not mean to change — a disk size someone rounded, or an accelerator string updated to a newer product name.
What is updatable in place is worth knowing for the opposite reason: the autoscaling bounds and the node count change without a replacement, as do the management flags. So resizing the pool is cheap and reshaping it is not.
The usual reflex — add create_before_destroy so the new pool exists before the old one goes — does not work as written, because the new pool would need the same name as the one still present. The provider’s documented escape is that name is optional and Terraform generates a unique name when it is left blank. There is also a name_prefix argument, but it is deprecated, so do not build on it.
resource "google_container_node_pool" "gpu" {
# name intentionally omitted: the provider generates a unique one,
# so a replacement can be created before the old pool is destroyed.
cluster = google_container_cluster.main.id
location = "us-central1"
lifecycle {
create_before_destroy = true
}
# ...
}Be honest about what that buys you. It removes the gap between old and new, at the cost of needing quota and physical capacity for both pools at once — which on GPUs is exactly the resource you are short of. On a large pool the more practical pattern is to create a second pool explicitly as a new resource, cordon and drain the first, then remove it in a later apply. That is three commits instead of one, and it is the difference between a migration you control and one that starts the moment somebody types yes.
One change arrives without any plan at all. With auto_upgrade enabled, GKE replaces nodes on its own schedule with a newer node version and potentially a newer driver. Terraform sees no drift, because the node pool resource is unchanged — but your pods moved and your CUDA userspace may not match what it did yesterday. Manage that with a maintenance window and maintenance exclusions on the cluster rather than by turning auto-upgrade off; nodes that never upgrade eventually fall outside GKE’s supported version skew, which is a worse problem than a scheduled reschedule.
Scaling to zero and what it costs you
min_node_count = 0 is the single biggest cost lever on this resource, and it is not free. A pool at zero must boot a node, pull a CUDA-sized image and install or load the driver before the first pod runs, which is a cold start measured in minutes rather than seconds.
The two honest options are: keep one node warm and pay for it, or accept the cold start and make it visible to whatever is queueing the work. What does not work is scaling to zero and then being surprised — see what an idle GPU node pool actually costs for the arithmetic on which side of that trade you are on, and the EKS equivalent if you are on AWS.