Requesting a Vertex AI Quota Increase
9 min read · updated August 11, 2026
Most quota requests fail for one of three reasons: they name the wrong metric, they are submitted by an identity that lacks the permission to submit them, or they ask for something that is not adjustable. All three are checkable in advance.
Find the metric before anything else
Google Cloud quotas are identified by a service and a metric, and Vertex AI has hundreds. The console surface is IAM & Admin → Quotas & System Limits, filtered by service to aiplatform.googleapis.com. That pane is the volatile part of this page — it gets redesigned — but the filter and the metric names are stable, and the metric name is the thing you actually need.
The names are descriptive once you have seen a few. custom_model_serving_nvidia_l4_gpus governs L4 accelerators for your own deployed models, with parallel metrics per accelerator family. Foundation-model serving uses metrics along the lines of online_prediction_requests_per_base_model, dimensioned by base model and region. Almost every one of them is regional, which means a limit raised in us-central1 does nothing for a deployment in europe-west4 — a surprisingly common way to raise a quota and see no change.
The same information is available without the console. Google’s Cloud Quotas documentation describes a gcloud quotas info command group for listing and describing quotas, and a gcloud quotas preferences group for requesting adjustments, backed by the QuotaInfo and QuotaPreference API resources. Scripting the read side is worthwhile if you manage many projects, because it turns “what are our limits” into a query rather than a click tour.
The permissions to make a request
A quota increase is an IAM-gated operation, and a developer who can deploy models frequently cannot request quota. Google’s documentation names roles/servicemanagement.quotaViewer for viewing quotas and roles/servicemanagement.quotaAdmin for requesting adjustments, with the underlying permissions serviceusage.quotas.update and cloudquotas.quotas.update.
Two consequences follow. First, if the increase button is greyed out or the API returns permission denied, the problem is the identity, not the quota. Second, this is a good thing to grant deliberately rather than by handing out Owner: the ability to raise quota is the ability to raise the ceiling on spend, so it belongs with whoever owns the budget. A budget alert is the natural companion control.
Submitting the increase
- Filter the Quotas page to the service, then to the metric name you identified from the error message.
- Select the specific rows you need. A metric usually has one row per region and sometimes per model; selecting the wrong row raises a limit you were not hitting.
- Enter the new value and the justification, then submit. Google’s documentation notes that adjustment requests are subject to review, and that you receive an acknowledgement email followed by a decision email.
- Group related requests. Google advises grouping requests by product and area, and warns that batching unrelated requests can extend processing time — so one request for four regions of the same metric is better than four requests, and better than one request mixing accelerators with API rates.
There is also an automated path. Google documents a quota adjuster that monitors usage and submits adjustment requests on your behalf, which suits steadily growing workloads and does not suit a project that is about to step change.
What the request should say
The review is at least partly human, and the justification is the only input you control. Public guidance and Google’s own advice converge on the same handful of things: ask for a proportionate increase rather than an enormous one, be specific about the workload, describe expected traffic including its shape rather than only its peak, and mention a launch date if one exists.
Concretely, a request that says “raising L4 serving quota in us-central1 from 1 to 8 to run two replicas of a 9B model with headroom for a blue-green deployment during a rollout on 3 September” gives a reviewer everything they need. “Need more quota” does not. Asking for ten times your current limit is a very different request from asking for a thousand times it, and the second invites scrutiny you do not want on a deadline.
Asking before you need it
A quota request is a review with an unpredictable turnaround, which makes it a terrible thing to discover on the critical path. The way to stop discovering it that way is to alert on quota utilisation rather than on quota errors.
Cloud Monitoring exposes quota usage and quota limit as metrics, so an alert at seventy or eighty percent of a limit fires days before the 429 does and while there is still time for a review. This is worth doing for the small number of quotas that actually bind — typically one accelerator metric per serving region and one request or token metric per base model — rather than for everything, because an alert policy over hundreds of Vertex AI metrics is noise nobody will read.
The other half is treating quota as part of an environment’s definition. A new project created for a new region starts with default limits, so a disaster-recovery plan that fails over to a region nobody has ever deployed to fails over into a project with zero accelerator quota — the plan is untested in the one way that matters. Request the quota in the standby region while it is quiet, and prove it by deploying something small there. Provisioning capacity you hope never to use is cheap; discovering during an incident that the failover region needs a multi-day review is not.
Quotas you cannot raise
- Anything served from a shared pool. Where a model is served from shared capacity rather than a per-project allowance, there is no counter to raise. The answer is retries, or Provisioned Throughput. See the 429 page for how to tell.
- Limits shown as zero for a restricted model. Some models — video generation has been a recurring example on Google’s developer forums — show a quota of zero because access is allowlisted rather than rate-limited. The request that unlocks them is an access request through your account team, not a quota adjustment.
- System limits. Hard platform maxima, such as the largest permitted value for a field, are documented alongside quotas but are not adjustable at all.
- Physical capacity. Quota is permission; capacity is availability. Holding quota for sixteen accelerators does not make them exist in a full zone, and only a reservation converts one into the other.