Skip to content

Triggering a Cloud Function From a Cloud Storage Upload

10 min read · updated August 11, 2026

A document lands in a bucket and something should read it. The wiring is three pieces — an event type, an IAM grant on a service agent you did not create, and a function that can be called twice for the same object without doing damage.

Which event you actually want

Google documents four Cloud Storage object event types for Eventarc triggers on the Cloud Storage triggers page:

  • google.cloud.storage.object.v1.finalized — an object was created, or an existing object was overwritten. This is the one you want for “a new document arrived”.
  • google.cloud.storage.object.v1.deleted — fires on deletion, including soft deletion.
  • google.cloud.storage.object.v1.archived — a live version became noncurrent, which only happens in a versioned bucket.
  • google.cloud.storage.object.v1.metadataUpdated — object metadata changed with no new data written.

Two things about finalized catch people. It fires on overwrite, not only on first write, so a client that re-uploads the same key re-triggers the whole pipeline. And it fires when the object becomes visible, which for a resumable upload is at the end — you will not get a partial object, but you also get nothing at all until the upload completes, which makes a stalled upload look like a broken trigger.

If you are moving from a 1st gen function, the old event names (google.storage.object.finalize) belong to a different trigger mechanism. The v1.finalized form is the Eventarc/CloudEvents one, and mixing them produces a deploy error rather than a silent mismatch, which is the merciful outcome.

The filter is on the bucket, and only on the bucket. There is no prefix or suffix condition at the trigger level, which means every object written anywhere in that bucket invokes your function, including the outputs your function itself writes. Two things follow. Write outputs to a different bucket, or you have built a loop that pays for itself twice and may not terminate. And accept that the endswith check at the top of the handler is a filter on work, not on invocations — you are still billed for the instance time of every ignored event, which for a bucket receiving a million small files is not nothing. Where the volume is high and the interesting objects are a minority, a dedicated bucket for the trigger is cheaper than a cheap check.

The grant the trigger needs before it exists

This is the step that produces a trigger which deploys cleanly and then never fires. Direct Cloud Storage events reach Eventarc through Pub/Sub, and the identity that publishes them is the Cloud Storage service agent for your project — an account in the gs-project-accounts.iam.gserviceaccount.com domain that you did not create and cannot see in the service-account list by default. Google’s Eventarc documentation requires it to hold roles/pubsub.publisher before the trigger is created.

PROJECT_ID=my-project

SERVICE_AGENT="$(gcloud storage service-agent --project="$PROJECT_ID")"

gcloud projects add-iam-policy-binding "$PROJECT_ID" \
  --member="serviceAccount:${SERVICE_AGENT}" \
  --role="roles/pubsub.publisher"

Use gcloud storage service-agent rather than constructing the address from the project number by hand. It is one command, it is correct, and the hand-built form is the kind of thing that works in one project and silently does not in another.

The function

A CloudEvent handler receives the event envelope; the object’s bucket and name are in the data payload. Note that the payload does not contain the object’s bytes — only its identity and metadata — so the function reads it back from Cloud Storage.

# main.py
import functions_framework
from cloudevents.http.event import CloudEvent
from google.cloud import storage
from google import genai

storage_client = storage.Client()
genai_client = genai.Client(vertexai=True, location="us-central1")


@functions_framework.cloud_event
def on_document(event: CloudEvent) -> None:
    data = event.data
    bucket_name = data["bucket"]
    object_name = data["name"]
    generation = data["generation"]

    if not object_name.endswith(".txt"):
        return

    blob = storage_client.bucket(bucket_name).blob(object_name)
    # Pin the generation: between the event and this read, the object
    # may already have been overwritten by a later upload.
    blob = storage_client.bucket(bucket_name).get_blob(
        object_name, generation=int(generation)
    )
    if blob is None:
        return  # object no longer exists at that generation

    text = blob.download_as_text()

    result = genai_client.models.generate_content(
        model="gemini-2.5-flash",
        contents=f"Extract the parties and the effective date:\n\n{text}",
    )

    out = storage_client.bucket(bucket_name).blob(f"extracted/{object_name}.json")
    # Only write if nothing is there yet — makes a duplicate delivery a no-op.
    out.upload_from_string(result.text, if_generation_match=0)

The generation pin is not paranoia. A finalized event refers to one specific version of an object, and by the time the function reads, a second upload may have replaced it. Without the pin the function processes whatever is there now, which for two rapid uploads means both invocations process the second document and the first is silently skipped.

Deploying the trigger

  1. Enable the APIs: gcloud services enable eventarc.googleapis.com run.googleapis.com cloudfunctions.googleapis.com pubsub.googleapis.com storage.googleapis.com.
  2. Grant the Cloud Storage service agent roles/pubsub.publisher as above. Do this first; the trigger creation can succeed without it and then deliver nothing.
  3. Give the function’s runtime service account read access to the bucket and roles/aiplatform.user on the project:
    gcloud storage buckets add-iam-policy-binding gs://incoming-docs \
      --member="serviceAccount:fn-docs@PROJECT_ID.iam.gserviceaccount.com" \
      --role="roles/storage.objectUser"
  4. Deploy with an event-filtered trigger. The bucket filter is exact match; there is no prefix filtering at the trigger level, so path selection happens in your code:
    gcloud functions deploy on-document \
      --gen2 \
      --runtime=python312 \
      --region=us-central1 \
      --source=. \
      --entry-point=on_document \
      --trigger-event-filters="type=google.cloud.storage.object.v1.finalized" \
      --trigger-event-filters="bucket=incoming-docs" \
      --service-account=fn-docs@PROJECT_ID.iam.gserviceaccount.com \
      --timeout=300s \
      --max-instances=10

The --max-instances is doing real work here. An event trigger has no natural back-pressure: copy ten thousand files into the bucket and ten thousand events arrive. Without a ceiling, that is ten thousand concurrent model calls, which will hit a Vertex AI quota before it hits anything else — see fixing RESOURCE_EXHAUSTED on Vertex AI for what that looks like from the other end.

Delivery is at-least-once

Eventarc delivers through Pub/Sub, and Pub/Sub is at-least-once. Your function will occasionally be called twice for one upload — not often, but often enough that a pipeline running for months will do it. If the function is not idempotent, that means a duplicate row, a duplicate charge, or a second model call billed for nothing.

The cheapest defence is the one in the code above: if_generation_match=0 on the output write, which makes the write fail rather than overwrite if the output already exists. It costs one argument and it turns a duplicate delivery into a caught exception instead of duplicated work. Where the output is not a Cloud Storage object, the equivalent is a key derived from bucket + name + generation, which is unique per version and stable across retries — the CloudEvent id is not, since a redelivery may carry a new one.

Retries are the other half of this and they are off by default. Deploying with --retry means an event whose handler raises is redelivered, which is what you want for a transient model error and exactly what you do not want for a malformed document: a permanently failing event retries on a schedule for as long as the platform’s retry window allows, spending a model call each time and, if it is one of many, holding instances that other events need. The standard defence is to make the handler give up on its own, based on the event’s age rather than on a counter you do not have:

from datetime import datetime, timedelta, timezone

MAX_EVENT_AGE = timedelta(hours=2)

created = datetime.fromisoformat(event["time"].replace("Z", "+00:00"))
if datetime.now(timezone.utc) - created > MAX_EVENT_AGE:
    # Too old to be worth another model call. Record it and return
    # cleanly so the platform stops redelivering.
    log_permanent_failure(bucket_name, object_name)
    return

Returning normally is what acknowledges the event; raising is what asks for another attempt. That single distinction is the whole retry contract, and getting it backwards — swallowing a transient error so it is never retried, or raising on a document that will never parse — is the most common way these pipelines misbehave. Watch the backlog rather than inferring it: Eventarc delivers through a Pub/Sub subscription in your project, and the oldest-unacked-message-age metric on that subscription is the number that tells you whether the function is keeping up.

Verifying end to end

echo "Contract between Acme Ltd and Beta GmbH, effective 2026-03-01." \
  > sample.txt
gcloud storage cp sample.txt gs://incoming-docs/sample.txt

gcloud logging read \
  'resource.type="cloud_run_revision"
   resource.labels.service_name="on-document"' \
  --limit=20 --freshness=10m --format='value(textPayload)'

gcloud storage cat gs://incoming-docs/extracted/sample.txt.json

If nothing appears in the logs at all, the event never reached the function, and the service-agent grant is the first thing to check. If the logs show an invocation that failed on permissions, it is the runtime service account and not the service agent. Those two look identical from the bucket’s side and have nothing to do with each other.