返回 Skills
google/skills· Apache-2.0 内容可用

cloud-run-basics

>-

安装

与 skills.sh 相同的 Command / Prompt 安装方式


name: cloud-run-basics metadata: category: Serverless description: >- Manages Cloud Run services, jobs, and worker pools. Use when you need to deploy applications responding to HTTP requests (services), run event-triggered or scheduled tasks (jobs), or handle always-on pull-based background processing (worker pools).

Cloud Run Basics

Cloud Run is a fully managed application platform for running your code, function, or container on top of Google's highly scalable infrastructure. It abstracts away infrastructure management, providing three primary resource types:

  1. Services: Responds to HTTP requests sent to a unique and stable endpoint, using stateless instances that autoscale based on a variety of key metrics, also responds to events and functions.
  2. Jobs: Executes parallelizable tasks that are executed manually, or on a schedule, and run to completion.
  3. Worker pools: Handles always-on background workloads such as pull-based workloads, for example, Kafka consumers, Pub/Sub pull queues, or RabbitMQ consumers.

Prerequisites

  1. Enable the Cloud Run Admin API and Cloud Build APIs:

    gcloud services enable run.googleapis.com cloudbuild.googleapis.com --quiet
    
  2. If you are under a domain restriction organization policy restricting unauthenticated invocations for your project, you will need to access your deployed service as described under Testing private services.

Required roles

You need the following roles to deploy your Cloud Run resource:

  • Cloud Run Admin (roles/run.admin) on the project
  • Cloud Run Source Developer (roles/run.sourceDeveloper) on the project
  • Service Account User (roles/iam.serviceAccountUser) on the service identity
  • Logs Viewer (roles/logging.viewer) on the project

Cloud Build automatically uses the Compute Engine default service account as the default Cloud Build service account to build your source code and Cloud Run resource, unless you override this behavior.

For Cloud Build to build your sources, grant the Cloud Build service account the Cloud Run Builder (roles/run.builder) role on your project:

gcloud projects add-iam-policy-binding PROJECT_ID \
    --member=serviceAccount:SERVICE_ACCOUNT_EMAIL_ADDRESS \
    --role=roles/run.builder \
    --quiet

Replace PROJECT_ID with your Google Cloud project ID and SERVICE_ACCOUNT_EMAIL_ADDRESS with the email address of the Cloud Build service account.

Deploy a Cloud Run service

You can deploy your service to Cloud Run by using a container image or deploy directly from source code using a single Google Cloud CLI command.

CRITICAL RULE: Any deployed code MUST listen on 0.0.0.0 (not 127.0.0.1) and use the injected $PORT environment variable (defaults to 8080), or it will crash on boot.

Deploy a container image to Cloud Run

Cloud Run imports your container image during deployment. Cloud Run keeps this copy of the container image as long as it is used by a serving revision. Container images are not pulled from their container repository when a new Cloud Run instance is started.

Supported container images

You can directly use container images stored in the Artifact Registry, or Docker Hub. Google recommends the use of Artifact Registry since Docker Hub images are cached for up to one hour.

You can use container images from other public or private registries (like JFrog Artifactory, Nexus, or GitHub Container Registry), by setting up an Artifact Registry remote repository.

You should only consider Docker Hub for deploying popular container images such as Docker Official Images or Docker Sponsored OSS images. For higher availability, Google recommends deploying these Docker Hub images using an Artifact Registry remote repository.

To deploy a container image, run the following command:

    gcloud run deploy SERVICE_NAME \
        --image IMAGE_URL \
        --region us-central1 \
        --allow-unauthenticated \
        --quiet

Replace the following:

  • SERVICE_NAME: the name of the service you want to deploy to. Service names must be 49 characters or less and must be unique per region and project. If the service does not exist yet, this command creates the service during the deployment. You can omit this parameter entirely, but you will be prompted for the service name if you omit it.
  • IMAGE_URL: a reference to the container image, for example, us-docker.pkg.dev/cloudrun/container/hello:latest. If you use Artifact Registry, the repository REPO_NAME must already be created. The URL follows the format of LOCATION-docker.pkg.dev/PROJECT_ID/REPO_NAME/PATH:TAG. Note that if you don't supply the --image flag, the deploy command will attempt to deploy from source code.

Deploy from source code

There are two different ways to deploy your service from source:

  • Deploy from source with build (default): This option uses Google Cloud's buildpacks and Cloud Build to automatically build container images from your source code without having to install Docker on your machine or set up buildpacks or Cloud Build. By default, Cloud Run uses the default machine type provided by Cloud Build.

    • To deploy from source with automatic base image updates enabled, run the following command:

      gcloud run deploy SERVICE_NAME --source . \
      --base-image BASE_IMAGE \
      --automatic-updates \
      --quiet
      

      Cloud Run only supports automatic base images that use Google Cloud's buildpacks base images.

      • To deploy from source using a Dockerfile, run the following command:
       gcloud run deploy SERVICE_NAME --source . --quiet
      
      When you provide a Dockerfile, Cloud Build runs it in the cloud, and
      deploys the service.
      
  • Deploy from source without build (Preview): This option deploys artifacts directly to Cloud Run, bypassing the Cloud Build step. This allows for rapid deployment times. To deploy from source without build, run the following command:

    gcloud beta run deploy SERVICE_NAME \
     --source APPLICATION_PATH \
     --no-build \
     --base-image=BASE_IMAGE \
     --command=COMMAND \
     --args=ARG \
     --quiet
    

    Replace the following:

    • SERVICE_NAME: the name of your Cloud Run service.
    • APPLICATION_PATH: the location of your application on the local file system.
    • BASE_IMAGE: the runtime base image you want to use for your application. For example, us-central1-docker.pkg.dev/serverless-runtimes/google-24-full/runtimes/nodejs24. You can also deploy a pre-compiled binary without configuring additional language-specific runtime components using the OS only base image, such as osonly24.
    • COMMAND: the command that the container starts up with.
    • ARG: an argument you send to the container command. If you use multiple arguments, specify each on its own line.

    For examples on deploying from source without build, see Examples of deploying from source without build.

Create and execute a Cloud Run job

To create a new job, run the following command:

gcloud run jobs create JOB_NAME --image IMAGE_URL OPTIONS --quiet

Alternatively, use the deploy command:

gcloud run jobs deploy JOB_NAME --image IMAGE_URL OPTIONS --quiet

Replace the following:

  • JOB_NAME: the name of the job you want to create. If you omit this parameter, you will be prompted for the job name when you run the command.

  • IMAGE_URL: a reference to the container image—for example, us-docker.pkg.dev/cloudrun/container/job:latest.

  • Optionally, replace OPTIONS with any of the following flags:

    • --tasks: Accepts integers greater or equal to 1. Defaults to 1; maximum is 10,000. Each task is provided the environment variables CLOUD_RUN_TASK_INDEX with a value between 0 and the number of tasks minus 1, along with CLOUD_RUN_TASK_COUNT, which is the number of tasks.
    • --max-retries: The number of times a failed task is retried. Once any task fails beyond this limit, the entire job is marked as failed. For example, if set to 1, a failed task will be retried once, for a total of two attempts. The default is 3. Accepts integers from 0 to 10.
    • --task-timeout: Accepts a duration like "2s". Defaults to 10 minutes; maximum is 168 hours (7 days). For tasks using GPUs, the maximum available timeout is 1 hour.
    • --parallelism: The maximum number of tasks that can execute in parallel. By default, tasks will be started as quickly as possible in parallel.
    • --execute-now: If set, immediately after the job is created, a job execution is started. Equivalent to calling gcloud run jobs create followed by gcloud run jobs execute.

    In addition to these preceding options, you also specify more configuration such as environment variables or memory limits.

For a full list of available options when creating a job, refer to the gcloud run jobs create command line documentation.

Wait for the job creation to finish. You'll see a success message upon a successful completion.

To execute an existing job, run the following command:

gcloud run jobs execute JOB_NAME --quiet

If you want the command to wait until the execution completes, run the following command:

gcloud run jobs execute JOB_NAME --wait --region=REGION --quiet

Replace the following:

  • JOB_NAME: the name of the job.
  • REGION: the region in which the resource can be found. For example, europe-west1. Alternatively, set the run/region property.

Deploy a worker pool

You can deploy a Cloud Run worker pool using container images or deploy directly from the source.

Deploy a container image

You can specify a container image with a tag (for example, us-docker.pkg.dev/my-project/container/my-image:latest) or with an exact digest (for example, us-docker.pkg.dev/my-project/container/my-image@sha256:41f34ab970ee...).

Supported container images

You can directly use container images stored in the Artifact Registry, or Docker Hub. Google recommends the use of Artifact Registry since Docker Hub images are cached for up to one hour.

You can use container images from other public or private registries (like JFrog Artifactory, Nexus, or GitHub Container Registry), by setting up an Artifact Registry remote repository.

You should only consider Docker Hub for deploying popular container images such as Docker Official Images or Docker Sponsored OSS images. For higher availability, Google recommends deploying these Docker Hub images using an Artifact Registry remote repository.

To deploy a container image, run the following command:

gcloud run worker-pools deploy WORKER_POOL_NAME --image IMAGE_URL --quiet

Replace the following:

  • WORKER_POOL_NAME: the name of the worker pool you want to deploy to. If the worker pool does not exist yet, this command creates the worker pool during the deployment. You can omit this parameter entirely, but you will be prompted for the worker pool name if you omit it.

  • IMAGE_URL: a reference to the container image that contains the worker pool, such as us-docker.pkg.dev/cloudrun/container/worker-pool:latest. Note that if you don't supply the --image flag, the deploy command attempts to deploy from source code.

Wait for the deployment to finish. Upon successful completion, Cloud Run displays a success message along with the revision information about the deployed worker pool.

Deploy a worker pool from source

You can deploy a new worker pool or worker pool revision to Cloud Run directly from source code using a single gcloud CLI command, gcloud run worker-pools deploy with the --source flag.

The deploy command defaults to source deployment if you don't supply the --image or --source flags.

Behind the scenes, this command uses Google Cloud's buildpacks and Cloud Build to automatically build container images from your source code without having to install Docker on your machine or set up buildpacks or Cloud Build. By default, Cloud Run uses the default machine type provided by Cloud Build.

To deploy a worker pool from source, run the following command:

gcloud run worker-pools deploy WORKER_POOL_NAME --source . --quiet

Replace WORKER_POOL_NAME with the name you want for your worker pool.

What to do if a deployment fails:

  1. IAM/Permission Error: Read iam-security.md.
  2. Crash on Boot / Healthcheck failed: Fetch the logs immediately using gcloud logging read "resource.labels.service_name=SERVICE_NAME" --limit=20 to find the exact runtime error.
  3. Native Dependency Error (Node/Python): If using --no-build, switch to --source . (Buildpacks) to compile native extensions properly for Linux.

Reference Directory

  • Core Concepts: Services vs. Jobs vs. Worker pools, resource model, and auto-scaling behavior for services.

  • CLI Usage: Essential gcloud run commands for deployment and management.

  • Client Libraries: Using Google Cloud client libraries to interact with Cloud Run.

  • MCP Usage: Using the Cloud Run remote MCP server.

  • Infrastructure as Code: Terraform examples for services, jobs, worker pools, and IAM bindings.

  • IAM & Security: Roles, service identities, and ingress/egress controls.

  • Networking Best Practices & Cost Optimization: Cost optimization strategies, Direct VPC egress, IP address and port exhaustion strategies, performance throughput tuning, and MTU settings.

If you need product information not found in these references, use the Developer Knowledge MCP server search_documents tool.

附带文件

references/cli-usage.md
# Cloud Run CLI

Use the `gcloud run` command to manage your Cloud Run applications.

## Basic Syntax

```bash
gcloud run [GROUP] [COMMAND] [FLAGS]
```

## Essential Commands

### Cloud Run service

-   **Deploy a service from an image:**

    ```bash
    gcloud run deploy my-service \
        --image us-docker.pkg.dev/cloudrun/container/hello:latest \
        --quiet
    ```

-   **Deploy from source code:**

    ```bash
    gcloud run deploy my-service --source . --quiet
    ```

-   **Deploy a Cloud Run function:** 

    ```bash
    gcloud run deploy my-service
    --source . --function example-hello --base-image go126 --region us-central1 --quiet
    ```

-   **List services:**

    ```bash
    gcloud run services list --quiet
    ```

-   **Update traffic split:**

    ```bash
    gcloud run services update-traffic my-service --to-revisions=REV1=50,REV2=50 --quiet
    ```

### Cloud Run job

-   **Create a job:**

    ```bash
    gcloud run jobs create my-job \
      --image us-docker.pkg.dev/cloudrun/container/job:latest \
      --quiet
    ```

-   **Execute a job:**

    ```bash
    gcloud run jobs execute my-job --quiet
    ```

-   **List jobs:** `gcloud run jobs list`

-   **List job executions:**

    ```bash
    gcloud run executions list --job my-job
    ```

### Cloud Run worker pools

-   **Deploy a worker pool from an image:**

    ```bash
    gcloud run worker-pools deploy my-workerpool \
      --image us-docker.pkg.dev/cloudrun/container/worker-pool:latest \
      --quiet
    ```

-   **Deploy from source code:**

    ```bash
    gcloud run worker-pools deploy my-workerpool --source . --quiet
    ```

-   **List worker pools:**

    ```bash
    gcloud run worker-pools list --region us-central1 --quiet
    ```

-   **Configure scaling (manual):**

    ```bash
    gcloud run worker-pools deploy my-workerpool --instances=5 \
      --image us-docker.pkg.dev/cloudrun/container/worker-pool:latest \
      --quiet
    ```

### Configuration and logs

-   **View more details about a service:** `gcloud run services describe my-service`

-   **View logs:**

    ```bash
    gcloud logging read "resource.type=cloud_run_revision AND \
      resource.labels.service_name=my-service" \
      --quiet
    ```

## Common Flags

-   `--region`: The region where the service or job is located.
-   `--allow-unauthenticated`: Makes the service publicly accessible.

-   `--no-allow-unauthenticated`: Restricts access to authenticated users only.
references/client-library-usage.md
# Cloud Run Client Libraries

Google Cloud client libraries provide an idiomatic way to manage Cloud Run
resources programmatically.

## Getting Started

Ensure you have the Google Cloud SDK installed and authenticated.
[Install Google Cloud SDK](https://cloud.google.com/sdk/docs/install)

### Python

- **Installation:**

  ```bash
  pip install --upgrade google-cloud-run
  ```

- **Usage Example:**

  ```python
  from google.cloud import run_v2
  client = run_v2.ServicesClient()
  request = run_v2.ListServicesRequest(
    parent="projects/my-project/locations/us-central1"
  )
  page_result = client.list_services(request=request)
  ```

- [Python Reference](https://docs.cloud.google.com/python/docs/reference/run/latest.md.txt)

### Java

- **Maven Dependency:**

  ```xml
  <dependencyManagement>
  <dependencies>
   <dependency>
      <groupId>com.google.cloud</groupId>
      <artifactId>libraries-bom</artifactId>
      <version>26.79.0</version>
      <type>pom</type>
      <scope>import</scope>
   </dependency>
  </dependencies>
  </dependencyManagement>
  <dependency>
    <groupId>com.google.cloud</groupId>
    <artifactId>google-cloud-run</artifactId>
  </dependency>
  ```

- **Usage Example:**

  ```java
  try (ServicesClient servicesClient = ServicesClient.create()) {
    ListServicesRequest request = ListServicesRequest.newBuilder()
        .setParent(LocationName.of("my-project", "us-central1").toString())
        .build();
    for (Service element : servicesClient.listServices(request).iterateAll()) {
      System.out.println(element.getName());
    }
  }
  ```

- [Java Reference](https://docs.cloud.google.com/java/docs/reference/google-cloud-run/latest/overview.md.txt)

### Node.js (TypeScript)

- **Installation:**

  ```bash
  npm install @google-cloud/run
  ```

- **Usage Example:**

  ```typescript
  import {ServicesClient} from '@google-cloud/run';
  const client = new ServicesClient();
  const [services] = await client.listServices({
    parent: 'projects/my-project/locations/us-central1',
  });
  ```

- [Node.js Reference](https://googleapis.dev/nodejs/run/latest/index.html)

### Go

- **Installation:**

  ```bash
  go get cloud.google.com/go/run/apiv2
  ```

- **Usage Example:**

  ```go
  package main

  import (
  	"context"
  	"fmt"
  	"log" // Import the log package

  	run "cloud.google.com/go/run/apiv2"
  	runpb "cloud.google.com/go/run/apiv2/runpb"
  	"google.golang.org/api/iterator"
  )

  func main() {
  	ctx := context.Background()
  	client, err := run.NewServicesClient(ctx)
  	if err != nil {
  		// Log the error and exit if the client can't be created
  		log.Fatalf("Failed to create Cloud Run Services client: %v", err)
  	}
  	defer client.Close()

  	req := &runpb.ListServicesRequest{
  		Parent: "projects/my-project/locations/us-central1", // Remember to replace my-project
  	}
  	it := client.ListServices(ctx, req)

  	fmt.Println("Cloud Run Services:")
  	for {
  		resp, err := it.Next()
  		if err == iterator.Done {
  			break // Finished iterating successfully
  		}
  		if err != nil {
  			// Log the error and exit if iteration fails
  			log.Fatalf("Error iterating services: %v", err)
  		}
  		fmt.Println(resp.GetName())
  	}
  }
  ```

- [Go Reference](https://docs.cloud.google.com/go/docs/reference/cloud.google.com/go/run/latest)

## Source Code Samples

For more examples across languages, visit the
[Cloud Run Code Samples](https://cloud.google.com/run/docs/samples).
references/core-concepts.md
# Cloud Run core concepts

Cloud Run is a fully managed application platform for running your code,
function, or container on top of Google's highly scalable infrastructure. On
Cloud Run, your code can run as a service, job, or worker pool. All of these
resource types are running sandboxed container instances in the same execution
environment and can integrate with Google Cloud services.

## Services vs. Jobs vs. Worker pools

-   **Cloud Run services:** Used for code that handles requests or events (e.g.,
    web apps, APIs). They provide an HTTPS endpoint and automatically scale
    based on traffic.

-   **Cloud Run jobs:** Used for code that performs a specific task and then
    exits (e.g., data processing, database migrations). They can run a single
    task or an array of parallel tasks.

-   **Cloud Run worker pools:** Designed for continuous, non-HTTP, pull-based
    background processing (e.g., Kafka consumers).

## Resource model

Cloud Run organizes resources as follows:

1.  **Service** The top-level resource. You can deploy a service from a
    container, repository, or source code.
                    1.  **Revision:** An immutable snapshot of a service's
                        configuration and container image. Each service
                        deployment creates a new revision.
                    1.  **Service instances:** The running container that
                        processes requests. Each service revision receiving
                        requests is automatically scaled to the number of
                        instances needed to handle all these requests.
                    1.  **Cloud Run functions**: Deploy functions as Cloud Run
                        services. You can deploy single-purpose functions that
                        respond to events emitted from your cloud infrastructure
                        and services

1.  **Job**: Executes one or more containers to completion. A job consists of
    one or multiple independent tasks that are executed in parallel in a given
    job execution.

1.  **Worker pools**: If your code processes workloads from an external source
    but not from an HTTP request, such as pulling work from a message queue, you
    can deploy it to a Cloud Run worker pool .

## Autoscaling for Cloud Run services

Cloud Run services scale automatically based on:

-   **Request concurrency:** The number of concurrent requests per instance.
-   **CPU utilization:**: The average CPU utilization of existing instances over
    a one minute window.
-   **Scale to zero:** Cloud Run autoscales from one to zero instances only
    after verifying that an instance is no longer processing requests. If you
    use instance-based billing, Cloud Run instances are charged for the entire
    lifecycle of instances, even when there are no incoming requests.

## Container contract 

Your container image can run code written in the programming language
of your choice and use any base image, provided that it respects the
constraints listed in the [Container runtime contract](https://docs.cloud.google.com/run/docs/container-contract.md.txt).

Executables in the container image must be compiled for
Linux 64-bit. Cloud Run specifically supports the Linux x86_64 ABI format.

Cloud Run accepts container images in the Docker Imag
 Manifest V2, Schema 1, Schema 2, and OCI image formats. Cloud Run
also accepts Zstd compressed container images.

If deploying a multi-architecture image, the manifest list must include
linux/amd64.

For functions deployed with Cloud Run, you can use one of the
Cloud Run runtime base images that are published by Google
Cloud's buildpacks to receive automatic security and maintenance updates.
For more information about the supported runtimes, see the [Runtime support schedule](https://docs.cloud.google.com/run/docs/runtime-support.md.txt).

### Container requirements

When deploying containers to Cloud Run, the following requirements must be met:

* Container deployed to services must listen for requests on the correct port
* A Cloud Run service starts Cloud Run instances to handle incoming
  requests. A Cloud Run instance always has one single ingress
  container that listens for requests, and optionally one or more
  sidecar containers. The following port configuration details
  apply only to the ingress container, not to sidecars.
* The ingress container within an instance must listen for
  requests on `0.0.0.0` on the port to which requests are sent. Notably,
  the ingress container should not listen on `127.0.0.1`. By default, request
  are sent to 8080, but you can configure Cloud Run to send requests to the port of your choice.
  Cloud Run injects the PORT environment variable into the ingress container.

## VPC network connectivity

Cloud Run services and jobs support Direct VPC egress. This means
that they can send traffic to private resources within your
configured VPC network, such as databases or internal services. Cloud Run
services and jobs don't support Direct VPC ingress.
Cloud Run worker pools support both Direct VPC egress and Direct VPC
ingress. When you configure Direct VPC for your Cloud Run worker pool
deployment, each worker instance receives a private IP address on the
configured network and subnet. Only resources from your VPC network can
connect to the worker pool private IP address endpoint. For more information
about obtaining the private IP addresses of your worker pool instance, see
[Retrieve the private IP addresses using the metadata server (MDS)](https://docs.cloud.google.com/run/docs/configuring/vpc-direct-vpc.md.txt).

For Cloud Run worker pools with Direct VPC ingress, such as database
connections or any other custom TCP-based protocol, the container must
listen for TCP connections on the port exposed in your container image
through the Dockerfile or specified by the PORT environment variable.

## AI and GPU support

Cloud Run supports hosting AI inference models. You can configure services with
GPUs (e.g., NVIDIA RTX PRO 6000 Blackwell GPU, NVIDIA L4) to accelerate
workloads like LLM inference using Gemma 3. For more information, see GPU
support for
[services](https://docs.cloud.google.com/run/docs/configuring/services/gpu.md.txt),
[jobs](https://docs.cloud.google.com/run/docs/configuring/jobs/gpu.md.txt), and [worker
pools](https://docs.cloud.google.com/run/docs/configuring/workerpools/gpu.md.txt).

## Pricing

Cloud Run uses a pay-as-you-go model:

-   **Request-based:** Charged for resources used during request processing.
-   **Instance-based:** Charged for the entire lifetime of an instance.

For the latest pricing, visit: [Cloud Run
pricing](https://cloud.google.com/run/pricing).
references/iac-usage.md
# Cloud Run Infrastructure as Code

Cloud Run resources can be provisioned and managed using Terraform and other IaC
tools.

## Terraform

The Google Cloud Terraform provider supports Cloud Run services, jobs, and worker pools.

### Cloud Run service example

```terraform
resource "google_cloud_run_v2_service" "default" {
  name     = "cloudrun-service"
  location = "us-central1"
  deletion_protection = false
  ingress  = "INGRESS_TRAFFIC_ALL"

  template {
    containers {
      image = "us-docker.pkg.dev/cloudrun/container/hello"
    }
  }
}

resource "google_cloud_run_v2_service_iam_member" "noauth" {
  location = google_cloud_run_v2_service.default.location
  name     = google_cloud_run_v2_service.default.name
  role     = "roles/run.invoker"
  member   = "allUsers"
}
```

### Cloud Run job example

```terraform
resource "google_cloud_run_v2_job" "default" {
  name     = "cloudrun-job"
  location = "us-central1"

  template {
    template {
      containers {
        image = "us-docker.pkg.dev/cloudrun/container/job"
      }
    }
  }
}
```

### Cloud Run worker pool example

```terraform
resource "google_cloud_run_v2_worker_pool" "default" {
  name     = "cloudrun-workerpool"
  location = "us-central1"

  template {
    containers {
      image = "us-docker.pkg.dev/cloudrun/container/worker-pool:latest"
    }
  }
}
```

### Reference documentation

- [Terraform Google Provider - Cloud Run v2 Service](https://registry.terraform.io/providers/hashicorp/google/latest/docs/resources/cloud_run_v2_service)

- [Terraform Google Provider - Cloud Run v2 Job](https://registry.terraform.io/providers/hashicorp/google/latest/docs/resources/cloud_run_v2_job)
- [Terraform Google Provider - Cloud Run v2 Worker pool](https://registry.terraform.io/providers/hashicorp/google/latest/docs/resources/cloud_run_v2_worker_pool)

## YAML

Cloud Run resources can also be defined using YAML. For more information, see
[Cloud Run YAML reference](https://docs.cloud.google.com/run/docs/reference/yaml/v1.md.txt).
references/iam-security.md
# Cloud Run IAM & security

Cloud Run uses Identity and Access Management (IAM) to secure your resources
and control who can deploy or invoke them.

## Predefined IAM Roles

| Predefined Role | Usage |
| :--- | :--- |
| `roles/run.admin` | Full control over all Cloud Run resources. |
| `roles/run.invoker` | Invoke Cloud Run services and execute jobs. |
| `roles/run.developer` | Read and write access; can't set IAM policies. |
| `roles/run.viewer` | Read-only access to Cloud Run resources. |

## Types of service accounts for service identity

Cloud Run resources run as a specific service account (the service
identity).

- **User-managed service account (recommended)**: You manually create this
  service account and determine the most minimal set of permissions that
  the service account needs to access specific Google Cloud resources. The
  user-managed service account follows the format of:
  `SERVICE_ACCOUNT_NAME@PROJECT_ID.iam.gserviceaccount.com`.

- **Compute Engine default service account:** Cloud Run automatically
  provides the Compute Engine default service account as the default
  service identity. The Compute Engine default service account
  follows the format of:
  `PROJECT_NUMBER-compute@developer.gserviceaccount.com`.

## Best practices

By default, the Compute Engine default service account is automatically
created. If you don't specify a service account when the Cloud Run resource is created, Cloud Run uses this service account.

Depending on your organization policy configuration, the default service
account might automatically be granted the Editor role on your project.
We strongly recommend that you disable the automatic role grant by enforcing
the `iam.automaticIamGrantsForDefaultServiceAccounts` organization policy
constraint. If you created your organization after May 3, 2024, this
constraint is enforced by default.

Create a user-managed service account with minimal permissions for each Cloud
Run resource.

To allow a service to access another GCP resource (e.g., Cloud SQL), grant the
service's identity the appropriate IAM role on that resource.

## Security controls

- **Ingress Settings:** Control whether your service is reachable from the
  internet (`all`), only from within the VPC (`internal`), or via a load
  balancer (`internal-and-cloud-load-balancing`).

- **VPC Egress:** Use a VPC connector or Direct VPC egress to allow Cloud Run
  to access resources in your VPC.

- **Binary Authorization:** Ensure only trusted container images are deployed.

- **Secrets Management:** Use Secret Manager to securely pass sensitive
  information (e.g., API keys, database passwords) to your containers as
  environment variables or volumes.

## Container security best practices

When packaging and deploying containerized applications to Cloud Run, follow
these security best practices:

-   **Build minimal container images:** Build minimal container images by working
    from lean images, such as
    Alpine or scratch, and utilize multi-stage builds to
    to keep your container light at run time.
-   **Run as a non-root user:** Avoid running your container as the root user.
    Configure your Dockerfile to run the container using a non-root user to
    limit privilege escalation and file access should the container be
    compromised.
-   **Keep base images updated:** Use actively maintained and secure base
    images such as Google base images or Docker Hub's official images.
-   **Enable vulnerability scanning:** Turn on automated vulnerability scanning
    in Artifact Registry to continuously scan your container images/packages
    for known CVEs. Apply
    the latest security updates by regularly rebuilding container images and
    redeploying your services.
-   **Implement deterministic builds:** Pin specific versions, tags, or digests
    for base images and package dependencies to ensure predictable and secure
    build outcomes. This also prevents unverified code from being included in
    your container.
-   **Control Preview features:** Prevent the use of Preview features by using
    custom organization policies.

## Public access

There are two ways to create a public Cloud Run service, you can either:

* Disable the Cloud Run Invoker IAM check (recommended).
* Assign the Cloud Run Invoker IAM role to the `allUsers` member type.

For more information, see:
[Cloud Run security overview](https://docs.cloud.google.com/run/docs/securing/managing-access.md.txt).

## Configure IAP to secure access

By enabling IAP on Cloud Run directly, you can secure traffic with a single
click from all ingress paths, including default `run.app` URLs and load
balancers.

When you integrate IAP with Cloud Run, you can manage user or group access in
the following ways:

* Inside the organization - configure access to users who are within the same
  organization as your Cloud Run service

* Outside the organization - configure access to users who are from
  organizations different than your Cloud Run service

* No organization - configure access in projects that are not part of any
  Google organization

Enabling IAP on a Cloud Run service can be as easy as deploying a new service
with the following flags:

```bash
gcloud run deploy SERVICE_NAME \
  --region=REGION \
  --image=IMAGE_URL \
  --no-allow-unauthenticated \
  --iap \
  --quiet
```
references/mcp-usage.md
# Cloud Run MCP Usage

Cloud Run is supported by a remote Model Context Protocol (MCP) server that
enables agents to deploy, manage, and monitor serverless applications.

## MCP Tools for Cloud Run

The Cloud Run MCP server typically includes tools for:

- `get_service`: Get info about a Cloud Run service, such as its URI and
  whether the deploy succeeded.
- `list_services`: List Cloud Run services in a given Google Cloud project and
  region.
- `deploy_service_from_image`: Deploy a container image from Artifact Registry
  or Docker Hub as a Cloud Run service.
- `deploy_service_from_archive`: Deploy a Cloud Run service directly from a
  self-contained source code archive (.tar.gz), skipping the container image
  build step for faster deployment. The archive must include all dependencies.
- `deploy_service_from_file_contents`: Deploys a Cloud Run service directly from
  local source files. This method is suitable for scripting languages like Python
  and Node.js, of which the source code can be embedded in the request. This is
  ideal for quick tests and development feedback loops. You must include all
  necessary dependencies within the source files because it skips the build step
  for faster deployment.

## Setup Instructions

To connect to the Cloud Run MCP server:

1.  Enable the Cloud Run API in your Google Cloud project.
2.  Configure the agent's MCP connection using the Gemini CLI extension.
3.  Follow the setup guide:
    [Setting up Cloud Run MCP](https://docs.cloud.google.com/run/docs/reference/mcp.md.txt).

## Supported Operations

Agents using the Cloud Run MCP can:

- Automate the rollout of new revisions.
- Troubleshoot failing deployments by inspecting logs and status.
- Manage scheduled jobs and verify their execution history.

Alternatively, use the [open source Cloud Run MCP server](https://github.com/GoogleCloudPlatform/cloud-run-mcp) which runs locally.
references/networking.md
# Cloud Run Networking best practices and cost optimization

When configuring networking options for Cloud Run services, apply these best
practices and cost optimization strategies.

## Optimize costs

When configuring networking options for your resources, consider the following:

*   **Co-locate your resources**: Deploy your Cloud Run resources in the same
    region as your backend databases (like Cloud SQL or Firestore) and Cloud
    Storage buckets. Data transfer between Google Cloud resources within the
    same region is free (`$0.00 / GiB`).
*   **Switch to Direct VPC egress**: If you are securely routing traffic to
    internal VPC network resources, switch to Direct VPC egress from Serverless
    VPC Access connectors. Direct VPC egress scales to zero, eliminating the
    baseline compute overhead and idle costs associated with connector
    instances.
*   **Offload static assets**: Use Cloud CDN in front of your Cloud Run
    resources to cache static assets and highly cacheable content. Serving data
    from the edge is significantly cheaper than paying for standard internet
    egress directly from Cloud Run.
*   **Monitor internet egress**: Inbound traffic (ingress) is always free, and 1
    GiB of free outbound internet data transfer is provided per month within
    North America. Focus your monitoring efforts on outbound traffic that
    crosses region boundaries or exceeds the free tier.

## Monitor IP address usage

If you're using Direct VPC egress, make sure that you have enough IP addresses
for your subnet. The number of IP addresses you use depends on the number of
instances that your workloads run, so we recommend monitoring your IP address
usage. Be sure that your IP usage over time stays within the bounds supported by
the subnet.

To estimate your IP address usage:

1.  Look up the number of instances in your project using the metric type
    `run.googleapis.com/container/instance_count`.
2.  Multiply the instance count metric's value by 2 to get an estimate of the
    number of IP addresses in use.

## IP address exhaustion strategies

Having a large number of Cloud Run workloads can cause IP exhaustion challenges
when using the RFC 1918 private IP address space with Direct VPC egress. The
following strategies can help you manage IP address exhaustion by using
alternative IP address ranges.

### Use non-RFC 1918 IPv4 addresses

Aside from the RFC 1918 IPv4 address ranges, Cloud Run also supports RFC 6598
(`100.64.0.0/10`) and Class E/RFC 5735 (`240.0.0.0/4`) ranges. All Google Cloud
services and features work with these non-RFC 1918 ranges, including VPC
networks, Cloud Load Balancing, and Private Service Connect. For best
compatibility, start with the RFC 6598 (`100.64.0.0/10`) range. If already in
use, consider using Class E/RFC 5735 (`240.0.0.0/4`).

### Use Cloud NAT or Private Service Connect

If your Cloud Run workload using a non-RFC 1918 range needs to reach an
on-premises destination that accepts only RFC 1918, use one of the following
solutions:

*   Use Hybrid NAT to perform address translation and egress using a small RFC
    1918 range.
*   Expose the on-premises service as a Private Service Connect hybrid service.

### Use IPv4 and IPv6 (dual-stack) subnets

Although it won't reduce IPv4 exhaustion, moving your apps to IPv6 is a good
first step. Set up dual-stack resources to avoid IPv4 exhaustion problems in the
future.

## Port exhaustion reduction strategies

When sending a large number of requests to a single destination IP address, use
connection pooling to maintain and reuse connections to the destination. High
connection rates to a single IP address can exhaust outbound ports and cause
connection refused errors.

### Use connection pooling and reuse connections

Use application-level connection pools (e.g., HTTP keep-alive or database
connection pooling libraries) and limit maximum pool size per container instance
to reuse active TCP connections rather than opening new sockets for every
request.

## Performance and throughput strategies

This section covers scalable options for improving network performance and
throughput towards the internet and Google services.

### Use the second generation execution environment

For the best networking performance for Cloud Run services, use the second
generation execution environment when routing traffic with Direct VPC egress.
The second generation environment provides faster network performance,
especially in the presence of packet loss.

### Use Direct VPC egress for faster network egress throughput

To achieve faster throughput across network egress connections, use Direct VPC
egress to route traffic through your VPC network. We recommend using this in
conjunction with the second generation execution environment.

#### Example 1: External traffic to the internet

If you're sending external traffic to the public internet, route all traffic
through the VPC network by setting `--vpc-egress=all-traffic`. With this
approach, you must set up Cloud NAT to reach the public internet.

#### Example 2: Internal traffic to a Google API

If you're using Direct VPC egress to send traffic to a Google API, such as Cloud
Storage, choose one of the following options:

*   Specify `private-ranges-only` (default) with Private Google Access:
    1.  Set the flag `--vpc-egress=private-ranges-only`.
    2.  Enable Private Google Access.
    3.  Configure DNS for Private Google Access so your target domain (such as
        `storage.googleapis.com`) maps to `199.36.153.8/30` or
        `199.36.153.4/30`.
*   Specify `all-traffic` with Private Google Access:
    1.  Set the flag `--vpc-egress=all-traffic`.
    2.  Enable Private Google Access.

## Use the default MTU setting for Cloud Run

Don't change the maximum transmission unit (MTU) setting of a VPC network when
using it with Cloud Run. Use the default MTU of 1,460 bytes instead.