返回 Skills
google/agents-cli· Apache-2.0 内容可用

google-agents-cli-observability

This skill should be used when the user wants to "set up tracing", "monitor my ADK agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring deployed ADK (Agent Development Kit) agents. Covers Cloud Trace, prompt-response logging, BigQuery Agent Analytics, third-party integrations (AgentOps, Phoenix, MLflow, etc.), and troubleshooting. Part of the Google ADK (Agent Development Kit) skills suite. Do NOT use for deployment setup (use google-agents-cli-deploy) or API code patterns (use google-agents-cli-adk-code).

安装

与 skills.sh 相同的 Command / Prompt 安装方式


name: google-agents-cli-observability description: > This skill should be used when the user wants to "set up tracing", "monitor my ADK agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring deployed ADK (Agent Development Kit) agents. Covers Cloud Trace, prompt-response logging, BigQuery Agent Analytics, third-party integrations (AgentOps, Phoenix, MLflow, etc.), and troubleshooting. Part of the Google ADK (Agent Development Kit) skills suite. Do NOT use for deployment setup (use google-agents-cli-deploy) or API code patterns (use google-agents-cli-adk-code). metadata: author: Google license: Apache-2.0 version: 1.1.0 requires: bins: - agents-cli install: "uv tool install google-agents-cli"

ADK Observability Guide

Cloud Trace works out of the box — no infrastructure needed. Prompt-response logging and BigQuery Agent Analytics require Terraform-provisioned infrastructure (service account, GCS bucket, BigQuery dataset). Run agents-cli infra single-project --project PROJECT_ID to provision these resources. See references/cloud-trace-and-logging.md for details, env vars, and verification commands. If your project isn't scaffolded yet, see /google-agents-cli-scaffold first.

Order of operations for agent_runtime deployments

For deployment_target = agent_runtime, run agents-cli infra single-project before the first agents-cli deploy. The Terraform module owns the entire Reasoning Engine resource (display_name, service account, deployment spec, env vars), so applying it after a SDK-based deploy creates a state mismatch — Terraform has no record of the SDK-deployed instance and cannot layer env vars onto it without taking ownership of the whole resource.

If you have already run agents-cli deploy, you have two options:

  1. Switch to Terraform-managed. Delete the SDK-deployed Reasoning Engine, then run agents-cli infra single-project followed by agents-cli deploy. Sessions and any in-flight state on the previous instance are lost.
  2. Keep the SDK-deployed instance. Skip infra single-project and set the observability env vars on the running instance directly via the vertexai client update API. You will also need to grant the instance's service account the IAM permissions required to emit telemetry — writing to the logs GCS bucket, BigQuery dataset access, log writer, etc. See deployment/terraform/single-project/iam.tf and telemetry.tf in your scaffolded project for the full set of bindings the Terraform module would otherwise provision. Terraform-managed env vars are not available in this mode.

Reference Files

FileContents
references/cloud-trace-and-logging.mdScaffolded project details — Terraform-provisioned resources, environment variables, verification commands, enabling/disabling locally
references/bigquery-agent-analytics.mdBQ Agent Analytics plugin — enabling, key features, GCS offloading, tool provenance

Observability Tiers

Choose the right level of observability based on your needs:

TierWhat It DoesScopeDefault StateBest For
Cloud TraceDistributed tracing — execution flow, latency, errors via OpenTelemetry spansAll templates, all environmentsAlways enabledDebugging latency, understanding agent execution flow
Prompt-Response LoggingGenAI interactions exported to GCS, BigQuery, and Cloud LoggingADK agents onlyDisabled locally, enabled when deployedAuditing LLM interactions, compliance
BigQuery Agent AnalyticsStructured agent events (LLM calls, tool use, outcomes) to BigQueryADK agents with plugin enabledOpt-in (--bq-analytics at scaffold time)Conversational analytics, custom dashboards, LLM-as-judge evals
Third-Party IntegrationsExternal observability platforms (AgentOps, Phoenix, MLflow, etc.)Any ADK agentOpt-in, per-provider setupTeam collaboration, specialized visualization, prompt management

Ask the user which tier(s) they need — they can be combined. Cloud Trace is always on; the others are additive.


Cloud Trace

ADK uses OpenTelemetry to emit distributed traces. Every agent invocation produces spans that track the full execution flow.

Span Hierarchy

invocation
  └── agent_run (one per agent in the chain)
        ├── call_llm (model request/response)
        └── execute_tool (tool execution)

Setup by Deployment Type

DeploymentSetup
Agent RuntimeAutomatic — traces are exported to Cloud Trace by default
Cloud Run (scaffolded)Automatic — setup_telemetry() configures Cloud Trace/Logging exporters
GKE (scaffolded)Automatic — setup_telemetry() configures Cloud Trace/Logging exporters
Cloud Run / GKE (manual)Configure OpenTelemetry exporter in your app
Local devWorks with agents-cli playground; traces visible in Cloud Console

View traces: Cloud Console → Trace → Trace explorer

For detailed setup instructions (Agent Runtime CLI/SDK, Cloud Run, custom deployments), fetch https://adk.dev/integrations/cloud-trace/index.md.


Prompt-Response Logging

Captures GenAI interactions (model name, tokens, timing) and exports to GCS (JSONL) and BigQuery (via direct log sinks and external tables). Privacy-preserving by default — only metadata is logged unless explicitly configured otherwise.

Key env var: OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT — OTel GenAI semantic-conventions standard (modes: span_only, event_only, span_and_event, no_content). The scaffolded setup_telemetry() collapses every non-false value to NO_CONTENT (metadata-only); false disables capture. Logging is disabled locally unless LOGS_BUCKET_NAME is set.

For scaffolded project details (Terraform resources, env vars, privacy modes, enabling/disabling, verification commands), see references/cloud-trace-and-logging.md.

For ADK logging docs (log levels, configuration, debugging), fetch https://adk.dev/observability/logging/index.md.


BigQuery Agent Analytics Plugin

Optional plugin that logs structured agent events to BigQuery. Enable with --bq-analytics at scaffold time. See references/bigquery-agent-analytics.md for details.


Third-Party Integrations

ADK supports several third-party observability platforms. Each uses OpenTelemetry or custom instrumentation to capture agent behavior.

PlatformKey DifferentiatorSetup ComplexitySelf-Hosted Option
AgentOpsSession replays, 2-line setup, replaces native telemetryMinimalNo (SaaS)
Arize AXCommercial platform, production monitoring, evaluation dashboardsLowNo (SaaS)
PhoenixOpen-source, custom evaluators, experiment testingLowYes
MLflowOTel traces to MLflow Tracking Server, span tree visualizationMedium (needs SQL backend)Yes
Monocle1-call setup, VS Code Gantt chart visualizerMinimalYes (local files)
WeaveW&B platform, team collaboration, timeline viewsLowNo (SaaS)
FreeplayPrompt management + evals + observability in one platformLowNo (SaaS)

Ask the user which platform they prefer — present the trade-offs and let them choose. For setup details, fetch the relevant ADK docs page from the Deep Dive table below.


Troubleshooting

IssueSolution
No traces in Cloud TraceVerify setup_telemetry() runs at startup and the service account has the cloudtrace.agent role
Prompt-response data not appearingCheck LOGS_BUCKET_NAME is set; verify SA has storage.objectCreator on the bucket; check app logs for telemetry setup warnings
Privacy mode misconfiguredCheck OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT value — use NO_CONTENT for metadata-only, false to disable
BigQuery Analytics not loggingVerify plugin is configured in app/agent.py; check BQ_ANALYTICS_DATASET_ID env var is set
Third-party integration not capturing spansCheck provider-specific env vars (API keys, endpoints); some providers (AgentOps) replace native telemetry
Traces missing tool spansTool execution spans appear under execute_tool — check trace explorer filters
High telemetry costsSwitch to NO_CONTENT mode; reduce BigQuery retention; disable unused tiers

Deep Dive: ADK Docs (WebFetch URLs)

For detailed documentation beyond what this skill covers, fetch these pages:

TopicURL
Observability overviewhttps://adk.dev/observability/index.md
Agent activity logginghttps://adk.dev/observability/logging/index.md
Cloud Trace integrationhttps://adk.dev/integrations/cloud-trace/index.md
BigQuery Agent Analyticshttps://adk.dev/integrations/bigquery-agent-analytics/index.md
AgentOpshttps://adk.dev/integrations/agentops/index.md
Arize AXhttps://adk.dev/integrations/arize-ax/index.md
Phoenix (Arize)https://adk.dev/integrations/phoenix/index.md
MLflow tracinghttps://adk.dev/integrations/mlflow-tracing/index.md
Monoclehttps://adk.dev/integrations/monocle/index.md
W&B Weavehttps://adk.dev/integrations/weave/index.md
Freeplayhttps://adk.dev/integrations/freeplay/index.md

Related Skills

  • /google-agents-cli-deploy — Deployment targets, CI/CD pipelines, and production workflows
  • /google-agents-cli-workflow — Development workflow, coding guidelines, and operational rules
  • /google-agents-cli-adk-code — ADK Python API quick reference for writing agent code

附带文件

references/bigquery-agent-analytics.md
# BigQuery Agent Analytics Plugin

> **Opt-in.** Enable with `--bq-analytics` at scaffold time, or add manually to `app/agent.py`.

An optional plugin that logs structured agent events directly to BigQuery via the Storage Write API. Enables:

- **Conversational analytics** — session flows, user interaction patterns
- **LLM-as-judge evals** — structured data for evaluation pipelines
- **Custom dashboards** — Looker Studio integration
- **Tool provenance tracking** — LOCAL, MCP, SUB_AGENT, A2A, TRANSFER_AGENT

## Enabling

| Method | How |
|--------|-----|
| **At scaffold time** | `agents-cli scaffold create <project-name> --bq-analytics` |
| **Post-scaffold** | Add the plugin manually to `app/agent.py` (see [ADK docs](https://adk.dev/integrations/bigquery-agent-analytics/index.md)) |

Infrastructure (BigQuery dataset, GCS offloading) is provisioned automatically by Terraform when enabled at scaffold time.

## Key Features

- Auto-schema upgrade (new fields added without migration)
- GCS offloading for multimodal content (images, audio)
- Distributed tracing via OpenTelemetry span context
- SQL-queryable event log for all agent interactions

For full schema, SQL query examples, and Looker Studio setup, fetch `https://adk.dev/integrations/bigquery-agent-analytics/index.md`.
references/cloud-trace-and-logging.md
# Cloud Trace & Prompt-Response Logging (Scaffolded Projects)

> **Assumes `/google-agents-cli-scaffold` scaffolding.** Observability infrastructure is provisioned by Terraform in scaffolded projects.

## Cloud Trace

Always-on distributed tracing: Cloud export is configured by `setup_telemetry()` (Cloud Run/GKE) and by `setup_agent_engine_telemetry()` (Agent Runtime); `get_fast_api_app` is called with `otel_to_cloud=False`. Tracks requests through LLM calls and tool executions with latency analysis and error visibility.

View traces: **Cloud Console → Trace → Trace explorer**

No configuration required. Works in local dev (`agents-cli playground`) and all deployed environments.

## Prompt-Response Logging Infrastructure

All provisioned automatically by `deployment/terraform/single-project/telemetry.tf` (and the `cicd/` variant):

- **Log sinks** — Route GenAI inference logs and feedback logs directly to BigQuery (partitioned tables)
- **BigQuery dataset** — Telemetry dataset with external tables over GCS data and pre-created log export table
- **Pre-created log export table** — `gen_ai_client_inference_operation_details` table with Cloud Logging BQ export schema (labels flattened: dots become underscores)
- **GCS logs bucket** — Stores completions as NDJSON
- **BigQuery connection** — Service account for GCS access from BigQuery
- **Completions view** — Joins BQ log export data with GCS-stored prompt/response data

Check `deployment/terraform/single-project/telemetry.tf` for exact configuration. IAM bindings grant log sink service accounts `roles/bigquery.dataEditor` on the telemetry dataset.

## Environment Variables

Set automatically by Terraform on the deployed service:

| Variable | Purpose |
|----------|---------|
| `LOGS_BUCKET_NAME` | GCS bucket for completions and logs. Required to enable prompt-response logging |
| `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT` | Controls logging state and content capture |
| `BQ_ANALYTICS_DATASET_ID` | BigQuery dataset for telemetry (only when scaffolded with `--bq-analytics`) |
| `BQ_ANALYTICS_CONNECTION_ID` | BigQuery connection for GCS access (only when scaffolded with `--bq-analytics`) |
| `BQ_ANALYTICS_GCS_BUCKET` | GCS bucket for BigQuery Analytics multimodal offloading (only when scaffolded with `--bq-analytics`) |
| `GENAI_TELEMETRY_PATH` | Optional: override upload path within bucket (default: `completions`) |

## Enabling / Disabling

### Enable Locally

Set these before running `agents-cli playground`:

```bash
export LOGS_BUCKET_NAME="your-bucket-name"
export OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT="NO_CONTENT"
```

### Disable in Deployed Environments

Set `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=false` in `deployment/terraform/single-project/service.tf` (or the `cicd/` variant) and re-apply Terraform.

## BigQuery Dataset Naming Convention

BigQuery dataset names **cannot contain hyphens**. Terraform automatically converts hyphens to underscores when creating dataset names from your project name:

- Project name `my-agent` → BQ dataset `my_agent_telemetry`

One dataset is created:
- **`{name}_telemetry`** — Contains external tables over GCS completions data (NDJSON), the pre-created log export table (`gen_ai_client_inference_operation_details`), and the `completions_view`

To discover the actual dataset name in your project:
```bash
bq ls --project_id=${PROJECT_ID}
```

## Verifying Telemetry

After deploying, verify prompt-response logging is working:

```bash
PROJECT_ID="your-dev-project-id"
PROJECT_NAME="your-app-name"  # The agents-cli project name (not the GCP project ID)

# Check GCS data
gsutil ls gs://${PROJECT_ID}-${PROJECT_NAME}-logs/completions/

# Check BigQuery log export table (logs arrive via sink, may take a few minutes)
bq query --use_legacy_sql=false \
  "SELECT COUNT(*) FROM \`${PROJECT_ID}.${PROJECT_NAME//-/_}_telemetry.gen_ai_client_inference_operation_details\`"

# Query completions external table
bq query --use_legacy_sql=false \
  "SELECT * FROM \`${PROJECT_ID}.${PROJECT_NAME//-/_}_telemetry.completions\` LIMIT 10"

# Query the completions view (joins log export with GCS data)
bq query --use_legacy_sql=false \
  "SELECT * FROM \`${PROJECT_ID}.${PROJECT_NAME//-/_}_telemetry.completions_view\` LIMIT 10"
```

If data is not appearing: check `LOGS_BUCKET_NAME` is set, verify SA has `storage.objectCreator` on the bucket, check application logs for telemetry setup warnings. Log export to BigQuery may take a few minutes to propagate.
    google-agents-cli-observability | Prompt Minder