Skip to main content

GCP Monitoring with OpenTelemetry - Architecture for base14 Scout

This is the architectural landing page for monitoring Google Cloud infrastructure with base14 Scout through the OpenTelemetry Collector. It covers the three paths telemetry can take out of GCP, the IAM each one needs, and the resource attributes every per-service guide sets. Read it once, then jump to the guide for the surface you are instrumenting.

Running this in production

Storing and querying this telemetry at production volume is what base14 Scout does. Check out Scout Metrics.

The three paths

Google Cloud emits telemetry in three shapes, and a production deployment usually runs all three.

┌──────────────────────── Google Cloud ─────────────────────────┐
│ │
│ Managed service (Cloud SQL, Pub/Sub, Cloud Run, LB, ...) │
│ │ │ │
│ │ metrics │ logs │
│ ▼ ▼ │
│ Cloud Monitoring Cloud Logging │
│ │ │ │
│ │ Log Router sink │
│ │ ▼ │
│ │ Pub/Sub topic │
│ │ │ │
└────────┼──────────────────────────┼───────────────────────────┘
│ Monitoring API (pull) │ subscription (push)
▼ ▼
googlecloudmonitoring googlecloudpubsub
receiver receiver
│ │
└──────────┬───────────────┘
│ ┌─────────────────────────────────┐
│ │ In-VPC collector │
│◀───────┤ postgresql / mysql / redis / │
│ │ nginx / prometheus receivers │
│ │ Cloud Run OTLP sidecar │
│ └─────────────────────────────────┘

OpenTelemetry Collector
│ OTLP

Scout
PathComponentSignalsFreshnessMain cost driver
Cloud Monitoring pullgooglecloudmonitoring receiverMetrics60s floor plus GCP's own export delay (often 3-5 min)Monitoring API read quota
Cloud Logging pushgooglecloudpubsub receiver + google_cloud_logentry_encodingLogsSecondsPub/Sub delivery and egress
Direct scrapepostgresql, mysql, redis, nginx, prometheus receiversMetricsYour collection_intervalCollector compute; no GCP API cost

The first two are documented once, as mechanism guides:

The per-service guides below assume you have read whichever of those two applies, and cover only what is specific to their surface.

What about traces?

No GCP managed service emits distributed traces. Cloud SQL, Pub/Sub, Cloud Load Balancing, API Gateway and VPC produce metrics and logs only. Traces in a GCP architecture come from two places:

  • Your application code, instrumented with an OpenTelemetry SDK. On Cloud Run, a collector sidecar is the cleanest way to get them out — see Cloud Run.
  • Google Cloud client libraries, which emit client spans for calls into managed services. Be aware that the Pub/Sub client libraries emit rpc.* attributes only and no messaging.* attributes at all; see Pub/Sub for what that means downstream.

Which collector image to run

Run otel/opentelemetry-collector-contrib. None of the GCP components — googlecloudmonitoring, googlecloudpubsub, the google_cloud_logentry_encoding extension — ship in the core collector distribution, and none of them are in the Scout collector distribution either. The same is true of the postgresql, mysql, redis and nginx receivers the alternative paths use.

If you already run a Scout collector for application telemetry, add a second collector for the GCP receivers rather than swapping the image on the first one. That also gives you the separate pipeline the next section requires.

Run the pull receiver on a single replica

googlecloudmonitoring polls on an interval. Every replica polls independently, so a three-replica Deployment triples your Monitoring API usage and produces duplicate series. Put it on a single-replica Deployment, never a DaemonSet.

Resource attributes every GCP pipeline sets

The googlecloudmonitoring receiver sets gcp.resource_type and the monitored-resource labels as resource attributes — and no service.name. Without one, every GCP metric arrives as unknown_service, which in the Scout data lake means it shares a sort-key prefix with everything else that went unnamed.

Set one per surface, following the convention already used for infra telemetry across the fleet (system-metrics, kubernetes-metrics):

Surfaceservice.namecloud.platform
Cloud SQLcloudsql-metricsgcp_cloud_sql
Memorystorememorystore-metricsgcp_memorystore
Cloud Load Balancingloadbalancing-metricsgcp_load_balancing
API Gateway (managed)apigateway-metricsgcp_api_gateway
nginx gateway (self-managed)nginx-gateway-metricsgcp_kubernetes_engine
Pub/Subpubsub-metricsgcp_pubsub
Cloud Runcloudrun-metricsgcp_cloud_run
VPCvpc-logs (flow logs), vpc-metrics (Cloud NAT)gcp_vpc

The block each guide repeats, with its own values substituted:

gcp-common.yaml
processors:
resource/cloudsql:
attributes:
- {key: service.name, value: cloudsql-metrics, action: insert}
- {key: cloud.provider, value: gcp, action: insert}
- {key: cloud.platform, value: gcp_cloud_sql, action: insert}
- {key: cloud.account.id, value: "${env:GCP_PROJECT_ID}", action: insert}
# Regional surfaces only — omit for Pub/Sub, Cloud Load Balancing and VPC
- {key: cloud.region, value: "${env:GCP_REGION}", action: insert}
- {key: deployment.environment.name, value: "${env:ENVIRONMENT}", action: upsert}
- {key: environment, value: "${env:ENVIRONMENT}", action: upsert}

cloud.region belongs only on surfaces that have one. Cloud SQL and Memorystore instances live in a region; Pub/Sub topics, global load balancers and VPC networks do not, and stamping a region on them invents a dimension that is not real.

Semconv version note

deployment.environment.name is the current OTel attribute (semantic conventions v1.27+, stable in v1.40.0). Scout's UI filters on the lowercase environment key, so emit it alongside the OTel-native deployment.environment.name. The legacy deployment.environment is still accepted for backward compatibility.

Give each surface its own pipeline

The insert action on service.name and the /cloudsql suffix on the processor are both load-bearing. A blanket resource processor that upserts a single service.name across a shared pipeline will stamp your application's name onto every GCP metric, and they become indistinguishable from application telemetry. Keep one receiver/processor/pipeline triple per surface, all suffix-keyed, so they coexist in one collector without overwriting each other.

Metric kinds and temporality

Each GCP metric has a kind (GAUGE, DELTA, CUMULATIVE) and a value type (INT64, DOUBLE, DISTRIBUTION). The receiver maps them like this:

GCP kind and typeOTel resultTemporality
GAUGE + scalarGauge
CUMULATIVE + scalarMonotonic sumCumulative
DELTA + scalarSumDelta
DELTA + DISTRIBUTIONHistogramDelta
GAUGE + DISTRIBUTIONUnsupported

Two consequences:

Most GCP counters are DELTA. request_count, ack_message_count, disk/read_ops_count and their siblings all arrive as delta sums, not cumulative ones. A panel or rollup that assumes a monotonically increasing counter — one that applies rate()-style logic, or subtracts consecutive points — will be wrong against them. Sum delta points over the window instead.

A GAUGE-kind DISTRIBUTION can drop everything. It yields an invalid data point that fails the entire scrape batch, so every metric from that receiver instance disappears, not just the offending one. If metrics vanish after a config change, remove the distribution-valued metric you just added. Latency metrics (*_latencies, *_times) are the usual culprits. Each per-service guide flags its distributions.

Authentication

Both GCP receivers use Application Default Credentials. Create one Google service account (GSA) and grant it what the paths you use need:

PathRoleScope
Cloud Monitoring pullroles/monitoring.viewerProject
Cloud Logging pushroles/pubsub.subscriberThe subscription
gcloud iam service-accounts create scout-telemetry-reader \
--display-name="base14 Scout telemetry reader"

gcloud projects add-iam-policy-binding PROJECT_ID \
--member="serviceAccount:scout-telemetry-reader@PROJECT_ID.iam.gserviceaccount.com" \
--role="roles/monitoring.viewer"
GKE Workload Identity

On GKE, bind the GSA to the collector's Kubernetes service account and skip key files entirely:

gcloud iam service-accounts add-iam-policy-binding \
scout-telemetry-reader@PROJECT_ID.iam.gserviceaccount.com \
--role="roles/iam.workloadIdentityUser" \
--member="serviceAccount:PROJECT_ID.svc.id.goog[NAMESPACE/KSA_NAME]"

Then annotate the KSA:

metadata:
annotations:
iam.gke.io/gcp-service-account: scout-telemetry-reader@PROJECT_ID.iam.gserviceaccount.com

Elsewhere, point GOOGLE_APPLICATION_CREDENTIALS at a service account key file. Both mechanism guides cover this in full.

One project per receiver

project_id is singular. To collect from several projects, add one receiver instance per project, each suffix-keyed:

multi-project.yaml
receivers:
googlecloudmonitoring/prod:
project_id: my-prod-project
metrics_list:
- metric_descriptor_filter: 'metric.type = starts_with("cloudsql.googleapis.com/")'
googlecloudmonitoring/staging:
project_id: my-staging-project
metrics_list:
- metric_descriptor_filter: 'metric.type = starts_with("cloudsql.googleapis.com/")'

Grant the GSA roles/monitoring.viewer in each project.

Verifying data has landed

GCP telemetry lands in the base otel_metrics_* and otel_logs tables, with the GCP labels in the ResourceAttributes and Attributes maps. The metric name is the raw GCP metric type, not a translated one — so MetricName is literally cloudsql.googleapis.com/database/cpu/utilization.

Every verification query in these guides filters on ServiceName, MetricName and a bounded one-hour window, and caps itself:

SELECT MetricName, count() AS points, max(Value) AS latest
FROM otel_metrics_gauge
WHERE ServiceName = 'cloudsql-metrics'
AND MetricName = 'cloudsql.googleapis.com/database/cpu/utilization'
AND TimeUnix >= now() - INTERVAL 1 HOUR
GROUP BY MetricName
SETTINGS max_execution_time = 30, max_rows_to_read = 50000000

Delta counters land in otel_metrics_sum and distributions in otel_metrics_histogram — if a metric is missing from one table, look for it in the other before concluding it failed.

Per-surface guides

GuideLead pathCovers
Cloud SQLCloud MonitoringHost and engine metrics, database logs, in-database scraping
MemorystoreCloud MonitoringRedis, Redis Cluster and Valkey engines
Cloud Load BalancingCloud Monitoring + LoggingRequest rates, latency distributions, access logs
API Gateway and nginxCloud Monitoring / PrometheusManaged API Gateway and self-managed nginx gateways
Pub/SubCloud MonitoringTopic and subscription health, backlog alerting
Cloud RunOTLP sidecarApplication traces, platform metrics, request logs
VPCCloud LoggingFlow logs, Cloud NAT, network metrics

FAQ

How does base14 Scout collect Google Cloud telemetry?

base14 Scout collects Google Cloud telemetry through the OpenTelemetry Collector, along three paths. The googlecloudmonitoring receiver pulls metrics from the Cloud Monitoring API, the googlecloudpubsub receiver consumes logs that a Log Router sink pushes into Pub/Sub, and standard receivers such as postgresql and redis scrape services directly from inside the VPC.

Which collector distribution do I need for GCP?

GCP telemetry needs the otel/opentelemetry-collector-contrib distribution. The GCP receivers and the Cloud Logging encoding extension are contrib components, present in neither the core collector nor the Scout collector distribution.

Do GCP managed services emit distributed traces?

No GCP managed service emits distributed traces. Cloud SQL, Pub/Sub, Cloud Load Balancing, API Gateway and VPC emit metrics and logs only. Traces come from your own application code, or from Google Cloud client libraries emitting client spans.

Why do my GCP metrics show up as unknown_service?

The googlecloudmonitoring receiver does not set service.name. Add a resource processor that inserts one per surface — cloudsql-metrics, pubsub-metrics, and so on. Because ServiceName is the leading sort key in the Scout data lake, leaving it unset makes queries far more expensive.

Why does my GCP counter graph look wrong?

Most GCP counters have DELTA kind and arrive as delta sums, not cumulative ones. Sum the points over your window rather than applying counter-rate logic that assumes a monotonically increasing series.

Can one collector serve several GCP projects?

One collector can serve several projects, but project_id is singular per receiver — add one receiver instance per project, and grant the service account roles/monitoring.viewer in each.

Reference

Was this page helpful?