Skip to main content

OpenTelemetry Collector deployment topology

Scout collectors run in three roles. An Agent Collector scrapes managed services, a Daemon Collector runs on each node (a VM Agent on each VM when you run on virtual machines), and a Gateway Collector tier receives OTLP from every other collector before anything leaves for Scout.

Diagram​

Component roles​

ComponentKubernetesVMsScalingWhy
Agent CollectorSingle replicaSingle instanceSingleton (one per managed service)Polls RDS, ElastiCache, and SQS. A second copy scrapes the same service twice. To grow, shard services so each has exactly one owner.
Daemon Collector / VM AgentDaemonSetOne VM Agent per VMSingleton per node/VMScrapes only its own host. The count grows with the fleet, never by adding replicas.
Gateway CollectorHorizontally scaled replicasHorizontally scaled instancesHorizontally scalableStateless. Runs PII redaction and other transforms in one place before export.

Why the gateway owns PII and transforms​

Every signal passes through the gateway, so redaction configured there applies to all of it. That gives you one policy to maintain and one place to audit, instead of a copy on every node that can fall out of step. The same holds for filtering, attribute renames, and other transforms. See Filters and Transformations for the processors and OTTL patterns.

Scaling the collector​

This section summarizes the OpenTelemetry guide Scaling the Collector.

What to scale. Treat each signal and receiver separately. A scraper owns its targets, so it scales differently from a receiver that accepts pushed OTLP.

When to scale. Watch for three things. The memory_limiter refuses data (otelcol_processor_refused_* climbs). The exporter queue fills: add capacity when otelcol_exporter_queue_size reaches about 60–70% of otelcol_exporter_queue_capacity. A scrape takes nearly as long as its interval, which means the targets need sharding.

When not to scale. If the queue stays near capacity after you enlarge it, or otelcol_exporter_send_failed_* keeps rising, the backend or network is the constraint. More collectors only add load to it.

How to scale. Put stateless collectors behind a gRPC-aware (L7) load balancer and add replicas, keeping at least three so one failure is survivable. An L4 balancer pins each long-lived OTLP/gRPC connection to one replica, leaving new ones idle. Split scrape targets across scrapers instead of duplicating them; on Kubernetes the Target Allocator does this. Stateful processing such as tail sampling or span-to-metrics needs a load-balancing exporter tier in front, routing by trace ID or service.

What that means here. The Agent Collector and the Daemon Collector or VM Agent are scrapers, so each stays a singleton. The Gateway Collector is stateless, so it scales out. Adding tail sampling to the gateway would make it stateful and call for a load-balancing exporter tier in front of it.

References​

FAQ​

Why can't I run two Agent Collectors for the same managed service?​

Two Agent Collectors pointed at the same managed service collect every metric twice. Each copy polls RDS, ElastiCache, or SQS on its own schedule, so Scout receives duplicate series and you pay the provider's API cost twice. To spread load, give each Agent Collector a different set of services so every service has exactly one owner.

How do I scale the Gateway Collector?​

Add Gateway Collector replicas behind a gRPC-aware (L7) load balancer. The gateway is stateless, so any replica can handle any request. An L4 balancer keeps each OTLP/gRPC connection on the replica it first reached, which leaves new replicas idle. Scale when the exporter queue passes about 60–70% of capacity, and keep at least three replicas.

Does the Daemon Collector scale horizontally?​

No, the Daemon Collector runs exactly one copy per node, or one VM Agent per VM. It scrapes only its own host, so a second copy on the same node would duplicate that host's data. The number of Daemon Collectors grows as you add nodes or VMs.

Was this page helpful?