OpenTelemetry Collector Cost Control: Cut Telemetry Spend Without Going Blind
Switching observability vendors moves the same volume to a cheaper rate. Cutting the volume in the Collector works on any backend, and here is the config that does it.
By VVV Ops ·
Your observability invoice climbed 40% and nobody renegotiated the contract. What changed was the fleet: more pods, more services, more spans per request, and agent frameworks that emit a span for every tool call. The usual reaction is a vendor bake-off, which takes two quarters and then moves the same volume to a slightly cheaper per-unit rate. OpenTelemetry Collector cost control is the other path, and we recommend it first on every engagement. Decide what telemetry is worth keeping, enforce that decision in the pipeline, and the backend you pick stops being the expensive part.
Where the telemetry bill actually comes from
Four line items, and each one scales on a different thing. Hosts scale with your infrastructure. Log ingestion scales with bytes. Log indexing scales with event count, which is not the same curve. Metrics scale with active series, which is cardinality, which is the one nobody models until it bites.
Here are the published list rates from two vendors, fetched this week, so you can do the arithmetic on your own volumes.
| Line item | Datadog (annual billing) | Grafana Cloud Pro | |---|---|---| | Infrastructure host | $15 per host, per month (Pro); $23 (Enterprise) | No per-host charge; $19/month platform fee | | APM host | $31 per host, per month | Billed as trace volume | | Log ingestion | $0.10 per ingested or scanned GB, per month | $0.050/GB process, $0.400/GB write, $0.100/GB retain | | Log retention | $1.70 per million log events, per month at 15-day retention | Included in the retain rate, 30 days | | Metrics | 100 custom metrics per host on Pro, 200 on Enterprise; overage is quote-only | $6.50 per 1k series, 13 months retention |
Take a 200-host estate on Datadog Pro with APM. Infrastructure is 200 × $15 = $3,000 a month. APM is 200 × $31 = $6,200. Say you ship 8,000 GB of logs and index 1,200 million events: ingestion is 8,000 × $0.10 = $800, indexing is 1,200 × $1.70 = $2,040. That is $12,040 a month, or $144,480 a year, before anyone opens a dashboard.
Notice which numbers you can move from the pipeline. The two host lines are fixed by your infrastructure. The $2,840 of log spend and the entire metrics bill are volume, and volume is a policy decision you have not written down yet.
Fix the pipeline before you shop for a cheaper backend
Migrating vendors is the most expensive way to cut an observability bill. You rebuild dashboards, retrain everyone on a new query language, rewire alert routing, and discover three months in that the incident runbooks all reference the old links. Do the pipeline work first. If you still want to migrate afterwards, you migrate a third of the volume and the quotes come back better anyway.
There is a second reason to put the policy in the Collector rather than in a vendor's UI. OpenTelemetry graduated from the CNCF on 21 May 2026, with over 12,000 contributors from more than 2,800 companies and the second-highest project velocity in the CNCF, behind only Kubernetes. Filter rules written as Collector config survive a backend change. Filter rules written in a vendor's exclusion-filter UI do not.
The market agrees on where the pressure is. In Grafana Labs' fourth annual survey of 1,363 practitioners across 76 countries, cost was the single most important tool selection criterion for the third year running, cited by 65%, and half of respondents expected to spend more on observability the following year.
Architecturally, run two layers. Agent Collectors sit next to the workload and do the cheap, local work: enrichment, redaction, and the obvious drops. A gateway tier does the work that needs a view across instances, which is tail sampling and aggregation. Do not try to do trace-level sampling in the agent tier. It cannot see the whole trace, so it will make the wrong call.
| Signal | What drives the cost | Collector lever | What you give up | |---|---|---|---| | Logs, by byte | Verbose access and debug logs | filter processor at the agent | Grep-ability of routine traffic | | Logs, by event count | Every line indexed by default | count connector, index only what you query | Ad hoc search over dropped classes | | Traces | One span per operation, plus agent tool calls | tail_sampling at the gateway | Exemplar traces for healthy requests | | Metrics | Per-pod labels on high-churn workloads | transform with keep_keys | Per-pod drill-down after the fact |
Drop at the edge with the filter processor
Most of what a service logs is a record that nothing happened. Health check responses, successful readiness probes, debug lines somebody left on after a 2 a.m. session. You pay to ship those bytes, then pay again to index the events.
The filter processor evaluates OTTL conditions and drops anything that matches. Run it in the agent tier so the bytes never leave the node.
processors:
filter:
error_mode: ignore
log_conditions:
- log.severity_number < SEVERITY_NUMBER_WARN
- IsMatch(log.body, ".*/healthz.*")
trace_conditions:
- span.attributes["http.request.method"] == nil
Two warnings before you copy that. The conditions are ORed, so any match drops the record, and a sloppy regex will quietly eat something you need. Second, error_mode: ignore means a condition that errors is logged and skipped rather than taking the pipeline down. That is the behaviour you want in production, and it is also the behaviour that hides your typo in staging. Test the conditions against a captured sample before you ship them.
Back to the 200-host example. Cut indexed log events in half, from 1,200 million to 600 million, and the indexing line goes from $2,040 to $1,020. Halve the bytes too, from 8,000 GB to 4,000 GB, and ingestion goes from $800 to $400. That is $1,420 a month, $17,040 a year, from one config file. In the estates we have worked on, half is the conservative figure. Health checks and debug output alone usually account for more.
Turn logs into metrics instead of paying to index them
Some log classes exist only so somebody can count them. Nobody reads the individual "payment declined" line. They read the graph of how many there were this hour versus last Tuesday.
Pay for the graph, not the lines. The count connector consumes logs and emits metrics, which you then export while the original log records go nowhere.
connectors:
count:
logs:
my.error.log.count:
description: Error+ logs.
conditions:
- severity_number >= SEVERITY_NUMBER_ERROR
attributes:
- key: env
service:
pipelines:
logs:
receivers: [otlp]
exporters: [count]
metrics:
receivers: [count]
exporters: [otlphttp/backend]
We use this for rate-shaped questions: error counts by service, retry counts, rejected requests by reason code. We do not use it for anything an engineer will need to read verbatim during an incident. Deciding which is which is a conversation with the on-call rotation, not a decision the platform team should make alone. If your incident response leans on automation, the signals that automation consumes are the ones to keep at full fidelity.
Cardinality is the metrics bill
Metrics look cheap per unit until you count series. One series is one unique combination of metric name and label values, and a single careless label multiplies everything.
Work an example. You have an http_server_duration histogram with 14 buckets, 80 route values, and it carries a pod attribute across 300 pods. That is 14 × 80 × 300 = 336,000 active series. At Grafana Cloud Pro's $6.50 per 1k series, that single metric costs 336 × $6.50 = $2,184 a month.
Now drop the pod attribute and aggregate at the service level, where you have 12 services. That is 14 × 80 × 12 = 13,440 series, or 13.44 × $6.50 = $87.36 a month. You saved $2,096.64 a month, roughly $25,000 a year, on one metric.
The transform processor does this with OTTL:
processors:
transform:
error_mode: ignore
metric_statements:
- keep_keys(resource.attributes, ["service.name", "service.namespace", "deployment.environment", "cloud.region"])
trace_statements:
- delete_key(span.attributes, "http.request.header.user_agent")
What you lose is real. You can no longer ask which pod was slow after the fact from this metric. Our position is that you should not be asking that question of a histogram anyway. Pod-level latency belongs in traces, where tail sampling has kept the slow ones, and in the pod's own logs for the window in question.
An allowlist beats a denylist here. keep_keys means a new attribute added by a library upgrade cannot silently double your series count next Tuesday. If your label schema is still ad hoc, fix that first; we wrote up a label schema that survives an audit in the context of GPU chargeback, and the same discipline applies to every metric you emit.
Tail sampling and the architecture it forces on you
Head sampling decides at the start of a trace, before anything interesting has happened, so it throws away errors at the same rate as successes. Tail sampling waits for the trace to complete and then decides. Keep every error, keep everything slow, keep a small random sample of the boring ones.
processors:
tail_sampling:
decision_wait: 10s
num_traces: 100000
expected_new_traces_per_sec: 2000
policies:
[
{
name: errors,
type: status_code,
status_code: {status_codes: [ERROR]}
},
{
name: slow,
type: latency,
latency: {threshold_ms: 2000}
},
{
name: baseline,
type: probabilistic,
probabilistic: {sampling_percentage: 5}
}
]
The catch is stateful. Every span of a trace has to reach the same Collector instance for the decision to be correct, which means you cannot put a round-robin load balancer in front of your gateway tier. The load_balancing exporter exists for exactly this, and its docs are explicit that it "ensures data from the same source is always routed to the same backend, enabling stateful downstream processing like log reduction, throttling, and tail-based sampling."
exporters:
load_balancing:
routing_key: traceID
protocol:
otlp:
timeout: 1s
resolver:
static:
hostnames:
- gateway-1:4317
- gateway-2:4317
Budget for the memory. num_traces is a buffer of in-flight traces held for decision_wait seconds, and setting it to 100,000 on a pod with a 512 MiB limit will get that pod OOM-killed during your next traffic spike. Size it from measured span rate, give the gateway tier its own node pool, and watch the eviction count for a week before you trust it.
On money: at Grafana Cloud Pro rates, 1,500 GB of trace data a month costs 1,500 × ($0.050 + $0.400 + $0.100) = $825. Keep 10% of traces and that becomes $82.50. The error traces you actually open during an incident are all still there.
What we would cut in the first two weeks
The order matters, because the first two steps pay for the rest.
- Instrument the bill before you change anything. Get a per-service breakdown of log bytes, indexed events, and active series. If you cannot get it from the vendor, put a
countconnector in a parallel pipeline for a week. - Drop debug and health check logs at the agent tier. This is the single largest reduction available and it carries almost no risk. Expect it in production within three days.
- Kill the top five metrics by series count. Not by name, by cardinality. In every estate we have looked at, the top five are 60% or more of the series total, and at least two of them are somebody's abandoned experiment.
- Move rate-shaped log classes to the
countconnector. One class per week, with the on-call rotation signing off on each. - Only now, stand up the gateway tier and turn on tail sampling. It is the highest-value change and also the one that will page you if you get it wrong, so it goes last.
Set a standing review. Volume creeps back the moment a new service ships with default instrumentation, so the pipeline config needs an owner and a monthly look at series count. That cadence belongs inside a real FinOps practice rather than in someone's calendar, and if you do not have one yet, start here.
One thing we will not do is cut telemetry to hit a number. If a team cannot debug a production incident because we dropped the wrong thing, the saving was negative. Every drop rule above is reversible in one deploy, and that is the property to preserve.
When to Get Help
Bring someone in when the bill is growing faster than the fleet, when nobody can tell you which service owns the top metric by cardinality, or when tail sampling has been on the roadmap for three quarters because the architecture question keeps stalling it. Those are two-week problems with the right pair of eyes and two-quarter problems without.
We do this work as a fixed-scope engagement: measure the current pipeline, land the drop rules with the on-call rotation in the room, and hand over Collector config your team owns rather than a vendor dashboard your team rents. Talk to us if that is the shape of your problem.