# How a B2B SaaS platform cut logs 24% and metrics 25% in six weeks, without touching a single customer log

A Latin American SaaS platform paid its observability backend per record. OllyGarden found where the volume came from and fixed it in their own OpenTelemetry Collector pipelines, with customer data off limits from day one.

> **Anonymized, but real.** This story comes from a real OllyGarden customer engagement. To protect the customer, we removed their name, people, products, and anything else that could identify them, and rounded some numbers. Everything else happened as described.

Published: 2026-10-03
Canonical URL: https://ollygarden.com/resources/case-studies/b2b-saas-cuts-logs-and-metrics-at-the-source

## At a glance

- **Company:** Mid-size, multi-tenant B2B SaaS platform that runs its customers' workloads
- **Region:** Latin America
- **Stack:** Kubernetes on two public clouds, upstream OpenTelemetry Collector, a backend built for OpenTelemetry, billed per record
- **OllyGarden:** Insights, telemetry assessments, and a six-week proof of value
- **Timeline:** November 2025 to September 2026

## Key numbers

- **24%** fewer log records (Target was 20%)
- **25%** fewer metric datapoints (Target was 10%)
- **61%** less volume on dedup-eligible logs (Counts preserved)
- **Zero** customer workload logs touched (Passed through one to one)

The bill grew with every pod. A Latin American B2B SaaS platform paid its observability backend for every log record and every metric datapoint, and nobody on the team could see, at the source, what was driving the volume. Six weeks later, log records reaching the backend were down 24% and metric datapoints about 25%, and not one customer workload log had been touched.

## A bill that grew with every pod

The company runs a multi-tenant platform that executes its customers' workloads on Kubernetes, across several production clusters on two public clouds. After moving away from a legacy APM vendor, it chose an observability SaaS built for OpenTelemetry that bills per ingested record. That model is fair, but unforgiving: every duplicate log line and every metric series nobody reads shows up on the invoice.

The platform team suspected waste. Nobody could say where the volume came from, and more than half of their metric data arrived without a service name, so they could not even say which team was paying for what.

> Once a log reaches the backend, the backend can't make it better. We want to fix telemetry where it's born.
>
> — Platform lead (paraphrased)

That conviction shaped the whole engagement. Rather than add more filtering in the backend, the team wanted control at the edge, in their own Collector pipelines, where they decide what leaves the cluster at all.

## One rule that could not bend

Most of the platform's log volume is customer workload output: the content of jobs that its customers run. That content may contain personal data, and for regulatory reasons it had to stay out of any reduction policy and out of any analysis. Whatever we changed, customer content had to pass through exactly as it came in.

The security review was just as strict. Work started in staging and pre-production, and production samples were shared only after the contract and the security review had cleared.

Then the team set the bar for a paid six-week proof of value, with pass/fail targets: at least 20% fewer log records and 10% fewer metric datapoints reaching the backend.

## What the first samples showed

OllyGarden Insights analyzed recurring telemetry samples, first from staging and then from production. The staging assessment, based on about 1.8 million records sampled over a week, was blunt:

- Metrics made up about 97% of all records, and more than half of the metric data had no service.name. Cost could not be attributed.
- About 89% of logs were duplicates. Deduplication alone could remove about 70% of log volume.
- About 29% of traces were single-span noise from cache health pings and liveness probes.
- No signal set deployment.environment.name, and SDK versions were fragmented across services.

Earlier in the pilot, something more urgent had already surfaced: passwords, usernames, API keys, and Authorization headers were leaking into logs from about five internal services. Insights flagged them, and the team fixed them at the source.

**Metrics that carry no signal.** About 57% of the Kubernetes volume-metric series came from static mounts: config maps, secrets, projected volumes, and the downward API. Those volumes never change size, and the same metric family was scraped three times more often than needed.

High-volume access logs from the API gateway and an HTTP trigger service, about 14% of log records, could become request-rate, error-rate, and latency metrics without losing anything an on-call engineer would look for.

## Six weeks, in waves

We delivered the fixes as numbered waves of drop-in Collector configurations. Every wave was replayed against real production samples before it went anywhere near production.

- Wave 1 cut the volume-metric scrape from three datapoints per minute to one.
- Wave 2 converted API gateway access logs into metrics.
- Wave 3 dropped the static-mount series and converted trigger access logs into metrics.
- Wave 5 added a log pipeline with template mining, attribute stripping, and count-preserving deduplication, which keeps how many times an event happened while dropping the copies.
- Wave 7 filtered pipeline self-noise on a newly onboarded cluster.

Customer workload content was routed around every one of these pipelines and passed through one to one.

## The day the replay paid for itself

Turning access logs into metrics is a well-known win, but it has a trap. One per-user path segment in the gateway routes was not templated, so a new log-to-metric rule copied real user email addresses into a metric route label. That was a privacy leak and an unbounded cardinality problem at the same time.

The next sample caught it. A fix shipped right away, and the following day's sample confirmed zero personal data in the metric. Without the replay-and-verify loop, those addresses would have sat in a backend that nobody was auditing for them.

## Noise that hid an outage

When a new cluster joined in June, about 78% of its log volume turned out to be the telemetry pipeline complaining about itself. An in-app agent could not reach its export endpoint and logged a stack trace on every pod, nonstop. Another 24% of the cluster's metrics was the logging stack observing itself.

The noise was hiding something worse: the cluster's application traces and metrics were not reaching the backend at all. Wave 7 removed about 70% of the cluster's log records and about 24% of its metrics, touched zero business logs, and made the real export failure impossible to miss.

## The results

In production, Wave 5 alone removed about 24% of log records overall and about 61% on the deduplication-eligible branch, consistently across two samples taken at midday and at midnight. Measured service by service, deduplication cut eligible log volume by 63% to 69%.

The proof of value closed at 24% fewer log records and a 24.6% average reduction in metric datapoints, against targets of 20% and 10%. The customer accepted it and kept OllyGarden on for continuous optimization. An engineering lead on their side called it a success for both teams.

## What happened next, honestly

A proof of value is a snapshot, and production keeps moving. After the six weeks, the production clusters moved off the vendor's operator and onto a self-managed upstream Collector, with OllyGarden checking each cutover.

Those checks found the hidden tax of a Collector migration. Service-name resolution had lived in the operator's configuration and was never carried over: after the cutover, 99.5% of logs and 77% of metrics reported service.name as unknown, which breaks dashboards and cost attribution alike.

Then, in August, the monthly backend bill still jumped by about a third, driven by incident-related metrics and deployments with hundreds to thousands of replicas. The analysis that followed showed that 10 metric groups made up about 48% of all metric volume, that self-monitoring was about 13% of metric datapoints, and that the median series sent about one datapoint per minute. The bill was driven by series count, not by scrape frequency.

So the relationship changed shape. Wave 8 moved 20 slow Kubernetes metric families to one datapoint per minute, added a selective log deduplication policy, and delivered a narrowly scoped email-redaction rule for cloud load-balancer logs. The team now brings OllyGarden in early whenever the bill starts to move, and the backlog is real: restoring service and cluster names, investigating single-span traces, and a filter that would remove about 43% of the in-scope logs.

## What other teams can take from this

- Decide what is off limits before you optimize. Routing customer content around every reduction made the rest of the work safe to do.
- Replay every change against real samples. That is how an email address in a metric label was caught within a day.
- Treat a Collector migration as a telemetry change. Logic that lived in a vendor operator, like service-name resolution, does not move by itself.
- Count series, not scrapes. Cardinality drives a per-datapoint bill far more than scrape intervals do.

## About these numbers

The proof-of-value results were measured on the pilot cluster with count-preserving pipeline measurements, not as a bill-to-bill comparison. Sample-based counts are subsampled, so this story reports ratios rather than absolute volumes. We do not publish dollar savings here: none have been confirmed, and the August bill increase shows why a single savings number would mislead.

## Listen to the story

Bianca and Florian, who you might know from the OTel Drops podcast, tell this story in a short conversation. The voices are AI-generated; the facts are the ones on this page.

[Episode (MP3)](https://ollygarden.com/audio/case-studies/b2b-saas-cuts-logs-and-metrics-at-the-source.mp3)

### Transcript

**Bianca:** Hi, I'm Bianca.

**Florian:** And I'm Florian. You might know us from the OTel Drops podcast, but today we're here to tell you a story.

**Bianca:** A real one. OllyGarden shared it with us, anonymized. So no company names and no people, but every number is real.

**Florian:** Here's the setup. A mid-size B2B SaaS platform in Latin America. Multi-tenant, running customer workloads on Kubernetes, across two public clouds.

**Bianca:** And an observability backend that bills per record. Every log line, every datapoint. So the bill grew with volume, and nobody could say what was driving it, because nobody had a view at the source.

**Florian:** What I like about this team is their platform lead's instinct. They didn't want more filtering in the backend. Their argument was simple: once a log reaches the backend, the backend can't make it better anymore. Fix it where it's born.

**Bianca:** That's the right instinct. But there was a hard constraint. Most of their log volume is customer workload content, the output of their customers' jobs. It can contain personal data.

**Florian:** So that content was off limits. Not reduced, not analyzed, not touched. Full stop.

**Bianca:** The security review was strict, too. They started in staging, and production data only flowed once the contract and the security approval were in place.

**Florian:** Then they set a bar. A six-week proof of value, with pass or fail targets. Twenty percent fewer log records, and ten percent fewer metric datapoints reaching the backend.

**Bianca:** So OllyGarden Insights started looking at samples. And the staging assessment was... a lot.

**Florian:** Metrics were about ninety-seven percent of all records. And more than half of the metric data had no service name. Just "unknown".

**Bianca:** Which means you can't attribute cost. You're paying for data, and you can't even say whose it is.

**Florian:** It gets better. About eighty-nine percent of the logs were duplicates.

**Bianca:** Eighty-nine?!

**Florian:** And almost a third of the traces were single spans. Cache health pings, liveness probes. Noise.

**Bianca:** And production?

**Florian:** Production had a Kubernetes detail I love. More than half of the volume metric series came from static mounts. Config maps, secrets, projected volumes. They never change size.

**Bianca:** So those series carry no signal at all.

**Florian:** None. And they were scraped three times more often than needed.

**Bianca:** So how do you fix all of that without breaking anything? OllyGarden shipped the work in numbered waves of drop-in OpenTelemetry Collector configurations.

**Florian:** And here's the part the operator in me appreciates. Every wave was replayed against real samples before it went anywhere near production.

**Bianca:** One wave cut those volume metrics from three datapoints a minute to one. Another turned high-volume access logs into request rate, error rate, and latency metrics.

**Florian:** And then the replays earned their keep. One of the new log-to-metric rules put real user email addresses into a metric label.

**Bianca:** Oof. How?

**Florian:** One per-user path segment wasn't templated. So every email became a label value. A privacy problem and a cardinality explosion, in one line of config.

**Bianca:** And the next sample caught it. A fix shipped, and the day after, zero personal data.

**Florian:** Then a new cluster came online. About seventy-eight percent of its log volume was the pipeline complaining about itself.

**Bianca:** Complaining about what?

**Florian:** An in-app agent couldn't reach its export endpoint, so it logged a stack trace on every pod, nonstop. And that noise was hiding a real outage. That cluster's traces and metrics weren't reaching the backend at all.

**Bianca:** That's the thing about noise. It doesn't just cost money. It hides the signal you actually need.

**Bianca:** So, the six weeks. Did they hit the targets?

**Florian:** They did. Twenty-four percent fewer log records, against a target of twenty. And about twenty-five percent fewer metric datapoints, against a target of ten.

**Bianca:** And the log deduplication is my favorite part. It's count-preserving, so you still know how many times something happened. On the eligible logs, it collapsed volume by about sixty-one percent.

**Florian:** And the customer workload logs?

**Bianca:** Passed through one to one. Untouched. Exactly as promised.

**Florian:** The customer accepted the proof of value, and the work kept going.

**Bianca:** Now, this is the part I respect most about this story. It doesn't end with a perfect bill.

**Florian:** No. After the proof of value, the clusters moved from the vendor's operator to a self-managed Collector. And service name resolution silently disappeared.

**Bianca:** How bad?

**Florian:** Ninety-nine and a half percent of logs, and seventy-seven percent of metrics, showed service name unknown. That logic lived in the operator's config, and nobody carried it over.

**Bianca:** The hidden tax of a Collector migration.

**Florian:** And in August, the monthly bill still jumped by about a third. Incident-driven metrics, and deployments with hundreds to thousands of replicas.

**Bianca:** Which is the real lesson. The cost driver was series count. Cardinality. Not scrape frequency.

**Florian:** So it stopped being a one-time project. The team now brings OllyGarden in early, whenever the bill starts to move.

**Bianca:** So, what would I take from this? First, look for telemetry that never changes. Static series and repeated logs are the cheapest wins you'll find.

**Florian:** Second, replay every change against real data before rollout. That's how you catch an email address in a metric label.

**Bianca:** Third, protect your customers' data by design. Decide what's off limits first, then optimize everything else.

**Florian:** And fourth, telemetry quality is a habit, not a project.

**Bianca:** The full written story, with all the numbers, is on the OllyGarden website. Thanks for listening.

**Florian:** Until next time.