Case study · Enterprise software · North America ·

Nearly half the spans came from a service called cli: what we found at a 500-service enterprise platform

OllyGarden Insights looked at staging OpenTelemetry® data from about 500 services at a large enterprise software company. One unnamed worker made 45% of all spans, and most logs were repeats.

  • 45%

    of all spans came from one unnamed worker

    Reported as cli

  • 32%

    of log records were exact duplicates

    About 60% reducible to patterns

  • 4.5%

    of logs were the Collector re-shipping its own output

    Every week

  • 9%

    of services set a deployment environment

    Across about 500 services

Listen to the story

Bianca and Florian, who you might know from the OTel Drops podcast, tell this story in a short conversation. The voices are AI-generated; the facts are the ones on this page.

Read the transcript

Bianca: Hi, I'm Bianca.

Florian: And I'm Florian. You might know us from the OTel Drops podcast. Today we're telling another real story from OllyGarden's work.

Bianca: Anonymized, so no names. And it's about what OllyGarden found in one company's telemetry.

Florian: Here's the setup. A large enterprise software company. Mostly Java, plus Python, Node.js, and PHP, on a multi-region internal Kubernetes platform.

Bianca: And one central team paying for all the telemetry, with no chargeback. OllyGarden Insights looked at their staging telemetry, across about five hundred services.

Florian: So, the biggest finding. Forty-five percent of all spans came from one service.

Bianca: Which one?

Florian: It was called C-L-I. The tracer's default name. A PHP background worker on two clusters.

Bianca: So the biggest trace producer in the whole fleet had no real identity. Nobody could tell from the telemetry who owned it.

Florian: And its traces showed why it was so busy. Every queued message made about fourteen spans. Including a brand new database connection, for every single message.

Bianca: Plus about six cache loads and four queries. Reuse the connection, and you cut latency and span volume at the same time.

Florian: And the spans were concentrated. Five services on a vendor tracer made forty-seven percent of all spans. The three hundred forty-three services on OpenTelemetry SDKs made thirty-eight.

Bianca: Now the logs. Thirty-two percent of log records were exact duplicates.

Florian: And once you normalize the IDs and numbers, about sixty percent collapse into repeating patterns.

Bianca: About a third of all logs were HTTP access logs. One service's ingress logs alone were sixteen and a half percent.

Florian: Request rate, errors, and latency by route. That's a metric. It's always been a metric.

Bianca: And thirteen percent were health checks, readiness checks, and pings.

Florian: Here's my favorite one. Every week, four and a half percent of all logs were the Collector printing records it had already shipped.

Bianca: And then?

Florian: And then those got collected and shipped again. The more it exports, the more it logs.

Bianca: A Collector that talks about its work, and then ships the transcript.

Florian: Then attribution. Twenty-six percent of log records had no service name at all.

Bianca: So no team to charge it back to. And forty-three percent of log bodies were unparsed JSON. INFO was spelled three different ways.

Florian: Only nine percent of services set a deployment environment.

Bianca: And thirty-five percent of traces were single spans. Only about five percent crossed a service boundary.

Florian: There were also pattern matches for bearer tokens and email addresses in a few services. Those still need a human review, so treat them as suspects for now.

Bianca: So, what would you check in your own telemetry after hearing this?

Florian: First, give every process a real service name. Defaults hide your biggest producers.

Bianca: Second, look at what your background workers do per message.

Florian: Third, turn access logs into metrics. And fourth, check whether your Collector is shipping its own logs twice.

Bianca: The full written story is on the OllyGarden website. Thanks for listening.

Florian: Until next time.

Across about 500 services at a large enterprise software company, one background worker with a default name produced 45% of all spans, a third of the log records were exact duplicates, and the Collector was logging its own exports and shipping them again.

The setup

The company runs a polyglot fleet, mostly Java with Python, Node.js, and PHP, on a multi-region internal Kubernetes platform. Telemetry flows through an in-house Collector, from a mix of OpenTelemetry SDKs and a vendor tracer. A central team pays for all of it, with no chargeback to the teams that produce it.

OllyGarden Insights analyzed the platform's staging telemetry in September 2026. This story covers what it showed.

The biggest producer had no name

45% of all spans came from a single PHP background worker running on two clusters. It reported under its tracer's default service name, cli. The biggest trace producer in the fleet had no real identity, so the telemetry could not say who owned it.

Its traces also showed why it was so busy. Each queued message produced about 14 spans: a new database connection for every message, about six cache loads, and about four queries. Reusing connections would cut both latency and span volume.

Span volume was concentrated in a few places. 5 services on a vendor tracer produced 47% of all spans, while 343 services on OpenTelemetry SDKs produced 38%.

Most logs were repeats

  • 32% of log records were exact duplicates, and about 60% were reducible once IDs and numbers were normalized into patterns.
  • About a third of logs were HTTP access logs. One service's ingress-request logs alone were 16.5% of all logs. Request rate, errors, and latency by route are better kept as metrics.
  • 13% of logs were health, readiness, or ping checks.
  • Stored-procedure timing messages, "executed in N ms", were metrics written as logs.

A Collector shipping its own logs

Every week, 4.5% of all logs were the Collector printing log records it had already shipped, which were then collected and shipped again. That loop grows with traffic: the more the Collector exports, the more it logs.

Logs nobody could attribute

26% of log records had no service name, so they could not be attributed to a team or charged back. 43% of log bodies were unparsed JSON, 16% carried no severity, and INFO was spelled three different ways.

Pattern detectors also matched a bearer-token format in about 17% of log records across 3 services, and email addresses in a few hundred records across 3 services. These are unverified pattern matches: the values were not inspected, and a human review has to confirm them.

Traces and metadata

  • 35% of traces were single-span, and only about 5% crossed a service boundary. The largest trace held about 11,600 spans.
  • Only 9% of services set a deployment environment, and 56% set service.instance.id.
  • kubelet and kube-state-metrics produced 48% of metric datapoints, data that cluster monitoring already collects.

What to check in your own telemetry

  • Give every process a real service name. Tracer defaults like cli hide your biggest producers.
  • Look at the per-message work in background workers. A new database connection per message costs latency and spans.
  • Turn access logs into request, error, and latency metrics.
  • Check whether your Collector's own logs are collected and shipped again.

Want a story like thisfor your telemetry?

Send us a sample and we will show you where your volume, cost, and risk really come from.