Back to Blog
October 3, 2026Author: Juraci Paixão Kröhling

Where Sensitive Data Hides in OpenTelemetry® Pipelines, and How to Find It

  • Security
  • PII
  • OpenTelemetry
  • Privacy

Trust store passwords in resource attributes, OAuth tokens in query strings, tax IDs in metrics: what we found in real pipelines, why tools cause most of it, and what privacy laws allow.

Cover illustration for “Where Sensitive Data Hides in OpenTelemetry Pipelines, and How to Find It”

The evening we first ran our personal-data detector against a customer's telemetry, I sent their platform lead a short message: "I have your trust store password. So does SiloLogic." Nobody at that bank had typed the password into a log statement. A Java process had started with the password as a command-line flag, the OpenTelemetry Java agent had copied the full command line into a resource attribute, and from there it rode along on every span the service produced, all the way into SiloLogic, the bank's observability vendor. The story is real, but we anonymized it: SiloLogic is a fictional name, and we left out every detail that could identify the bank.

Sensitive data in OpenTelemetry pipelines is more frequent than you might think, and more diverse in shape than you can imagine. Sensitive data in telemetry means any personal data, credential, or secret that ends up in a span, metric, or log record: a customer's email address, a national tax ID, an OAuth access token, a database password. Since then, we have found it at banks, fintechs, health platforms, and IoT companies, in traces, metrics, and logs alike. This post walks through what we found, why most of it is caused by misconfigured tools rather than careless developers, what the main data privacy frameworks allow and prohibit, and where to look in your own telemetry.

What we found in real OpenTelemetry pipelines

The table below lists findings from organizations we work with. We anonymized every one of them, and most came from staging or development environments, which matters for reasons we cover later.

What leaked Where it showed up What caused it
Java trust store passwords process.command_line resource attribute on every span Password passed as a -D flag, captured by the Java agent's process resource detector
OAuth authorization codes, access tokens, and JSON Web Tokens (JWTs) url.query and url.full span attributes, sometimes URL-encoded twice HTTP auto-instrumentation records the query string by default
Authorization headers with bearer tokens Log records and span attributes HTTP header capture enabled for all headers
Brazilian tax IDs: individual (CPF) and company (CNPJ) numbers Log messages, url.full, and metric attributes Developer log statements and identifiers in REST paths
Email addresses Metric attributes, span attributes, and logs Identifiers chosen as dimensions, logged objects
Names, emails, phone numbers, and street addresses of clinicians Log records An object-relational mapper (ORM) configured to log every SQL statement with its bound values
Date of birth A span attribute Request object recorded as attributes
Customer phone numbers Prometheus metric labels The phone number doubled as a money-transfer key
Wi-Fi passwords Span attributes A decorator that recorded every method parameter
Database connection strings with embedded credentials A span attribute on hundreds of thousands of spans Client library instrumentation recording the connection string
Signed storage URLs Span attributes The signature in the URL granted 15 minutes of access to signed contracts
Japanese tax numbers Logs Regional identifiers no existing redaction rule covered

Two things stand out. The first is how often the leak sits in a field nobody thinks about. Teams that review their log statements carefully never look at process.command_line, because nobody wrote it. The second is shape. A Brazilian CPF looks nothing like a JWT, which looks nothing like a street address, which looks nothing like a Wi-Fi password. A single regular expression, or a single list of forbidden attribute names, catches a fraction of it.

A password nobody logged

The trust store password is the story we tell most often, because we have seen it at more than one company. Java applications that talk to services over mutual Transport Layer Security (TLS) need a trust store, and the quickest way to configure one is -Djavax.net.ssl.trustStorePassword=... on the command line. The OpenTelemetry Java agent, like other process resource detectors, can record the full command line as process.command_line or process.command_args. Resource attributes are attached to every span, metric, and log the process emits, so one flag turns into millions of copies of the password in the backend.

Someone might argue that anyone who can read a process's command line already owns the host. That is not true once the command line leaves the host. Telemetry backends have far broader access than production machines: developers, support engineers, vendors, and now AI agents can query them.

Tokens in the query string

Our most common finding is an OAuth token or authorization code in url.query. Every OAuth and OpenID Connect flow ends with a redirect back to the application, and that redirect often carries a code, an access_token, or an id_token in the query string. HTTP auto-instrumentation records the query string by default. The semantic conventions do ask instrumentations to redact query parameters by default, but the list covers signed cloud storage URLs: AWSAccessKeyId, Signature, sig, and X-Goog-Signature. An OAuth code is not on it.

The sneakiest variant we found was a token inside a redirect parameter, URL-encoded inside another URL-encoded value. To read it, you have to decode the attribute twice. Most users we show this to do not remember that their login flow involves an OAuth exchange at all.

Most leaks come from tools, not careless developers

When teams find personal data in their telemetry, the first reaction is often to remind developers to be careful. In our experience, that is aimed at the wrong target. Most of what we find is a tool doing what it was configured to do.

The Java agent records the command line because a resource detector is enabled. HTTP instrumentation records the query string because that is the default. At a European digital-health platform, one shared ORM setting logged every SQL query with its values at the INFO level, which produced about 100,000 log lines a week carrying clinicians' personal data from a single service; we wrote about it in a case study. Header capture configured with a wildcard records Authorization. A log bridge that serializes whole request objects records whatever those objects contain. None of these required a developer to write the word "password" anywhere.

Developers do make mistakes as well. We found a CPF in a log message that announced a successful "CPF validation", Wi-Fi passwords recorded by a decorator that logged every method parameter, a customer email three levels deep inside a data transfer object whose toString method was logged, and a redaction rule that masked the field the team was worried about while an API key in the same log line went through untouched. These are real, but they are the minority. Fixing the shared setting removes far more personal data than asking each developer to review each log line.

Why your SIEM did not catch it

Some of these findings came from organizations that pay good money for specialized security information and event management (SIEM) tools, with teams whose job is finding this kind of leak. The tools were not broken. They were looking in places where OpenTelemetry data does not look like what they expect.

SIEM tools are built around text logs. OpenTelemetry data travels as OpenTelemetry Protocol (OTLP) messages, usually Protocol Buffers, and it is structured: a value can sit in a resource attribute, a span attribute, an event, a metric data point attribute, or a log body. A SIEM rule that greps for password= never sees a resource attribute that was never written to a log file, and it does not decode a token that has been URL-encoded twice inside url.full. It also does not know that url.query is where OAuth redirects end up, or that process.command_args exists.

We found these leaks because we know OpenTelemetry deeply. We know which instrumentation libraries record which attributes by default, which semantic conventions carry user input, and which resource detectors copy host details into every signal. Knowing where to look matters more than the sophistication of the pattern matching. That knowledge is what the Suspicious PII Leakage insight in OllyGarden Insights encodes: it checks attribute names and values across traces and logs, URL query strings included.

Where to look in your own telemetry

If you want to start a review today, these are the places that produced the most findings for us:

  • process.command_line and process.command_args resource attributes, especially on Java services
  • url.query and url.full span attributes, decoded at least twice
  • http.request.header.* and http.response.header.* attributes, especially authorization, cookie, and set-cookie
  • db.query.text (formerly db.statement) and ORM or driver logs that include bound parameter values
  • exception.message and exception.stacktrace, which often embed the input that caused the error
  • Log bodies from frameworks that serialize whole objects, such as request bodies, DTOs, and HTTP client debug logs
  • Metric attributes, where an email address or tax ID also causes a cardinality problem
  • Messaging payloads and generative AI prompts recorded as span attributes or events

Attribute names help, but values matter more. A key called customer.reference can hold an email address, and a key called id can hold a tax ID.

What data privacy frameworks allow and prohibit

Privacy law does not ban personal data in telemetry. It requires that every piece of personal data you process has a reason, and that you process no more than that reason needs. The frameworks below are the ones our customers ask about most. This is a practitioner's summary, not legal advice; your privacy officer or counsel makes the call for your organization.

Framework Where it applies What it says about data like this
General Data Protection Regulation (GDPR) European Union and European Economic Area Personal data needs a lawful basis (Article 6), must be limited to what is necessary (data minimisation, Article 5), kept no longer than needed, and secured. Health data is a special category (Article 9). Breaches must usually be reported within 72 hours (Article 33).
UK GDPR and Data Protection Act 2018 United Kingdom Keeps the EU GDPR's principles after Brexit: lawful basis, data minimisation, storage limitation, and security. Breaches go to the Information Commissioner's Office (ICO) within 72 hours. Fines reach £17.5 million or 4% of worldwide turnover, whichever is higher.
Privacy Act 1988 and the Australian Privacy Principles (APPs) Australia Collect personal information only when reasonably necessary for your functions (APP 3), protect it and destroy or de-identify it when no longer needed (APP 11). Tax file numbers have their own rule that limits their use to tax purposes. The Notifiable Data Breaches scheme requires an assessment within 30 days. Penalties for serious interferences reach the greater of 50 million Australian dollars, three times the benefit gained, or 30% of adjusted turnover.
Lei Geral de Proteção de Dados (LGPD) Brazil Modeled on the GDPR, with similar legal bases and necessity principle. A CPF is personal data. Fines reach 2% of revenue in Brazil, up to 50 million reais per infraction.
California Consumer Privacy Act (CCPA), as amended by the CPRA California IP addresses and email addresses are personal information. Social Security numbers and account logins with passwords are "sensitive personal information". Collection must be reasonably necessary and proportionate to its purpose.
Health Insurance Portability and Accountability Act (HIPAA) United States health care Protected health information includes names, dates of birth, email addresses, and IP addresses linked to health data. The minimum necessary standard applies, and a vendor that stores it needs a business associate agreement.
Payment Card Industry Data Security Standard (PCI DSS) 4.0 Anyone handling payment cards Card verification codes and other sensitive authentication data must never be stored after authorization. Card numbers must be unreadable wherever they are stored, and a telemetry backend that holds them is in scope for the audit.

Credentials such as passwords, tokens, and connection strings are a separate category. They are rarely personal data, but every security standard and audit, from ISO 27001 to SOC 2, expects you to keep them out of logs. They are also the leak an attacker can use immediately.

Client IP addresses are fine, until you can tie them to a user

IP addresses are the most common debate we have with customers. Under the GDPR, and the UK GDPR that mirrors it, an IP address is usually personal data: the Court of Justice of the European Union ruled in 2016 (Breyer, C-582/14) that even a dynamic IP address is personal data when the site operator has legal means to identify the person behind it. The same regulation also recognizes network and information security as a legitimate interest (Recital 49). Keeping client.address on a span to investigate abuse, rate limiting, or an attack is usually easy to justify, especially with a short retention period.

The picture changes when the same span, or the same trace, also carries enduser.id, an email address, or an account number. At that point the telemetry is a record of which known person did what, from where, and when. That record needs its own justification, and it is a much harder one to write.

Every piece of personal data needs a justification

Frameworks allow some personal data when it is necessary for the business operation. A payment service may need an account identifier to trace a failed transfer. A support team may need a customer ID to find a user's request. What they do not allow is personal data because it was convenient, or because nobody turned the default off. In practice, this means each personal attribute in your telemetry should have an answer to three questions: why it is there, who can see it, and how long it stays. If nobody can answer them, the attribute should go.

There is also a timing aspect. A leak you do not know about is a problem. A leak you do know about, and leave in place, is a decision, and auditors treat it as one. Finding sensitive data creates an obligation to act on it.

We rarely see real sensitive data, and we do not need to

There is an irony in this work. We have found sensitive data at almost every organization we have worked with, but we have seen very little real sensitive data ourselves. We are happy to analyze staging data. We do not need production telemetry to find a production leak.

If it looks like sensitive data, it might be sensitive data. A span attribute named customer.name with the value "Jane Doe" is clearly test data. The application that records it, though, will record real customer names the moment it runs in production. The same code path, the same instrumentation, and the same defaults produce the same attribute. A trust store password from a staging environment usually means the production deployment passes its own password the same way.

This also means detection has to tolerate false positives. In one assessment, millisecond timestamps looked like phone numbers, and cloud service-account addresses looked like personal email addresses. Every hit needs a human or an agent to confirm it before anyone acts on it.

Summary

Sensitive data in OpenTelemetry pipelines is common and comes in many shapes: trust store passwords in resource attributes, OAuth tokens in query strings, tax IDs and email addresses in metric attributes, and personal data in SQL logs. Most of it comes from tools doing what they were configured to do, such as auto-instrumentation defaults, ORM logging, and header capture, and a smaller share from developer mistakes. SIEM tools tend to miss it because they read text logs, not structured OTLP data. Privacy frameworks such as the GDPR, the UK GDPR, the Australian Privacy Act, LGPD, CCPA, HIPAA, and PCI DSS allow personal data only when it is necessary and justified, and an IP address that can be tied to a user ID needs a stronger justification than one that cannot. Staging data is enough to find these leaks, because the same code records the same fields in production.

Other Posts

Keep up to date with OllyGarden news and insights.

Check Blog
Cover illustration for “Observability for Humans and Agents in the AI Age”
Sep 28, 2026Authors: Nicolas Wörner and Severin Neumann

Observability for Humans and Agents in the AI Age

OllyGarden and Bronto on where AI agents help with instrumentation and analysis today, and why they only work when the telemetry is intentional and reachable.

  • AI
  • Instrumentation
  • Observability
Read Article
Cover illustration for “What's changed about OpenTelemetry vendor lock-in since 2024”
Aug 21, 2026Author: Juraci Paixão Kröhling

What's changed about OpenTelemetry vendor lock-in since 2024

OpenTelemetry still does not make dashboards portable, but agents, CLIs, and MCP servers have made switching observability backends noticeably less painful.

  • OpenTelemetry
  • Observability
  • Vendor Lock-in
Read Article
Cover illustration for “The Scrape Interval Nobody Chose”
Jul 31, 2026Author: Juraci Paixão Kröhling

The Scrape Interval Nobody Chose

A wrong metric scrape interval inflates ingest, egress, and compute. Learn how to measure DPM and match intervals to the decisions each metric supports.

  • Metrics
  • OpenTelemetry
  • Observability
Read Article

Keep your observability backendSend it better telemetry

Get Started

You’re Heading to the App

Create your OllyGarden account over at ollygarden.app. You can get started for free, no credit card needed.