Where Sensitive Data Hides in OpenTelemetry® Pipelines, and How to Find It
- Security
- PII
- OpenTelemetry
- Privacy
Trust store passwords in resource attributes, OAuth tokens in query strings, tax IDs in metrics: what we found in real pipelines, why tools cause most of it, and what privacy laws allow.

The evening we first ran our personal-data detector against a customer's telemetry, I sent their platform lead a short message: "I have your trust store password. So does SiloLogic." Nobody at that bank had typed the password into a log statement. A Java process had started with the password as a command-line flag, the OpenTelemetry Java agent had copied the full command line into a resource attribute, and from there it rode along on every span the service produced, all the way into SiloLogic, the bank's observability vendor. The story is real, but we anonymized it: SiloLogic is a fictional name, and we left out every detail that could identify the bank.
Sensitive data in OpenTelemetry pipelines is more frequent than you might think, and more diverse in shape than you can imagine. Sensitive data in telemetry means any personal data, credential, or secret that ends up in a span, metric, or log record: a customer's email address, a national tax ID, an OAuth access token, a database password. Since then, we have found it at banks, fintechs, health platforms, and IoT companies, in traces, metrics, and logs alike. This post walks through what we found, why most of it is caused by misconfigured tools rather than careless developers, what the main data privacy frameworks allow and prohibit, and where to look in your own telemetry.
What we found in real OpenTelemetry pipelines
The table below lists findings from organizations we work with. We anonymized every one of them, and most came from staging or development environments, which matters for reasons we cover later.
| What leaked | Where it showed up | What caused it |
|---|---|---|
| Java trust store passwords | process.command_line resource attribute on every span |
Password passed as a -D flag, captured by the Java agent's process resource detector |
| OAuth authorization codes, access tokens, and JSON Web Tokens (JWTs) | url.query and url.full span attributes, sometimes URL-encoded twice |
HTTP auto-instrumentation records the query string by default |
| Authorization headers with bearer tokens | Log records and span attributes | HTTP header capture enabled for all headers |
| Brazilian tax IDs: individual (CPF) and company (CNPJ) numbers | Log messages, url.full, and metric attributes |
Developer log statements and identifiers in REST paths |
| Email addresses | Metric attributes, span attributes, and logs | Identifiers chosen as dimensions, logged objects |
| Names, emails, phone numbers, and street addresses of clinicians | Log records | An object-relational mapper (ORM) configured to log every SQL statement with its bound values |
| Date of birth | A span attribute | Request object recorded as attributes |
| Customer phone numbers | Prometheus metric labels | The phone number doubled as a money-transfer key |
| Wi-Fi passwords | Span attributes | A decorator that recorded every method parameter |
| Database connection strings with embedded credentials | A span attribute on hundreds of thousands of spans | Client library instrumentation recording the connection string |
| Signed storage URLs | Span attributes | The signature in the URL granted 15 minutes of access to signed contracts |
| Japanese tax numbers | Logs | Regional identifiers no existing redaction rule covered |
Two things stand out. The first is how often the leak sits in a field nobody thinks about. Teams that review their log statements carefully never look at process.command_line, because nobody wrote it. The second is shape. A Brazilian CPF looks nothing like a JWT, which looks nothing like a street address, which looks nothing like a Wi-Fi password. A single regular expression, or a single list of forbidden attribute names, catches a fraction of it.
A password nobody logged
The trust store password is the story we tell most often, because we have seen it at more than one company. Java applications that talk to services over mutual Transport Layer Security (TLS) need a trust store, and the quickest way to configure one is -Djavax.net.ssl.trustStorePassword=... on the command line. The OpenTelemetry Java agent, like other process resource detectors, can record the full command line as process.command_line or process.command_args. Resource attributes are attached to every span, metric, and log the process emits, so one flag turns into millions of copies of the password in the backend.
Someone might argue that anyone who can read a process's command line already owns the host. That is not true once the command line leaves the host. Telemetry backends have far broader access than production machines: developers, support engineers, vendors, and now AI agents can query them.
Tokens in the query string
Our most common finding is an OAuth token or authorization code in url.query. Every OAuth and OpenID Connect flow ends with a redirect back to the application, and that redirect often carries a code, an access_token, or an id_token in the query string. HTTP auto-instrumentation records the query string by default. The semantic conventions do ask instrumentations to redact query parameters by default, but the list covers signed cloud storage URLs: AWSAccessKeyId, Signature, sig, and X-Goog-Signature. An OAuth code is not on it.
The sneakiest variant we found was a token inside a redirect parameter, URL-encoded inside another URL-encoded value. To read it, you have to decode the attribute twice. Most users we show this to do not remember that their login flow involves an OAuth exchange at all.
Most leaks come from tools, not careless developers
When teams find personal data in their telemetry, the first reaction is often to remind developers to be careful. In our experience, that is aimed at the wrong target. Most of what we find is a tool doing what it was configured to do.
The Java agent records the command line because a resource detector is enabled. HTTP instrumentation records the query string because that is the default. At a European digital-health platform, one shared ORM setting logged every SQL query with its values at the INFO level, which produced about 100,000 log lines a week carrying clinicians' personal data from a single service; we wrote about it in a case study. Header capture configured with a wildcard records Authorization. A log bridge that serializes whole request objects records whatever those objects contain. None of these required a developer to write the word "password" anywhere.
Developers do make mistakes as well. We found a CPF in a log message that announced a successful "CPF validation", Wi-Fi passwords recorded by a decorator that logged every method parameter, a customer email three levels deep inside a data transfer object whose toString method was logged, and a redaction rule that masked the field the team was worried about while an API key in the same log line went through untouched. These are real, but they are the minority. Fixing the shared setting removes far more personal data than asking each developer to review each log line.
Why your SIEM did not catch it
Some of these findings came from organizations that pay good money for specialized security information and event management (SIEM) tools, with teams whose job is finding this kind of leak. The tools were not broken. They were looking in places where OpenTelemetry data does not look like what they expect.
SIEM tools are built around text logs. OpenTelemetry data travels as OpenTelemetry Protocol (OTLP) messages, usually Protocol Buffers, and it is structured: a value can sit in a resource attribute, a span attribute, an event, a metric data point attribute, or a log body. A SIEM rule that greps for password= never sees a resource attribute that was never written to a log file, and it does not decode a token that has been URL-encoded twice inside url.full. It also does not know that url.query is where OAuth redirects end up, or that process.command_args exists.
We found these leaks because we know OpenTelemetry deeply. We know which instrumentation libraries record which attributes by default, which semantic conventions carry user input, and which resource detectors copy host details into every signal. Knowing where to look matters more than the sophistication of the pattern matching. That knowledge is what the Suspicious PII Leakage insight in OllyGarden Insights encodes: it checks attribute names and values across traces and logs, URL query strings included.
Where to look in your own telemetry
If you want to start a review today, these are the places that produced the most findings for us:
process.command_lineandprocess.command_argsresource attributes, especially on Java servicesurl.queryandurl.fullspan attributes, decoded at least twicehttp.request.header.*andhttp.response.header.*attributes, especiallyauthorization,cookie, andset-cookiedb.query.text(formerlydb.statement) and ORM or driver logs that include bound parameter valuesexception.messageandexception.stacktrace, which often embed the input that caused the error- Log bodies from frameworks that serialize whole objects, such as request bodies, DTOs, and HTTP client debug logs
- Metric attributes, where an email address or tax ID also causes a cardinality problem
- Messaging payloads and generative AI prompts recorded as span attributes or events
Attribute names help, but values matter more. A key called customer.reference can hold an email address, and a key called id can hold a tax ID.
What data privacy frameworks allow and prohibit
Privacy law does not ban personal data in telemetry. It requires that every piece of personal data you process has a reason, and that you process no more than that reason needs. The frameworks below are the ones our customers ask about most. This is a practitioner's summary, not legal advice; your privacy officer or counsel makes the call for your organization.
| Framework | Where it applies | What it says about data like this |
|---|---|---|
| General Data Protection Regulation (GDPR) | European Union and European Economic Area | Personal data needs a lawful basis (Article 6), must be limited to what is necessary (data minimisation, Article 5), kept no longer than needed, and secured. Health data is a special category (Article 9). Breaches must usually be reported within 72 hours (Article 33). |
| UK GDPR and Data Protection Act 2018 | United Kingdom | Keeps the EU GDPR's principles after Brexit: lawful basis, data minimisation, storage limitation, and security. Breaches go to the Information Commissioner's Office (ICO) within 72 hours. Fines reach £17.5 million or 4% of worldwide turnover, whichever is higher. |
| Privacy Act 1988 and the Australian Privacy Principles (APPs) | Australia | Collect personal information only when reasonably necessary for your functions (APP 3), protect it and destroy or de-identify it when no longer needed (APP 11). Tax file numbers have their own rule that limits their use to tax purposes. The Notifiable Data Breaches scheme requires an assessment within 30 days. Penalties for serious interferences reach the greater of 50 million Australian dollars, three times the benefit gained, or 30% of adjusted turnover. |
| Lei Geral de Proteção de Dados (LGPD) | Brazil | Modeled on the GDPR, with similar legal bases and necessity principle. A CPF is personal data. Fines reach 2% of revenue in Brazil, up to 50 million reais per infraction. |
| California Consumer Privacy Act (CCPA), as amended by the CPRA | California | IP addresses and email addresses are personal information. Social Security numbers and account logins with passwords are "sensitive personal information". Collection must be reasonably necessary and proportionate to its purpose. |
| Health Insurance Portability and Accountability Act (HIPAA) | United States health care | Protected health information includes names, dates of birth, email addresses, and IP addresses linked to health data. The minimum necessary standard applies, and a vendor that stores it needs a business associate agreement. |
| Payment Card Industry Data Security Standard (PCI DSS) 4.0 | Anyone handling payment cards | Card verification codes and other sensitive authentication data must never be stored after authorization. Card numbers must be unreadable wherever they are stored, and a telemetry backend that holds them is in scope for the audit. |
Credentials such as passwords, tokens, and connection strings are a separate category. They are rarely personal data, but every security standard and audit, from ISO 27001 to SOC 2, expects you to keep them out of logs. They are also the leak an attacker can use immediately.
Client IP addresses are fine, until you can tie them to a user
IP addresses are the most common debate we have with customers. Under the GDPR, and the UK GDPR that mirrors it, an IP address is usually personal data: the Court of Justice of the European Union ruled in 2016 (Breyer, C-582/14) that even a dynamic IP address is personal data when the site operator has legal means to identify the person behind it. The same regulation also recognizes network and information security as a legitimate interest (Recital 49). Keeping client.address on a span to investigate abuse, rate limiting, or an attack is usually easy to justify, especially with a short retention period.
The picture changes when the same span, or the same trace, also carries enduser.id, an email address, or an account number. At that point the telemetry is a record of which known person did what, from where, and when. That record needs its own justification, and it is a much harder one to write.
Every piece of personal data needs a justification
Frameworks allow some personal data when it is necessary for the business operation. A payment service may need an account identifier to trace a failed transfer. A support team may need a customer ID to find a user's request. What they do not allow is personal data because it was convenient, or because nobody turned the default off. In practice, this means each personal attribute in your telemetry should have an answer to three questions: why it is there, who can see it, and how long it stays. If nobody can answer them, the attribute should go.
There is also a timing aspect. A leak you do not know about is a problem. A leak you do know about, and leave in place, is a decision, and auditors treat it as one. Finding sensitive data creates an obligation to act on it.
We rarely see real sensitive data, and we do not need to
There is an irony in this work. We have found sensitive data at almost every organization we have worked with, but we have seen very little real sensitive data ourselves. We are happy to analyze staging data. We do not need production telemetry to find a production leak.
If it looks like sensitive data, it might be sensitive data. A span attribute named customer.name with the value "Jane Doe" is clearly test data. The application that records it, though, will record real customer names the moment it runs in production. The same code path, the same instrumentation, and the same defaults produce the same attribute. A trust store password from a staging environment usually means the production deployment passes its own password the same way.
This also means detection has to tolerate false positives. In one assessment, millisecond timestamps looked like phone numbers, and cloud service-account addresses looked like personal email addresses. Every hit needs a human or an agent to confirm it before anyone acts on it.
Summary
Sensitive data in OpenTelemetry pipelines is common and comes in many shapes: trust store passwords in resource attributes, OAuth tokens in query strings, tax IDs and email addresses in metric attributes, and personal data in SQL logs. Most of it comes from tools doing what they were configured to do, such as auto-instrumentation defaults, ORM logging, and header capture, and a smaller share from developer mistakes. SIEM tools tend to miss it because they read text logs, not structured OTLP data. Privacy frameworks such as the GDPR, the UK GDPR, the Australian Privacy Act, LGPD, CCPA, HIPAA, and PCI DSS allow personal data only when it is necessary and justified, and an IP address that can be tied to a user ID needs a stronger justification than one that cannot. Staging data is enough to find these leaks, because the same code records the same fields in production.






