# The PII leak was a logger setting, not a developer problem: what we found at a European digital-health platform

OllyGarden Insights and Rose looked at staging OpenTelemetry data and code from a European digital-health platform. Most of the personal data in its logs came from one shared ORM setting.

> **Anonymized, but real.** This story comes from a real OllyGarden customer engagement. To protect the customer, we removed their name, people, products, and anything else that could identify them, and rounded some numbers. Everything else happened as described.

Published: 2026-10-03
Canonical URL: https://ollygarden.com/resources/case-studies/digital-health-pii-leak-was-a-logger-setting

## At a glance

- **Company:** Small-to-mid-size digital-health platform with a small engineering team
- **Region:** Europe
- **Stack:** Node.js and TypeScript microservices with shared OpenTelemetry auto-instrumentation, a self-hosted open-source observability stack
- **Scale:** 50 to 120 services; in staging, about 15 million log records and 28 million metric datapoints a day
- **OllyGarden:** Insights, with weekly written assessments, and Rose reviewing the code
- **Timeline:** April to June 2026

## Key numbers

- **17%** of services flagged for personal data in week one (12 of about 70)
- **100k** log lines a week with personal data, from one service (One ORM logging setting)
- **100%** of logs arrived with no severity (In every sample)
- **1 day** for Rose to find auth tokens in logs (Confirmed by the customer)

A European digital-health platform knew personal data might be in its logs, but not where it came from. Five weekly staging samples and a code review later, the answer was mostly a single shared setting: an ORM configured to log every SQL query with its values. On one service, that came to about 100,000 log lines a week carrying clinicians' personal data.

## The setup

The platform runs a fleet of Node.js and TypeScript microservices, between 50 and 120 depending on the sample, instrumented through a shared internal auto-instrumentation package. Telemetry goes to a self-hosted open-source observability stack. The engineering team is small, and nobody owned observability full time. With security and privacy audits ahead, personal data in logs was the engineering lead's top priority.

> However careful a team is, it never feels sure it isn't leaking secrets or personal data.
>
> — CTO (paraphrased)

OllyGarden Insights analyzed short staging samples and sent five written assessments between April and June 2026, and Rose reviewed the code. This story covers what they found.

## Personal data in 1 in 6 services

In the first weekly sample, Insights flagged personal data in 12 of about 70 services, or 17% of the fleet. As more services started sending traces, the count grew to 18 and then 25 over the next two weeks.

The source was not careless developers. One shared ORM configuration logged every SQL query with its bound values at INFO level. On a service that handles clinician records, about 60% of the query logs in a 30-minute window touched personal-data columns: email, phone, full name, and address. Over a full week, that is roughly 170,000 query logs from one service, about 100,000 of them carrying personal data. Two more services used the same configuration.

**Fix the setting, not the people.** When one shared setting causes most of a leak, the fix belongs in that setting, not in a reminder to developers to be more careful.

Later samples found more: a bundle of customer personal data (email, name, phone, street, and postal code) in one service, authorization headers in three services, a date of birth on a trace attribute, and OAuth tokens captured verbatim in spans.

## Tokens in the headers

Within about a day of being connected to the repositories, Rose found authentication tokens being written to logs inside HTTP request headers. The customer's DevOps lead checked the finding against their own logs, confirmed it, and called it critical.

## The hits we ruled out

A scanner that cries wolf gets ignored. Several hits were checked and ruled out as false positives: git SSH URLs, millisecond timestamps that looked like phone numbers, and cloud service-account addresses. In the first week, traces were clean, so only the log pipeline needed work.

One trace-side issue did show up: medical-image and signature file fetches were traced as one span per file, which created high-cardinality spans and exposed storage bucket paths.

## Logs with no severity

Every log record in every sample, about 50,000 to 60,000 per sample, arrived with no severity. The bridge that converted the logs to OTLP was dropping the level. That one cause accounted for about 70% of all active insights in the first week, 90 of 128, and a single collector transform would close almost all of them.

## Where the volume came from

- 16% of logs were exact duplicates in the first week, mostly from a secrets-cache loop that logged on every check. By June, the staging application logs were 84% to 99% reducible by deduplication.
- One background worker produced about 60% to 65% of all log volume, three weeks in a row.
- Once metrics arrived, the monitoring stack's own self-monitoring was about 40% of all metric volume. Scrape intervals were already efficient, so the lever was dropping series, not slowing scrapes.

## Gaps in the picture

Tracing barely existed. In the first week, about 3 of 72 services sent traces, and every trace was a single span, so there was no distributed tracing at all. There were no metrics for the first three weeks, and deployment.environment.name was set on 2 of 72 services, so staging and production could not be told apart.

Some checks came back clean: no metric attribute exceeded 1,000 distinct values, and no attribute values were oversized.

## What to check in your own telemetry

- Before looking for careless code, look at shared settings. ORM and HTTP logging that records query values or headers can put personal data in every service that uses it.
- Check that your log pipeline keeps severity. A bridge that drops it makes every log look the same.
- Find the one service or worker that writes most of your logs.
- Review PII hits before acting on them. Timestamps and service accounts can look like personal data.

## About these numbers

All numbers come from short staging samples, about five minutes each, analyzed between April and June 2026, plus Rose's review of the code. Figures are rounded, and the weekly figures for the clinician-records service are extrapolated from a 30-minute window. Production was not analyzed. Several PII hits were checked and ruled out as false positives before reporting. This story reports what OllyGarden found; it does not claim any outcome.

## Listen to the story

Bianca and Florian, who you might know from the OTel Drops podcast, tell this story in a short conversation. The voices are AI-generated; the facts are the ones on this page.

[Episode (MP3)](https://ollygarden.com/audio/case-studies/digital-health-pii-leak-was-a-logger-setting.mp3)

### Transcript

**Bianca:** Hi, I'm Bianca.

**Florian:** And I'm Florian. You might know us from the OTel Drops podcast. Today we're telling another real story from OllyGarden's work.

**Bianca:** Anonymized, so no names. And this one is about what OllyGarden found, start to finish.

**Florian:** Here's the setup. A European digital-health platform. A small engineering team, a fleet of Node.js and TypeScript microservices, and nobody owning observability full time.

**Bianca:** And security and privacy audits coming up. So personal data in logs was their top priority.

**Florian:** Their CTO put it well, and I'm paraphrasing: however careful a team is, it never feels sure it isn't leaking secrets or personal data.

**Bianca:** That's honest. And it's true for almost every team I know.

**Florian:** So OllyGarden Insights looked at their staging samples every week. And in the very first one?

**Bianca:** Personal data in twelve of about seventy services. Roughly one in six.

**Florian:** That sounds like a lot of careless code.

**Bianca:** That's the interesting part. It wasn't. Most of it came from one shared ORM setting. It logged every SQL query, with its values, at INFO level.

**Florian:** So every query, values included, went straight into the logs.

**Bianca:** On one service that handles clinician records, about sixty percent of the query logs touched personal data columns. Email, phone, full name, address.

**Florian:** And over a week?

**Bianca:** Roughly a hundred thousand log lines carrying personal data. From one service. And two more services used the same setting.

**Florian:** So the fix isn't a training session for developers. It's one configuration.

**Bianca:** Exactly. Fix the setting, not the people.

**Florian:** Then Rose got connected to the repositories.

**Bianca:** And within about a day, it found authentication tokens being written to logs, inside HTTP request headers.

**Florian:** The customer's DevOps lead checked it against their own logs, confirmed it, and called it critical.

**Bianca:** I also like what OllyGarden didn't report. Several hits turned out to be false positives. Git SSH URLs, service-account addresses, and millisecond timestamps that looked like phone numbers.

**Florian:** A timestamp with a phone number's haircut.

**Bianca:** Right. But a scanner that cries wolf gets ignored. So those were ruled out before they went into a report.

**Florian:** Then the less dramatic, but very expensive, findings. Every single log record, in every sample, arrived with no severity.

**Bianca:** None at all?

**Florian:** None. The bridge converting their logs to OTLP was dropping the level. That one cause was about seventy percent of all their open insights in the first week.

**Bianca:** And one collector transform would close almost all of them.

**Florian:** Then volume. One background worker wrote about sixty percent of all the logs, three weeks in a row.

**Bianca:** And a secrets-cache loop logged on every single check. That was most of their exact duplicates.

**Florian:** When metrics showed up, about forty percent of them were the monitoring stack watching itself.

**Bianca:** And tracing barely existed. A handful of services, and every trace was a single span.

**Florian:** So, what would you check in your own telemetry after hearing this?

**Bianca:** First, look at shared settings before you blame anyone's code. ORM and HTTP logging can put personal data in every service at once.

**Florian:** Second, check that your log pipeline keeps severity. And third, find the one worker that writes most of your logs.

**Bianca:** And fourth, review your PII hits before acting on them. Not everything that looks like a phone number is one.

**Florian:** The full written story is on the OllyGarden website.

**Bianca:** Thanks for listening.

**Florian:** Until next time.