# The Instrumentation Score

Open specification · Apache 2.0

# The Instrumentation Score

The Instrumentation Score is a 0–100 measure of how well a service's OpenTelemetry® telemetry follows best practices and semantic conventions. It checks the traces, metrics, logs, and resource attributes a service already sends against 25 open rules, and weighs each failed rule by its impact.

[Get your score](https://ollygarden.app)

## You’re Heading to the App

Create your OllyGarden account over at ollygarden.app. You can get started for free, no credit card needed.

Stay Here[Continue to the App](https://ollygarden.app)

[Read the specification](https://github.com/instrumentation-score/spec)

How it is calculated

## Weighted rules, one number

The score is a meta-signal: a signal about the quality of your signals. Every rule passes or fails for a service, and failed rules cost points in proportion to their impact, so one Critical failure costs four times as much as one Low failure. Nothing is capped, and fixing the most severe problems first always moves the score the most.

- 01

### Analyze the telemetry

An implementation reads the OTLP data a service already sends, grouped by service.name, over a sliding window that defaults to 30 days. The score follows the service as it changes: spans that helped during development can become noise once it is stable.

- 02

### Check every rule

Each of the 25 rules is a pass or fail condition on resources, spans, metrics, logs, or the SDK. A rule with several conditions passes only when all of them hold.

- 03

### Weigh the results

The score is the weighted share of passed rules: the sum of passed rules times their weight, divided by the sum of all rules times their weight, times 100.

Score = Σ (passed rules × weight) ÷ Σ (all rules × weight) × 100

### Impact levels and weights



| Impact | Weight | Rules |
| --- | --- | --- |
| Critical | 40 | 5 |
| Important | 30 | 13 |
| Normal | 20 | 6 |
| Low | 10 | 1 |



### What a score means



| Score | Category | Meaning |
| --- | --- | --- |
| 90–100 | Excellent | A high standard of instrumentation quality. |
| 75–89 | Good | Solid quality; minor improvements are possible. |
| 50–74 | Needs Improvement | Tangible issues that need attention. |
| 0–49 | Poor | Significant problems that need urgent action. |



The design borrows from scores engineers already trust: CVSS for vulnerabilities, Google Lighthouse for web pages, and SSL Labs for TLS configuration. Like them, it is strict about comparability. An implementation must not add its own rules to the score, and must tell users when it does not implement every rule, so the same telemetry gets the same score in any tool.

The rules

## All 25 rules

The rules come from OpenTelemetry semantic conventions and community practice. This table is generated from the specification repository, most severe first; each rule ID links to its full definition and criteria.

Synced from the specification on 2026-09-24 (commit 0ef3cc6).



| Rule | Applies to | Impact | Why it matters |
| --- | --- | --- | --- |
| LOG-003: Log records do not contain sensitive data such as PII, financial identifiers, credentials, or health information. | Log | Criticalweight 40 pts | Backends give more people access than production databases do, so personal data in logs becomes a compliance problem. How-to → |
| MET-008: Metrics do not contain sensitive data such as PII, financial identifiers, credentials, or health information. | Metric | Criticalweight 40 pts | Personal data in metric names, labels, or exemplars reaches everyone with dashboard access. How-to → |
| RES-005: service.name is present. | Resource | Criticalweight 40 pts | service.name is required by the semantic conventions; without it, telemetry belongs to nobody. |
| RES-010: Resource attribute values do not contain sensitive data such as PII, financial identifiers, credentials, or health information. | Resource | Criticalweight 40 pts | Credentials in resource attributes, often from command-line arguments, are copied onto every record a process sends. How-to → |
| SPA-006: Spans do not contain sensitive data such as PII, financial identifiers, credentials, or health information. | Span | Criticalweight 40 pts | Spans record query text, URLs, and headers, which is where personal data and tokens leak. How-to → |
| LOG-001: Debug-level logs are not enabled in production environments for longer than 14 days. | Log | Importantweight 30 pts | Debug logs left on in production cost storage, can carry sensitive values, and bury the records you need. How-to → |
| LOG-002: Log records have their severityNumber set. | Log | Importantweight 30 pts | Without a severity number every record looks the same, so you cannot filter or alert on errors. |
| MET-001: Metric attributes have bound cardinality. | Metric | Importantweight 30 pts | Unbounded attribute values multiply time series, which slows queries and raises storage cost. How-to → |
| MET-002: Metrics have useful metric units. | Metric | Importantweight 30 pts | A value without a unit cannot be read: 1% of memory, 1 MB, and 1 GB look the same. How-to → |
| MET-003: Metric names are consistently associated with the same metric unit. | Metric | Importantweight 30 pts | When one metric name carries different units, every query across services mixes them. How-to → |
| MET-006: Metric names do not equal semantic convention attribute keys. | Metric | Importantweight 30 pts | A metric named like a semantic convention attribute, such as http.response.status_code, confuses queries and readers. |
| RES-002: service.instance.id is unique across logical resources within a given service.name. | Resource | Importantweight 30 pts | An instance ID shared by several pods is worse than none: it merges their telemetry. How-to → |
| RES-003: k8s.pod.uid is present in telemetry collected from applications running on a Kubernetes cluster, or from the control pane of the Kubernetes cluster itself. | Resource | Importantweight 30 pts | k8s.pod.uid lets the Collector attach Kubernetes metadata reliably, even behind a service mesh. |
| RES-004: Semantic conventions attributes are used at the right level. | Resource, Log, Span | Importantweight 30 pts | Attributes at the wrong level, such as service.name on a span, break queries that look where the conventions say they live. |
| RES-007: deployment.environment.name is present. | Resource | Importantweight 30 pts | Without an environment, staging and production data mix in the same views and alerts. |
| SPA-003: Span names have bound cardinality.Criteria still being defined in the specification. | Span | Importantweight 30 pts | Span names with IDs or literal URLs defeat grouping and can blow up backend indexes. How-to → |
| SPA-004: Root spans are not CLIENT spans. | Span | Importantweight 30 pts | A trace that starts with a CLIENT span is missing the work that caused the call, or lost its context. |
| SPA-005: Traces do not contain a high number of short duration spans. | Span | Importantweight 30 pts | Dozens of spans shorter than 5 ms in one trace add overhead and hide the operations that matter. How-to → |
| MET-004: Histogram metrics consistently use the same histogram buckets per metric name. | Metric | Normalweight 20 pts | Different buckets for the same histogram make aggregated quantiles less precise. |
| MET-005: Metric names do not contain the name of the metric unit. | Metric | Normalweight 20 pts | The unit belongs in the unit field; a name that repeats it breaks when the unit changes. How-to → |
| RES-001: service.instance.id is present. | Resource | Normalweight 20 pts | service.instance.id tells instances of a service apart without combining other attributes. How-to → |
| RES-006: service.criticality uses a valid enum value. | Resource | Normalweight 20 pts | Standard criticality values let tooling rank alerts and incidents consistently across teams. |
| SPA-001: Traces contain a limited number of INTERNAL spans per service. | Span | Normalweight 20 pts | Many INTERNAL spans in one trace usually mean over-instrumented code paths that add noise and cost. How-to → |
| SPA-002: Traces do not contain orphan spans. | Span | Normalweight 20 pts | A span whose parent never arrives means context propagation broke and the trace is incomplete. How-to → |
| SDK-001: Dependencies (language and runtime) are supported by the SDK. | SDK | Lowweight 10 pts | Languages and runtimes outside the SDK's support window miss fixes and current semantic conventions. |



Worked example

## Scoring a checkout service

A checkout service sends traces, metrics, and logs and fails 4 of the 25 rules. Here is how those findings become its score.

service.name = checkout

- [SPA-006](#SPA-006)Critical−40 pts

url.full on password-reset spans carries the reset token in its query string.

- [RES-007](#RES-007)Important−30 pts

No deployment.environment.name, so staging and production share dashboards.

- [MET-002](#MET-002)Important−30 pts

Custom queue metrics are exported without a unit.

- [RES-001](#RES-001)Normal−20 pts

No service.instance.id, so the three replicas cannot be told apart.

Instrumentation Score

83.3

Good

(720 − 120) ÷ 720 × 100 = 83.3

With every rule passing, the service would hold all 720 weighted points. The 4 failures cost 120, which leaves 83.3: Good.Fixing the Critical token leak alone raises the score to 88.9. Adding the two Important fixes reaches 97.2, which is Excellent.That ordering is the point of the weights: the score tells a team which fix to make first.

Get your score

## Compute it yourself, or let Insights do it

The specification is open, so anyone can implement it. OllyGarden Insights is a hosted implementation that keeps the score current for every service.

### Use the open specification

Read the rules and the formula on GitHub, run them against your own OTLP data, and propose changes. Discussion happens in the #instrumentation-score channel on the CNCF Slack.

- [instrumentation-score/spec on GitHub](https://github.com/instrumentation-score/spec)
- [Join #instrumentation-score →](https://cloud-native.slack.com/archives/C090FEG5R0F)

### Get it continuously with Insights

Send a sample of your OTLP data to OllyGarden Insights and get a score for every service, the findings behind it, and remediation guidance. The Free plan includes the Instrumentation Score, with no credit card.

[Get your score](https://ollygarden.app)

## You’re Heading to the App

Create your OllyGarden account over at ollygarden.app. You can get started for free, no credit card needed.

Stay Here[Continue to the App](https://ollygarden.app)

[Learn more about Insights →](/products/insights)

In real fleets

## What the rules catch

These findings come from our anonymized case studies. Each one is a rule that failed in production or staging telemetry, and a real cost behind it.

[Read the case studies →](/resources/case-studies)

- 9%

of about 500 services set a deployment environment; 56% set service.instance.id.

[RES-007](#RES-007)[RES-001](#RES-001)

[Read the case study →](/resources/case-studies/enterprise-platform-half-the-spans-had-no-name)
- 12 of ~70

services were flagged for personal data in logs in the first week.

[LOG-003](#LOG-003)

[Read the case study →](/resources/case-studies/digital-health-pii-leak-was-a-logger-setting)
- 100%

of log records arrived with no severity, because the log bridge dropped the level.

[LOG-002](#LOG-002)

[Read the case study →](/resources/case-studies/digital-health-pii-leak-was-a-logger-setting)
- 1 day

to catch real user email addresses copied into a metric label by a new log-to-metric rule.

[MET-008](#MET-008)

[Read the case study →](/resources/case-studies/b2b-saas-cuts-logs-and-metrics-at-the-source)

## Frequently Asked Questions

What teams ask us about the Instrumentation Score, the specification, and how to use it.

### What is the Instrumentation Score?

The Instrumentation Score is a 0–100 measure of how well a service's OpenTelemetry telemetry follows best practices and semantic conventions. It checks the traces, metrics, logs, and resource attributes a service already sends against 25 open rules, and weighs each failed rule by its impact.

### How is the Instrumentation Score calculated?

Each rule has an impact level with a weight: Critical 40, Important 30, Normal 20, and Low 10. The score is the sum of the weights of the rules a service passes, divided by the sum of the weights of all rules, times 100. The current specification has 5 Critical, 13 Important, 6 Normal, and 1 Low rules.

### What is a good Instrumentation Score?

The specification maps scores to four categories: 90 to 100 is Excellent, 75 to 89 is Good, 50 to 74 is Needs Improvement, and 0 to 49 is Poor. Do not chase 100; a score of 75 or more is a sound target for most services. Fix Critical failures first: one of them, such as personal data in spans, costs 40 of 720 points and leaves 94.4, and two leave 88.9, which is Good.

### Is the Instrumentation Score vendor-neutral?

Yes. The specification is open source under the Apache 2.0 license, and its maintainers work at Splunk, New Relic, Dash0, and OllyGarden. Implementations may not add rules of their own, so the same telemetry gets the same score in any tool.

### Who maintains the Instrumentation Score specification?

OllyGarden started the specification in 2025 and hosts it today under an open governance model. Maintainers from Splunk, New Relic, Dash0, and OllyGarden review changes, and the goal is to move it to a neutral foundation such as the CNCF or the OpenTelemetry project.

### How is the Instrumentation Score different from SLOs?

An SLO measures how a service behaves for its users, such as latency or error rate. The Instrumentation Score measures the telemetry itself: whether the data your SLOs, dashboards, and alerts are built on is complete, consistent, and safe. A service can meet its SLOs while its telemetry hides problems.

### How is the Instrumentation Score different from data quality checks?

Generic data quality checks validate schemas, freshness, or volumes. The Instrumentation Score checks telemetry against OpenTelemetry semantic conventions and instrumentation practice, such as service identity, units, cardinality, broken traces, and sensitive data, and summarizes the result in one number you can compare across services and tools.

### Do I have to change application code to improve my score?

Not always. Many failed rules can be fixed in the OpenTelemetry Collector: the transform processor, written in the OpenTelemetry Transformation Language (OTTL), can add missing resource attributes, move attributes to the right level, or redact sensitive values, and the filter processor can drop debug logs and noisy spans. Fixing the instrumentation at the source is still the better long-term answer.

### Has the scoring formula changed since the 2025 launch?

Yes. The first draft, published in June 2025, started every service at 80 and added or subtracted fixed points per rule. The community replaced it with the weighted share of passed rules described on this page, which keeps the score between 0 and 100 without clamping and makes every rule's effect proportional to its impact.

### Which signals does the Instrumentation Score cover?

The rules cover resource attributes, spans, metrics, and logs, plus one rule on SDK support for the language and runtime. New rules are added through the specification repository.

### How do I get an Instrumentation Score for my services?

Implement the open specification against your own OTLP data, or use OllyGarden Insights, which computes the score continuously for every service and shows the findings behind it. The Insights Free plan includes the Instrumentation Score, with no credit card.

### Can I propose a new rule?

Yes. Open an issue or pull request in the instrumentation-score/spec repository on GitHub, or start a discussion in the #instrumentation-score channel on the CNCF Slack. Every rule needs a rationale, pass or fail criteria, a target, and an impact level.

## Keep reading

- [BlogIntroducing the Instrumentation ScoreThe Instrumentation Score is an open, standardized measure of OpenTelemetry instrumentation quality, scoring OTLP data against best-practice rules.](/resources/blog/instrumentation-score)
- [BlogYou don't have too much telemetry. You have bad telemetry.High observability costs are rarely a volume problem. Learn how to fix bad telemetry at the source before relying on sampling.](/resources/blog/you-dont-have-too-much-telemetry-you-have-bad-telemetry)
- [BlogWhere Sensitive Data Hides in OpenTelemetry Pipelines, and How to Find ItTrust store passwords in resource attributes, OAuth tokens in query strings, tax IDs in metrics: what we found in real pipelines, why tools cause most of it, and what privacy laws allow.](/resources/blog/where-sensitive-data-hides-in-opentelemetry-pipelines)[What is telemetry quality?](/resources/telemetry-quality)[Watch the KubeCon talk on the score](/resources/webinars-conferences)[Read the FAQ](/resources/faq)

## Get yourInstrumentation Score

Send a sample of your OpenTelemetry data to Insights and see the score for every service, with the findings behind it.

[Get Started](https://ollygarden.app)

## You’re Heading to the App

Create your OllyGarden account over at ollygarden.app. You can get started for free, no credit card needed.

Stay Here[Continue to the App](https://ollygarden.app)

[Book a Demo](/contact)

Canonical URL: https://ollygarden.com/resources/instrumentation-score