Open specification · Apache 2.0

The Instrumentation Score

The Instrumentation Score is a 0–100 measure of how well a service's OpenTelemetry® telemetry follows best practices and semantic conventions. It checks the traces, metrics, logs, and resource attributes a service already sends against 25 open rules, and weighs each failed rule by its impact.

Get your score

You’re Heading to the App

Create your OllyGarden account over at ollygarden.app. You can get started for free, no credit card needed.

Read the specification

How it is calculated

Weighted rules, one number

The score is a meta-signal: a signal about the quality of your signals. Every rule passes or fails for a service, and failed rules cost points in proportion to their impact, so one Critical failure costs four times as much as one Low failure. Nothing is capped, and fixing the most severe problems first always moves the score the most.

  1. 01

    Analyze the telemetry

    An implementation reads the OTLP data a service already sends, grouped by service.name, over a sliding window that defaults to 30 days. The score follows the service as it changes: spans that helped during development can become noise once it is stable.

  2. 02

    Check every rule

    Each of the 25 rules is a pass or fail condition on resources, spans, metrics, logs, or the SDK. A rule with several conditions passes only when all of them hold.

  3. 03

    Weigh the results

    The score is the weighted share of passed rules: the sum of passed rules times their weight, divided by the sum of all rules times their weight, times 100.

Score = Σ (passed rules × weight) ÷ Σ (all rules × weight) × 100

Impact levels and weights

ImpactWeightRules
Critical405
Important3013
Normal206
Low101

What a score means

ScoreCategoryMeaning
90–100ExcellentA high standard of instrumentation quality.
75–89GoodSolid quality; minor improvements are possible.
50–74Needs ImprovementTangible issues that need attention.
0–49PoorSignificant problems that need urgent action.

The design borrows from scores engineers already trust: CVSS for vulnerabilities, Google Lighthouse for web pages, and SSL Labs for TLS configuration. Like them, it is strict about comparability. An implementation must not add its own rules to the score, and must tell users when it does not implement every rule, so the same telemetry gets the same score in any tool.

The rules

All 25 rules

The rules come from OpenTelemetry semantic conventions and community practice. This table is generated from the specification repository, most severe first; each rule ID links to its full definition and criteria.

Synced from the specification on 2026-09-24 (commit 0ef3cc6).

Instrumentation Score rules with their target, impact level, weight, and rationale
RuleApplies toImpactWhy it matters
LOG-003: Log records do not contain sensitive data such as PII, financial identifiers, credentials, or health information.LogCriticalweight 40 ptsBackends give more people access than production databases do, so personal data in logs becomes a compliance problem. How-to →
MET-008: Metrics do not contain sensitive data such as PII, financial identifiers, credentials, or health information.MetricCriticalweight 40 ptsPersonal data in metric names, labels, or exemplars reaches everyone with dashboard access. How-to →
RES-005: service.name is present.ResourceCriticalweight 40 ptsservice.name is required by the semantic conventions; without it, telemetry belongs to nobody.
RES-010: Resource attribute values do not contain sensitive data such as PII, financial identifiers, credentials, or health information.ResourceCriticalweight 40 ptsCredentials in resource attributes, often from command-line arguments, are copied onto every record a process sends. How-to →
SPA-006: Spans do not contain sensitive data such as PII, financial identifiers, credentials, or health information.SpanCriticalweight 40 ptsSpans record query text, URLs, and headers, which is where personal data and tokens leak. How-to →
LOG-001: Debug-level logs are not enabled in production environments for longer than 14 days.LogImportantweight 30 ptsDebug logs left on in production cost storage, can carry sensitive values, and bury the records you need. How-to →
LOG-002: Log records have their severityNumber set.LogImportantweight 30 ptsWithout a severity number every record looks the same, so you cannot filter or alert on errors.
MET-001: Metric attributes have bound cardinality.MetricImportantweight 30 ptsUnbounded attribute values multiply time series, which slows queries and raises storage cost. How-to →
MET-002: Metrics have useful metric units.MetricImportantweight 30 ptsA value without a unit cannot be read: 1% of memory, 1 MB, and 1 GB look the same. How-to →
MET-003: Metric names are consistently associated with the same metric unit.MetricImportantweight 30 ptsWhen one metric name carries different units, every query across services mixes them. How-to →
MET-006: Metric names do not equal semantic convention attribute keys.MetricImportantweight 30 ptsA metric named like a semantic convention attribute, such as http.response.status_code, confuses queries and readers.
RES-002: service.instance.id is unique across logical resources within a given service.name.ResourceImportantweight 30 ptsAn instance ID shared by several pods is worse than none: it merges their telemetry. How-to →
RES-003: k8s.pod.uid is present in telemetry collected from applications running on a Kubernetes cluster, or from the control pane of the Kubernetes cluster itself.ResourceImportantweight 30 ptsk8s.pod.uid lets the Collector attach Kubernetes metadata reliably, even behind a service mesh.
RES-004: Semantic conventions attributes are used at the right level.Resource, Log, SpanImportantweight 30 ptsAttributes at the wrong level, such as service.name on a span, break queries that look where the conventions say they live.
RES-007: deployment.environment.name is present.ResourceImportantweight 30 ptsWithout an environment, staging and production data mix in the same views and alerts.
SPA-003: Span names have bound cardinality.Criteria still being defined in the specification.SpanImportantweight 30 ptsSpan names with IDs or literal URLs defeat grouping and can blow up backend indexes. How-to →
SPA-004: Root spans are not CLIENT spans.SpanImportantweight 30 ptsA trace that starts with a CLIENT span is missing the work that caused the call, or lost its context.
SPA-005: Traces do not contain a high number of short duration spans.SpanImportantweight 30 ptsDozens of spans shorter than 5 ms in one trace add overhead and hide the operations that matter. How-to →
MET-004: Histogram metrics consistently use the same histogram buckets per metric name.MetricNormalweight 20 ptsDifferent buckets for the same histogram make aggregated quantiles less precise.
MET-005: Metric names do not contain the name of the metric unit.MetricNormalweight 20 ptsThe unit belongs in the unit field; a name that repeats it breaks when the unit changes. How-to →
RES-001: service.instance.id is present.ResourceNormalweight 20 ptsservice.instance.id tells instances of a service apart without combining other attributes. How-to →
RES-006: service.criticality uses a valid enum value.ResourceNormalweight 20 ptsStandard criticality values let tooling rank alerts and incidents consistently across teams.
SPA-001: Traces contain a limited number of INTERNAL spans per service.SpanNormalweight 20 ptsMany INTERNAL spans in one trace usually mean over-instrumented code paths that add noise and cost. How-to →
SPA-002: Traces do not contain orphan spans.SpanNormalweight 20 ptsA span whose parent never arrives means context propagation broke and the trace is incomplete. How-to →
SDK-001: Dependencies (language and runtime) are supported by the SDK.SDKLowweight 10 ptsLanguages and runtimes outside the SDK's support window miss fixes and current semantic conventions.

Worked example

Scoring a checkout service

A checkout service sends traces, metrics, and logs and fails 4 of the 25 rules. Here is how those findings become its score.

service.name = checkout

  • SPA-006Critical−40 pts

    url.full on password-reset spans carries the reset token in its query string.

  • RES-007Important−30 pts

    No deployment.environment.name, so staging and production share dashboards.

  • MET-002Important−30 pts

    Custom queue metrics are exported without a unit.

  • RES-001Normal−20 pts

    No service.instance.id, so the three replicas cannot be told apart.

Instrumentation Score

83.3

Good

(720 − 120) ÷ 720 × 100 = 83.3

With every rule passing, the service would hold all 720 weighted points. The 4 failures cost 120, which leaves 83.3: Good.Fixing the Critical token leak alone raises the score to 88.9. Adding the two Important fixes reaches 97.2, which is Excellent.That ordering is the point of the weights: the score tells a team which fix to make first.

Get your score

Compute it yourself, or let Insights do it

The specification is open, so anyone can implement it. OllyGarden Insights is a hosted implementation that keeps the score current for every service.

Get it continuously with Insights

Send a sample of your OTLP data to OllyGarden Insights and get a score for every service, the findings behind it, and remediation guidance. The Free plan includes the Instrumentation Score, with no credit card.

Get your score

You’re Heading to the App

Create your OllyGarden account over at ollygarden.app. You can get started for free, no credit card needed.

Learn more about Insights →

In real fleets

What the rules catch

These findings come from our anonymized case studies. Each one is a rule that failed in production or staging telemetry, and a real cost behind it.

Read the case studies →

Frequently Asked Questions

What teams ask us about the Instrumentation Score, the specification, and how to use it.

What is the Instrumentation Score?

The Instrumentation Score is a 0–100 measure of how well a service's OpenTelemetry telemetry follows best practices and semantic conventions. It checks the traces, metrics, logs, and resource attributes a service already sends against 25 open rules, and weighs each failed rule by its impact.

How is the Instrumentation Score calculated?

Each rule has an impact level with a weight: Critical 40, Important 30, Normal 20, and Low 10. The score is the sum of the weights of the rules a service passes, divided by the sum of the weights of all rules, times 100. The current specification has 5 Critical, 13 Important, 6 Normal, and 1 Low rules.

What is a good Instrumentation Score?

The specification maps scores to four categories: 90 to 100 is Excellent, 75 to 89 is Good, 50 to 74 is Needs Improvement, and 0 to 49 is Poor. Do not chase 100; a score of 75 or more is a sound target for most services. Fix Critical failures first: one of them, such as personal data in spans, costs 40 of 720 points and leaves 94.4, and two leave 88.9, which is Good.

Is the Instrumentation Score vendor-neutral?

Yes. The specification is open source under the Apache 2.0 license, and its maintainers work at Splunk, New Relic, Dash0, and OllyGarden. Implementations may not add rules of their own, so the same telemetry gets the same score in any tool.

Who maintains the Instrumentation Score specification?

OllyGarden started the specification in 2025 and hosts it today under an open governance model. Maintainers from Splunk, New Relic, Dash0, and OllyGarden review changes, and the goal is to move it to a neutral foundation such as the CNCF or the OpenTelemetry project.

How is the Instrumentation Score different from SLOs?

An SLO measures how a service behaves for its users, such as latency or error rate. The Instrumentation Score measures the telemetry itself: whether the data your SLOs, dashboards, and alerts are built on is complete, consistent, and safe. A service can meet its SLOs while its telemetry hides problems.

How is the Instrumentation Score different from data quality checks?

Generic data quality checks validate schemas, freshness, or volumes. The Instrumentation Score checks telemetry against OpenTelemetry semantic conventions and instrumentation practice, such as service identity, units, cardinality, broken traces, and sensitive data, and summarizes the result in one number you can compare across services and tools.

Do I have to change application code to improve my score?

Not always. Many failed rules can be fixed in the OpenTelemetry Collector: the transform processor, written in the OpenTelemetry Transformation Language (OTTL), can add missing resource attributes, move attributes to the right level, or redact sensitive values, and the filter processor can drop debug logs and noisy spans. Fixing the instrumentation at the source is still the better long-term answer.

Has the scoring formula changed since the 2025 launch?

Yes. The first draft, published in June 2025, started every service at 80 and added or subtracted fixed points per rule. The community replaced it with the weighted share of passed rules described on this page, which keeps the score between 0 and 100 without clamping and makes every rule's effect proportional to its impact.

Which signals does the Instrumentation Score cover?

The rules cover resource attributes, spans, metrics, and logs, plus one rule on SDK support for the language and runtime. New rules are added through the specification repository.

How do I get an Instrumentation Score for my services?

Implement the open specification against your own OTLP data, or use OllyGarden Insights, which computes the score continuously for every service and shows the findings behind it. The Insights Free plan includes the Instrumentation Score, with no credit card.

Can I propose a new rule?

Yes. Open an issue or pull request in the instrumentation-score/spec repository on GitHub, or start a discussion in the #instrumentation-score channel on the CNCF Slack. Every rule needs a rationale, pass or fail criteria, a target, and an impact level.

Get yourInstrumentation Score

Send a sample of your OpenTelemetry data to Insights and see the score for every service, with the findings behind it.

Get Started

You’re Heading to the App

Create your OllyGarden account over at ollygarden.app. You can get started for free, no credit card needed.

Book a Demo