Turn a slow request into a visible path
A report that checkout is slow does not reveal whether inventory, payment or the database is responsible. A trace splits one request into related operations called spans. Parent-child relationships reveal where time was spent. OpenTelemetry provides common APIs and SDKs for producing and exporting this telemetry. This guide builds a small Python starting point. More telemetry is not automatically better: each signal should help the team make a specific troubleshooting decision.
Logs describe individual events, metrics summarize measurements and traces show a particular execution path. Use them together. An error-rate metric can lead to a failed checkout trace, whose trace ID connects the relevant logs. Set service, environment and version consistently so test traffic does not mix with production. The example exports to the console for inspection. That output is a teaching tool, not a production storage backend or an observability dashboard.
Collect data that supports a decision
Instrument one meaningful business path: accepting an order, checking stock and persisting the result. Use stable span names such as inventory.lookup rather than embedding a customer ID in the name. Add small, useful attributes. Tokens, phone numbers, complete private documents and payment payloads require deliberate handling. Our editorial recommendation is an attribute allowlist: every new field needs a documented purpose, a sensitivity review and an owner who can justify collecting it.
Across service boundaries, propagation must carry context or the execution becomes several disconnected traces. Automatic instrumentation can cover framework calls; manual spans describe business boundaries. Combining them carelessly produces duplicates. For queue processing and retries, decide how attempts relate to the original task and how failures are recorded. A dependency error should be diagnosable without printing credentials or request bodies into the telemetry stream.
Manage telemetry as a workload with costs and failure modes. Sampling, attribute limits and retention should follow measured volume. Excessive sampling can hide rare but important failures. A production Collector can centralize routing and processing, including removing fields, but its buffering and network-failure behavior need testing. The support team should know whether exporter outages affect requests, which data can be lost and how to identify a stalled telemetry pipeline.
Code example and verification
This educational example demonstrates the implementation path. Check the stated runtime and prerequisites in a test environment; the notes explain what remains before production use.
# python -m pip install opentelemetry-api opentelemetry-sdk
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace.export import (
SimpleSpanProcessor, ConsoleSpanExporter,
)
provider = TracerProvider(resource=Resource.create({
"service.name": "order-demo",
"deployment.environment.name": "test",
}))
provider.add_span_processor(SimpleSpanProcessor(ConsoleSpanExporter()))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("liyan.order-demo")
with tracer.start_as_current_span("order.prepare"):
with tracer.start_as_current_span("inventory.lookup") as span:
span.set_attribute("inventory.available", True)
provider.shutdown()Expect two spans sharing a trace ID with a parent-child relationship. Shutdown finishes export. This demo has no network or database dependency. Production commonly uses BatchSpanProcessor and an appropriate exporter; configure and test timeouts, sampling and sensitive-field handling first.
Build a useful troubleshooting pilot
Verify a successful request and an intentional test failure. Identify the slow span, service version and related logs, then compare resource use and response latency before and after instrumentation. A useful acceptance criterion is that the team can locate the cause of a delayed order from the recorded path. Define who can view the data and who handles alerts. A busy dashboard by itself does not establish operational value.
Implementation checklist
- Separate the API, SDK and exporter roles; a provider is required for meaningful output.
- Keep names stable; avoid high-cardinality metric dimensions.
- Test a controlled failure, slow operation and exporter outage.
- Remove sensitive fields before export and define access to traces.
Practical explanations and recommendations are Liyan Knowledge editorial analysis.Sources: OpenTelemetry — Python getting started · OpenTelemetry — Python exporters · OpenTelemetry — Context propagation
This Liyan Knowledge article is an editorial synthesis based on the original source.View original source





