For decades, knowing whether your infrastructure was working was relatively simple: you installed an agent, defined CPU, memory and disk thresholds, and waited for an alert to fire. That is monitoring, and it is still essential. But when an application stops being a single process on a single server and becomes dozens of microservices spread across a Kubernetes cluster, those alerts start telling you "something is wrong" without telling you "where or why".
That is where observability comes in. It is not a trendy synonym for monitoring, nor a new tool that replaces the old one: it is its natural evolution. Monitoring tells you the system is broken; observability gives you the data to understand why it is broken, even for failures you never anticipated. This article picks up where classic 24/7 monitoring leaves off and explains its three pillars, how they correlate and when it is worth making the leap.
From Monitoring to Observability
The essential difference comes down to a single distinction: known-unknowns versus unknown-unknowns. Monitoring is built around failures you already know and anticipate: you know the disk can fill up, the CPU can saturate or a service can stop responding, so you define metrics and alerts for those specific cases. It is a status-dashboard approach that answers the question "is it broken?" very well.
The problem appears with failures you did not anticipate. In a distributed system, high latency at checkout can be caused by a slow third-party service, an exhausted connection pool three hops down, a cascade of retries or a single cluster node with a degraded disk. No predefined alert covers that exact combination, because you had never seen it before. Observability answers "why is it broken?": instead of checking fixed hypotheses, it lets you ask new questions about the system's behaviour based on the data it already emits.
Put another way, monitoring is a finite set of questions with predefined answers; observability is the ability to ask questions you had not foreseen without having to deploy new code. And that ability is built on three types of telemetry data that complement each other.
The Three Pillars: Metrics, Logs and Traces
Observability rests on three signals, each with its strengths and its cost. They do not compete: they are used together. Understanding what each one answers is the key to designing a strategy that becomes neither blind nor ruinously expensive.
1. Metrics. These are numeric values aggregated over time: requests per second, p99 latency, memory usage, error rate. Because they are time series, they are cheap to store and query, aggregable and perfect for alerting and for trend dashboards. Their limit is that they lack individual context: they tell you the error rate rose to 5%, but not which specific requests failed or why. The reference tool is Prometheus, with Grafana as the visualization layer.
2. Logs. These are discrete events with context: each line records what happened, when and with what detail (error message, user ID, stack trace). They are irreplaceable for fine-grained diagnosis, but expensive to store and index at scale: a high-traffic system generates terabytes of logs a day. The typical platforms are Loki (more economical, indexes only labels) and the ELK / Elasticsearch stack (full indexing, more powerful for search but more costly).
3. Distributed traces. This is the pillar classic monitoring lacks. A trace reconstructs the complete journey of a request as it travels through all the microservices that handle it. Each leg of the journey is a span, with its own duration and metadata, and the whole set is linked by a shared identifier. That way, for one specific slow request, you can see where the time went: whether 90% of the latency was in a database query or in a call to an external service. The reference tools are Jaeger and Grafana Tempo.
| Criterion | Metrics | Logs | Traces |
|---|---|---|---|
| What does it answer? | What is happening and how much? | What exactly happened? | Where did the time go? |
| Data type | Aggregated time series | Discrete event with context | Request journey (spans) |
| Storage cost | Low | High | Medium (with sampling) |
| Cardinality risk | High | Medium | Low |
| Best for | Alerting and trends | Fine-grained diagnosis | Latency in distributed systems |
| Typical tool | Prometheus | Loki / ELK | Jaeger / Tempo |
Correlation: From Symptom to Root Cause
The value of observability is not in each pillar on its own, but in being able to jump from one to another without friction. An isolated pillar is an island; the three correlated together form a map. The ideal investigation flow during an incident almost always follows the same path:
- arrow_right You start with the metric: an alert shows that the payment service's p99 latency has spiked. You know what is happening and how much, but not why.
- arrow_right You jump to the trace: you filter the slow requests in the affected window and examine their spans. You discover that almost all the time is lost in a call to an inventory service.
- arrow_right You land on the log: from that specific span you open the correlated logs for that service at that instant and find the exact cause: database timeouts due to an exhausted connection pool.
That path — from the metric to the trace and from the trace to the log — is what cuts MTTR (mean time to resolution) from hours to minutes. For it to work, all three signals must share context: the same trace identifiers and the same service labels propagated throughout the whole system. And this is where the standard that makes it possible comes in.
OpenTelemetry: Unified Instrumentation Without Lock-in
Historically, each observability vendor had its own agent and its own SDK. Instrumenting your code for one meant getting tied to it: switching tools meant reinstrumenting the entire application. OpenTelemetry (OTel) solves exactly that: it is an open, vendor-neutral standard for generating metrics, logs and traces.
The idea is simple and powerful: you instrument once with the OTel libraries, and you decide afterwards which backend you send the data to. You can send your metrics to Prometheus, your traces to Tempo and your logs to Loki; or migrate to a commercial platform tomorrow without touching your application code. The OpenTelemetry Collector acts as the central piece that receives, processes and re-exports telemetry to one or several destinations.
For a business, this is above all a strategic decision: avoiding vendor lock-in. Telemetry is an asset that should not depend on whichever vendor you happen to use. Adopting OTel as your instrumentation layer gives you the freedom to choose — and change — backend based on cost, performance or project needs, without redoing the instrumentation work.
Cardinality, Cost and Tiered Retention
Observability has a silent enemy that can send the bill through the roof: cardinality. It refers to the number of unique label combinations a metric generates. Adding a label with few values (for example, the HTTP method) is harmless. Adding a high-cardinality label — user ID, session ID, full URL with parameters — multiplies the number of time series into the millions and can bring down a Prometheus instance or make the storage cost unviable.
Practical rule on labels:
High-cardinality data (user, request or session IDs) belongs in logs and traces, not in metric labels. Metrics should only be labelled with low-cardinality dimensions (service, endpoint, status code). Getting this wrong is the number-one cause of runaway observability bills.
With traces, cost control is called sampling. Storing 100% of the traces from a high-traffic system is expensive and unnecessary: it is enough to keep a representative fraction, or better still, tail sampling that always keeps traces of slow or errored requests and discards most of the fast, correct ones. That way you retain what matters for diagnosis without paying for the noise.
Finally, tiered retention balances usefulness and budget: keep recent data in fast, detailed storage (days or weeks), gradually downsample to aggregated resolutions for the medium term, and archive to cheap object storage for the long term. Not all telemetry data deserves the same cost or the same access speed.
SLI, SLO and Error Budgets: the Reliability Framework
Observability is not an end in itself: it exists to sustain measurable reliability goals. The framework popularised by SRE practice rests on three linked concepts. An SLI (Service Level Indicator) is a concrete metric of the user experience: for example, the percentage of requests served correctly under 300 ms. An SLO (Service Level Objective) is the target you set for that indicator: for example, meeting it 99.9% of the time.
The direct consequence of the SLO is the error budget: if your target is 99.9%, you have a 0.1% of "allowed" failure over the period. That budget is a decision-making tool: as long as you have margin, you can deploy and experiment; if you burn through it, you freeze changes and prioritise stability. This framework turns reliability into something negotiable and quantifiable, rather than a vague aspiration to "100% always". To understand how this translates into a contractual commitment, see our article on the 99.99% SLA explained and how much downtime each availability level really represents.
When Is Monitoring Enough and When Do You Need Observability?
Adopting full observability takes effort and money, so the important question is not "do I want it?" but "do I need it?". The answer depends on the system's architecture, not the size of the company.
Simple monitoring is enough when you have a monolith or a few servers with clear responsibilities. If a request is handled within a single process, when something goes wrong the search space is small: the application logs and the system metrics are usually enough. Adding distributed traces and the whole OTel machinery would be over-engineering.
You need observability when the system is distributed: microservices, Kubernetes, message queues, multiple databases and calls between services. In that scenario a single request may touch ten different components, and locating where something fails without traces is like searching in the dark. The more network hops there are between the user and the response, the more justified the investment.
- check_circle A sign you need observability: you spend more time locating which service a problem is in than actually fixing it.
- check_circle Another sign: incidents keep reproducing combinations of failures nobody had foreseen, and existing alerts do not explain them.
- check_circle Recommended approach: start with a solid base of monitoring and metrics, and add traces and correlation as distributed complexity justifies it.
EasyDataHost: Managed Monitoring as the Foundation
Observability is built on solid foundations, and those foundations are well-done monitoring. At EasyDataHost we offer managed monitoring (Monitoring as a Service) as the base layer on which your team can develop a full observability strategy:
- arrow_right Collection and retention of infrastructure metrics (Prometheus and compatible) with dashboards, thresholds and alerts managed by our team.
- arrow_right Aggregation and retention of logs with the tiered policy suited to your volume, avoiding cost surprises at scale.
- arrow_right Infrastructure ready for OpenTelemetry, so you can instrument your applications without tying yourself to a specific vendor.
- arrow_right Integration with our managed services, with 24/7 support and incident response from our own datacenter in Spain.
Our infrastructure operates with ISO 27001 certification and ENS compliance. If you want to move from "just monitoring" to genuinely understanding why your platform behaves the way it does, talk to our team to design the right strategy for your architecture.
Frequently Asked Questions
What is the difference between monitoring and observability?
Monitoring answers "is it broken?" using metrics and alerts defined in advance for problems you already anticipate (known-unknowns). Observability lets you answer "why is it broken?" by correlating metrics, logs and traces to investigate failures you never predicted (unknown-unknowns), which is essential in distributed systems and microservices.
Do I need observability if I have a monolith and a few servers?
Not necessarily. For a monolith with a few servers, classic monitoring with system metrics and alerts is usually enough to operate reliably. Observability delivers real value when a request travels across many independent services: microservices, Kubernetes and distributed architectures where pinpointing where something fails is not obvious.
What is OpenTelemetry and why does it matter?
It is the open, vendor-neutral instrumentation standard for metrics, logs and traces. You instrument your application once and can send the data to any backend (Prometheus, Jaeger, Tempo, Loki and others), which avoids vendor lock-in and unifies telemetry under a single format.
Conclusion
Observability does not replace monitoring: it extends it. Going from "is it broken?" to "why is it broken?" is what separates operating a simple system from operating a distributed one with confidence:
- arrow_right Monitoring covers the failures you anticipate (known-unknowns); observability lets you investigate the ones you did not (unknown-unknowns).
- arrow_right Its three pillars complement each other: metrics to alert and see trends, logs for detail, and traces to know where the time goes in a distributed system.
- arrow_right Correlation between the three (metric → trace → log) and OpenTelemetry as a neutral instrumentation layer are what turn scattered data into answers.
- arrow_right Control cardinality and cost, lean on SLI/SLO/error budgets, and adopt observability when your distributed architecture genuinely justifies it, not before.