Search

Elasticsearch Infrastructure Monitoring and APM: A Practical Guide

Elasticsearch Infrastructure Monitoring and APM: A Practical Guide

Elasticsearch Infrastructure Monitoring and APM: A Practical Guide

No headings found on page

Written by

Dina Bridge

|

Subscribe

Subscribe to get the latest insights straight in your inbox

Learn how Elasticsearch connects infrastructure monitoring, APM, logs, metrics, and traces to help teams investigate application performance problems.

Infrastructure health does not explain application behavior

A host reaches 90% CPU utilization. A Kubernetes pod restarts. Application response time increases. Customers report failed transactions. These events may be connected, but infrastructure monitoring alone cannot prove the relationship. Infrastructure metrics show what is happening across hosts, containers, cloud services, and Kubernetes environments. Application performance monitoring shows how requests move through services, where transactions slow down, and which dependencies contribute to latency or errors.

An effective Elasticsearch application performance monitoring architecture connects these perspectives. Engineers should be able to begin with an infrastructure symptom, move to the affected service, examine representative traces, and inspect related logs without rebuilding the investigation in several disconnected tools. Elasticsearch provides the searchable data layer behind this workflow. Elastic Agent, Elastic APM, OpenTelemetry, integrations, and Kibana contribute different collection, instrumentation, processing, and analysis capabilities. The result is not automatic correlation. The platform becomes useful only when the telemetry shares enough consistent context to support a defensible investigation.

Infrastructure monitoring and APM answer different questions

Infrastructure monitoring and application performance monitoring overlap, but they are not interchangeable.

What infrastructure monitoring tells you

Infrastructure monitoring focuses on the systems that run applications. Depending on the environment, this can include:

  • CPU utilization

  • Memory consumption

  • Filesystem capacity

  • Disk latency and throughput

  • Network traffic

  • Host availability

  • Container resource usage

  • Kubernetes node, pod, and workload health

  • Cloud-service metrics

  • Process behavior

Elastic’s infrastructure monitoring capabilities can visualize host and infrastructure metrics, help identify high resource utilization, track Kubernetes pods, and connect metrics with logs and APM data. These signals help answer questions such as:

  • Which hosts are under pressure?

  • Did resource usage change before the incident?

  • Are failures concentrated in one availability zone?

  • Is a Kubernetes workload repeatedly restarting?

  • Are several services competing for the same infrastructure?

  • Did storage or network behavior change?

Infrastructure monitoring is essential, but it mainly describes the operating environment. It does not automatically show which customer request was affected or which application operation caused the pressure.

What application performance monitoring tells you

Application performance monitoring focuses on the behavior of applications and services. Typical APM data includes:

  • Transaction duration

  • Request rate

  • Error rate

  • Spans within a distributed trace

  • Database calls

  • External service calls

  • Service dependencies

  • Runtime metrics

  • Exception details

  • Application-specific labels and attributes

Distributed tracing follows a request across instrumented services. Elastic describes traces as linked transactions that show how a request was served and which services participated in it. APM helps answer different questions:

  • Which service is slow?

  • Which transaction is failing?

  • Where is time spent inside the request?

  • Is the delay in application code, a database, or another service?

  • Did a new deployment change latency or error behavior?

  • Are failures limited to one environment, version, or customer group?

Why teams need both

Imagine that an API becomes slow while CPU utilization rises on several Kubernetes nodes. Infrastructure monitoring shows the resource pressure. APM shows that one service is executing unusually expensive database calls. Logs reveal repeated retries. Deployment metadata shows that the behavior began after a new version was released.

No individual signal provides the full explanation. The useful investigation comes from connecting:

Infrastructure condition → affected service → slow transaction → dependency behavior → relevant logs → deployment context

That connection is the real value of combining infrastructure monitoring and APM.

How Elasticsearch supports infrastructure and application monitoring

Elasticsearch is not the infrastructure agent, application instrumentation library, or visualization interface. Its primary role is to store, index, search, aggregate, and retrieve the telemetry collected from those systems. A simplified Elastic monitoring architecture has five stages:

  1. Instrument applications and collect infrastructure data

  2. Process and enrich the telemetry

  3. Send the data to the Elastic platform

  4. Index the data in Elasticsearch

  5. Investigate and visualize it through Elastic Observability and Kibana

1. Collect infrastructure telemetry

Elastic Agent can collect logs, metrics, and other operational data from hosts. Elastic integrations provide collection and visualization assets for common technologies and services. For infrastructure monitoring, teams may collect data from:

  • Linux and Windows hosts

  • Kubernetes clusters

  • Docker environments

  • AWS, Microsoft Azure, and Google Cloud

  • NGINX and web servers

  • Databases

  • Message brokers

  • Network services

  • Elasticsearch clusters

  • Prometheus-compatible endpoints

Elastic currently recommends Elastic Agent for infrastructure metric collection. Its System integration can collect host metrics and logs, while other integrations address Kubernetes, cloud platforms, databases, and services. Existing environments may also use Beats, Logstash, upstream OpenTelemetry Collectors, or other supported collection paths. The right design depends on the current environment, deployment model, processing requirements, and migration constraints.

2. Instrument application services

Applications must be instrumented before APM can explain their internal behavior. Instrumentation can collect:

  • Transactions

  • Spans

  • Errors

  • Service names

  • Environment information

  • Runtime metrics

  • Database activity

  • External calls

  • Trace context

Teams can use Elastic-supported instrumentation or OpenTelemetry instrumentation, depending on their languages, architecture, support requirements, and portability goals. OpenTelemetry provides APIs, SDKs, and collection components for telemetry. Elastic supports OpenTelemetry-based ingestion, including paths that use Elastic’s distribution, upstream-compatible components, or managed OTLP endpoints where available.

Instrumentation should be treated as an engineering design decision. Installing an agent does not guarantee useful traces. Service naming, environment attributes, propagation, sampling, sensitive-data controls, and ownership must also be defined.

3. Process and enrich the data

Raw telemetry often needs processing before it becomes consistently searchable. Processing may include:

  • Parsing log messages

  • Normalizing field names

  • Adding service and environment context

  • Enriching events with cloud or Kubernetes metadata

  • Redacting sensitive information

  • Filtering low-value data

  • Sampling traces

  • Routing data to appropriate destinations

  • Converting source-specific fields into a common schema

This stage may happen in the application SDK, Elastic Agent, an OpenTelemetry Collector, Logstash, Elasticsearch ingest pipelines, or a combination of components. More processing layers are not automatically better. Each additional component creates configuration, capacity, failure, and troubleshooting requirements. The architecture should use the smallest number of stages that satisfies the organization’s collection, governance, routing, and resiliency needs.

4. Index telemetry in Elasticsearch

Once telemetry reaches Elasticsearch, its field structure determines how effectively engineers can search and correlate it. Useful fields commonly describe:

  • Service

  • Host

  • Container

  • Kubernetes workload

  • Cloud account and region

  • Deployment environment

  • Trace and transaction identifiers

  • Application version

  • Event dataset

  • Timestamp

  • Log level

  • Error information

Consistent identifiers allow engineers to move between different telemetry signals. For example:

  • A trace identifies a slow service.

  • The service record identifies its Kubernetes workload.

  • Infrastructure data shows resource pressure on the relevant pod and node.

  • A trace identifier connects the request to related logs.

  • Version metadata connects the failure to a recent deployment.

Elasticsearch can store all these signals, but correlation depends on the fields being populated correctly. Storing data in the same cluster does not make it automatically related.

5. Investigate through Elastic Observability and Kibana

Elastic Observability provides purpose-built views for infrastructure, hosts, services, traces, logs, alerts, and related operational workflows. Kibana also gives teams flexible search, visualization, dashboard, and analytical capabilities. This matters when a predefined application view does not answer the investigation question.

Engineers may begin with:

  • A host showing high utilization

  • An application service with increased latency

  • A failed synthetic check

  • An alert

  • A trace

  • A log pattern

  • A Kubernetes workload

  • A custom Kibana dashboard

The platform is most valuable when the engineer can preserve useful context while moving between those views.

The main components of an Elastic monitoring architecture

Several products and standards appear in an Elastic observability environment. Their roles should not be confused.

Component

Primary role

Elastic Agent

Collects logs, metrics, and other data from hosts and integrations

Fleet

Centrally manages Elastic Agents and their policies

Elastic integrations

Provide collection configurations, field mappings, ingest assets, and dashboards for supported technologies

Elastic APM instrumentation

Captures application transactions, spans, errors, and runtime data

OpenTelemetry SDKs

Instrument applications using the OpenTelemetry standard

OpenTelemetry Collector

Receives, processes, and exports telemetry

Logstash

Provides flexible pipeline processing and routing where needed

Elasticsearch ingest pipelines

Transform and enrich documents during ingestion

Elasticsearch

Stores, indexes, searches, and aggregates telemetry

Elastic Observability

Provides infrastructure, APM, logs, alerting, and investigation workflows

Kibana

Provides visualization, dashboards, search, and analysis interfaces

Not every deployment needs every component. A small environment might use Elastic Agent for host telemetry, application instrumentation for traces, and a managed Elastic deployment. A larger enterprise might use gateway collectors, several processing pipelines, multiple regions, strict data controls, tiered retention, and separate production and monitoring environments. Architecture should reflect actual requirements—not the maximum number of available components.

A practical infrastructure-to-application investigation

Consider an online service with an unexpected increase in checkout latency.

Step 1: Confirm the symptom

Begin with the user-facing or service-level symptom:

  • Which transaction is slow?

  • When did the problem start?

  • Is the increase visible at p50, p95, or p99?

  • Is it limited to one region, environment, or version?

  • Did request volume change?

  • Are errors increasing as well?

An average can hide a severe experience affecting a smaller percentage of requests. Segment the data before concluding that the entire service behaves the same way.

Step 2: Identify the affected service and transaction

Use APM data to determine:

  • Which service owns the slow transaction

  • Which transaction types are affected

  • Whether the time is spent inside the service or a dependency

  • Whether the problem affects all requests or a subset

  • Whether a deployment or configuration change aligns with the increase

Inspect representative traces instead of relying only on aggregate charts.

Step 3: Follow the distributed trace

A distributed trace can show the path of the request across instrumented services and dependencies. Inspect:

  • The longest spans

  • Database operations

  • External calls

  • Repeated requests

  • Synchronous dependencies

  • Error events

  • Missing sections of the trace

Elastic’s service map is based on instrumented services and propagated trace context. If a service is not instrumented, or trace context is not propagated, the relationship may not appear. A missing connection should therefore be treated as an instrumentation question, not proof that the dependency does not exist.

Step 4: Correlate the service with infrastructure

Once the affected service is known, examine the infrastructure supporting it:

  • Pod CPU and memory

  • Container restarts

  • Node pressure

  • Network behavior

  • Filesystem or storage latency

  • Replica availability

  • Autoscaling events

  • Resource limits

  • Cloud-service metrics

The goal is not to find any infrastructure anomaly. It is to determine whether an infrastructure condition aligns with the affected service, time period, and requests. High CPU alone is not a root cause. It may be the cause of slow transactions, the consequence of expensive application behavior, or an unrelated workload occurring on the same infrastructure.

Step 5: Inspect related logs

Logs can provide evidence that traces and metrics do not contain:

  • Timeout messages

  • Retry behavior

  • Connection failures

  • Application exceptions

  • Resource-limit messages

  • Dependency errors

  • Configuration changes

  • Deployment events

Where trace and log correlation is configured, trace identifiers can help locate logs associated with a particular request. Without consistent identifiers, engineers may have to approximate the relationship using time windows, service names, hosts, or Kubernetes metadata. That makes the investigation slower and less reliable.

Step 6: Test the leading explanation

After correlating the signals, state the hypothesis precisely. For example:

Checkout latency increased because version 4.2 generated repeated inventory requests, which increased database activity and CPU consumption on the affected pods.

That hypothesis can be tested. A vague conclusion such as “the cluster was overloaded” is much harder to validate and may encourage the wrong response.

Step 7: Verify the remediation

After applying a fix, verify the result using the same measurements that established the problem:

  • Transaction latency

  • Error rate

  • Trace duration

  • Dependency behavior

  • Infrastructure utilization

  • Log patterns

  • User-facing outcome

A graph returning to normal is useful evidence, but the validation should cover the complete service outcome—not only the metric that first triggered the investigation.

Why correlation depends on telemetry design

Observability platforms are often presented as though correlation appears automatically once logs, metrics, and traces enter the same system. In practice, correlation is an architectural outcome.

Establish consistent service identity

Every service should have a stable and governed identity. Common problems include:

  • The same service using several names

  • Different naming conventions across languages

  • Environment names embedded inconsistently

  • Temporary deployment identifiers replacing logical service names

  • Teams using application, repository, container, and Kubernetes names interchangeably

A service-name standard should be defined before instrumentation expands across the organization.

Preserve environment and deployment context

Production, staging, test, and development telemetry must be distinguishable. Useful context includes:

  • Deployment environment

  • Application version

  • Cloud region

  • Availability zone

  • Kubernetes namespace

  • Cluster name

  • Service instance

  • Team or ownership metadata

Without this context, engineers can find an error without knowing where it occurred or which deployment introduced it.

Propagate trace context

Distributed tracing depends on trace context moving between services. If propagation fails at a proxy, queue, asynchronous process, or uninstrumented service, a single request can appear as several unrelated traces. Teams should test propagation across:

  • HTTP services

  • Message queues

  • Background workers

  • Serverless functions

  • API gateways

  • Third-party calls

  • Asynchronous workflows

Align logs with traces

Application logs should contain enough structured context to support correlation. Where appropriate, include:

  • Trace ID

  • Transaction or span ID

  • Service name

  • Environment

  • Application version

  • Request or correlation ID

  • Error type

  • Relevant business-safe identifiers

Do not add sensitive information simply because it may help troubleshooting. Logging standards must account for security, privacy, retention, and access-control requirements.

Control cardinality

Telemetry fields with a very large number of unique values can increase storage, indexing, query, and metric-processing costs. Common high-cardinality sources include:

  • Customer IDs

  • Session IDs

  • Request IDs

  • Container IDs

  • Kubernetes labels

  • URLs containing dynamic values

  • Unbounded exception messages

  • User-defined attributes

Some high-cardinality fields are operationally valuable. The goal is not to eliminate them indiscriminately. Teams should decide which values belong in indexed fields, logs, traces, metric dimensions, or another controlled representation.

Common implementation mistakes

Mistake 1: Collecting everything before defining questions

More telemetry does not automatically produce better observability. Before expanding collection, define the operational questions the platform must answer:

  • Which services affect a critical customer journey?

  • What evidence is needed during an incident?

  • Which logs are required for diagnosis?

  • Which infrastructure metrics influence capacity decisions?

  • Which traces must be retained?

  • What data must be excluded or redacted?

Collection should follow operational value, not fear of missing data.

Mistake 2: Treating Elasticsearch as the only system to monitor

An Elasticsearch-based monitoring platform must itself be observable. Teams should monitor:

  • Ingestion delay

  • Pipeline failures

  • Rejected work

  • Storage growth

  • JVM memory pressure

  • Query performance

  • Shard behavior

  • Collector health

  • Agent health

  • Sampling behavior

  • Data completeness

If the monitoring platform becomes unhealthy during an incident, engineers may lose the evidence they need most.

Mistake 3: Using dashboards as the complete investigation model

Dashboards are useful for detecting patterns and summarizing known conditions. They cannot anticipate every failure. A mature workflow must also support:

  • Ad hoc search

  • Trace inspection

  • Log exploration

  • Field-level filtering

  • Time comparison

  • Service and infrastructure drilldowns

  • Validation of unusual hypotheses

A dashboard should lead engineers toward evidence, not trap them inside a fixed view. For additional guidance, read Kibana Dashboard Best Practices for Elasticsearch Teams.

Mistake 4: Ignoring data retention by signal

Logs, metrics, and traces do not necessarily need the same retention period. Retention should reflect:

  • Investigation value

  • Compliance requirements

  • Data volume

  • Query frequency

  • Storage tier

  • Sampling strategy

  • Incident history

  • Recovery requirements

Keeping every signal in an expensive tier for the same duration is rarely the most defensible design.

Mistake 5: Assuming OpenTelemetry removes backend decisions

OpenTelemetry can improve instrumentation consistency and portability, but it does not eliminate architectural choices. Teams must still decide:

  • Which signals to collect

  • Where collectors run

  • How data is sampled

  • Which attributes are permitted

  • How telemetry is enriched

  • What happens during backpressure

  • Which destinations receive the data

  • How schemas are governed

  • How costs are controlled

Portability should be tested with real telemetry and a realistic export path.

Mistake 6: Scaling before locating the bottleneck

Additional infrastructure may be justified, but it should not be the first answer to every performance problem. A slow investigation experience could be caused by:

  • Excessive telemetry volume

  • Poor field mappings

  • Uncontrolled shard growth

  • Expensive queries

  • High-cardinality aggregations

  • Oversized dashboard time ranges

  • Pipeline delays

  • Hot nodes

  • Slow storage

  • Application-side behavior

  • Incomplete correlation fields

Identify the limiting component before paying to scale it.

How to evaluate an Elasticsearch monitoring architecture

A proof of concept should recreate real operational workflows—not merely demonstrate that data reaches a dashboard.

Use representative data

Include:

  • Normal and peak ingestion

  • Realistic host and Kubernetes cardinality

  • Representative application traces

  • Expected log formats

  • Production-like retention

  • Concurrent searches

  • Failure conditions

  • Schema changes

A short test with clean sample data may conceal the problems that emerge at production scale.

Reproduce real investigations

Choose two or three incidents that the engineering team has previously experienced. Measure whether the proposed architecture helps engineers:

  1. Detect the symptom

  2. Identify the affected service

  3. Locate representative traces

  4. Find related infrastructure

  5. Retrieve relevant logs

  6. Form a defensible hypothesis

  7. Validate the remediation

Record the number of steps, time required, missing context, query performance, and points where engineers must leave the platform.

Test failure behavior

Observability pipelines also fail. Test:

  • Collector interruption

  • Network loss

  • Backpressure

  • Endpoint unavailability

  • Partial telemetry loss

  • Invalid data

  • Mapping conflicts

  • Sudden ingestion growth

  • Agent-policy errors

  • Credential or certificate expiration

Document buffering, retry, data-loss, and recovery behavior.

Model cost at realistic growth

Calculate the cost of:

  • Collection

  • Processing

  • Ingestion

  • Compute

  • Storage

  • Replication

  • Retention

  • Data transfer

  • Support

  • Administration

  • Engineering ownership

Run the model at current, 2x, 5x, and 10x telemetry volumes. Separate logs, metrics, and traces because their volumes, retention requirements, and collection strategies differ.

When Elasticsearch is the right fit

An Elasticsearch-based monitoring architecture can be a strong fit when an organization:

  • Needs flexible search across large operational datasets

  • Relies heavily on logs during investigations

  • Wants infrastructure, application, and log context in one platform

  • Already operates Elasticsearch or Kibana successfully

  • Requires deployment flexibility

  • Has engineers capable of governing mappings, lifecycle policies, ingestion, and platform performance

  • Wants to build custom analytical workflows beyond predefined dashboards

  • Needs to retain and interrogate detailed operational events

It may be a weaker fit when an organization:

  • Wants minimal platform engineering responsibility

  • Does not have the skills to operate or govern the environment

  • Needs a highly guided experience with little customization

  • Cannot control telemetry growth

  • Expects correlation without investing in instrumentation and schema standards

  • Has a small, simple environment that does not justify the operational model

The question is not whether Elasticsearch can store monitoring data. It can. The more important question is whether the organization can design and operate the surrounding collection, instrumentation, correlation, retention, and investigation system effectively.

Frequently asked questions

Is Elasticsearch an application performance monitoring tool?
Elasticsearch is the search and analytics data layer used by Elastic’s APM and observability capabilities. Application instrumentation, telemetry collection, processing, and purpose-built investigation interfaces are provided through other components such as Elastic APM, Elastic Agent, OpenTelemetry, Elastic integrations, and Kibana.

What is the difference between Elastic APM and infrastructure monitoring?
Elastic APM focuses on application transactions, traces, dependencies, errors, and service performance. Infrastructure monitoring focuses on hosts, containers, Kubernetes environments, cloud systems, and their resource behavior. They become more useful when service, host, container, trace, deployment, and environment context can be correlated.

Can Elastic monitor Kubernetes infrastructure and applications?
Yes. Elastic provides Kubernetes integrations for collecting cluster logs and metrics, while instrumented applications can send traces and application-performance data. The quality of the combined investigation depends on consistent Kubernetes metadata, service identity, environment fields, and trace propagation.

Can OpenTelemetry send application data to Elastic?
Yes. OpenTelemetry-compatible SDKs and Collectors can send telemetry to supported Elastic ingestion endpoints. The exact architecture depends on the Elastic deployment model, language support, chosen Collector distribution, and processing requirements.

Does putting logs, metrics, and traces in Elasticsearch automatically correlate them?No. The signals need shared, consistently populated identifiers and metadata. Service names, environments, trace identifiers, infrastructure metadata, timestamps, and deployment information must be designed and validated.

Should every application transaction be traced?
Not necessarily. The appropriate strategy depends on transaction volume, diagnostic value, compliance requirements, storage cost, and the sampling capabilities in use.

Sampling decisions should preserve the traces needed to investigate errors, high latency, and important business transactions.

Can Elastic replace separate infrastructure and APM tools?
Potentially, but replacement should be proven through a production-like evaluation.

Confirm telemetry coverage, investigation workflows, integrations, instrumentation support, alerting, governance, failure behavior, retention, cost, and operational ownership before consolidating existing tools.

Final takeaway

Infrastructure monitoring tells engineers what is happening across the systems running an application. APM shows how application requests and dependencies behave. Logs add detailed event evidence. Traces connect work across services. Elasticsearch makes these signals searchable at scale, but successful correlation depends on the architecture around it.

Teams must establish consistent service identity, propagate trace context, preserve infrastructure metadata, structure logs, govern schemas, control cardinality, and design retention by signal. Without that foundation, the organization may collect large amounts of telemetry while still struggling to explain incidents. The goal is not to place every operational event in one database. It is to create a monitoring system that helps engineers move from symptom to evidence, test a hypothesis, and verify the result.

DinaBridge provides Elasticsearch Consulting Services for organizations designing, improving, or scaling production search and observability environments. If your infrastructure metrics, application traces, and logs remain disconnected—or your Elastic environment is becoming difficult to operate, discuss your platform challenge with DinaBridge.

References