Search

Written by
Dina Bridge
|
Subscribe
Subscribe to get the latest insights straight in your inbox
Learn how Elasticsearch connects infrastructure monitoring, APM, logs, metrics, and traces to help teams investigate application performance problems.
Infrastructure health does not explain application behavior
A host reaches 90% CPU utilization. A Kubernetes pod restarts. Application response time increases. Customers report failed transactions. These events may be connected, but infrastructure monitoring alone cannot prove the relationship. Infrastructure metrics show what is happening across hosts, containers, cloud services, and Kubernetes environments. Application performance monitoring shows how requests move through services, where transactions slow down, and which dependencies contribute to latency or errors.
An effective Elasticsearch application performance monitoring architecture connects these perspectives. Engineers should be able to begin with an infrastructure symptom, move to the affected service, examine representative traces, and inspect related logs without rebuilding the investigation in several disconnected tools. Elasticsearch provides the searchable data layer behind this workflow. Elastic Agent, Elastic APM, OpenTelemetry, integrations, and Kibana contribute different collection, instrumentation, processing, and analysis capabilities. The result is not automatic correlation. The platform becomes useful only when the telemetry shares enough consistent context to support a defensible investigation.
Infrastructure monitoring and APM answer different questions
Infrastructure monitoring and application performance monitoring overlap, but they are not interchangeable.
What infrastructure monitoring tells you
Infrastructure monitoring focuses on the systems that run applications. Depending on the environment, this can include:
CPU utilization
Memory consumption
Filesystem capacity
Disk latency and throughput
Network traffic
Host availability
Container resource usage
Kubernetes node, pod, and workload health
Cloud-service metrics
Process behavior
Elastic’s infrastructure monitoring capabilities can visualize host and infrastructure metrics, help identify high resource utilization, track Kubernetes pods, and connect metrics with logs and APM data. These signals help answer questions such as:
Which hosts are under pressure?
Did resource usage change before the incident?
Are failures concentrated in one availability zone?
Is a Kubernetes workload repeatedly restarting?
Are several services competing for the same infrastructure?
Did storage or network behavior change?
Infrastructure monitoring is essential, but it mainly describes the operating environment. It does not automatically show which customer request was affected or which application operation caused the pressure.
What application performance monitoring tells you
Application performance monitoring focuses on the behavior of applications and services. Typical APM data includes:
Transaction duration
Request rate
Error rate
Spans within a distributed trace
Database calls
External service calls
Service dependencies
Runtime metrics
Exception details
Application-specific labels and attributes
Distributed tracing follows a request across instrumented services. Elastic describes traces as linked transactions that show how a request was served and which services participated in it. APM helps answer different questions:
Which service is slow?
Which transaction is failing?
Where is time spent inside the request?
Is the delay in application code, a database, or another service?
Did a new deployment change latency or error behavior?
Are failures limited to one environment, version, or customer group?
Why teams need both
Imagine that an API becomes slow while CPU utilization rises on several Kubernetes nodes. Infrastructure monitoring shows the resource pressure. APM shows that one service is executing unusually expensive database calls. Logs reveal repeated retries. Deployment metadata shows that the behavior began after a new version was released.
No individual signal provides the full explanation. The useful investigation comes from connecting:
Infrastructure condition → affected service → slow transaction → dependency behavior → relevant logs → deployment context
That connection is the real value of combining infrastructure monitoring and APM.
How Elasticsearch supports infrastructure and application monitoring
Elasticsearch is not the infrastructure agent, application instrumentation library, or visualization interface. Its primary role is to store, index, search, aggregate, and retrieve the telemetry collected from those systems. A simplified Elastic monitoring architecture has five stages:
Instrument applications and collect infrastructure data
Process and enrich the telemetry
Send the data to the Elastic platform
Index the data in Elasticsearch
Investigate and visualize it through Elastic Observability and Kibana
1. Collect infrastructure telemetry
Elastic Agent can collect logs, metrics, and other operational data from hosts. Elastic integrations provide collection and visualization assets for common technologies and services. For infrastructure monitoring, teams may collect data from:
Linux and Windows hosts
Kubernetes clusters
Docker environments
AWS, Microsoft Azure, and Google Cloud
NGINX and web servers
Databases
Message brokers
Network services
Elasticsearch clusters
Prometheus-compatible endpoints
Elastic currently recommends Elastic Agent for infrastructure metric collection. Its System integration can collect host metrics and logs, while other integrations address Kubernetes, cloud platforms, databases, and services. Existing environments may also use Beats, Logstash, upstream OpenTelemetry Collectors, or other supported collection paths. The right design depends on the current environment, deployment model, processing requirements, and migration constraints.
2. Instrument application services
Applications must be instrumented before APM can explain their internal behavior. Instrumentation can collect:
Transactions
Spans
Errors
Service names
Environment information
Runtime metrics
Database activity
External calls
Trace context
Teams can use Elastic-supported instrumentation or OpenTelemetry instrumentation, depending on their languages, architecture, support requirements, and portability goals. OpenTelemetry provides APIs, SDKs, and collection components for telemetry. Elastic supports OpenTelemetry-based ingestion, including paths that use Elastic’s distribution, upstream-compatible components, or managed OTLP endpoints where available.
Instrumentation should be treated as an engineering design decision. Installing an agent does not guarantee useful traces. Service naming, environment attributes, propagation, sampling, sensitive-data controls, and ownership must also be defined.
3. Process and enrich the data
Raw telemetry often needs processing before it becomes consistently searchable. Processing may include:
Parsing log messages
Normalizing field names
Adding service and environment context
Enriching events with cloud or Kubernetes metadata
Redacting sensitive information
Filtering low-value data
Sampling traces
Routing data to appropriate destinations
Converting source-specific fields into a common schema
This stage may happen in the application SDK, Elastic Agent, an OpenTelemetry Collector, Logstash, Elasticsearch ingest pipelines, or a combination of components. More processing layers are not automatically better. Each additional component creates configuration, capacity, failure, and troubleshooting requirements. The architecture should use the smallest number of stages that satisfies the organization’s collection, governance, routing, and resiliency needs.
4. Index telemetry in Elasticsearch
Once telemetry reaches Elasticsearch, its field structure determines how effectively engineers can search and correlate it. Useful fields commonly describe:
Service
Host
Container
Kubernetes workload
Cloud account and region
Deployment environment
Trace and transaction identifiers
Application version
Event dataset
Timestamp
Log level
Error information
Consistent identifiers allow engineers to move between different telemetry signals. For example:
A trace identifies a slow service.
The service record identifies its Kubernetes workload.
Infrastructure data shows resource pressure on the relevant pod and node.
A trace identifier connects the request to related logs.
Version metadata connects the failure to a recent deployment.
Elasticsearch can store all these signals, but correlation depends on the fields being populated correctly. Storing data in the same cluster does not make it automatically related.
5. Investigate through Elastic Observability and Kibana
Elastic Observability provides purpose-built views for infrastructure, hosts, services, traces, logs, alerts, and related operational workflows. Kibana also gives teams flexible search, visualization, dashboard, and analytical capabilities. This matters when a predefined application view does not answer the investigation question.
Engineers may begin with:
A host showing high utilization
An application service with increased latency
A failed synthetic check
An alert
A trace
A log pattern
A Kubernetes workload
A custom Kibana dashboard
The platform is most valuable when the engineer can preserve useful context while moving between those views.
The main components of an Elastic monitoring architecture
Several products and standards appear in an Elastic observability environment. Their roles should not be confused.
Component | Primary role |
Elastic Agent | Collects logs, metrics, and other data from hosts and integrations |
Fleet | Centrally manages Elastic Agents and their policies |
Elastic integrations | Provide collection configurations, field mappings, ingest assets, and dashboards for supported technologies |
Elastic APM instrumentation | Captures application transactions, spans, errors, and runtime data |
OpenTelemetry SDKs | Instrument applications using the OpenTelemetry standard |
OpenTelemetry Collector | Receives, processes, and exports telemetry |
Logstash | Provides flexible pipeline processing and routing where needed |
Elasticsearch ingest pipelines | Transform and enrich documents during ingestion |
Elasticsearch | Stores, indexes, searches, and aggregates telemetry |
Elastic Observability | Provides infrastructure, APM, logs, alerting, and investigation workflows |
Kibana | Provides visualization, dashboards, search, and analysis interfaces |
Not every deployment needs every component. A small environment might use Elastic Agent for host telemetry, application instrumentation for traces, and a managed Elastic deployment. A larger enterprise might use gateway collectors, several processing pipelines, multiple regions, strict data controls, tiered retention, and separate production and monitoring environments. Architecture should reflect actual requirements—not the maximum number of available components.
A practical infrastructure-to-application investigation
Consider an online service with an unexpected increase in checkout latency.
Step 1: Confirm the symptom
Begin with the user-facing or service-level symptom:
Which transaction is slow?
When did the problem start?
Is the increase visible at p50, p95, or p99?
Is it limited to one region, environment, or version?
Did request volume change?
Are errors increasing as well?
An average can hide a severe experience affecting a smaller percentage of requests. Segment the data before concluding that the entire service behaves the same way.
Step 2: Identify the affected service and transaction
Use APM data to determine:
Which service owns the slow transaction
Which transaction types are affected
Whether the time is spent inside the service or a dependency
Whether the problem affects all requests or a subset
Whether a deployment or configuration change aligns with the increase
Inspect representative traces instead of relying only on aggregate charts.
Step 3: Follow the distributed trace
A distributed trace can show the path of the request across instrumented services and dependencies. Inspect:
The longest spans
Database operations
External calls
Repeated requests
Synchronous dependencies
Error events
Missing sections of the trace
Elastic’s service map is based on instrumented services and propagated trace context. If a service is not instrumented, or trace context is not propagated, the relationship may not appear. A missing connection should therefore be treated as an instrumentation question, not proof that the dependency does not exist.
Step 4: Correlate the service with infrastructure
Once the affected service is known, examine the infrastructure supporting it:
Pod CPU and memory
Container restarts
Node pressure
Network behavior
Filesystem or storage latency
Replica availability
Autoscaling events
Resource limits
Cloud-service metrics
The goal is not to find any infrastructure anomaly. It is to determine whether an infrastructure condition aligns with the affected service, time period, and requests. High CPU alone is not a root cause. It may be the cause of slow transactions, the consequence of expensive application behavior, or an unrelated workload occurring on the same infrastructure.
Step 5: Inspect related logs
Logs can provide evidence that traces and metrics do not contain:
Timeout messages
Retry behavior
Connection failures
Application exceptions
Resource-limit messages
Dependency errors
Configuration changes
Deployment events
Where trace and log correlation is configured, trace identifiers can help locate logs associated with a particular request. Without consistent identifiers, engineers may have to approximate the relationship using time windows, service names, hosts, or Kubernetes metadata. That makes the investigation slower and less reliable.
Step 6: Test the leading explanation
After correlating the signals, state the hypothesis precisely. For example:
Checkout latency increased because version 4.2 generated repeated inventory requests, which increased database activity and CPU consumption on the affected pods.
That hypothesis can be tested. A vague conclusion such as “the cluster was overloaded” is much harder to validate and may encourage the wrong response.
Step 7: Verify the remediation
After applying a fix, verify the result using the same measurements that established the problem:
Transaction latency
Error rate
Trace duration
Dependency behavior
Infrastructure utilization
Log patterns
User-facing outcome
A graph returning to normal is useful evidence, but the validation should cover the complete service outcome—not only the metric that first triggered the investigation.
Why correlation depends on telemetry design
Observability platforms are often presented as though correlation appears automatically once logs, metrics, and traces enter the same system. In practice, correlation is an architectural outcome.
Establish consistent service identity
Every service should have a stable and governed identity. Common problems include:
The same service using several names
Different naming conventions across languages
Environment names embedded inconsistently
Temporary deployment identifiers replacing logical service names
Teams using application, repository, container, and Kubernetes names interchangeably
A service-name standard should be defined before instrumentation expands across the organization.
Preserve environment and deployment context
Production, staging, test, and development telemetry must be distinguishable. Useful context includes:
Deployment environment
Application version
Cloud region
Availability zone
Kubernetes namespace
Cluster name
Service instance
Team or ownership metadata
Without this context, engineers can find an error without knowing where it occurred or which deployment introduced it.
Propagate trace context
Distributed tracing depends on trace context moving between services. If propagation fails at a proxy, queue, asynchronous process, or uninstrumented service, a single request can appear as several unrelated traces. Teams should test propagation across:
HTTP services
Message queues
Background workers
Serverless functions
API gateways
Third-party calls
Asynchronous workflows
Align logs with traces
Application logs should contain enough structured context to support correlation. Where appropriate, include:
Trace ID
Transaction or span ID
Service name
Environment
Application version
Request or correlation ID
Error type
Relevant business-safe identifiers
Do not add sensitive information simply because it may help troubleshooting. Logging standards must account for security, privacy, retention, and access-control requirements.
Control cardinality
Telemetry fields with a very large number of unique values can increase storage, indexing, query, and metric-processing costs. Common high-cardinality sources include:
Customer IDs
Session IDs
Request IDs
Container IDs
Kubernetes labels
URLs containing dynamic values
Unbounded exception messages
User-defined attributes
Some high-cardinality fields are operationally valuable. The goal is not to eliminate them indiscriminately. Teams should decide which values belong in indexed fields, logs, traces, metric dimensions, or another controlled representation.
Common implementation mistakes
Mistake 1: Collecting everything before defining questions
More telemetry does not automatically produce better observability. Before expanding collection, define the operational questions the platform must answer:
Which services affect a critical customer journey?
What evidence is needed during an incident?
Which logs are required for diagnosis?
Which infrastructure metrics influence capacity decisions?
Which traces must be retained?
What data must be excluded or redacted?
Collection should follow operational value, not fear of missing data.
Mistake 2: Treating Elasticsearch as the only system to monitor
An Elasticsearch-based monitoring platform must itself be observable. Teams should monitor:
Ingestion delay
Pipeline failures
Rejected work
Storage growth
JVM memory pressure
Query performance
Shard behavior
Collector health
Agent health
Sampling behavior
Data completeness
If the monitoring platform becomes unhealthy during an incident, engineers may lose the evidence they need most.
Mistake 3: Using dashboards as the complete investigation model
Dashboards are useful for detecting patterns and summarizing known conditions. They cannot anticipate every failure. A mature workflow must also support:
Ad hoc search
Trace inspection
Log exploration
Field-level filtering
Time comparison
Service and infrastructure drilldowns
Validation of unusual hypotheses
A dashboard should lead engineers toward evidence, not trap them inside a fixed view. For additional guidance, read Kibana Dashboard Best Practices for Elasticsearch Teams.
Mistake 4: Ignoring data retention by signal
Logs, metrics, and traces do not necessarily need the same retention period. Retention should reflect:
Investigation value
Compliance requirements
Data volume
Query frequency
Storage tier
Sampling strategy
Incident history
Recovery requirements
Keeping every signal in an expensive tier for the same duration is rarely the most defensible design.
Mistake 5: Assuming OpenTelemetry removes backend decisions
OpenTelemetry can improve instrumentation consistency and portability, but it does not eliminate architectural choices. Teams must still decide:
Which signals to collect
Where collectors run
How data is sampled
Which attributes are permitted
How telemetry is enriched
What happens during backpressure
Which destinations receive the data
How schemas are governed
How costs are controlled
Portability should be tested with real telemetry and a realistic export path.
Mistake 6: Scaling before locating the bottleneck
Additional infrastructure may be justified, but it should not be the first answer to every performance problem. A slow investigation experience could be caused by:
Excessive telemetry volume
Poor field mappings
Uncontrolled shard growth
Expensive queries
High-cardinality aggregations
Oversized dashboard time ranges
Pipeline delays
Hot nodes
Slow storage
Application-side behavior
Incomplete correlation fields
Identify the limiting component before paying to scale it.
How to evaluate an Elasticsearch monitoring architecture
A proof of concept should recreate real operational workflows—not merely demonstrate that data reaches a dashboard.
Use representative data
Include:
Normal and peak ingestion
Realistic host and Kubernetes cardinality
Representative application traces
Expected log formats
Production-like retention
Concurrent searches
Failure conditions
Schema changes
A short test with clean sample data may conceal the problems that emerge at production scale.
Reproduce real investigations
Choose two or three incidents that the engineering team has previously experienced. Measure whether the proposed architecture helps engineers:
Detect the symptom
Identify the affected service
Locate representative traces
Find related infrastructure
Retrieve relevant logs
Form a defensible hypothesis
Validate the remediation
Record the number of steps, time required, missing context, query performance, and points where engineers must leave the platform.
Test failure behavior
Observability pipelines also fail. Test:
Collector interruption
Network loss
Backpressure
Endpoint unavailability
Partial telemetry loss
Invalid data
Mapping conflicts
Sudden ingestion growth
Agent-policy errors
Credential or certificate expiration
Document buffering, retry, data-loss, and recovery behavior.
Model cost at realistic growth
Calculate the cost of:
Collection
Processing
Ingestion
Compute
Storage
Replication
Retention
Data transfer
Support
Administration
Engineering ownership
Run the model at current, 2x, 5x, and 10x telemetry volumes. Separate logs, metrics, and traces because their volumes, retention requirements, and collection strategies differ.
When Elasticsearch is the right fit
An Elasticsearch-based monitoring architecture can be a strong fit when an organization:
Needs flexible search across large operational datasets
Relies heavily on logs during investigations
Wants infrastructure, application, and log context in one platform
Already operates Elasticsearch or Kibana successfully
Requires deployment flexibility
Has engineers capable of governing mappings, lifecycle policies, ingestion, and platform performance
Wants to build custom analytical workflows beyond predefined dashboards
Needs to retain and interrogate detailed operational events
It may be a weaker fit when an organization:
Wants minimal platform engineering responsibility
Does not have the skills to operate or govern the environment
Needs a highly guided experience with little customization
Cannot control telemetry growth
Expects correlation without investing in instrumentation and schema standards
Has a small, simple environment that does not justify the operational model
The question is not whether Elasticsearch can store monitoring data. It can. The more important question is whether the organization can design and operate the surrounding collection, instrumentation, correlation, retention, and investigation system effectively.
Frequently asked questions
Is Elasticsearch an application performance monitoring tool?
Elasticsearch is the search and analytics data layer used by Elastic’s APM and observability capabilities. Application instrumentation, telemetry collection, processing, and purpose-built investigation interfaces are provided through other components such as Elastic APM, Elastic Agent, OpenTelemetry, Elastic integrations, and Kibana.
What is the difference between Elastic APM and infrastructure monitoring?
Elastic APM focuses on application transactions, traces, dependencies, errors, and service performance. Infrastructure monitoring focuses on hosts, containers, Kubernetes environments, cloud systems, and their resource behavior. They become more useful when service, host, container, trace, deployment, and environment context can be correlated.
Can Elastic monitor Kubernetes infrastructure and applications?
Yes. Elastic provides Kubernetes integrations for collecting cluster logs and metrics, while instrumented applications can send traces and application-performance data. The quality of the combined investigation depends on consistent Kubernetes metadata, service identity, environment fields, and trace propagation.
Can OpenTelemetry send application data to Elastic?
Yes. OpenTelemetry-compatible SDKs and Collectors can send telemetry to supported Elastic ingestion endpoints. The exact architecture depends on the Elastic deployment model, language support, chosen Collector distribution, and processing requirements.
Does putting logs, metrics, and traces in Elasticsearch automatically correlate them?No. The signals need shared, consistently populated identifiers and metadata. Service names, environments, trace identifiers, infrastructure metadata, timestamps, and deployment information must be designed and validated.
Should every application transaction be traced?
Not necessarily. The appropriate strategy depends on transaction volume, diagnostic value, compliance requirements, storage cost, and the sampling capabilities in use.
Sampling decisions should preserve the traces needed to investigate errors, high latency, and important business transactions.
Can Elastic replace separate infrastructure and APM tools?
Potentially, but replacement should be proven through a production-like evaluation.
Confirm telemetry coverage, investigation workflows, integrations, instrumentation support, alerting, governance, failure behavior, retention, cost, and operational ownership before consolidating existing tools.
Final takeaway
Infrastructure monitoring tells engineers what is happening across the systems running an application. APM shows how application requests and dependencies behave. Logs add detailed event evidence. Traces connect work across services. Elasticsearch makes these signals searchable at scale, but successful correlation depends on the architecture around it.
Teams must establish consistent service identity, propagate trace context, preserve infrastructure metadata, structure logs, govern schemas, control cardinality, and design retention by signal. Without that foundation, the organization may collect large amounts of telemetry while still struggling to explain incidents. The goal is not to place every operational event in one database. It is to create a monitoring system that helps engineers move from symptom to evidence, test a hypothesis, and verify the result.
DinaBridge provides Elasticsearch Consulting Services for organizations designing, improving, or scaling production search and observability environments. If your infrastructure metrics, application traces, and logs remain disconnected—or your Elastic environment is becoming difficult to operate, discuss your platform challenge with DinaBridge.
References
Next article
