Observability

Written by
Dina Bridge
|
Subscribe
Subscribe to get the latest insights straight in your inbox
A Kibana dashboard should do more than display Elasticsearch data. During an incident, it should help an engineer recognize what changed, narrow the affected scope, and reach the underlying evidence without losing context.
That is a higher standard than arranging charts on a page. A dashboard can look polished and still fail when the on-call engineer cannot tell whether an error spike is global, isolated to one service, or caused by a recent deployment. This guide explains how to build Kibana dashboards for real operational use. It focuses on investigation speed, consistent data, useful context, dashboard performance, and maintainability across teams.
1. Start with an investigation question
Do not begin by asking which charts to add. Begin with the decisions the dashboard must support. A useful dashboard usually answers one primary question, such as:
Is the service currently healthy?
Which endpoint is responsible for the latency increase?
Is the failure limited to one region, version, host, or customer segment?
Did the error rate change after a deployment?
Are infrastructure symptoms correlated with application failures?
Which logs or traces contain the evidence needed for diagnosis?
Write the question at the top of the dashboard description. Then define the intended user and action. An SRE service overview, an application-owner dashboard, and an executive availability summary should not be the same page. This discipline prevents the “one dashboard for everyone” problem. When a page tries to serve every audience, it usually becomes crowded, slow, and difficult to interpret.
Define the dashboard contract
Before building, document five items:
Item | Example |
Primary user | On-call SRE |
Question | Why did checkout latency increase? |
Decision | Escalate, roll back, or continue investigating |
Required dimensions | Service, environment, region, version, endpoint |
Evidence path | Overview → trace sample → related logs |
If the team cannot define that contract, it is not ready to choose visualizations.
2. Use a layered dashboard structure
Operational dashboards work best when they reveal information in layers. The first screen should explain the current state. More detailed views should help isolate the cause. A practical structure is:
Layer 1: Current health
Place the few signals that indicate whether the service is operating normally at the top:
Request rate
Error rate
Latency, preferably relevant percentiles rather than an average alone
Availability or SLO status
Active alerts
These panels should answer, “Is there a problem, and when did it begin?”
Layer 2: Scope and correlation
Use the middle of the dashboard to segment the problem by dimensions such as service, endpoint, region, availability zone, host, pod, deployment version, or error type. This layer should answer, “Where is the problem concentrated, and what changed with it?”
Layer 3: Evidence
The lower section should provide a path to detailed events: representative errors, slow transactions, trace samples, affected hosts, and relevant log records. Do not force users to inspect every row on the overview page. Provide enough evidence to select the next investigative path, then use a drilldown to continue.
3. Choose visualizations that support decisions
The correct visualization depends on the question, not on what looks sophisticated.
Question | Useful visualization |
Is a metric changing over time? | Line or area chart |
What is the current value? | Metric panel with clear unit |
Which services or error types contribute most? | Sorted bar chart or table |
Is latency degrading across percentiles? | Time series with p50, p95, and p99 |
Where are failures occurring? | Map only when geography is operationally relevant |
What events require inspection? | Table with timestamp and diagnostic fields |
Avoid pie charts with many slices, gauges without meaningful thresholds, and metrics that lack units. A number such as “12.4” is not useful unless the viewer knows whether it means milliseconds, seconds, percent, or thousands of events. Use consistent colors across dashboards. For example, red should not mean errors on one dashboard and ordinary traffic on another. Do not rely on color alone; titles, labels, and values must communicate the state as well.
Kibana Lens is the default visualization editor and can be used to build charts, tables, and metrics from Elasticsearch data. Elastic also supports ES|QL-powered visualizations. Use the simplest option that expresses the operational question clearly; complexity in the query or chart must earn its place.
4. Design filters around operational dimensions
A useful Kibana dashboard lets an engineer narrow the scope without rewriting a query. Common controls include:
Environment
Service name
Region or availability zone
Kubernetes cluster and namespace
Host or pod
Deployment version
Transaction or endpoint name
Log level
Customer or tenant identifier, when access controls permit it
Elastic supports options-list, range, and time-based controls. Its current documentation also allows controls to be populated from a data-view field or an ES|QL query. This can be useful when a field has high cardinality and returning every value would be impractical. Keep the control set deliberate. Ten poorly ordered filters create work instead of reducing it. Put the filters most likely to divide an incident near the top, use human-readable labels, and set sensible defaults. Filters only work reliably when telemetry uses consistent fields. If one team records the production environment as prod, another as production, and a third omits it, the dashboard will provide a fragmented view. Fix the schema or ingestion pipeline rather than hiding the inconsistency in the visualization.
5. Preserve context with drilldowns
An overview dashboard should not contain every detail. It should provide reliable navigation to the next level. Kibana supports dashboard, Discover, and URL drilldowns. A dashboard drilldown can carry the selected value, time range, query, and filters into a more specific dashboard. A Discover drilldown can open the underlying documents associated with a visualization selection. URL drilldowns can connect an operational value to another approved system.
Useful paths include:
Service overview → endpoint detail
Region overview → host or pod detail
Error category → filtered log records in Discover
Latency spike → slow transaction or trace view
Deployment marker → approved deployment record
Host selection → infrastructure investigation dashboard
Preserving context is critical. If an engineer selects a service, region, and 15-minute incident window, the destination should not reset to all services over the last 24 hours. Drilldowns also impose a design requirement: important operational dimensions should exist as indexed fields. Elastic notes that query-time computed values do not support every filtering or drilldown action because no matching field exists in the underlying index. Plan the telemetry model with the investigation workflow in mind.
6. Connect dashboards to raw evidence
Aggregations identify patterns; individual events confirm what happened. Every incident dashboard should provide a short path from a visual symptom to the supporting documents. For logs, include fields that help the user decide whether to open a record:
@timestampservice.nameservice.versionenvironmentor the organization’s standardized equivalenthost.name,container.id, orkubernetes.pod.nametrace.idandtransaction.id, when availableError type and message
HTTP route and status code
Cloud region or availability zone
Do not add every available field. Choose the fields that help distinguish one failure from another. When traces, logs, and metrics share consistent resource attributes and correlation identifiers, engineers can move between signals with less manual searching. This is one reason telemetry design must be treated as part of dashboard design rather than as a separate ingestion task.
7. Make time and comparison explicit
Time is often the most important incident filter. A dashboard should make the selected range obvious and should behave predictably when users change it. Apply these practices:
Use a global time range unless a panel has a documented reason to differ.
Label panels that use a custom time range.
Select an automatic refresh interval that matches the operational need and system cost.
Add deployment or change annotations when the data supports them.
Compare the incident window with an appropriate baseline.
Use percentiles and rates when totals would be distorted by traffic changes.
A prior-period comparison can be useful, but it is not always valid. Traffic by hour of day, weekday, season, or release cycle may make “the previous hour” a misleading baseline. State what is being compared. Do not save a dashboard with an unexpectedly narrow time range unless that behavior is intentional. Elastic allows the current time filter to be stored with a dashboard; use that setting carefully so future viewers do not mistake an old window for current health.
8. Protect dashboard performance
A slow dashboard damages the investigation it is supposed to support. Performance problems commonly come from expensive queries, excessive panels, broad time ranges, high-cardinality aggregations, and inefficient field or index design. Review:
The number of panels loaded at once
Query latency at normal and peak ingestion
Default and maximum time ranges
Aggregation cardinality and bucket size
Wildcard and unbounded searches
Runtime or computed fields used repeatedly
Data-view scope
Refresh frequency
Shard, mapping, and index design underneath the dashboard
Elastic permits many panels on a dashboard, but a supported maximum is not a design target. If users need dozens of panels to answer unrelated questions, split the page by workflow and connect the views with drilldowns. Measure dashboard load time using production-like data and realistic user concurrency. A dashboard that performs well on a sample index may degrade substantially after retention, cardinality, and data volume increase.
9. Standardize ownership and governance
Dashboards become unreliable when no one owns their data, thresholds, or meaning. For every production dashboard, record:
Business or service purpose
Technical owner
Data sources and required fields
Alert and threshold definitions
Intended audience
Last validation date
Expected review cadence
Related runbook or investigation dashboard
Use clear titles, descriptions, and tags so users can find the correct dashboard. Define naming conventions for service, environment, team, and lifecycle status. Separate draft or experimental content from approved operational dashboards. Reusable library panels can improve consistency, but changes may affect every dashboard that uses them. Elastic distinguishes between panels saved in the Visualize Library and panels created locally on a dashboard. Decide which elements should be centrally maintained and which should remain specific to one workflow. For repeatable environments, consider managing approved dashboards and visualizations as code. Elastic documents APIs and command-line workflows for programmatic dashboard management. Whatever method you choose, test promotion across development, staging, and production rather than editing critical content informally in each environment.
10. Test the dashboard with a real incident
Dashboard review should be task-based, not aesthetic. Reproduce a known incident or create a controlled scenario. Give the dashboard to an SRE, application developer, and platform engineer who did not build it. Ask each person to identify:
What changed?
When did it begin?
Which component or segment is affected?
How large is the impact?
What evidence supports the conclusion?
What should be investigated next?
Measure time to first useful hypothesis, time to supporting evidence, failed clicks, query rewrites, and requests for help. Watch where users hesitate. A title that seems obvious to the creator may be ambiguous to everyone else. Repeat the test after schema changes, new service versions, and significant growth. Dashboards are operational software: they require validation and maintenance.
Kibana dashboard review checklist
Before approving a dashboard for production use, confirm:
The dashboard has one primary operational purpose.
The intended user and action are documented.
Health, scope, and evidence are presented in a logical order.
Every metric has a clear unit and definition.
Filters use consistent, governed fields.
Time range and refresh behavior are obvious.
Important selections preserve context through drilldowns.
Users can reach the underlying logs or events.
Colors and thresholds have consistent meanings.
The dashboard performs at production volume and concurrency.
Access controls prevent inappropriate data exposure.
An owner and review date are recorded.
The workflow has been tested with a realistic incident.
Frequently asked questions
What should a Kibana dashboard include?
A production Kibana dashboard should include the minimum signals needed to establish current health, isolate the affected scope, and reach supporting evidence. The exact panels depend on the user and investigation question; more panels do not automatically create better visibility.
How many visualizations should a Kibana dashboard have?
There is no universal correct number. Use only the panels required for one operational workflow. If the dashboard mixes unrelated questions or becomes slow and difficult to scan, split it into overview and detail dashboards connected by drilldowns.
What is the difference between a filter and a drilldown in Kibana?
A filter narrows the data shown on the current dashboard. A drilldown opens another dashboard, Discover, or a URL while carrying relevant context such as the selected value, filters, query, and time range.
Why is a Kibana dashboard slow?
Common causes include expensive queries, too many panels, long time ranges, high-cardinality aggregations, repeated computed fields, frequent refreshes, and inefficient Elasticsearch mappings or shard design. Diagnose the queries and data model instead of treating the layout alone.
Should every team build its own dashboards?
Teams should own the views needed for their services, but shared field conventions, naming standards, access controls, reusable panels, and review requirements should be governed centrally. This balances local usefulness with organization-wide consistency.
Final takeaway
The best Kibana dashboard is not the one with the most data. It is the one that helps the intended user move from a visible symptom to a defensible next action with minimal delay. Build around an investigation question. Structure the page from health to scope to evidence. Preserve time and filter context. Test performance with realistic data. Then validate the experience with people who will actually use it during an incident. DinaBridge provides Elasticsearch Consulting Services for teams designing, improving, and governing production Elastic environments. If your dashboards are slow, difficult to maintain, or disconnected from real incident workflows, discuss your platform challenge with DinaBridge.
References
Next article
