Observability

Kibana Dashboard Best Practices for Incident Investigation

Kibana Dashboard Best Practices for Incident Investigation

Kibana Dashboard Best Practices for Incident Investigation

Green light beam illuminating a futuristic circuit board
No headings found on page

Written by

Dina Bridge

|

Subscribe

Subscribe to get the latest insights straight in your inbox

A Kibana dashboard should do more than display Elasticsearch data. During an incident, it should help an engineer recognize what changed, narrow the affected scope, and reach the underlying evidence without losing context.

That is a higher standard than arranging charts on a page. A dashboard can look polished and still fail when the on-call engineer cannot tell whether an error spike is global, isolated to one service, or caused by a recent deployment. This guide explains how to build Kibana dashboards for real operational use. It focuses on investigation speed, consistent data, useful context, dashboard performance, and maintainability across teams.

1. Start with an investigation question

Do not begin by asking which charts to add. Begin with the decisions the dashboard must support. A useful dashboard usually answers one primary question, such as:

  • Is the service currently healthy?

  • Which endpoint is responsible for the latency increase?

  • Is the failure limited to one region, version, host, or customer segment?

  • Did the error rate change after a deployment?

  • Are infrastructure symptoms correlated with application failures?

  • Which logs or traces contain the evidence needed for diagnosis?

Write the question at the top of the dashboard description. Then define the intended user and action. An SRE service overview, an application-owner dashboard, and an executive availability summary should not be the same page. This discipline prevents the “one dashboard for everyone” problem. When a page tries to serve every audience, it usually becomes crowded, slow, and difficult to interpret.

Define the dashboard contract

Before building, document five items:

Item

Example

Primary user

On-call SRE

Question

Why did checkout latency increase?

Decision

Escalate, roll back, or continue investigating

Required dimensions

Service, environment, region, version, endpoint

Evidence path

Overview → trace sample → related logs

If the team cannot define that contract, it is not ready to choose visualizations.

2. Use a layered dashboard structure

Operational dashboards work best when they reveal information in layers. The first screen should explain the current state. More detailed views should help isolate the cause. A practical structure is:

Layer 1: Current health
Place the few signals that indicate whether the service is operating normally at the top:

  • Request rate

  • Error rate

  • Latency, preferably relevant percentiles rather than an average alone

  • Availability or SLO status

  • Active alerts

These panels should answer, “Is there a problem, and when did it begin?”

Layer 2: Scope and correlation
Use the middle of the dashboard to segment the problem by dimensions such as service, endpoint, region, availability zone, host, pod, deployment version, or error type. This layer should answer, “Where is the problem concentrated, and what changed with it?”

Layer 3: Evidence
The lower section should provide a path to detailed events: representative errors, slow transactions, trace samples, affected hosts, and relevant log records. Do not force users to inspect every row on the overview page. Provide enough evidence to select the next investigative path, then use a drilldown to continue.

3. Choose visualizations that support decisions

The correct visualization depends on the question, not on what looks sophisticated.

Question

Useful visualization

Is a metric changing over time?

Line or area chart

What is the current value?

Metric panel with clear unit

Which services or error types contribute most?

Sorted bar chart or table

Is latency degrading across percentiles?

Time series with p50, p95, and p99

Where are failures occurring?

Map only when geography is operationally relevant

What events require inspection?

Table with timestamp and diagnostic fields

Avoid pie charts with many slices, gauges without meaningful thresholds, and metrics that lack units. A number such as “12.4” is not useful unless the viewer knows whether it means milliseconds, seconds, percent, or thousands of events. Use consistent colors across dashboards. For example, red should not mean errors on one dashboard and ordinary traffic on another. Do not rely on color alone; titles, labels, and values must communicate the state as well.

Kibana Lens is the default visualization editor and can be used to build charts, tables, and metrics from Elasticsearch data. Elastic also supports ES|QL-powered visualizations. Use the simplest option that expresses the operational question clearly; complexity in the query or chart must earn its place.

4. Design filters around operational dimensions

A useful Kibana dashboard lets an engineer narrow the scope without rewriting a query. Common controls include:

  • Environment

  • Service name

  • Region or availability zone

  • Kubernetes cluster and namespace

  • Host or pod

  • Deployment version

  • Transaction or endpoint name

  • Log level

  • Customer or tenant identifier, when access controls permit it

Elastic supports options-list, range, and time-based controls. Its current documentation also allows controls to be populated from a data-view field or an ES|QL query. This can be useful when a field has high cardinality and returning every value would be impractical. Keep the control set deliberate. Ten poorly ordered filters create work instead of reducing it. Put the filters most likely to divide an incident near the top, use human-readable labels, and set sensible defaults. Filters only work reliably when telemetry uses consistent fields. If one team records the production environment as prod, another as production, and a third omits it, the dashboard will provide a fragmented view. Fix the schema or ingestion pipeline rather than hiding the inconsistency in the visualization.

5. Preserve context with drilldowns

An overview dashboard should not contain every detail. It should provide reliable navigation to the next level. Kibana supports dashboard, Discover, and URL drilldowns. A dashboard drilldown can carry the selected value, time range, query, and filters into a more specific dashboard. A Discover drilldown can open the underlying documents associated with a visualization selection. URL drilldowns can connect an operational value to another approved system.

Useful paths include:

  • Service overview → endpoint detail

  • Region overview → host or pod detail

  • Error category → filtered log records in Discover

  • Latency spike → slow transaction or trace view

  • Deployment marker → approved deployment record

  • Host selection → infrastructure investigation dashboard

Preserving context is critical. If an engineer selects a service, region, and 15-minute incident window, the destination should not reset to all services over the last 24 hours. Drilldowns also impose a design requirement: important operational dimensions should exist as indexed fields. Elastic notes that query-time computed values do not support every filtering or drilldown action because no matching field exists in the underlying index. Plan the telemetry model with the investigation workflow in mind.

6. Connect dashboards to raw evidence

Aggregations identify patterns; individual events confirm what happened. Every incident dashboard should provide a short path from a visual symptom to the supporting documents. For logs, include fields that help the user decide whether to open a record:

  • @timestamp

  • service.name

  • service.version

  • environment or the organization’s standardized equivalent

  • host.name, container.id, or kubernetes.pod.name

  • trace.id and transaction.id, when available

  • Error type and message

  • HTTP route and status code

  • Cloud region or availability zone

Do not add every available field. Choose the fields that help distinguish one failure from another. When traces, logs, and metrics share consistent resource attributes and correlation identifiers, engineers can move between signals with less manual searching. This is one reason telemetry design must be treated as part of dashboard design rather than as a separate ingestion task.

7. Make time and comparison explicit

Time is often the most important incident filter. A dashboard should make the selected range obvious and should behave predictably when users change it. Apply these practices:

  • Use a global time range unless a panel has a documented reason to differ.

  • Label panels that use a custom time range.

  • Select an automatic refresh interval that matches the operational need and system cost.

  • Add deployment or change annotations when the data supports them.

  • Compare the incident window with an appropriate baseline.

  • Use percentiles and rates when totals would be distorted by traffic changes.

A prior-period comparison can be useful, but it is not always valid. Traffic by hour of day, weekday, season, or release cycle may make “the previous hour” a misleading baseline. State what is being compared. Do not save a dashboard with an unexpectedly narrow time range unless that behavior is intentional. Elastic allows the current time filter to be stored with a dashboard; use that setting carefully so future viewers do not mistake an old window for current health.

8. Protect dashboard performance

A slow dashboard damages the investigation it is supposed to support. Performance problems commonly come from expensive queries, excessive panels, broad time ranges, high-cardinality aggregations, and inefficient field or index design. Review:

  • The number of panels loaded at once

  • Query latency at normal and peak ingestion

  • Default and maximum time ranges

  • Aggregation cardinality and bucket size

  • Wildcard and unbounded searches

  • Runtime or computed fields used repeatedly

  • Data-view scope

  • Refresh frequency

  • Shard, mapping, and index design underneath the dashboard

Elastic permits many panels on a dashboard, but a supported maximum is not a design target. If users need dozens of panels to answer unrelated questions, split the page by workflow and connect the views with drilldowns. Measure dashboard load time using production-like data and realistic user concurrency. A dashboard that performs well on a sample index may degrade substantially after retention, cardinality, and data volume increase.

9. Standardize ownership and governance

Dashboards become unreliable when no one owns their data, thresholds, or meaning. For every production dashboard, record:

  • Business or service purpose

  • Technical owner

  • Data sources and required fields

  • Alert and threshold definitions

  • Intended audience

  • Last validation date

  • Expected review cadence

  • Related runbook or investigation dashboard

Use clear titles, descriptions, and tags so users can find the correct dashboard. Define naming conventions for service, environment, team, and lifecycle status. Separate draft or experimental content from approved operational dashboards. Reusable library panels can improve consistency, but changes may affect every dashboard that uses them. Elastic distinguishes between panels saved in the Visualize Library and panels created locally on a dashboard. Decide which elements should be centrally maintained and which should remain specific to one workflow. For repeatable environments, consider managing approved dashboards and visualizations as code. Elastic documents APIs and command-line workflows for programmatic dashboard management. Whatever method you choose, test promotion across development, staging, and production rather than editing critical content informally in each environment.

10. Test the dashboard with a real incident

Dashboard review should be task-based, not aesthetic. Reproduce a known incident or create a controlled scenario. Give the dashboard to an SRE, application developer, and platform engineer who did not build it. Ask each person to identify:

  1. What changed?

  2. When did it begin?

  3. Which component or segment is affected?

  4. How large is the impact?

  5. What evidence supports the conclusion?

  6. What should be investigated next?

Measure time to first useful hypothesis, time to supporting evidence, failed clicks, query rewrites, and requests for help. Watch where users hesitate. A title that seems obvious to the creator may be ambiguous to everyone else. Repeat the test after schema changes, new service versions, and significant growth. Dashboards are operational software: they require validation and maintenance.

Kibana dashboard review checklist

Before approving a dashboard for production use, confirm:

  • The dashboard has one primary operational purpose.

  • The intended user and action are documented.

  • Health, scope, and evidence are presented in a logical order.

  • Every metric has a clear unit and definition.

  • Filters use consistent, governed fields.

  • Time range and refresh behavior are obvious.

  • Important selections preserve context through drilldowns.

  • Users can reach the underlying logs or events.

  • Colors and thresholds have consistent meanings.

  • The dashboard performs at production volume and concurrency.

  • Access controls prevent inappropriate data exposure.

  • An owner and review date are recorded.

  • The workflow has been tested with a realistic incident.

Frequently asked questions

What should a Kibana dashboard include?
A production Kibana dashboard should include the minimum signals needed to establish current health, isolate the affected scope, and reach supporting evidence. The exact panels depend on the user and investigation question; more panels do not automatically create better visibility.

How many visualizations should a Kibana dashboard have?
There is no universal correct number. Use only the panels required for one operational workflow. If the dashboard mixes unrelated questions or becomes slow and difficult to scan, split it into overview and detail dashboards connected by drilldowns.

What is the difference between a filter and a drilldown in Kibana?
A filter narrows the data shown on the current dashboard. A drilldown opens another dashboard, Discover, or a URL while carrying relevant context such as the selected value, filters, query, and time range.

Why is a Kibana dashboard slow?
Common causes include expensive queries, too many panels, long time ranges, high-cardinality aggregations, repeated computed fields, frequent refreshes, and inefficient Elasticsearch mappings or shard design. Diagnose the queries and data model instead of treating the layout alone.

Should every team build its own dashboards?
Teams should own the views needed for their services, but shared field conventions, naming standards, access controls, reusable panels, and review requirements should be governed centrally. This balances local usefulness with organization-wide consistency.

Final takeaway

The best Kibana dashboard is not the one with the most data. It is the one that helps the intended user move from a visible symptom to a defensible next action with minimal delay. Build around an investigation question. Structure the page from health to scope to evidence. Preserve time and filter context. Test performance with realistic data. Then validate the experience with people who will actually use it during an incident. DinaBridge provides Elasticsearch Consulting Services for teams designing, improving, and governing production Elastic environments. If your dashboards are slow, difficult to maintain, or disconnected from real incident workflows, discuss your platform challenge with DinaBridge.

References