Search

Running Elasticsearch on GKE at Scale: When Reliability Becomes an Architecture Problem

Running Elasticsearch on GKE at Scale: When Reliability Becomes an Architecture Problem

Running Elasticsearch on GKE at Scale: When Reliability Becomes an Architecture Problem

No headings found on page

Written by

Dina Bridge

|

Subscribe

Subscribe to get the latest insights straight in your inbox

Google Kubernetes Engine can make infrastructure easier to provision, automate, and scale. But for teams already running Elasticsearch on GKE, infrastructure elasticity does not automatically translate into a scalable Elasticsearch architecture. That distinction often becomes visible only after the environment grows. More data arrives. Search or observability workloads increase. More teams depend on the cluster. Recovery becomes slower. Infrastructure costs rise. Engineers add capacity, adjust Kubernetes resources, or increase storage, and the environment stabilizes temporarily before another bottleneck appears.

At that point, the question for a VP of Engineering, Head of Platform, or Director of SRE is no longer simply: “Do we need more GKE capacity?”

It becomes: “Is the Elasticsearch architecture still appropriate for the workload we are asking it to support?”

That is an architecture question, and answering it requires looking at Elasticsearch and Kubernetes as two interconnected distributed systems.

Why Elasticsearch on GKE Gets Harder as It Scales

Running Elasticsearch on Kubernetes can initially look straightforward. Google Kubernetes Engine manages the Kubernetes control plane and provides the orchestration layer. Kubernetes gives teams mechanisms for scheduling, persistent storage, resource allocation, upgrades, and workload management. Google documents StatefulSets and persistent volumes specifically for stateful workloads, while Elastic Cloud on Kubernetes (ECK) provides Kubernetes-native orchestration for Elasticsearch and other Elastic components. The difficult part begins when the workload becomes significant. An Elasticsearch cluster that supported a modest workload six months ago may now be handling substantially more:

  • indexed data

  • queries

  • concurrent users

  • ingestion pipelines

  • retention requirements

  • dashboards

  • application dependencies

  • recovery expectations

The Kubernetes infrastructure may still be functioning exactly as designed. Elasticsearch may still report a functioning cluster. But the system as a whole can become increasingly expensive and difficult to operate. That is where platform leaders need to distinguish infrastructure scaling from Elasticsearch scaling. They are related, but they are not the same problem.

The Two Systems Your Team Is Actually Operating

An Elasticsearch deployment on GKE combines two systems with their own scheduling, resource, availability, and recovery behavior.

GKE manages infrastructure orchestration

At the Kubernetes layer, the team is dealing with concepts such as:

  • nodes

  • pods

  • CPU and memory requests

  • persistent volumes

  • StatefulSets

  • scheduling

  • node maintenance

  • disruption budgets

  • autoscaling

Elasticsearch manages the distributed data system

At the Elasticsearch layer, the team must consider:

  • indices

  • primary and replica shards

  • node roles

  • heap and JVM pressure

  • data tiers

  • ingestion

  • query workload

  • shard allocation

  • cluster state

  • recovery

  • snapshots

  • retention

A decision that appears sensible from the Kubernetes side can therefore have very different consequences for Elasticsearch. This is one reason scaling Elasticsearch on GKE should not be treated like scaling a stateless application.

Five Signs Capacity Is No Longer the Main Problem

Adding capacity is appropriate when the underlying architecture is sound and the workload simply requires more resources. But repeatedly adding resources without understanding the bottleneck can turn architectural inefficiency into infrastructure spend. For a VP or Director, five patterns should trigger a deeper review.

1. GKE capacity keeps increasing without proportional performance improvement

The team adds nodes, memory, CPU, or storage, but search latency, indexing performance, or stability does not improve proportionally. The immediate pressure may disappear. The underlying constraint remains.

2. Routine maintenance creates disproportionate operational risk

Node maintenance, Kubernetes upgrades, Elasticsearch upgrades, or pod rescheduling require unusually careful intervention. Operations that should be controlled events start becoming production concerns.

3. Recovery takes longer as the environment grows

A failed or rescheduled Elasticsearch node triggers significant shard movement and recovery activity. The cluster technically recovers, but recovery consumes enough resources to affect production workloads.

4. Engineers spend increasing time managing cluster behavior

Platform or SRE engineers repeatedly investigate:

  • allocation problems

  • JVM pressure

  • storage saturation

  • slow queries

  • indexing backlogs

  • unassigned shards

  • resource contention

  • unstable nodes

This is an important organizational signal. The cost of the platform is no longer limited to the Google Cloud bill. It now includes increasing engineering attention.

5. Nobody can confidently explain the next scaling threshold

Perhaps the clearest warning sign is uncertainty. If another 30% or 50% increase in workload arrived, would the team know:

  • which Elasticsearch tier needs capacity?

  • how many additional nodes are required?

  • whether shard distribution remains appropriate?

  • whether storage can sustain the workload?

  • what happens during a node failure?

  • how recovery time changes?

  • how much the additional capacity should cost?

If the answer is unclear, the organization has a capacity-planning problem—not merely a capacity problem.

Why Adding GKE Nodes Can Hide Elasticsearch Problems

Kubernetes makes horizontal expansion accessible.

That can create a dangerous operational habit:
performance problem → add infrastructure → temporary improvement → repeat.

Sometimes that is exactly the correct response. Sometimes it is expensive symptom management. Consider an Elasticsearch cluster with inefficient shard distribution. Adding data nodes gives Elasticsearch additional resources, but it does not necessarily address why the workload is inefficiently distributed in the first place. The same applies to poorly matched node roles, inappropriate storage, excessive shard counts, inefficient indexing patterns, or resource contention.

Elastic's ECK autoscaling capabilities demonstrate this distinction clearly. ECK can adjust pod counts and allocated resources for supported Elasticsearch tiers within defined policies, but autoscaling still operates within architectural boundaries established by the deployment.

Autoscaling is a capacity mechanism. It is not a substitute for capacity architecture.

For a platform leader, that distinction matters because infrastructure can continue scaling, and billing, while the underlying efficiency of the Elasticsearch environment deteriorates.

Storage Can Become a Reliability Problem

Elasticsearch is stateful. That makes storage architecture fundamental to performance and recovery. GKE supports persistent storage for stateful applications, and Kubernetes associates persistent volumes with StatefulSet workloads.

But simply having persistent storage does not mean the storage architecture is appropriate for Elasticsearch. Elastic recommends evaluating storage based on the expected Elasticsearch workload and notes that storage options have different performance characteristics.

At scale, platform teams need to understand:

  • latency

  • throughput

  • IOPS requirements

  • volume expansion behavior

  • failure characteristics

  • recovery behavior

  • snapshot strategy

The question is not: Does Elasticsearch have persistent storage?

It is: Does the storage architecture support Elasticsearch under normal load, peak load, and recovery conditions?

Those are very different standards.

Shard Architecture Still Determines Elasticsearch Behavior

Kubernetes does not eliminate Elasticsearch fundamentals. Shards remain one of the most consequential architectural decisions in an Elasticsearch environment. Too many shards can create unnecessary overhead. Poorly sized shards can complicate recovery and resource utilization.

Inappropriate index and retention strategies can gradually turn what was once a healthy deployment into a difficult cluster to operate. This is particularly important because shard problems often develop incrementally. The environment works. Data grows. New indices appear. Retention increases. More applications begin using Elasticsearch. Eventually the architecture that worked at the original scale is still running, but the workload around it has changed substantially. Adding Kubernetes capacity may provide more room for that architecture to operate. It does not automatically make the architecture appropriate again.

Kubernetes Scheduling and Elasticsearch Availability Must Align

High availability is not simply a matter of having multiple Elasticsearch pods. Those pods must also be placed and disrupted in ways that preserve the availability assumptions of the Elasticsearch cluster. ECK supports Kubernetes scheduling controls and manages PodDisruptionBudgets for Elasticsearch resources. PodDisruptionBudgets limit voluntary disruption during operations such as Kubernetes node maintenance. But platform teams still need to understand the failure domains they are designing around.

For example:

  • Where are master-eligible nodes scheduled?

  • Where are replica shards located?

  • Can multiple critical pods disappear during the same infrastructure event?

  • What happens during node maintenance?

  • What happens when storage becomes unavailable?

  • What does recovery look like under production traffic?

These questions become especially important as uptime requirements increase. A Kubernetes deployment can be highly available from an orchestration perspective while the Elasticsearch architecture still contains concentrated failure risk.

ECK Helps With Operations, but It Does Not Design the Architecture

Elastic Cloud on Kubernetes is extremely useful for teams operating Elasticsearch on Kubernetes. ECK provides Kubernetes-native management capabilities for Elasticsearch and other Elastic applications and handles areas such as orchestration, configuration, certificates, updates, and integrations with Kubernetes lifecycle mechanisms.

That reduces significant operational work. But adopting ECK does not remove the need to make Elasticsearch architecture decisions. ECK can orchestrate what you define. It cannot decide the business requirements behind that definition. Your team still needs to determine:

  • appropriate node roles

  • resource allocation

  • shard strategy

  • data tiers

  • storage characteristics

  • resilience requirements

  • recovery objectives

  • workload separation

  • capacity thresholds

  • scaling boundaries

This distinction is important for leadership. A well-operated deployment is not necessarily a well-designed deployment.

The Cost of an Inefficient Elasticsearch-on-GKE Architecture

Infrastructure inefficiency becomes particularly visible in cloud environments because capacity has a direct recurring cost. Suppose Elasticsearch begins experiencing performance pressure. The team increases:

  • GKE nodes

  • CPU

  • memory

  • persistent storage

Performance improves. Three months later, the same pattern returns.

More capacity is added. The question leadership should ask is not whether the additional resources helped.

It is: Did workload growth justify the increase in infrastructure, or are we paying to compensate for architectural inefficiency?

The distinction can materially affect the economics of the platform. There is also a second cost: engineering time. A cluster that requires continuous manual intervention consumes SRE and platform engineering capacity that could otherwise support product delivery, reliability improvements, or platform modernization.

The true operating cost therefore becomes: Cloud infrastructure + Elastic resources + engineering effort + reliability risk.

Optimizing only the first line item misses most of the problem.

A Diagnostic Framework for Platform Leaders

A VP or Director does not need to troubleshoot individual Elasticsearch nodes. They do need a reliable way to determine whether the platform requires deeper architectural work. Start with four questions.

1. Reliability

Can the environment tolerate expected failures without material service degradation? Look at:

  • node failures

  • pod rescheduling

  • maintenance

  • upgrades

  • recovery duration

  • shard availability

2. Scalability

Can the team explain how Elasticsearch should scale for the next stage of workload growth? Look at:

  • data growth

  • ingestion growth

  • query concurrency

  • shard growth

  • storage requirements

  • node capacity

3. Operational complexity

Is the platform becoming easier or harder to operate as it grows? Look at:

  • incident frequency

  • manual interventions

  • upgrade complexity

  • troubleshooting time

  • recurring cluster-health issues

4. Cost efficiency

Is infrastructure growth broadly proportional to workload growth? Look at:

  • GKE node growth

  • CPU and memory utilization

  • storage growth

  • Elasticsearch resource utilization

  • engineering time required to operate the platform

The goal is not to minimize every metric. The goal is to understand whether reliability, performance, cost, and operational effort are scaling predictably together. If they are not, another infrastructure increase may not be the right next move.

Can Your Team Fix It Internally?

Not every Elasticsearch scaling problem requires external specialists. A capable internal platform team should generally continue internally when:

  • the bottleneck is clearly identified

  • Elasticsearch expertise exists inside the organization

  • workload growth is predictable

  • recovery behavior is understood

  • shard and index strategies are documented

  • the team understands its capacity thresholds

  • upgrades and maintenance are controlled

  • scaling decisions produce predictable results

In that situation, the organization may simply need focused tuning and disciplined capacity planning. The decision changes when the team can see the symptoms but cannot confidently identify the architectural cause. External Elasticsearch expertise becomes more useful when:

  • adding GKE resources no longer produces predictable improvements

  • recurring reliability incidents have different immediate causes

  • shard architecture has evolved without deliberate planning

  • recovery behavior creates production risk

  • storage performance is difficult to characterize

  • Kubernetes and Elasticsearch scaling strategies are poorly aligned

  • upgrades have become high-risk events

  • cloud spending is increasing faster than expected

  • the organization lacks senior Elasticsearch architecture expertise internally

At that point, the objective should not be to outsource day-to-day platform ownership. It should be to establish what is actually limiting the environment and create an architecture the internal team can operate confidently.

When Elasticsearch Consulting Services Make Sense

The best time to engage Elasticsearch Consulting Services is not necessarily when Elasticsearch is failing. It is often when the platform is still functioning but its behavior is becoming increasingly difficult to predict. A focused Elasticsearch architecture assessment should answer questions such as:

  • Is the current cluster topology appropriate for the workload?

  • Are node roles and resources aligned with actual usage?

  • Is shard architecture creating unnecessary overhead?

  • Does the storage design support performance and recovery requirements?

  • Are Kubernetes scheduling and disruption policies aligned with Elasticsearch availability?

  • What happens during realistic failure scenarios?

  • Where are the current scaling limits?

  • Which capacity increases are justified?

  • Which infrastructure costs are symptoms of architectural inefficiency?

  • What needs to change before the next stage of growth?

The deliverable should not simply be a list of Elasticsearch settings to change. Leadership should leave with a clear view of current risk, target architecture, remediation priorities, capacity requirements, and the operational model required to sustain the environment. That turns Elasticsearch scaling from reactive infrastructure management into an engineering decision.

Final Takeaway

Google Kubernetes Engine provides a strong platform for running stateful workloads, and ECK provides valuable orchestration capabilities for operating Elasticsearch on Kubernetes. Neither eliminates the architectural characteristics of Elasticsearch. As an Elasticsearch-on-GKE environment grows, the warning sign is not simply higher CPU utilization or another storage expansion.

It is loss of predictability. When more infrastructure produces less predictable improvement, failures become harder to recover from, routine operations carry greater risk, engineers spend increasing time stabilizing the cluster, and cloud costs grow faster than the workload, the organization should reassess the architecture before adding another layer of capacity.

The question is no longer: How do we give Elasticsearch more GKE resources?

It is: What architecture will let Elasticsearch scale reliably from here?

DinaBridge provides senior-led Elasticsearch Consulting Services for organizations operating complex search and observability environments. We help platform and engineering teams assess Elasticsearch architecture, identify scaling and reliability constraints, and build a practical remediation path around performance, resilience, and cost.

If your Elasticsearch environment on GKE is becoming harder to scale, recover, or predict, contact DinaBridge to evaluate whether the next step is more capacity—or a better architecture.

Technical References

Elastic Cloud on Kubernetes documentation

Google Cloud: Deploying stateful applications on GKE