Search

Written by
Dina Bridge
|
Subscribe
Subscribe to get the latest insights straight in your inbox
A production Elasticsearch cluster has been running on virtual machines for years. It works. Your teams understand it. Applications depend on it. But the infrastructure around it has changed. Kubernetes is now becoming the standard platform for application deployment. Platform engineering owns more of the infrastructure lifecycle. Provisioning is increasingly declarative. Deployment patterns are standardized. Application teams expect infrastructure capabilities to fit into the Kubernetes operating model. Elasticsearch is becoming an exception.
For a VP or Director of Platform Engineering, that creates a reasonable question: Should we move our self-managed Elasticsearch environment from VMs to Kubernetes?
The answer is not automatically yes. Moving Elasticsearch to Kubernetes with Elastic Cloud on Kubernetes (ECK) can improve standardization and automate parts of the operational lifecycle. But it also means moving a stateful distributed system that may already contain years of architecture decisions, accumulated data, performance dependencies, and operational knowledge. The decision should therefore be based on whether Kubernetes creates a better Elasticsearch operating model, not simply whether Kubernetes has become the organizational standard.
Why Elasticsearch Becomes the Infrastructure Exception
This decision rarely begins because Elasticsearch suddenly stops working. It usually begins because the platform around Elasticsearch evolves. A company that originally deployed Elasticsearch onto VMs may now operate most application workloads through Kubernetes. Platform engineering has standardized:
Compute
Networking
Secrets
Deployment automation
Infrastructure as code
Observability
Resource governance
Availability patterns
Environment provisioning
Elasticsearch remains outside that model. That creates operational questions. Should the platform team maintain separate infrastructure processes for Elasticsearch? Should Elasticsearch provisioning become declarative? Could scaling and upgrades fit into existing Kubernetes workflows? Could development, staging, and production environments become more consistent? Would moving Elasticsearch reduce infrastructure exceptions? These are legitimate reasons to investigate Elasticsearch on Kubernetes. But reducing infrastructure exceptions is valuable only when the resulting platform is easier to operate.
What Are You Actually Trying to Improve?
Before evaluating ECK, establish why the existing VM architecture is no longer sufficient. This is the first decision gate. If Elasticsearch on VMs is stable, appropriately sized, recoverable, well understood, and inexpensive to operate, migration needs a compelling justification. Moving infrastructure simply because Kubernetes is newer is not platform strategy. Instead, identify a specific operational problem.
Perhaps provisioning additional Elasticsearch infrastructure requires too much manual work. Perhaps environments have diverged because configurations are managed differently. Perhaps upgrades require coordination across infrastructure teams. Perhaps capacity changes are slow. Perhaps the organization wants infrastructure definitions stored and managed through the same declarative workflows used elsewhere. Or perhaps the platform organization is deliberately reducing infrastructure models to decrease operational overhead. Those are measurable problems. The migration should be judged against them.
Before asking: Can Elasticsearch run on Kubernetes?
ask: What becomes materially better if Elasticsearch runs on Kubernetes?
If there is no clear answer, the migration case is already weak.
What Changes When Elasticsearch Moves to Kubernetes?
It is easy to describe this migration as: VMs → Kubernetes.
Operationally, much more changes. On VMs, infrastructure teams may provision hosts, attach storage, configure networking, install Elasticsearch, manage services, and maintain the underlying operating system. On Kubernetes, Elasticsearch becomes part of a different control model. Compute becomes Kubernetes resources. Storage is integrated through PersistentVolumes and storage classes. Placement is influenced by scheduling rules and Kubernetes topology. Infrastructure configuration becomes declarative.
Maintenance interacts with Kubernetes node lifecycle. Scaling interacts with both Kubernetes capacity and Elasticsearch shard allocation. Availability needs to account for both Kubernetes failure domains and Elasticsearch cluster topology. Monitoring now needs to answer questions about two systems. The migration is therefore not simply a change in compute. It is a change in the Elasticsearch operating model. That distinction should influence how the project is funded, staffed, designed, and measured.
What ECK Solves
If Elasticsearch moves to Kubernetes, Elastic Cloud on Kubernetes (ECK) should be central to the evaluation. ECK is Elastic's Kubernetes operator for orchestrating Elasticsearch and other Elastic components. Instead of treating Elasticsearch like a generic containerized application, ECK understands parts of the Elastic lifecycle and allows teams to manage Elastic resources through Kubernetes-native configuration.
For platform teams, ECK can help with areas including:
Elasticsearch deployment
Cluster configuration
Scaling
TLS certificates
Node configuration
Rolling changes
Version upgrades
Kubernetes-aware orchestration
This matters. Without an operator that understands Elasticsearch, the platform team would need to build significantly more lifecycle logic around a stateful distributed system. ECK reduces that burden.
But there is a critical distinction: ECK automates parts of Elasticsearch operations. It does not make Elasticsearch architecture decisions for you.
That is where migration projects can become more difficult than expected.
What ECK Does Not Solve
Suppose the existing VM-based Elasticsearch cluster has:
Too many shards
Poorly sized shards
Inefficient index structures
Excessive retention
Storage bottlenecks
Heap pressure
Unnecessary replicas
Weak lifecycle policies
Slow recovery
Inconsistent snapshots
Poor capacity planning
Moving that cluster to Kubernetes does not inherently correct those problems. It can reproduce them. The infrastructure may become more automated while the underlying Elasticsearch architecture remains inefficient. This is particularly important because Kubernetes and Elasticsearch manage different things. Kubernetes understands infrastructure state. Elasticsearch understands data and cluster state. A Kubernetes environment can therefore look healthy while Elasticsearch is experiencing:
Search latency
Indexing pressure
Shard allocation problems
Disk pressure
Long recoveries
JVM pressure
Cluster-state overhead
For platform leadership, the important principle is: Kubernetes can modernize Elasticsearch infrastructure management. It cannot replace Elasticsearch engineering.
Assess the Existing Elasticsearch Cluster First
One of the biggest mistakes in an Elasticsearch-to-Kubernetes migration is designing the target environment before understanding the source environment. The VM cluster is your evidence. Use it. Before defining the Kubernetes architecture, establish a baseline across the existing Elasticsearch deployment.
Evaluate:
Data volume
Daily ingestion
Retention requirements
Search workload
Query latency
Indexing throughput
Node utilization
CPU utilization
Memory utilization
JVM pressure
Disk utilization
Index count
Primary shard count
Replica strategy
Average shard size
Peak workloads
Recovery duration
Snapshot behavior
Historical growth
Then ask a more difficult question: Which parts of the current architecture should not migrate?
A cluster that has existed for years often contains decisions made for requirements that no longer exist. Old indices remain. Retention expands. New workloads are added. Shard strategies become inconsistent. Hardware is added incrementally. Migration provides an opportunity to reset some of that technical debt. Do not waste it by recreating the VM cluster node-for-node inside Kubernetes.
Five Questions Platform Leaders Should Answer
The decision does not require a VP or Director of Platform Engineering to become an Elasticsearch administrator. It does require leadership to challenge the architecture in the right places.
1. Is Kubernetes Ready for Elasticsearch?
Your Kubernetes platform may already be mature for application workloads. That does not automatically mean it is ready for Elasticsearch. Elasticsearch is stateful and storage-intensive. Evaluate whether the platform already handles production stateful workloads reliably.
Specifically:
Is persistent storage mature?
Are failure domains properly designed?
Is capacity predictable?
Can nodes be maintained without unexpected disruption?
Are resource requests and limits governed?
Are Kubernetes upgrades predictable?
Is stateful workload recovery understood?
If the Kubernetes platform itself remains operationally unstable, moving Elasticsearch onto it compounds risk.
2. What Does Kubernetes Improve?
The migration should produce specific improvements. Potential benefits include:
Faster infrastructure provisioning
More repeatable deployments
Declarative configuration
Standardized platform workflows
Reduced manual infrastructure administration
Consistent environment creation
Better integration with existing Kubernetes operations
Define those benefits before migration. Then measure them afterward. Otherwise, platform engineering can complete a technically successful migration without improving the operating model.
3. Who Owns Elasticsearch After Migration?
This needs an explicit answer. Before migration, infrastructure engineering may own the VMs while another team owns Elasticsearch. After migration, the ownership boundary can become less obvious.
Platform engineering may own: Kubernetes.
But who owns: ECK?
Who owns: Elasticsearch topology?
Who owns: Shard architecture?
Who owns: Index lifecycle?
Who owns: Cluster performance?
Who responds when Kubernetes says every Pod is healthy but Elasticsearch performance is deteriorating? Those responsibilities should be defined before production moves. Otherwise, the migration creates an organizational dependency that architecture diagrams will not reveal.
4. Can the Target Architecture Survive Real Failure?
Do not validate the architecture only under normal conditions. Model failure. What happens if:
A Kubernetes worker fails?
A data node disappears?
An availability zone becomes unavailable?
A PersistentVolume has problems?
Multiple Pods restart?
Kubernetes drains a node?
Elasticsearch needs to relocate large shards?
Disk utilization approaches a watermark?
An upgrade fails halfway through?
The cluster needs to be restored from snapshots?
Kubernetes restarting an Elasticsearch Pod is not the same as Elasticsearch recovering safely. Platform leadership should understand both sides of the recovery process.
5. Is the Migration Worth the Operational Change?
There is a cost to maintaining infrastructure exceptions. There is also a cost to migrations. Moving Elasticsearch from VMs to Kubernetes can require:
Architecture work
Migration engineering
Parallel infrastructure
Testing
Data movement
Validation
Operational training
Runbook changes
Monitoring changes
Recovery testing
That investment should produce a meaningful return. The benefit may not be direct infrastructure savings. It may instead come from standardization, deployment speed, operational consistency, resilience, or reduced manual administration. But the expected return should be explicit.
When Moving Elasticsearch from VMs to Kubernetes Makes Sense
A migration becomes compelling when several conditions align.
Kubernetes is already a mature internal platform
The organization has reliable production patterns for networking, storage, monitoring, upgrades, capacity, and stateful workloads.
Elasticsearch is important enough to standardize
Multiple applications, teams, or operational capabilities depend on it.
The existing VM operating model creates friction
Provisioning, scaling, upgrades, or environment management require unnecessary manual work.
Platform engineering wants declarative infrastructure
Elasticsearch can become part of the organization's broader infrastructure-as-code and platform-management model.
Elasticsearch architecture is understood
The organization knows its workloads, capacity requirements, shard behavior, storage requirements, and recovery characteristics.
Ownership is clear
There is an explicit operating model across Kubernetes, ECK, and Elasticsearch. Under those conditions, Kubernetes can improve how Elasticsearch infrastructure is delivered and operated.
When Keeping Elasticsearch on VMs Makes More Sense
Platform engineering should also be willing to decide not to migrate. Keeping Elasticsearch on VMs can be the better decision when:
The existing environment is stable.
Operating costs are reasonable.
Provisioning requirements are infrequent.
The Kubernetes platform is still maturing.
Stateful Kubernetes experience is limited.
Storage architecture is uncertain.
Elasticsearch expertise is limited.
Ownership after migration is unclear.
Migration risk is high.
The expected operational benefit is small.
There is no requirement for every infrastructure service to share the same deployment model. Standardization is valuable because it reduces complexity. If standardization increases complexity, it has defeated its purpose. A mature platform organization should be comfortable maintaining an exception when the exception is economically and operationally justified.
The Migration Risk Platform Leaders Should Not Underestimate
The technical challenge is not deploying an empty Elasticsearch cluster with ECK. The challenge is moving a production workload into it. Existing Elasticsearch environments may contain terabytes, or considerably more, of data.
Applications may depend on:
Existing endpoints
Index names
Mappings
Templates
Pipelines
Security configuration
Dashboards
Integrations
Data streams
Application behavior
A migration therefore needs to address more than infrastructure provisioning.
Platform teams need a deliberate approach to:
Data movement
How does existing data reach the new environment?
Synchronization
What happens to new writes while historical data moves?
Application cutover
When and how do producers and consumers switch?
Validation
How do you establish that the new cluster behaves correctly?
Rollback
What happens if the target environment fails validation?
Downtime
How much interruption can the business tolerate?
Performance
Does the Kubernetes environment perform at least as well under production load?
The target architecture should be proven before it becomes the only production architecture.
A Decision Framework for Platform Engineering
Before approving the migration, score the proposal across six dimensions.
Decision Area | Question Platform Leadership Should Answer |
Platform maturity | Can Kubernetes reliably support a business-critical stateful workload? |
Operational benefit | What specifically becomes easier or better after migration? |
Elasticsearch health | Is the current cluster architecture understood and worth carrying forward? |
Ownership | Who owns Kubernetes, ECK, Elasticsearch architecture, and production incidents? |
Resilience | Can the target design survive realistic infrastructure and Elasticsearch failures? |
Economics | Is the operational benefit worth the migration and ongoing platform complexity? |
Strong answers across all six create a credible migration case. Weak answers do not necessarily mean the organization should abandon Kubernetes. They mean the architecture is not ready for approval. And repeated answers of “we don't know” indicate that the next step should probably be an assessment rather than a migration.
Assess Before You Migrate
A strong Elasticsearch-to-Kubernetes project should therefore begin before ECK is installed. The assessment should establish four things.
Current State
Document the VM environment, workload characteristics, performance baseline, data growth, topology, storage, and operational dependencies.
Target State
Define the Kubernetes and ECK architecture based on workload requirements rather than copying the existing VM topology.
Migration Path
Determine how data, applications, integrations, and traffic can move while remaining within acceptable risk and downtime boundaries.
Operating Model
Define who owns the environment after migration and how capacity, upgrades, incidents, backups, recovery, and performance will be managed.
This gives platform leadership something much more useful than an architecture diagram. It creates a basis for deciding whether the migration should happen at all.
Where Elasticsearch Consulting Services Fit
For a production Elasticsearch environment, the highest-value external support often comes before the migration begins.
The difficult question is not: How do we install ECK?
It is: Should this Elasticsearch environment move to Kubernetes, what should the target architecture look like, and how do we move it without carrying existing problems into the new platform?
DinaBridge's Elasticsearch Consulting Services can support platform engineering organizations with:
Current-state Elasticsearch assessments
VM-to-Kubernetes migration readiness
ECK architecture
Target cluster topology
Shard and index analysis
Storage architecture
Capacity planning
Resilience and recovery design
Migration strategy
Cutover planning
Production validation
Post-migration optimization
The objective is not Kubernetes adoption. The objective is a better Elasticsearch operating model.
Final Takeaway
If Kubernetes has become your organization's platform standard, questioning whether Elasticsearch should remain on VMs is reasonable. But platform standardization alone is not sufficient justification for migration.
Start with the existing Elasticsearch environment.
Is it difficult to provision?
Is it difficult to scale?
Is it expensive to maintain?
Are environments inconsistent?
Are upgrades unnecessarily manual?
Would moving to Kubernetes materially improve those conditions?
Then evaluate the Kubernetes platform itself.
Is storage mature?
Are stateful workloads already reliable?
Are failure domains understood?
Can the platform support Elasticsearch recovery behavior?
Finally, define ownership.
Because after the migration, your organization is not simply operating Elasticsearch.
It is operating Elasticsearch, ECK, and Kubernetes as one production system.
If that combined operating model is more standardized, resilient, repeatable, and supportable than the VM environment it replaces, the migration has a strong case. If it is not, keeping Elasticsearch on VMs may be the more mature platform decision.
DinaBridge helps VP and Director-level platform engineering teams evaluate that decision before migration begins. Our Elasticsearch Consulting Services cover existing-cluster assessment, ECK architecture, VM-to-Kubernetes migration planning, capacity, resilience, and production optimization. Contact DinaBridge to determine whether moving your Elasticsearch environment to Kubernetes will actually improve the platform you operate.
