OpenSearch Kubernetes Operator 3.0 went GA on September 25, 2026. From the team that maintains it: what changed from 2.x, which defaults break, and how to upgrade in place.
OpenSearch Kubernetes Operator 3.0 went GA on September 25, 2026, and the 2.x line is end of life. BigData Boutique maintains the operator: two of its three active maintainers work here, and our team drove the 3.0 rewrite. So this is a maintainers' view of the release, not a summary of the changelog.
The short version: 2.x could remove a node that still held data during scale-down, never rotated node certificates, and let one failing reconciler freeze all maintenance on a cluster. 3.0 fixes those, changes several defaults on the way, and upgrades in place with one rolling restart.
This post covers what was fixed, which defaults change under you, and a step-by-step upgrade sequence. It builds on the alpha announcement, which covered the feature list. The official GA post has the project view. If you would rather not run the migration alone, we offer an enterprise version of the operator and enterprise support for it, and for OpenSearch on Kubernetes in general. Details are at the end.
Why did 2.x reach end of life?
The OpenSearch Kubernetes Operator is a Kubernetes controller that creates, scales, upgrades and secures OpenSearch clusters and Dashboards from a single OpenSearchCluster custom resource. Version 3.0 rewrites its core reconcilers around one goal: no operation may put data or quorum at risk.
Since v2.8.0, the project merged more than 200 pull requests from more than 20 contributors, and over 100 of them landed after the alpha. Most are fixes for bugs found by a new test harness that builds the operator from source, installs it on a real cluster and replays scenarios that simulate a year of cluster life. About 60 playbooks caught close to 50 bugs between alpha and GA.
Six of those bugs mattered most, and all of them existed in 2.8:
| Bug in 2.x | What happened | Tracking |
|---|---|---|
| Scale-down drain failed open | An API error or a still-relocating shard could get a node removed while it held data | #1447, fixed in #1500 |
| Certificates never renewed | Generated node certs expired after a year and were not reissued. Renewed certs were not loaded by nodes | #1451, fixed in #1460 |
| Two pods down at once | A restart could delete a second pod while the one removed by the scaler was still terminating | #1572, fixed in #1577 |
| One failure froze everything | A failing sub-reconciler blocked scaling, upgrades and restarts, and requeue signals were dropped | #1454, fixed in #1469 |
| emptyDir recovery | A transient readiness blip could tear down every StatefulSet | #1455, fixed in #1467 |
| Aborted upgrade | Reverting left shard allocation set to primaries indefinitely while the CR reported RUNNING | #1571, fixed in #1575 |
A nuance on the first row: in 2.8 SmartScaler, the draining scaler, was off by default. So by default there was no drain at all. Undrained scale-down is the one to check in your own clusters today.
What does 3.0 do differently?
Three behaviors matter most in production.
SmartScaler is on by default. spec.confMgmt.smartScaler now defaults to true, even when the CR has no confMgmt block. Scale-down drains a node and confirms it holds no data before removal, and master-eligible nodes always go through voting-config exclusions. It is not mandatory: smartScaler: false still works, and the operator then removes nodes without draining and emits a Warning event each time.
Certificates rotate on their own. Generated TLS certs now renew 30 days before expiry, and expired or CA-mismatched certs are always reissued. On OpenSearch 2.19.1 and later, nodes hot-reload the new cert. Older versions get a rolling restart.
spec:
security:
tls:
transport:
generate: true
perNode: true
rotateDaysBeforeExpiry: 30
http:
generate: true
rotateDaysBeforeExpiry: 30
Failures stay local. One failing sub-reconciler no longer blocks the rest, and admission webhooks reject bad configs at create time: missing cluster-manager pool, unknown roles, storage class changes, duplicate pool names, downgrades and multi-major version jumps.
New capabilities
Shard allocation awareness now reads Kubernetes node labels. An init container copies the label into a node attribute:
spec:
general:
nodeAttributes:
- name: zone
nodeLabel: topology.kubernetes.io/zone
additionalConfig:
cluster.routing.allocation.awareness.attributes: zone
Each node pool can also set its own image, imagePullPolicy and imagePullSecrets, which is handy for ML nodes that need a different image. And spec.general.persistentVolumeClaimRetentionPolicy takes the standard Kubernetes whenDeleted and whenScaled values. Note that it applies to every node pool, not per pool.
What breaks when you upgrade from 2.x?
Most upgrade surprises come from defaults that changed, and from CRs that keep their stored old values because CRD defaulting does not rewrite existing objects. Check each of these before you start.
| Area | 2.x | 3.0 |
|---|---|---|
| SmartScaler | Off by default | On by default. Existing CRs with a stored false stay false |
| Cert rotation | Node certs never rotated | Renew 30 days before expiry. Existing CRs keep a stored -1 until you set 30 |
| Admin password | admin:admin |
Random, in secret <cluster>-admin-password, unless you set adminCredentialsSecret |
setVMMaxMapCount |
false |
true |
| API group | opensearch.opster.io/v1 |
opensearch.org/v1, legacy deprecated |
| Config validation | None | Admission webhooks, failurePolicy: Fail |
| StatefulSets | OrderedReady |
Parallel pod management, new selector labels |
Two corrections to things you may have read. TLS is not simply "on by default": the operator enables it when a security.tls.transport or http block exists, and a CR with no security block gets no TLS and no security plugin. And the operator authenticates with basic auth from your admin credentials secret by default. Client-certificate auth is opt-in through operatorClientCert.
Two hard requirements. The webhooks need a serving certificate, so install cert-manager 1.0 or later first, or set webhook.certManager.enabled=false and supply your own secret. Without cert-manager, helm install fails with no matches for kind "Certificate" in version "cert-manager.io/v1". And OpenSearch must already be at 2.19.2 or later. The operator supports 2.19.2 through the latest 3.x.
How do you upgrade in place?
The migration controller creates an opensearch.org/v1 twin of every legacy resource, copies the spec, transfers certificate secret ownership and keeps status in sync. It only migrates resources that are healthy: clusters must be RUNNING, and users, roles, tenants and policies must be CREATED.
- Confirm OpenSearch is on 2.19.2 or later, the cluster is green, and every index has at least one replica.
- Take a snapshot. The migration guide does not require one, and we would not skip it.
- Check that every legacy CR is healthy:
kubectl get opensearchclusters.opensearch.opster.io <name> -o jsonpath='{.status.phase}' - Record
spec.confMgmt.smartScalerandspec.general.setVMMaxMapCounton each cluster, and set the values you actually want. - Install cert-manager, or prepare a webhook cert.
- Upgrade the chart. Pass a values file, and avoid
--reuse-valuesfrom a 2.x release because it drops the newwebhook.*defaults:helm repo update helm upgrade opensearch-operator opensearch-operator/opensearch-operator -n <ns> -f values.yaml - Watch the operator logs, the new
opensearch.orgresources and theStatefulSetRecreatedandRollingRestartevents until the cluster is green. - Switch GitOps manifests to
apiVersion: opensearch.org/v1. The legacy webhook rejects creation and spec changes on old-group resources. - Delete the legacy CRs, confirm none remain, then run the chart with
--set legacyAPI.enabled=false.
The one rolling restart
The first reconcile of a migrated cluster changes the StatefulSet selector and the pod management policy, and both fields are immutable. The operator relabels the pods, deletes each node-pool StatefulSet with orphan propagation, recreates it, and restarts pods one at a time. Expect yellow health during the restart, and hours for large pools. PV-backed data is kept. Nodes on emptyDir lose their data, so every index needs a replica first.
Pitfalls to know
- Image tag. At the time of writing, the chart pulls
opensearchproject/opensearch-operator:3.0.0, but the image was pushed asv3.0.0(#1610, open). If your pods hitImagePullBackOff, setmanager.image.tag=v3.0.0. Check whether a chart fix has shipped first. - Uninstall is destructive. The chart templates its CRDs without a keep policy, so
helm uninstalldeletes them and cascades to everyOpenSearchCluster(#1539). The documented rollback of "uninstall 3.x, install 2.x" is therefore not safe as written. Do not treat operator rollback as a plan. - Chart and app versions differ. App
v3.0.0ships as chart3.0.14. Chart 2.8.3 and 2.8.4 even carry the alpha app version. Do not pin by guessing. - User-supplied TLS keys must be PKCS#8. cert-manager's PKCS#1 default crash-loops nodes (#1527).
Who maintains the operator, and how we can help
BigData Boutique maintains the OpenSearch Kubernetes Operator. The 3.0 work was driven by our team, an OpenSearch Software Foundation member and the first accredited Long-Term Support provider in the Foundation's LTS program. Two of the operator's three active maintainers, Jose Barato and Itamar Syn-Hershko, are from BigData Boutique. Ryan Patterson and Lior Friedler contributed heavily to the reconciler rework and the test harness, and Prudhvi Godithi co-maintained throughout. The operator stays open source, and we keep contributing upstream.
On top of that, we offer two things for teams that run OpenSearch themselves:
- An enterprise version of the operator. It ships with our OpenSearch Enterprise Distribution, the enterprise-hardened OpenSearch distribution with long-term support, so the operator and the OpenSearch versions it manages come from one supported source.
- Enterprise support for the operator and for OpenSearch on Kubernetes. OpenSearch enterprise support covers the operator itself, the clusters it runs, and the Kubernetes layer around them, with 24/7 coverage and contractual SLAs. That includes upgrade planning, so a 2.x to 3.0 migration is a normal conversation for us.
Key takeaways
- 3.0 GA shipped on September 25, 2026 as app
v3.0.0(chart3.0.14). 2.x is end of life. - Scale-down data loss, non-rotating node certs, double pod restarts and reconciler freezes are fixed. Six major bugs, all present in 2.8.
- SmartScaler is on by default but not mandatory. Cert rotation defaults to 30 days. Existing CRs may still hold the old stored values, so set them explicitly.
- The admin password is random, TLS depends on the presence of a
tlsblock, and the webhooks need cert-manager or your own cert. - The in-place upgrade causes one rolling restart per cluster. Have replicas, a snapshot and a maintenance window ready.
- Treat operator rollback as unsafe until the CRD uninstall behavior is fixed. Test the whole path in a non-production cluster first.