OpenSearch Kubernetes Operator 3.0 went GA on September 25, 2026. From the team that maintains it: what changed from 2.x, which defaults break, and how to upgrade in place.

OpenSearch Kubernetes Operator 3.0: What Changed and How to Upgrade from 2.x

OpenSearch Kubernetes Operator 3.0 went GA on September 25, 2026, and the 2.x line is end of life. BigData Boutique maintains the operator: two of its three active maintainers work here, and our team drove the 3.0 rewrite. So this is a maintainers' view of the release, not a summary of the changelog.

The short version: 2.x could remove a node that still held data during scale-down, never rotated node certificates, and let one failing reconciler freeze all maintenance on a cluster. 3.0 fixes those, changes several defaults on the way, and upgrades in place with one rolling restart.

This post covers what was fixed, which defaults change under you, and a step-by-step upgrade sequence. It builds on the alpha announcement, which covered the feature list. The official GA post has the project view. If you would rather not run the migration alone, we offer an enterprise version of the operator and enterprise support for it, and for OpenSearch on Kubernetes in general. Details are at the end.

Why did 2.x reach end of life?

The OpenSearch Kubernetes Operator is a Kubernetes controller that creates, scales, upgrades and secures OpenSearch clusters and Dashboards from a single OpenSearchCluster custom resource. Version 3.0 rewrites its core reconcilers around one goal: no operation may put data or quorum at risk.

Since v2.8.0, the project merged more than 200 pull requests from more than 20 contributors, and over 100 of them landed after the alpha. Most are fixes for bugs found by a new test harness that builds the operator from source, installs it on a real cluster and replays scenarios that simulate a year of cluster life. About 60 playbooks caught close to 50 bugs between alpha and GA.

Six of those bugs mattered most, and all of them existed in 2.8:

Bug in 2.x What happened Tracking
Scale-down drain failed open An API error or a still-relocating shard could get a node removed while it held data #1447, fixed in #1500
Certificates never renewed Generated node certs expired after a year and were not reissued. Renewed certs were not loaded by nodes #1451, fixed in #1460
Two pods down at once A restart could delete a second pod while the one removed by the scaler was still terminating #1572, fixed in #1577
One failure froze everything A failing sub-reconciler blocked scaling, upgrades and restarts, and requeue signals were dropped #1454, fixed in #1469
emptyDir recovery A transient readiness blip could tear down every StatefulSet #1455, fixed in #1467
Aborted upgrade Reverting left shard allocation set to primaries indefinitely while the CR reported RUNNING #1571, fixed in #1575

A nuance on the first row: in 2.8 SmartScaler, the draining scaler, was off by default. So by default there was no drain at all. Undrained scale-down is the one to check in your own clusters today.

What does 3.0 do differently?

Three behaviors matter most in production.

SmartScaler is on by default. spec.confMgmt.smartScaler now defaults to true, even when the CR has no confMgmt block. Scale-down drains a node and confirms it holds no data before removal, and master-eligible nodes always go through voting-config exclusions. It is not mandatory: smartScaler: false still works, and the operator then removes nodes without draining and emits a Warning event each time.

Certificates rotate on their own. Generated TLS certs now renew 30 days before expiry, and expired or CA-mismatched certs are always reissued. On OpenSearch 2.19.1 and later, nodes hot-reload the new cert. Older versions get a rolling restart.

spec:
    security:
      tls:
        transport:
          generate: true
          perNode: true
          rotateDaysBeforeExpiry: 30
        http:
          generate: true
          rotateDaysBeforeExpiry: 30
  

Failures stay local. One failing sub-reconciler no longer blocks the rest, and admission webhooks reject bad configs at create time: missing cluster-manager pool, unknown roles, storage class changes, duplicate pool names, downgrades and multi-major version jumps.

New capabilities

Shard allocation awareness now reads Kubernetes node labels. An init container copies the label into a node attribute:

spec:
    general:
      nodeAttributes:
        - name: zone
          nodeLabel: topology.kubernetes.io/zone
      additionalConfig:
        cluster.routing.allocation.awareness.attributes: zone
  

Each node pool can also set its own image, imagePullPolicy and imagePullSecrets, which is handy for ML nodes that need a different image. And spec.general.persistentVolumeClaimRetentionPolicy takes the standard Kubernetes whenDeleted and whenScaled values. Note that it applies to every node pool, not per pool.

What breaks when you upgrade from 2.x?

Most upgrade surprises come from defaults that changed, and from CRs that keep their stored old values because CRD defaulting does not rewrite existing objects. Check each of these before you start.

Area 2.x 3.0
SmartScaler Off by default On by default. Existing CRs with a stored false stay false
Cert rotation Node certs never rotated Renew 30 days before expiry. Existing CRs keep a stored -1 until you set 30
Admin password admin:admin Random, in secret <cluster>-admin-password, unless you set adminCredentialsSecret
setVMMaxMapCount false true
API group opensearch.opster.io/v1 opensearch.org/v1, legacy deprecated
Config validation None Admission webhooks, failurePolicy: Fail
StatefulSets OrderedReady Parallel pod management, new selector labels

Two corrections to things you may have read. TLS is not simply "on by default": the operator enables it when a security.tls.transport or http block exists, and a CR with no security block gets no TLS and no security plugin. And the operator authenticates with basic auth from your admin credentials secret by default. Client-certificate auth is opt-in through operatorClientCert.

Two hard requirements. The webhooks need a serving certificate, so install cert-manager 1.0 or later first, or set webhook.certManager.enabled=false and supply your own secret. Without cert-manager, helm install fails with no matches for kind "Certificate" in version "cert-manager.io/v1". And OpenSearch must already be at 2.19.2 or later. The operator supports 2.19.2 through the latest 3.x.

How do you upgrade in place?

The migration controller creates an opensearch.org/v1 twin of every legacy resource, copies the spec, transfers certificate secret ownership and keeps status in sync. It only migrates resources that are healthy: clusters must be RUNNING, and users, roles, tenants and policies must be CREATED.

  1. Confirm OpenSearch is on 2.19.2 or later, the cluster is green, and every index has at least one replica.
  2. Take a snapshot. The migration guide does not require one, and we would not skip it.
  3. Check that every legacy CR is healthy:
    kubectl get opensearchclusters.opensearch.opster.io <name> -o jsonpath='{.status.phase}'
      
  4. Record spec.confMgmt.smartScaler and spec.general.setVMMaxMapCount on each cluster, and set the values you actually want.
  5. Install cert-manager, or prepare a webhook cert.
  6. Upgrade the chart. Pass a values file, and avoid --reuse-values from a 2.x release because it drops the new webhook.* defaults:
    helm repo update
      helm upgrade opensearch-operator opensearch-operator/opensearch-operator -n <ns> -f values.yaml
      
  7. Watch the operator logs, the new opensearch.org resources and the StatefulSetRecreated and RollingRestart events until the cluster is green.
  8. Switch GitOps manifests to apiVersion: opensearch.org/v1. The legacy webhook rejects creation and spec changes on old-group resources.
  9. Delete the legacy CRs, confirm none remain, then run the chart with --set legacyAPI.enabled=false.

The one rolling restart

The first reconcile of a migrated cluster changes the StatefulSet selector and the pod management policy, and both fields are immutable. The operator relabels the pods, deletes each node-pool StatefulSet with orphan propagation, recreates it, and restarts pods one at a time. Expect yellow health during the restart, and hours for large pools. PV-backed data is kept. Nodes on emptyDir lose their data, so every index needs a replica first.

Pitfalls to know

  • Image tag. At the time of writing, the chart pulls opensearchproject/opensearch-operator:3.0.0, but the image was pushed as v3.0.0 (#1610, open). If your pods hit ImagePullBackOff, set manager.image.tag=v3.0.0. Check whether a chart fix has shipped first.
  • Uninstall is destructive. The chart templates its CRDs without a keep policy, so helm uninstall deletes them and cascades to every OpenSearchCluster (#1539). The documented rollback of "uninstall 3.x, install 2.x" is therefore not safe as written. Do not treat operator rollback as a plan.
  • Chart and app versions differ. App v3.0.0 ships as chart 3.0.14. Chart 2.8.3 and 2.8.4 even carry the alpha app version. Do not pin by guessing.
  • User-supplied TLS keys must be PKCS#8. cert-manager's PKCS#1 default crash-loops nodes (#1527).

Who maintains the operator, and how we can help

BigData Boutique maintains the OpenSearch Kubernetes Operator. The 3.0 work was driven by our team, an OpenSearch Software Foundation member and the first accredited Long-Term Support provider in the Foundation's LTS program. Two of the operator's three active maintainers, Jose Barato and Itamar Syn-Hershko, are from BigData Boutique. Ryan Patterson and Lior Friedler contributed heavily to the reconciler rework and the test harness, and Prudhvi Godithi co-maintained throughout. The operator stays open source, and we keep contributing upstream.

On top of that, we offer two things for teams that run OpenSearch themselves:

  • An enterprise version of the operator. It ships with our OpenSearch Enterprise Distribution, the enterprise-hardened OpenSearch distribution with long-term support, so the operator and the OpenSearch versions it manages come from one supported source.
  • Enterprise support for the operator and for OpenSearch on Kubernetes. OpenSearch enterprise support covers the operator itself, the clusters it runs, and the Kubernetes layer around them, with 24/7 coverage and contractual SLAs. That includes upgrade planning, so a 2.x to 3.0 migration is a normal conversation for us.

Key takeaways

  • 3.0 GA shipped on September 25, 2026 as app v3.0.0 (chart 3.0.14). 2.x is end of life.
  • Scale-down data loss, non-rotating node certs, double pod restarts and reconciler freezes are fixed. Six major bugs, all present in 2.8.
  • SmartScaler is on by default but not mandatory. Cert rotation defaults to 30 days. Existing CRs may still hold the old stored values, so set them explicitly.
  • The admin password is random, TLS depends on the presence of a tls block, and the webhooks need cert-manager or your own cert.
  • The in-place upgrade causes one rolling restart per cluster. Have replicas, a snapshot and a maintenance window ready.
  • Treat operator rollback as unsafe until the CRD uninstall behavior is fixed. Test the whole path in a non-production cluster first.