Skip to content

OpenBao init Job fails: manual raft join conflicts with retry_join and cannot recover on retry #2313

Description

@sunilthorat09

Describe the bug

The OpenBao init Job (deploy/helm/openbao, helm/scripts/deploy.sh, helm install method) runs a manual bao operator raft join for the follower pods in unseal_cluster. The server config (server.ha.raft.config in values.yaml) also configures retry_join for all three peers.

When retry_join auto-joins the followers first, the manual raft join fails with failed to get raft challenge, so the init Job fails. The Job is not idempotent, so its retry re-runs bao operator init against the now-initialized cluster and fails with Vault is already initialized. The Job never succeeds and helm install never completes - even though the raft cluster actually formed and every pod is unsealed.

Steps or code to reproduce bug

  1. Install the chart (default Shamir seal) with image nvcf-openbao:2.6.2-nv-1.3.4.
  2. Follow the init Job (openbao-server-initialize-cluster) logs: operator init and the pod-0 unseal succeed, then operator raft join of pod-1 returns Code: 500 ... failed to join raft cluster: failed to get raft challenge.
  3. The Job retries; the new attempt fails at operator init with Vault is already initialized.
  4. Meanwhile bao operator raft list-peers shows all three nodes joined and bao status shows every pod sealed=false: the cluster is healthy, but the Job and helm install report failure.

Reproduced on a kind cluster (via podman) with the image above. The manual-join vs retry_join ordering looks timing-sensitive.

Expected behavior

The init Job completes. Either drop the manual raft join and rely on the config's retry_join, or make the init path idempotent in the helm method too: skip operator init and the manual join when the cluster is already initialized (the helm method currently skips pre_checks, which is the existing already-initialized guard, so retries re-init).

Additional context

Observed while testing #2310 (which does not change this code path). No data loss - the raft cluster forms correctly via retry_join; the problem is only that the Job and release report failure and the Job cannot recover on retry.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions