Describe the bug
The OpenBao init Job (deploy/helm/openbao, helm/scripts/deploy.sh, helm install method) runs a manual bao operator raft join for the follower pods in unseal_cluster. The server config (server.ha.raft.config in values.yaml) also configures retry_join for all three peers.
When retry_join auto-joins the followers first, the manual raft join fails with failed to get raft challenge, so the init Job fails. The Job is not idempotent, so its retry re-runs bao operator init against the now-initialized cluster and fails with Vault is already initialized. The Job never succeeds and helm install never completes - even though the raft cluster actually formed and every pod is unsealed.
Steps or code to reproduce bug
- Install the chart (default Shamir seal) with image
nvcf-openbao:2.6.2-nv-1.3.4.
- Follow the init Job (
openbao-server-initialize-cluster) logs: operator init and the pod-0 unseal succeed, then operator raft join of pod-1 returns Code: 500 ... failed to join raft cluster: failed to get raft challenge.
- The Job retries; the new attempt fails at
operator init with Vault is already initialized.
- Meanwhile
bao operator raft list-peers shows all three nodes joined and bao status shows every pod sealed=false: the cluster is healthy, but the Job and helm install report failure.
Reproduced on a kind cluster (via podman) with the image above. The manual-join vs retry_join ordering looks timing-sensitive.
Expected behavior
The init Job completes. Either drop the manual raft join and rely on the config's retry_join, or make the init path idempotent in the helm method too: skip operator init and the manual join when the cluster is already initialized (the helm method currently skips pre_checks, which is the existing already-initialized guard, so retries re-init).
Additional context
Observed while testing #2310 (which does not change this code path). No data loss - the raft cluster forms correctly via retry_join; the problem is only that the Job and release report failure and the Job cannot recover on retry.
Describe the bug
The OpenBao init Job (
deploy/helm/openbao,helm/scripts/deploy.sh, helm install method) runs a manualbao operator raft joinfor the follower pods inunseal_cluster. The server config (server.ha.raft.configinvalues.yaml) also configuresretry_joinfor all three peers.When
retry_joinauto-joins the followers first, the manualraft joinfails withfailed to get raft challenge, so the init Job fails. The Job is not idempotent, so its retry re-runsbao operator initagainst the now-initialized cluster and fails withVault is already initialized. The Job never succeeds andhelm installnever completes - even though the raft cluster actually formed and every pod is unsealed.Steps or code to reproduce bug
nvcf-openbao:2.6.2-nv-1.3.4.openbao-server-initialize-cluster) logs:operator initand the pod-0 unseal succeed, thenoperator raft joinof pod-1 returnsCode: 500 ... failed to join raft cluster: failed to get raft challenge.operator initwithVault is already initialized.bao operator raft list-peersshows all three nodes joined andbao statusshows every podsealed=false: the cluster is healthy, but the Job andhelm installreport failure.Reproduced on a kind cluster (via podman) with the image above. The manual-join vs retry_join ordering looks timing-sensitive.
Expected behavior
The init Job completes. Either drop the manual
raft joinand rely on the config'sretry_join, or make the init path idempotent in the helm method too: skipoperator initand the manual join when the cluster is already initialized (the helm method currently skipspre_checks, which is the existing already-initialized guard, so retries re-init).Additional context
Observed while testing #2310 (which does not change this code path). No data loss - the raft cluster forms correctly via
retry_join; the problem is only that the Job and release report failure and the Job cannot recover on retry.