Problem
When an InferenceReplica instance's pod-0 crashloops and SurgeThenDrain brings up a healthy pod-1 on the new revision, the old pod-0 is never cleaned up: on moirai-dev we observed router instance pod-0s crashlooping for 7h49m with 96 restarts each, running a dead revision, alongside the serving pod-1s. They also hold their (host) ports and scheduler capacity.
Observed on latest main (dev-bf797d6), deploymentMode OMENative, single-pod instances (router component, 2 replicas). Expectation: once the surge replacement is Ready and the instance converges, the superseded ordinal-0 pod should be deleted.
Repro sketch
- OMENative ISVC with a router whose command crashes (bad flag).
- Fix the runtime → new revision surges pod-1, becomes Ready.
- pod-0 keeps crashlooping on the old revision indefinitely.
From the moirai PD campaign; manual kubectl delete pod was the workaround.
Problem
When an InferenceReplica instance's pod-0 crashloops and SurgeThenDrain brings up a healthy pod-1 on the new revision, the old pod-0 is never cleaned up: on moirai-dev we observed router instance pod-0s crashlooping for 7h49m with 96 restarts each, running a dead revision, alongside the serving pod-1s. They also hold their (host) ports and scheduler capacity.
Observed on latest main (dev-bf797d6), deploymentMode OMENative, single-pod instances (router component, 2 replicas). Expectation: once the surge replacement is Ready and the instance converges, the superseded ordinal-0 pod should be deleted.
Repro sketch
From the moirai PD campaign; manual
kubectl delete podwas the workaround.