test(move): make the reverse-window waits say why they timed out - #1265
Conversation
TestMoveReverseWindowNMResumesAfterKill fails in CI at reversewindow_test.go:581 with "timed out waiting for the reverse window to open" and nothing else. On the most recent occurrence the runner had logged the window open on the very database the test polls, one second into a 30-second wait — so the checkpoint row it was looking for was there for the whole budget and the helper still never saw it. That narrows the cause to the read side, and the helper is built so that no read-side failure can be told apart from any other: it discards every query error, has no bound on the individual query, and never looks at what the run itself is doing. waitForReverseWindow is also the first thing in that test to use the control connection at all, which makes a control-connection failure — refused, out of connections, wrong schema — the cheapest explanation for every poll failing identically, and exactly the one the message cannot distinguish from "not there yet". So give the waits the same shape pkg/datasync got in block#1261: - Every wait selects on the run result. A move that dies before the window fails the test with its own error instead of an opaque wait. - Every query is bounded. An unbounded read can spend the whole budget inside one call, which is the 30.2s no-information failure above. - Only sql.ErrNoRows and ER_NO_SUCH_TABLE mean "not yet" for a checkpoint read. Everything else fails immediately, naming the error. - The timeout message carries the last phase seen, the last read error, how many reads were blocked, and the runner's state. awaitReverseWindow now waits on the runner's state as well as the checkpoint phase, state first: the state is in-process and cannot be blocked by the server, so a checkpoint read that then fails is reported against a window we already know is open. Two tests had that state wait open-coded after the phase wait; they lose the duplicate. The budget drops from 30s to 20s, under the 30s reverse window these tests configure — when the window elapses the terminal action drops the checkpoint, so a phase wait that outlives the window can never be satisfied. The handle also owns teardown, which fixes a hazard the old shape had: `defer utils.CloseAndLog(runner)` runs while a fatal wait leaves Run still in flight, and Runner.Close is safe only once Run has returned. Cleanup now cancels, drains, and only then closes — and reports the run as stuck rather than closing underneath it. Measured against injected failures: an unreachable target now fails in 0.04s quoting the connect error instead of waiting out the budget, a bad schema fails in 0.12s quoting error 1049, and the timeout path prints the phase and state it last saw. Full package green, and -race green over three runs of the reverse-window tests. This does not claim the root cause of block#1239 — it makes the next occurrence name it. Refs block#1239 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Address the two moderate timeout-budget issues in pkg/move/reversewindow_test.go.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
Open (1)
What changed in this PR
Improves reverse-window test diagnostics, bounded polling, runner error reporting, and asynchronous cleanup.
Changes:
- Adds diagnostic, runner-aware polling.
- Bounds checkpoint queries and reports read errors.
- Centralizes cancellation, draining, and teardown.
| File | Summary |
|---|---|
pkg/move/reversewindow_test.go |
Adds lifecycle-safe wait helpers and diagnostics. Requires a shared composite deadline and preservation of the 30-second acquisition budget. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
awaitReverseWindow called two waits that each started their own waitTimeout, so the composite could run to 40s. That is past the 30s reverse window these tests configure, and the second half polls the checkpoint phase — which the terminal action drops when the window elapses. A slow first half therefore left the second half waiting for a row that could no longer exist, timing out with a message that blames the phase. That is the misleading timeout this harness exists to remove, reintroduced by composing two waits that each looked safe alone. poll now takes the deadline instead of starting its own budget, and awaitReverseWindow passes one deadline to both halves. awaitTable keeps its own: nothing it waits for expires, because the renames it watches for are made by the cutover and by completing the window forward, so a window elapsing underneath it can only make the table more likely to appear. A caller that also needs the window still open — the kill tests — learns otherwise from its own assertion on Run's error. The per-wait timing text moves into poll, which now reports how long it actually waited against the budget, so the two halves cannot disagree about what "within 20s" meant. Verified with a deadline deliberately aged 15s: the second half gives up after the remaining 4.9s and the composite ends at 20.05s, where before it would have run 35s. Addresses Copilot review feedback on block#1265. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
🤖 Adversarial correctness review — 0 blocking, 4 non-blockingReviewed Everything below ran locally against the repo's MySQL on The headline claim holds, measured. Injecting a failure immediately after the start-up log in
Thirty seconds and a bare sentence becomes forty milliseconds, the error, and the state the runner was actually in. The Non-blocking1. Probe run at
|
aparajon
left a comment
There was a problem hiding this comment.
🤖 Approving: 0 blocking, 4 non-blocking — see the correctness review. The headline claim is verified by measurement: an injected setup failure takes 30.06s and names nothing at 164d2ee0, and 0.04s with the error and the runner's last state at this head. The four non-blocking items are poll printing the constant instead of the budget it actually had (measured, probe included), no final condition check when the timer fires, awaitTable's own budget letting the N:M kill test's waits outlast the window it kills inside, and five completion waits left as bare literals outside the new timing block.
This stamp was left by Claude Code (claude-opus-5).
|
🤖 Review findings - created by Kiran's code review agent - for spirit/pull/1265, 3796b2b. Non-blocking
The one thing that could have broken, verifiedThe wait ordering inside Verified correct
This review was generated by Claude Code (claude-opus-5). |
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- poll now reports the deadline it was actually given, not the waitTimeout constant, and re-checks cond once more when the timer fires so a condition that flips true during a bounded query isn't missed (aparajon). - awaitTable takes a shared deadline like its siblings instead of starting its own, so the N:M kill test's two retire-waits after awaitReverseWindow can no longer sum past the window they run inside (aparajon). - awaitTable now treats a query timeout as "not yet" and retries, instead of hard-failing the test on a slow information_schema read (Kiran01bm). - awaitReverseWindow gives the state wait its own generous budget (reverseWindowOpenTimeout, matching waitForMoveStatus's slow-CI allowance) since none of that time counts against the window's countdown; only the phase check and whatever a caller shares its deadline with are bounded by waitTimeout (Kiran01bm). - named the five remaining bare 30s/60s awaitDone timeouts (reverseCutoverTimeout, nmReverseCutoverTimeout) (aparajon). All non-blocking; verified against the real MySQL, package + -race x3 green, golangci-lint clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
🤖 Reply from Morgan's AI agent. Addressed all four non-blocking findings from the correctness review and both from Kiran's, in f14b3d4:
Verified against the real MySQL: |

TestMoveReverseWindowNMResumesAfterKillfails intermittently in CI atreversewindow_test.go:581withtimed out waiting for the reverse window to openand nothing else (#1239, most recently on #1225's 8.0.45-with-replicas job).What the last failure actually shows
The runner logged the window open one second into a 30-second wait, on the very database the test polls:
That log line comes after
persistReverseWindow, andpersistReverseWindowreturns its write error intoPostSwitch— a failed write aborts the cutover and the window never opens. It also setsr.cutoverAt, without which the window would have elapsed on its first tick and logged so. So the checkpoint row was present, with the right phase, in the right database, for the entire budget — and the helper never saw it.That puts the fault on the read side, and the helper is built so that no read-side failure can be distinguished from any other: it discards every query error, puts no bound on the individual query, and never looks at what the run is doing.
waitForReverseWindowis also the first thing in that test to use the control connection at all, which makes a control-connection failure the cheapest explanation for every poll failing identically — and precisely the one the message cannot tell apart from "not there yet".Change
The same shape
pkg/datasyncgot in #1261:sql.ErrNoRowsandER_NO_SUCH_TABLEmean "not yet" for a checkpoint read. Everything else fails immediately, naming the error.awaitReverseWindownow waits on the runner's state as well as the checkpoint phase, state first — the state is in-process and cannot be blocked by the server, so a checkpoint read that then fails is reported against a window we already know is open. Two tests had that state wait open-coded after the phase wait; they lose the duplicate.The budget drops 30s → 20s, under the 30s reverse window these tests configure: once the window elapses the terminal action drops the checkpoint, so a phase wait outliving the window can never be satisfied.
The handle also owns teardown, fixing a hazard in the old shape —
defer utils.CloseAndLog(runner)runs while a fatal wait leavesRunin flight, andRunner.Closeis safe only onceRunhas returned. Cleanup now cancels, drains, then closes, and reports a stuck run rather than closing underneath it.Verification
Measured against injected failures:
ping failed: dial tcp 127.0.0.1:1: connect: connection refusedError 1049 (42000): Unknown database ...timed out waiting for the reverse window to openlast phase="reverse_window", last read error=<nil>, blocked reads=0, runner state=reverseWindow./pkg/move/...green,-racegreen over three runs of the reverse-window tests,golangci-lintclean.Scope
This does not claim a root cause for #1239 — the evidence narrows it to the read side but does not identify which read-side failure. It makes the next occurrence name it, and removes the helper's ability to burn its entire budget in one blocked query. Leaving #1239 open.
Refs #1239
🤖 Generated with Claude Code