Repository navigation
Conversation
|
Do you have a test/repro case that would trigger this race condition? I see |
It will be hard to write a deterministic test for this precisely because it is a race condition and the window is pretty narrow (a couple of assembly instructions). But it works like this:
Yes, this can only happen while pg_net bg worker is shutting down. The fix closes that window between two loads of if condition and |
Maybe we could inject some sleep via postgres injection points to trigger this failure and prove the fix? It would be a custom build but we could do it via some config in xpg. |
Possibly this can be used in some way? https://www.postgresql.org/docs/devel/xfunc-c.html#XFUNC-ADDIN-INJECTION-POINTS |
Possibly, with a few caveats. First, injection points are available in PG 17 and later only, so we won't be able to test it on older versions. Second, it needs support in xpg to add the I've let AI write a potential test using injection points in this commit on a separate branch (rs/segfault-repro-with-sleeps) but haven't run in locally due to missing I was also thinking if this is a scalable testing strategy. I asked AI find more race conditions and it came up with the following. I've not verified them all, but assuming there are more race conditions would we want to add more injection points to test them all? My worry is we might litter the code with them hurting readability. |
In the old code SetLatch was called on the shared_latch only after checking that it was not NULL. But there was a narrow race condition in which the shared_latch variable might be assigned to NULL after the if condition succeeded but before the SetLatch was called if the worker exited in that narrow window and net_on_exit function set the shared_latch to NULL. The fix is to make the shared_latch pointer volatile and then to load it once before checking it for NULL and calling SetLatch. This commit also removes the pg_write_barrier calls because these are redundant: SetLatch and ConditionVariableBroadcast functions already have a pg_write_barrier call inside them.
7f7daea to
2c3ae39
Compare
|
Tbh for this PR all I really need is a way to repro what ur talking about locally so I can validate the issue. Ill take a look at the above. |
Just my 2 cents. |
In the old code
SetLatchwas called on theshared_latchonly after checking that it was notNULL. But there was a narrow race condition in which theshared_latchvariable might be assigned toNULLafter the if condition succeeded but before theSetLatchwas called if the worker exited in that narrow window andnet_on_exitfunction set theshared_latchtoNULL. The fix is to make theshared_latchpointer volatile and then to load it once before checking it forNULLand callingSetLatch.This PR also removes the
pg_write_barriercalls because these are redundant:SetLatchandConditionVariableBroadcastfunctions already have apg_write_barriercall inside them.