Testing.SyncReverseProxy.SendReceiveServiceTests.SendNewProject_Big fails intermittently in the GHA integration tests because the hgweb container is OOMKilled while unbundling the push. The test is not at fault — it sits right at the memory ceiling by design — but it fails on unrelated PRs and gets re-rolled, so it costs review time.
Evidence
Two runs about 50 minutes apart, on unrelated branches, with an identical signature:
| Run |
Branch |
unbundle body |
| 31696184703 |
chore/document-dotnet8-chorusmerge-pin |
189,695,076 bytes → 502 |
| 31700197936 |
pt/missing-lex-code (#2547) |
189,695,051 bytes → 502 |
Chain, from each run's dotnet-k8s-logs artifact:
logs-ingress.txt — nginx buffers the request body to a temp file, then POST /send-new-project-<id>-dev-flex?cmd=unbundle returns 502 after ~1.1 s.
logs-lexbox.txt — the API proxies to hg:8088 and gets System.Net.Http.HttpIOException: The response ended prematurely. (ResponseEnded), followed by a run of HttpRequestException: Connection refused (hg:8088). Collateral: Job DataUpdate.UpdateProjectMetadataJob threw a JobExecutionException: Connection refused (hg:8088).
describe-hg.txt — the hgweb container's Last State is Terminated / Reason: OOMKilled / Exit Code: 137, finishing at the same second as the 502. Restart Count: 1; the pod comes back and the rest of the suite passes.
- Test-side: Chorus reports
abort: HTTP Error 502: Bad Gateway, and the assertion fails with Did not expect srOutput "Sync failure: ... Failed to send to LexBox".
Cause
hgweb is capped at 400Mi:
https://github.com/sillsdev/languageforge-lexbox/blob/develop/deployment/base/hg-deployment.yaml#L86-L90
The test pushes 180 MB across 10 commits:
https://github.com/sillsdev/languageforge-lexbox/blob/develop/backend/Testing/SyncReverseProxy/SendReceiveServiceTests.cs#L145-L148
The existing comment in SendNewProject — "must be in separate commits otherwise hg runs out of memory" — is the same tuning acknowledgement from the client side. A 400Mi ceiling for a 180 MB unbundle leaves no headroom, so whether it OOMs depends on runner memory pressure at that moment. Note ENABLE_DEMAND_IMPORT: false in the same deployment, which raises Mercurial's baseline footprint; that is longstanding config, not a recent change.
Options
- Raise the
hgweb memory limit in the base deployment (or just in the test overlay) enough to clear a 180 MB unbundle. Most faithful to what the test is meant to cover, but it also reflects a real production limit worth thinking about — a 180 MB user push would hit the same wall.
- Shrink the test payload, which weakens what the test proves.
- Split the push into more commits, which only moves the ceiling.
The first seems right, but the number should come from someone who knows what production actually needs to survive.
Also seen
The same suite failed elsewhere on unrelated branches the same day, so SendNewProject_Big may not be the only instability: UnauthorizedUserCloneHgWeb (31695546591), InitFwDataProject_PopulatesTheEmptyRepoWithTheRequestedWritingSystems (31680119938), and whole-block ApiTests failures (31673009635, 31667434391). Those were not investigated here and may well have separate causes.
Testing.SyncReverseProxy.SendReceiveServiceTests.SendNewProject_Bigfails intermittently in the GHA integration tests because thehgwebcontainer is OOMKilled while unbundling the push. The test is not at fault — it sits right at the memory ceiling by design — but it fails on unrelated PRs and gets re-rolled, so it costs review time.Evidence
Two runs about 50 minutes apart, on unrelated branches, with an identical signature:
chore/document-dotnet8-chorusmerge-pinpt/missing-lex-code(#2547)Chain, from each run's
dotnet-k8s-logsartifact:logs-ingress.txt— nginx buffers the request body to a temp file, thenPOST /send-new-project-<id>-dev-flex?cmd=unbundlereturns502after ~1.1 s.logs-lexbox.txt— the API proxies tohg:8088and getsSystem.Net.Http.HttpIOException: The response ended prematurely. (ResponseEnded), followed by a run ofHttpRequestException: Connection refused (hg:8088). Collateral:Job DataUpdate.UpdateProjectMetadataJob threw a JobExecutionException: Connection refused (hg:8088).describe-hg.txt— thehgwebcontainer'sLast StateisTerminated / Reason: OOMKilled / Exit Code: 137, finishing at the same second as the 502.Restart Count: 1; the pod comes back and the rest of the suite passes.abort: HTTP Error 502: Bad Gateway, and the assertion fails withDid not expect srOutput "Sync failure: ... Failed to send to LexBox".Cause
hgwebis capped at 400Mi:https://github.com/sillsdev/languageforge-lexbox/blob/develop/deployment/base/hg-deployment.yaml#L86-L90
The test pushes 180 MB across 10 commits:
https://github.com/sillsdev/languageforge-lexbox/blob/develop/backend/Testing/SyncReverseProxy/SendReceiveServiceTests.cs#L145-L148
The existing comment in
SendNewProject— "must be in separate commits otherwise hg runs out of memory" — is the same tuning acknowledgement from the client side. A 400Mi ceiling for a 180 MB unbundle leaves no headroom, so whether it OOMs depends on runner memory pressure at that moment. NoteENABLE_DEMAND_IMPORT: falsein the same deployment, which raises Mercurial's baseline footprint; that is longstanding config, not a recent change.Options
hgwebmemory limit in the base deployment (or just in the test overlay) enough to clear a 180 MB unbundle. Most faithful to what the test is meant to cover, but it also reflects a real production limit worth thinking about — a 180 MB user push would hit the same wall.The first seems right, but the number should come from someone who knows what production actually needs to survive.
Also seen
The same suite failed elsewhere on unrelated branches the same day, so
SendNewProject_Bigmay not be the only instability:UnauthorizedUserCloneHgWeb(31695546591),InitFwDataProject_PopulatesTheEmptyRepoWithTheRequestedWritingSystems(31680119938), and whole-blockApiTestsfailures (31673009635, 31667434391). Those were not investigated here and may well have separate causes.