Summary
We instrumented the async conformance client (conformance/test/client.py + src/connectrpc/_client_async.py) on an otherwise idle 64-core Linux box (Rocky 10, Python 3.13.14, no coverage tracing) and measured, per run:
- Baseline
pytest -k client_async: 196 Deadline Propagation + 21 Client Cancellation scenario failures.
- The client's single event-loop thread is CPU-saturated for ~90% of the run: a 10 ms ticker task logged 128 stall bursts (median 658 ms wall lag, max 1.9 s), and
time.thread_time() deltas show the lag is CPU-bound (median thread-CPU 636 ms ≈ wall 658 ms; totals 78.4 s CPU vs 80.9 s wall in an 86 s run). It is not GIL/futex blocking (0 events where thread-CPU < 20% of wall), not cgroup throttling (cpu.stat nr_throttled=0, loadavg 0.44/0.45/0.76 1m/5m/15m at run time, box idle), and not a single slow Python callback (asyncio debug mode logs only 7 steps >100 ms, max 0.46 s). The specific CPU consumer is not yet attributed — candidates include per-call harness/codec work and synchronous sections in the pyqwest bridge — but the loop thread being ~90% busy while the driver pipelines many concurrent cases into the process is measured directly.
- A harness-level experiment supports a second, independent contributor: in the async
BidiStream path of conformance/test/client.py, the same task sleeps request_delay_ms between sends and is the only task advancing the response generator. Moving response pumping into a separate task (non-cancel scenarios only) reduced Deadline Propagation failures 196 → 139 in one run. (Caveat: that patch variant also changed cancel-path semantics, so the cancel-family number from that run is not comparable.)
- A minimal micro-test of the library's streaming shape:
asyncio.timeout() entered inside an async generator cancels the consuming task; if the consumer is parked at another await, CancelledError surfaces at the consumer's await point unconverted, and the generator's TimeoutError only surfaces at generator finalization. This matches the observed mix of late DEADLINE_EXCEEDED surfacing and CANCELED conversions.
Hypotheses we tested and ruled out (with evidence)
- Premature deadline timer: no ~1.1 s abort cluster exists; per-call pairing shows aborts at/after budget, header
connect-timeout-ms equals the full budget.
- Uncancellable transport reads (pyqwest): micro-benchmark cancel-to-
CancelledError = 1 ms, 4/4 runs.
- Coverage tracing overhead: runs used plain pytest;
cov fixture resolves to None without --cov, so the client subprocess is untraced.
- The modified 3.10
asyncio_timeout backport: Python 3.13 uses stdlib asyncio.timeout; the backport never executes.
- Environment load: another suite (sync client) passes on the same box in the same session; cgroup throttling zero.
Environment
Open questions
- Is the reference async Python client expected to execute driver-pipelined concurrent cases on a single loop thread, given the 200 ms/2 s budgets in
Deadline Propagation? On this box the thread saturates and budgets blow.
- Would the recommended shape be a dedicated reader task per streaming call (mirroring the Go reference client's structure), or is the CPU burn itself (per-call compression/codec work on the loop) the thing to address first?
- We could not run py-spy against this interpreter build (0.4.2 does not detect the uv cpython-3.13.14 binary); if useful we can produce a cProfile breakdown of the loop thread's CPU next.
Happy to iterate on instrumentation or a PR direction — not claiming a specific library defect yet; the harness single-task interleave looks like the most actionable piece.
Summary
We instrumented the async conformance client (
conformance/test/client.py+src/connectrpc/_client_async.py) on an otherwise idle 64-core Linux box (Rocky 10, Python 3.13.14, no coverage tracing) and measured, per run:pytest -k client_async: 196Deadline Propagation+ 21Client Cancellationscenario failures.time.thread_time()deltas show the lag is CPU-bound (median thread-CPU 636 ms ≈ wall 658 ms; totals 78.4 s CPU vs 80.9 s wall in an 86 s run). It is not GIL/futex blocking (0 events where thread-CPU < 20% of wall), not cgroup throttling (cpu.stat nr_throttled=0, loadavg 0.44/0.45/0.76 1m/5m/15m at run time, box idle), and not a single slow Python callback (asyncio debug mode logs only 7 steps >100 ms, max 0.46 s). The specific CPU consumer is not yet attributed — candidates include per-call harness/codec work and synchronous sections in the pyqwest bridge — but the loop thread being ~90% busy while the driver pipelines many concurrent cases into the process is measured directly.BidiStreampath ofconformance/test/client.py, the same task sleepsrequest_delay_msbetween sends and is the only task advancing the response generator. Moving response pumping into a separate task (non-cancel scenarios only) reducedDeadline Propagationfailures 196 → 139 in one run. (Caveat: that patch variant also changed cancel-path semantics, so the cancel-family number from that run is not comparable.)asyncio.timeout()entered inside an async generator cancels the consuming task; if the consumer is parked at another await,CancelledErrorsurfaces at the consumer's await point unconverted, and the generator'sTimeoutErroronly surfaces at generator finalization. This matches the observed mix of lateDEADLINE_EXCEEDEDsurfacing andCANCELEDconversions.Hypotheses we tested and ruled out (with evidence)
connect-timeout-msequals the full budget.CancelledError= 1 ms, 4/4 runs.covfixture resolves toNonewithout--cov, so the client subprocess is untraced.asyncio_timeoutbackport: Python 3.13 uses stdlibasyncio.timeout; the backport never executes.Environment
2179d4a(post-Queue docs deployments instead of racing them #355), driver pinVERSION_CONFORMANCE = "v1.0.5"(conformance/test/_util.py:11)pytest -k client_asyncOpen questions
Deadline Propagation? On this box the thread saturates and budgets blow.Happy to iterate on instrumentation or a PR direction — not claiming a specific library defect yet; the harness single-task interleave looks like the most actionable piece.