Skip to content

fix: ak.from_iter of np.clongdouble, 0-d arrays, and np.void scalars - #4392

Open
lgray wants to merge 5 commits into
scikit-hep:mainfrom
graphed-org:fix/from-iter-clongdouble
Open

lgray wants to merge 5 commits into
scikit-hep:mainfrom
graphed-org:fix/from-iter-clongdouble

Conversation

@lgray

@lgray lgray commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

🤖 AI text below 🤖

Bug

ak.from_iter segfaults on np.clongdouble:

>>> import numpy as np, awkward as ak
>>> ak.from_iter([np.clongdouble(1 + 1j)])
Segmentation fault

builder_fromiter (awkward-cpp/src/python/content.cpp) has no branch for complex numpy scalars other than complex128 (a Python complex subclass), so np.clongdouble reaches the tolist branch. np.clongdouble.tolist() returns the scalar itself, and the builder recurses until the stack overflows. The same happens for a clongdouble ndarray, and for ak.from_iter(ak.to_list(<clongdouble array>)).

The same function also recurses without a bound on anything that never bottoms out: a tolist()/to_list() that returns another such object, a 0-d object array that contains itself, or very deep nesting with a raised sys.setrecursionlimit. Each of these segfaults on main.

Fix

In builder_fromiter:

  • np.complexfloating scalars are built as complex128, the same widening complex64 already gets.
  • A 0-d np.ndarray recurses on obj[()]. Before, tolist() dropped the datetime64/timedelta64 unit (datetime64[ns] gave int64) and the field names of a structured value.
  • A structured np.void scalar becomes a record over dtype.names; an unstructured one becomes bytes. Before, a structured scalar was iterated as a list (var * float64, ints turned into floats).
  • The recursion depth is capped at min(sys.getrecursionlimit(), 1000) and exceeding it raises RecursionError.

In ErrorContext.format_argument, arguments are formatted with a reprlib.Repr bounded in depth. With a raised recursion limit on CPython 3.10 and 3.11, the plain repr() of the deep argument overflowed the stack while the RecursionError message was being built. The text is unchanged for any argument that fits in the message.

Behavior changes

  • A 1-d structured ndarray or recarray now gives a record. ak.from_iter([np.array([(1, 2.5)], dtype=[("a", "i8"), ("b", "f8")])]) gave 1 * var * var * float64 and now gives 1 * var * {a: int64, b: float64}; the same for .view(np.recarray). This is what ak.from_numpy gives for the same array.
  • Nesting deeper than about 1000 levels raises RecursionError whatever sys.getrecursionlimit() is. At the default limit, main already raises RecursionError from Python-level code for a list nested 987 to 20,000 deep inside [x], but segfaults at 200,000 deep (macOS arm64, and in the linux/amd64 test run below); this branch raises RecursionError at every one of those depths. With a raised limit, main built deeper nesting until the stack overflowed; this branch raises from 999 levels. The cap is not tied to sys.getrecursionlimit() because a raised limit brings the segfault back, and neither CPython's C recursion limit (Py_EnterRecursiveCall) nor the limit stays below the stack's capacity for these frames. Nesting that deep could not be used anyway: ak.forms.from_json of the resulting form raises RecursionError at depth 1000.

Tests

tests/test_4392_from_iter_numpy_scalars.py: values checked against ak.from_numpy, clongdouble scalars/0-d/ndarrays, the non-terminating members and the depth cap. The cases that crash on main run in a subprocess (skipped on emscripten, as in test_2682), so on main they fail rather than take down the test session.

On main, 19 of the first 29 tests fail and 10 pass, on macOS arm64 and on linux/amd64 with awkward-cpp built from the tree. The later commits add a dict member to the non-terminating cases, test_argument_text_matches_repr (including a dict nested past the depth cap, which found that reprlib.Repr.fillvalue does not exist on 3.10) and test_argument_whose_repr_raises. On this branch all 38 tests pass on macOS arm64 with CPython 3.13 and 3.10; with the format_argument change reverted, the two raised-limit members fail on 3.10.

Found by the numpy-dtype survey behind #4390.

builder_fromiter had no branch for complex numpy scalars other than
complex128, so np.clongdouble reached the tolist() branch; its tolist()
returns itself and the builder recursed until the stack overflowed.

- np.complexfloating scalars are built as complex128.
- 0-d ndarrays are built from obj[()], keeping datetime64/timedelta64
  units and structured field names that tolist() dropped.
- structured np.void scalars are built as records; unstructured ones
  as bytes.
- an object whose tolist() returns its own type raises TypeError
  instead of recursing.
builder_fromiter recursed on nested containers and on tolist()/to_list()
results with nothing to stop it, so a tolist() 2-cycle, a to_list() that
returns its own type, a 0-d object array containing itself, or a list,
tuple or dict nested 200000 deep overflowed the stack.

- A thread-local depth counter raises RecursionError once the nesting
  reaches sys.getrecursionlimit(). Py_EnterRecursiveCall is not enough:
  on Python 3.13 linux/amd64 its C-level limit lets tuples and dicts
  nested 200000 deep still overflow the stack.
- The tolist() same-type TypeError is removed; the bound covers it.
- The np.void branch moves after the dict branch and skips lists, so
  from_iter of records is not slowed by the numpy lookup.
- Subprocess tests are skipped on emscripten, which has no subprocess.
The depth bound was sys.getrecursionlimit(), which users raise. Past about
7600 levels (tuples, linux/amd64) or 4800 (nested 0-d object arrays) the
builder overflows the stack first, so setrecursionlimit(10_000) or higher
brought the segfault back.

The bound is now min(sys.getrecursionlimit(), 1000). 1000 is CPython's
default limit, so default behavior is unchanged and no platform recurses
deeper than it already did at the default; Windows gets roughly a third of
the linux C-stack budget (Py_C_RECURSION_LIMIT 3000 vs 10000). A list
nested deeper than 1000 now raises RecursionError even with a raised limit.
@github-actions github-actions Bot added the type/fix PR title type: fix (set automatically) label Sep 26, 2026
@codecov

codecov Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 84.10%. Comparing base (1f04d6f) to head (55f2271).

Additional details and impacted files
Files with missing lines Coverage Δ
src/awkward/_errors.py 85.82% <100.00%> (+1.24%) ⬆️

@lgray
lgray force-pushed the fix/from-iter-clongdouble branch from d8c74c0 to 55f2271 Compare September 26, 2026 20:21
@lgray

lgray commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

py3.14t failing for unrelated reasons (build itself fails, PR does not touch CI config)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type/fix PR title type: fix (set automatically)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant