Skip to content

perf: cache backend and nplike lookups for unrecognised types - #4341

Open
ikrommyd wants to merge 4 commits into
scikit-hep:mainfrom
ikrommyd:perf-improve-dispatch-cache
Open

ikrommyd wants to merge 4 commits into
scikit-hep:mainfrom
ikrommyd:perf-improve-dispatch-cache

Conversation

@ikrommyd

Copy link
Copy Markdown
Member

backend_of_obj and nplike_of_obj re-scanned every registered lookup factory on every call for types that are not arrays at all, for example the field name string passed to SlicingErrorContext. We cache the misses too and clear them when a new backend or nplike is registered.

@github-actions github-actions Bot added the type/perf PR title type: perf (set automatically) label Sep 12, 2026
@codecov

codecov Bot commented Sep 12, 2026 •

Copy link
Copy Markdown

⚠️ JUnit XML file not found

The CLI was unable to find any JUnit XML files to upload.
For more help, visit our troubleshooting guide.

@ikrommyd
ikrommyd marked this pull request as ready for review September 13, 2026 10:33
@ikrommyd

Copy link
Copy Markdown
Member Author

🤖 AI text below 🤖

CPU timings, main vs this branch, median of 7 repeats per case.

case base (us) branch (us) change
to_list roundtrip [n=9000] 14,874.8 1,946.0 -86.9%
to_list roundtrip [n=3] 81.71 67.83 -17.0%
getitem python slice [n=9000] 17.37 15.43 -11.2%
getitem python slice [n=3] 17.04 15.35 -9.9%
fill_none python str [n=9000] 297.1 294.6 -0.8%

Median across all cases: -3.8% (28 cases total, showing the four best and the worst).

benchmark script
# Benchmark for PR #4341: caching backend/nplike lookups for unrecognised types.
# Every statement below mixes plain Python objects (int/float/list/str) into
# awkward operations, so the "type nothing recognises" lookup path is hit often.
import statistics, timeit
import awkward as ak

def report(label, stmt, glb, *, number, repeats=7):
    ts = timeit.repeat(stmt, number=number, repeat=repeats, globals=glb)
    us = sorted(t / number * 1e6 for t in ts)
    print(f"{label:<45} {statistics.median(us):10.3f} us  min {us[0]:10.3f}  sd {statistics.stdev(us):8.3f}  (n={number}x{repeats})")

print(f"awkward {ak.__version__} from {ak.__file__}")


def make(n):
    return ak.Array([[1.1, 2.2, None], [], [3.3]] * n)


CASES = [
    ("add python int", "array + 1"),
    ("add python float", "array + 1.5"),
    ("multiply python int", "array * 2"),
    ("compare python float", "array > 2.0"),
    ("fill_none python int", "ak.fill_none(array, 0)"),
    ("fill_none python str", "ak.fill_none(array, 'x')"),
    ("getitem python int", "array[0]"),
    ("getitem python slice", "array[1:]"),
    ("getitem nested python ints", "array[0, 0]"),
    ("getitem python list", "array[[0, 1, 2]]"),
    ("with_field python scalar", "ak.zip({'x': array, 'y': array}) "),
    ("concatenate with python list", "ak.concatenate([array, [[9.9]]])"),
    ("where python scalars", "ak.where(array > 2.0, 1.0, 0.0)"),
    ("to_list roundtrip", "ak.Array(pylist)"),
]

for n in (1, 3000):
    array = make(n)
    pylist = array.to_list()
    glb = {"ak": ak, "array": array, "pylist": pylist}
    number = 1000 if n == 1 else 50
    print(f"\n--- {len(array)} entries ---")
    for label, stmt in CASES:
        report(f"{label} [n={len(array)}]", stmt, glb, number=number)

github-actions Bot added a commit that referenced this pull request Sep 14, 2026
@github-actions

Copy link
Copy Markdown
Contributor

The documentation preview is ready to be viewed at https://awkward-array.org/doc/pr/4341/

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type/perf PR title type: perf (set automatically)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant