Skip to content

perf: cache compiled kernels and call them through void pointers - #4338

Open
ikrommyd wants to merge 10 commits into
scikit-hep:mainfrom
ikrommyd:perf-improve-kernel-dispatch
Open

ikrommyd wants to merge 10 commits into
scikit-hep:mainfrom
ikrommyd:perf-improve-kernel-dispatch

Conversation

@ikrommyd

@ikrommyd ikrommyd commented Sep 12, 2026 •

Copy link
Copy Markdown
Member

Kernels were constructed anew on every dispatch, so they are cached on the backend now. Building a typed ctypes pointer for each buffer also costs several times more than the kernel call itself, so the same function is re-prototyped with void * parameters and handed plain addresses instead.

The numpy and jax kernels call the same awkward-cpp functions (which is why the jax backend needs its buffers on the cpu), so they now share that calling convention in a common base and differ only in how a buffer's address is taken. The cache sits in Backend.__getitem__ with each backend providing _new_kernel, so cupy and typetracer get it as well. It matters most for jax, whose kernel constructor was importing jax and parsing a version string on every lookup.

Kernel lookup goes from 143 ns to 44 ns, and a small kernel call from 9.2 us to 5.2 us.

@github-actions github-actions Bot added the type/perf PR title type: perf (set automatically) label Sep 12, 2026
@ikrommyd
ikrommyd marked this pull request as ready for review September 13, 2026 10:33
@ikrommyd
ikrommyd marked this pull request as draft September 13, 2026 11:32
@ikrommyd
ikrommyd marked this pull request as ready for review September 13, 2026 13:27
@codecov

codecov Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 83.90%. Comparing base (60390ac) to head (2287776).

Additional details and impacted files
Files with missing lines Coverage Δ
src/awkward/_backends/backend.py 93.75% <100.00%> (+1.64%) ⬆️
src/awkward/_backends/cupy.py 97.72% <100.00%> (+0.05%) ⬆️
src/awkward/_backends/jax.py 100.00% <100.00%> (ø)
src/awkward/_backends/numpy.py 100.00% <100.00%> (ø)
src/awkward/_backends/typetracer.py 100.00% <100.00%> (ø)
src/awkward/_kernels.py 95.80% <100.00%> (+4.41%) ⬆️

... and 1 file with indirect coverage changes

@ikrommyd

Copy link
Copy Markdown
Member Author

🤖 AI text below 🤖

CPU timings, main vs this branch, median of 7 repeats per case.

case base (us) branch (us) change
ak.drop_none(small_opt) 109.4 86.95 -20.5%
small[:, 1:] 83.11 66.89 -19.5%
ak.concatenate([small, small]) 108.6 92.96 -14.4%
ak.sum(small, axis=1) 52.28 46.40 -11.3%
ak.num(small) 22.87 23.68 +3.5%

Median across all cases: -0.4% (15 cases total, showing the four best and the worst).

benchmark script
"""Benchmark for PR #4338: cached kernel lookup + void-pointer ctypes calls.

The win is per-kernel-call overhead, so most cases are many operations on
*small* arrays; the last case is large so that the kernel body dominates and
shows there is no regression there.
"""

from __future__ import annotations

import statistics
import timeit

import numpy as np

import awkward as ak


def report(label, stmt, glb, *, number, repeats=7):
    ts = timeit.repeat(stmt, number=number, repeat=repeats, globals=glb)
    us = sorted(t / number * 1e6 for t in ts)
    print(
        f"{label:<45} {statistics.median(us):10.3f} us  min {us[0]:10.3f}  sd {statistics.stdev(us):8.3f}  (n={number}x{repeats})"
    )


print(f"awkward {ak.__version__} from {ak.__file__}")

rng = np.random.default_rng(12345)

small = ak.Array([[1, 2, 3], [], [4, 5], [6], [7, 8, 9, 10]])
small_rec = ak.Array([{"x": [1, 2], "y": 1.5}, {"x": [], "y": 2.5}, {"x": [3], "y": 3.5}])
small_opt = ak.Array([[1, None, 3], [], [None], [4, 5]])

counts = rng.integers(0, 8, size=1_000_000)
big = ak.unflatten(ak.Array(rng.random(int(counts.sum()))), counts)

glb = globals()

# --- small arrays: dispatch overhead dominates ---------------------------------
report("ak.num(small)", "ak.num(small)", glb, number=2000)
report("ak.flatten(small)", "ak.flatten(small)", glb, number=2000)
report("small[:, 1:]", "small[:, 1:]", glb, number=2000)
report("small[[0, 2, 4]]", "small[[0, 2, 4]]", glb, number=2000)
report("small[small_mask]", "small[np.array([True, False] * 2 + [True])]", glb, number=2000)
report("ak.to_list(small)", "ak.to_list(small)", glb, number=1000)
report("ak.is_none(small_opt, axis=1)", "ak.is_none(small_opt, axis=1)", glb, number=1000)
report("ak.drop_none(small_opt)", "ak.drop_none(small_opt)", glb, number=500)
report("ak.sum(small, axis=1)", "ak.sum(small, axis=1)", glb, number=1000)
report("ak.cartesian([small, small], axis=1)", "ak.cartesian([small, small], axis=1)", glb, number=200)
report("small_rec.x", "small_rec.x", glb, number=5000)
report("ak.concatenate([small, small])", "ak.concatenate([small, small])", glb, number=500)

# --- large array: kernel body dominates, check for no regression ---------------
report("ak.num(big)  [1M lists]", "ak.num(big)", glb, number=20)
report("ak.flatten(big)  [1M lists]", "ak.flatten(big)", glb, number=20)
report("ak.sum(big, axis=1)  [1M lists]", "ak.sum(big, axis=1)", glb, number=10)

github-actions Bot added a commit that referenced this pull request Sep 15, 2026
@github-actions

Copy link
Copy Markdown
Contributor

The documentation preview is ready to be viewed at https://awkward-array.org/doc/pr/4338/

github-actions Bot added a commit that referenced this pull request Sep 15, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type/perf PR title type: perf (set automatically)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant