docs: read the Chicago taxi dataset from an R2 mirror - #4327
Conversation
|
The documentation preview is ready to be viewed at https://awkward-array.org/doc/pr/4327/ |
|
The PR preview looks like this:
The 2.13 doc is
The page is a hands-on tutorial for new users. I think that the current doc is easier to follow. metadata = ak.metadata_from_parquet(
"https://zenodo.org/records/14537442/files/chicago-taxi.parquet"
)There should be a way to use caching without changing the visible rendered doc pages. |
|
Maybe we should copy this data to somewhere else? Zenodo is a permanent archive site with DOIs. It might not be designed for serving data to CI. We can copy the data somewhere more appropriate for hosting data, and then put the Zenodo DOI as a citation. |
I put the file in an R2 bucket. @ianna could you try changing this PR to use https://test-files.awkward-array.org/chicago-taxi.parquet ? |
@ariostas - thanks! Feel free to push the update. |
The 10-minutes and how-to-examine-single-item notebooks fetched chicago-taxi.parquet from Zenodo during sphinx-build, so the docs build failed whenever Zenodo was slow or returned a 5xx. Point the reads at https://test-files.awkward-array.org/chicago-taxi.parquet, a byte-identical copy (640173859 bytes, 25 row groups) on an R2 bucket we control. Range requests work, so `ak.metadata_from_parquet` and the `row_groups=[0]` read still download only what they need, and both are noticeably faster than Zenodo. The prose keeps a link to the Zenodo record for provenance. Co-authored-by: Ianna Osborne <ianna.osborne@cern.ch> Assisted-by: claude-code:claude-opus-5
4153a1c to
433388d
Compare
|
That worked well. So @TaiSakuma, this is ready for re-review when |
TaiSakuma
left a comment
There was a problem hiding this comment.
Thanks. It is nice to have the data file served from R2 and accessible under awkward-array.org.
🤖 The text below was written by Claude.
Verification of the R2 mirror and the rendered pages, done on 2026-09-15 against commit 433388d:
- The mirror is byte-identical to the Zenodo copy. A full download of
https://test-files.awkward-array.org/chicago-taxi.parquetis 640173859 bytes with MD5e7bde64b9e87f41b27edfa2da7424a23, equal to the checksum in the Zenodo record's API metadata. The mirror's ETag is a multipart one, so the identity had to be confirmed by hashing rather than from headers. - Range requests work as the notebooks need. Both endpoints answer partial reads with 206; the mirror served 1 MB ranges in 0.2 to 0.3 s against 1.1 s from Zenodo. Zenodo's file endpoint also advertises a rate limit (
x-ratelimit-limit: 133), which supports moving build traffic off it. - The two notebook cells reproduce with awkward 2.13.0 in a fresh venv:
ak.metadata_from_parquetreports 25 row groups and 7728 rows, andak.from_parquet(..., row_groups=[0], columns=[...])returns 353 elements withak.sum(taxi.trip.km)= 7097739.0, matching the table in the description. - The rendered PR preview pages differ from
doc/mainanddoc/2.13only in the URL string in the code cells, the'paths'entry of the metadata output, the added prose sentence linking the Zenodo record, and theHTTPFileSystemobject address that changes on every build. All other outputs on both pages are identical. - The Build Docs job executed both notebooks live against the mirror (14.15 s and 3.57 s) rather than serving them from the Jupyter cache.
- No other reference to the Chicago taxi file exists in the repository outside these two notebooks.
Public API and tests are untouched.
🤖 Generated with Claude Code


🤖 AI text below 🤖
The 10-minutes and how-to-examine-single-item notebooks fetched the 611 MB
chicago-taxi.parquetlive from Zenodo duringsphinx-build, so the docs build failed whenever Zenodo was slow or returned a 5xx (e.g. a 504 Gateway Time-out).This replaces the earlier caching approach in this PR with a much smaller change: point the reads at
https://test-files.awkward-array.org/chicago-taxi.parquet, a copy of the dataset on an R2 bucket we control. Nodocs/_cache/, noAWKWARD_CHICAGO_TAXI_PARQUETenv var, noactions/cachestep — the notebooks stay runnable standalone with no fallback logic.The mirror is byte-identical to the Zenodo copy, and Cloudflare honours range requests, so
ak.metadata_from_parquetand therow_groups=[0]read still download only what they need rather than the whole 611 MB:metadata_from_parquetfrom_parquet(row_groups=[0], columns=...)sum(trip.km)= 7097739.0The prose in 10 minutes to Awkward Array keeps a link to the Zenodo record for provenance, since that is the citable source.
The tradeoff versus the caching approach: docs builds still need the network at build time, so an R2 outage breaks the build the way a Zenodo outage did. The difference is that this bucket is ours rather than a shared service under load. If we want belt-and-braces later, a caching layer can sit on top of this URL.