Skip to content

docs: read the Chicago taxi dataset from an R2 mirror - #4327

Merged
ianna merged 1 commit into
mainfrom
ianna/cache-chicago-taxi-parquet-docs
Sep 16, 2026
Merged

ianna merged 1 commit into
mainfrom
ianna/cache-chicago-taxi-parquet-docs

Conversation

@ianna

@ianna ianna commented Sep 9, 2026 •

Copy link
Copy Markdown
Member

🤖 AI text below 🤖

The 10-minutes and how-to-examine-single-item notebooks fetched the 611 MB chicago-taxi.parquet live from Zenodo during sphinx-build, so the docs build failed whenever Zenodo was slow or returned a 5xx (e.g. a 504 Gateway Time-out).

This replaces the earlier caching approach in this PR with a much smaller change: point the reads at https://test-files.awkward-array.org/chicago-taxi.parquet, a copy of the dataset on an R2 bucket we control. No docs/_cache/, no AWKWARD_CHICAGO_TAXI_PARQUET env var, no actions/cache step — the notebooks stay runnable standalone with no fallback logic.

The mirror is byte-identical to the Zenodo copy, and Cloudflare honours range requests, so ak.metadata_from_parquet and the row_groups=[0] read still download only what they need rather than the whole 611 MB:

R2 Zenodo
size 640173859 B 640173859 B
metadata_from_parquet 1.1 s 1.0 s
from_parquet(row_groups=[0], columns=...) 1.3 s 8.5 s
num_row_groups / num_rows 25 / 7728 25 / 7728
resulting array 353 elements, sum(trip.km) = 7097739.0 same

The prose in 10 minutes to Awkward Array keeps a link to the Zenodo record for provenance, since that is the citable source.

The tradeoff versus the caching approach: docs builds still need the network at build time, so an R2 outage breaks the build the way a Zenodo outage did. The difference is that this bucket is ours rather than a shared service under load. If we want belt-and-braces later, a caching layer can sit on top of this URL.

@github-actions github-actions Bot added the type/docs PR title type: docs (set automatically) label Sep 9, 2026
@ianna
ianna requested a review from TaiSakuma September 9, 2026 04:16
github-actions Bot added a commit that referenced this pull request Sep 9, 2026
@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

The documentation preview is ready to be viewed at https://awkward-array.org/doc/pr/4327/

@codecov

codecov Bot commented Sep 9, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 62.45%. Comparing base (08dcbdc) to head (433388d).
✅ All tests successful. No failed tests found.

Additional details and impacted files

@TaiSakuma

Copy link
Copy Markdown
Member

The PR preview looks like this:

Screenshot 2026-09-09 at 8 27 26 AM

The 2.13 doc is

Screenshot 2026-09-09 at 8 29 13 AM

The page is a hands-on tutorial for new users. I think that the current doc is easier to follow.

metadata = ak.metadata_from_parquet(
    "https://zenodo.org/records/14537442/files/chicago-taxi.parquet"
)

There should be a way to use caching without changing the visible rendered doc pages.

@TaiSakuma

Copy link
Copy Markdown
Member

Maybe we should copy this data to somewhere else?

Zenodo is a permanent archive site with DOIs. It might not be designed for serving data to CI. We can copy the data somewhere more appropriate for hosting data, and then put the Zenodo DOI as a citation.

@ariostas

Copy link
Copy Markdown
Member

Maybe we should copy this data to somewhere else?

I put the file in an R2 bucket.

@ianna could you try changing this PR to use https://test-files.awkward-array.org/chicago-taxi.parquet ?

@ianna

ianna commented Sep 15, 2026

Copy link
Copy Markdown
Member Author

Maybe we should copy this data to somewhere else?

I put the file in an R2 bucket.

@ianna could you try changing this PR to use https://test-files.awkward-array.org/chicago-taxi.parquet ?

@ariostas - thanks! Feel free to push the update.

The 10-minutes and how-to-examine-single-item notebooks fetched
chicago-taxi.parquet from Zenodo during sphinx-build, so the docs build
failed whenever Zenodo was slow or returned a 5xx.

Point the reads at https://test-files.awkward-array.org/chicago-taxi.parquet,
a byte-identical copy (640173859 bytes, 25 row groups) on an R2 bucket we
control. Range requests work, so `ak.metadata_from_parquet` and the
`row_groups=[0]` read still download only what they need, and both are
noticeably faster than Zenodo. The prose keeps a link to the Zenodo record
for provenance.

Co-authored-by: Ianna Osborne <ianna.osborne@cern.ch>
Assisted-by: claude-code:claude-opus-5
@ariostas
ariostas force-pushed the ianna/cache-chicago-taxi-parquet-docs branch from 4153a1c to 433388d Compare September 15, 2026 15:58
@ariostas ariostas changed the title docs: cache the Chicago taxi Parquet dataset for the docs build docs: read the Chicago taxi dataset from an R2 mirror Sep 15, 2026
@ariostas

Copy link
Copy Markdown
Member

That worked well. So @TaiSakuma, this is ready for re-review when

github-actions Bot added a commit that referenced this pull request Sep 15, 2026

@TaiSakuma TaiSakuma left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. It is nice to have the data file served from R2 and accessible under awkward-array.org.

🤖 The text below was written by Claude.


Verification of the R2 mirror and the rendered pages, done on 2026-09-15 against commit 433388d:

  • The mirror is byte-identical to the Zenodo copy. A full download of https://test-files.awkward-array.org/chicago-taxi.parquet is 640173859 bytes with MD5 e7bde64b9e87f41b27edfa2da7424a23, equal to the checksum in the Zenodo record's API metadata. The mirror's ETag is a multipart one, so the identity had to be confirmed by hashing rather than from headers.
  • Range requests work as the notebooks need. Both endpoints answer partial reads with 206; the mirror served 1 MB ranges in 0.2 to 0.3 s against 1.1 s from Zenodo. Zenodo's file endpoint also advertises a rate limit (x-ratelimit-limit: 133), which supports moving build traffic off it.
  • The two notebook cells reproduce with awkward 2.13.0 in a fresh venv: ak.metadata_from_parquet reports 25 row groups and 7728 rows, and ak.from_parquet(..., row_groups=[0], columns=[...]) returns 353 elements with ak.sum(taxi.trip.km) = 7097739.0, matching the table in the description.
  • The rendered PR preview pages differ from doc/main and doc/2.13 only in the URL string in the code cells, the 'paths' entry of the metadata output, the added prose sentence linking the Zenodo record, and the HTTPFileSystem object address that changes on every build. All other outputs on both pages are identical.
  • The Build Docs job executed both notebooks live against the mirror (14.15 s and 3.57 s) rather than serving them from the Jupyter cache.
  • No other reference to the Chicago taxi file exists in the repository outside these two notebooks.

Public API and tests are untouched.

🤖 Generated with Claude Code

@ianna
ianna merged commit 5342e8d into main Sep 16, 2026
19 checks passed
@ianna
ianna deleted the ianna/cache-chicago-taxi-parquet-docs branch September 16, 2026 03:41
github-actions Bot added a commit that referenced this pull request Sep 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type/docs PR title type: docs (set automatically)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants