From 433388d79792e91c649544bc96d71adbd089f2e3 Mon Sep 17 00:00:00 2001 From: Andres Rios Tascon Date: Tue, 15 Sep 2026 11:55:58 -0400 Subject: [PATCH] docs: read the Chicago taxi dataset from an R2 mirror The 10-minutes and how-to-examine-single-item notebooks fetched chicago-taxi.parquet from Zenodo during sphinx-build, so the docs build failed whenever Zenodo was slow or returned a 5xx. Point the reads at https://test-files.awkward-array.org/chicago-taxi.parquet, a byte-identical copy (640173859 bytes, 25 row groups) on an R2 bucket we control. Range requests work, so `ak.metadata_from_parquet` and the `row_groups=[0]` read still download only what they need, and both are noticeably faster than Zenodo. The prose keeps a link to the Zenodo record for provenance. Co-authored-by: Ianna Osborne Assisted-by: claude-code:claude-opus-5 --- docs/getting-started/10-minutes-to-awkward-array.md | 6 +++--- docs/user-guide/how-to-examine-single-item.md | 2 +- 2 files changed, 4 insertions(+), 4 deletions(-) diff --git a/docs/getting-started/10-minutes-to-awkward-array.md b/docs/getting-started/10-minutes-to-awkward-array.md index 37fe5638bb..5260444003 100644 --- a/docs/getting-started/10-minutes-to-awkward-array.md +++ b/docs/getting-started/10-minutes-to-awkward-array.md @@ -24,7 +24,7 @@ In this guide, we'll look at how to manipulate a jagged dataset to plot taxi rou ## Loading the dataset -Our dataset is formatted as a 611 MB [Apache Parquet](https://parquet.apache.org/) file, provided [here](https://zenodo.org/records/14537442/files/chicago-taxi.parquet). Alongside JSON, and raw buffers, Awkward can also read Parquet files and Arrow tables. +Our dataset is formatted as a 611 MB [Apache Parquet](https://parquet.apache.org/) file, provided [here](https://test-files.awkward-array.org/chicago-taxi.parquet) (a mirror of [this Zenodo record](https://zenodo.org/records/14537442)). Alongside JSON, and raw buffers, Awkward can also read Parquet files and Arrow tables. Given that this file is so large, let's first look at the *metadata* with `ak.metadata_from_parquet` to see what we're working with: @@ -43,7 +43,7 @@ import numpy as np import awkward as ak metadata = ak.metadata_from_parquet( - "https://zenodo.org/records/14537442/files/chicago-taxi.parquet" + "https://test-files.awkward-array.org/chicago-taxi.parquet" ) ``` @@ -59,7 +59,7 @@ There are a lot of different columns here (`trip.sec`, `trip.begin.lon`, `trip.p ```{code-cell} ipython3 taxi = ak.from_parquet( - "https://zenodo.org/records/14537442/files/chicago-taxi.parquet", + "https://test-files.awkward-array.org/chicago-taxi.parquet", row_groups=[0], columns=["trip.km", "trip.begin.l*", "trip.end.l*", "trip.path.*"], ) diff --git a/docs/user-guide/how-to-examine-single-item.md b/docs/user-guide/how-to-examine-single-item.md index fa763f75c1..f7003dc691 100644 --- a/docs/user-guide/how-to-examine-single-item.md +++ b/docs/user-guide/how-to-examine-single-item.md @@ -27,7 +27,7 @@ First, let's load the dataset using the {func}`ak.from_parquet` function. We wil ```{code-cell} ipython3 import awkward as ak -url = "https://zenodo.org/records/14537442/files/chicago-taxi.parquet" +url = "https://test-files.awkward-array.org/chicago-taxi.parquet" taxi = ak.from_parquet( url, row_groups=[0],