Skip to content

PLAN: record the P3 data half, and retract a box that is now wrong - #64

Merged
mmcky merged 2 commits into
mainfrom
docs/plan-p3-status
Aug 10, 2026
Merged

PLAN: record the P3 data half, and retract a box that is now wrong#64
mmcky merged 2 commits into
mainfrom
docs/plan-p3-status

Conversation

@mmcky

@mmcky mmcky commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Docs only, no code. Housekeeping after #62 and #63, while the consuming repoints wait on a zh-cn PR audit.

The retraction is the part worth reviewing. Repoint generating_mini.md's input URL was written while #14 was open and hedged it as required "either way". #14 then settled the question the other way: the builder lands as provenance, committed-frozen, not edited at all — including its input URL. Left as a bare unticked box, that line reads as outstanding work whose completion would destroy the artifact's value as provenance. It is struck with its reasoning intact rather than deleted, because the underlying worry was legitimate: an archived repo keeps serving that URL. The answer is that the rule requiring a builder to read from sources/ binds builders that run, and the substitution is recorded as prose in sources/README.md.

One figure corrected. The fold's consuming reads are 28 across four repos, not 21 across three — QuantEcon/test-actions-lecture-intro was found after that line was written. The old figure is named in the line rather than silently replaced, since it has already been mistaken for the total once.

Two additions that are not status. The headline now explains why CATALOG.md says 24 datasets while audit.json says 18 migrated — one counts what lives here, the other what is read from here, and a wave landing ahead of its repoints shows that gap by design. Without it the two documents read as contradicting each other. And P3 now records three things it proved that were not on its test list, the load-bearing one being that a constructed dataset's builder must land in the same PR as its data, because check_consumed_files.py asserts the builder: path resolves.

Ticked with outcomes rather than bare marks: sources/ exists and its README is load-bearing rather than documentary, and the 0.88% blob-limit margin is recorded where it explains why git check-attr is a gate.

Verified: consumed-file-check 24 manifests / 25 files / 0 errors, strict audit exit 0.

🤖 Generated with Claude Code

Housekeeping after #62 and #63. The retraction matters more than the
ticks.

`Repoint generating_mini.md's input URL` was written while #14 was open
and said to do it "either way". #14 then settled the question the other
way: the builder lands as provenance, committed-frozen, not edited at
all. The box is struck rather than deleted, with the reasoning, because
its worry was legitimate — an archived repo keeps serving that URL — and
the answer is that the rule binds builders that RUN. Left as a bare
unticked box it reads as outstanding work that would destroy the
artifact's value as provenance.

Corrected while here: the fold's consuming reads are 28 across FOUR
repos, not 21 across three. `test-actions-lecture-intro` was found after
that line was written and has been mistaken for out-of-scope once
already, so the old figure is named rather than silently replaced.

Ticked with their outcomes: sources/ exists and its README is now
load-bearing rather than documentary (the hash gate), and the 0.88%
blob-limit margin is recorded where it explains why check-attr is a gate.

Two additions that are not just status. The headline now explains why
CATALOG.md says 24 datasets while audit.json says 18 migrated — one
counts what lives here, the other what is read from here, and a wave
that lands ahead of its repoints shows the gap by design. And P3 records
three things it proved that were not on its test list, including that a
constructed dataset's builder must land in the same PR as its data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings August 10, 2026 05:04

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Updates PLAN.md to reflect post-P3 realities (after #62/#63) by clarifying metric discrepancies, recording completed Phase 3 “data half” outcomes, and explicitly retracting an outdated checklist item whose completion would now be harmful (editing a provenance-only frozen builder).

Changes:

  • Explain why CATALOG.md can report more datasets than audit.json reports as “migrated” (landed-without-consumers gap during repoint waves).
  • Replace Phase 3 checklist items with outcome-recording text (including the retraction of the generating_mini.md input-URL repoint box and updated consumer-read counts).
  • Add Phase 3 notes capturing additional proven constraints/observations discovered during the fold.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread PLAN.md Outdated
Comment thread PLAN.md Outdated
Copilot review on #64.

"one row per committed file" predates the format and is now wrong in a
way that costs someone a red build. check_sources() splits the README on
`## ` headings and reads the first 64-hex token in each filename-shaped
section, so a second sources/ file documented as a row in a shared table
has no recorded hash and fails the required check. AGENTS.md already
described the enforced shape correctly; PLAN.md was the last place
carrying the old one, which is exactly the drift this document's own
header warns about.

The retracted sentence is named rather than silently replaced, on the
same principle as the generating_mini.md box in the previous commit: a
reader who remembers the old wording should learn it was wrong, not
wonder whether they misread it.

Also named the two paths behind "both sources/ paths return 404 on
Pages", and re-measured them: sources/SCF_plus.dta and
sources/README.md are 404, lectures/cities_us.csv is 200. Every other
measurement in this document names its target; that one did not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@mmcky
mmcky merged commit 29289ef into main Aug 10, 2026
1 check passed
@mmcky
mmcky deleted the docs/plan-p3-status branch August 10, 2026 05:52
mmcky added a commit that referenced this pull request Aug 11, 2026
…cn lines (#68)

* PLAN: rule 6's enumeration omitted the canary and carried drifted zh-cn lines

Rule 6 is the section a repointer consults for "the enumerated reads, the
acceptance check, and why CI does not cover it". Its tables were the last
place in this document still saying 21 reads across three repos.

#64 corrected the P3 checklist line to 28 across four
repos but not this section, so the corrected total and the stale enumeration
sat in the same file. Anyone following the pointer got a missing consumer.

- 21/three -> 28/four throughout, with test-actions-lecture-intro's seven
  reads enumerated (heavy_tails 822/849/850/874 and data.ipynb:37 on media,
  mle:93 and inequality:249 on the redirect form)
- lecture-intro.zh-cn's heavy_tails lines corrected 810/837/838/862 ->
  811/838/839/863, the +1 drift from its 2026-08-10 sync PRs, and the whole
  table date-stamped with a re-derive instruction rather than presented as
  fact
- the acceptance grep gains the canary path, which lives under the other
  clone root and so was invisible to a repos/-shaped command; plus a second
  grep for `high_dim_data`, since the media-host grep passes trivially on a
  tree nobody repointed
- the audit-coverage paragraph recounted: the strict audit sees 12 of 28,
  the two data-url-guards recover 2 more, and 14 reads across zh-cn and the
  canary have no automated check of any kind

Both superseded figures are now recorded in the section as superseded, since
each has been mistaken for the total once.

Measured against each repo's main on 2026-08-11.

Plan: QuantEcon/workspace-lectures#23

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PLAN: say why rule 6's acceptance grep is deliberately broad

Copilot read the bare `high_dim_data` needle as an imprecision that
would false-positive on prose, and suggested narrowing both greps to
`QuantEcon/high_dim_data`. The needle is deliberate, but nothing said
so, which is a fair documentation gap.

Records the reasoning rather than the conclusion: an acceptance gate's
error costs are asymmetric (a false positive costs five seconds, a
false negative ships a broken fold); the broad form also catches a fork
reference and a URL that lost its org prefix; and inside a lecture tree
a non-URL hit is a finding to sweep, not noise -- which is the opposite
of the org-wide gate, worded as zero *executable data reads* precisely
because prose mentions there are permanent.

Measured while checking: all 28 hits across the four lecture trees are
org-qualified URL reads, so the two forms are equivalent today. The
broad one differs only on cases you want to hear about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants