Skip to content

Alignment machinery design: production-mode reconciliation with displacement disclosure (draft for referee)#217

Draft
MaxGhenis wants to merge 14 commits into
masterfrom
sol/alignment-design
Draft

Alignment machinery design: production-mode reconciliation with displacement disclosure (draft for referee)#217
MaxGhenis wants to merge 14 commits into
masterfrom
sol/alignment-design

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Summary

  • designs the missing producer for a Trustees-aligned after projection around the existing build_alignment_displacement accounting function
  • keeps the immutable unaligned projection as the scientific and gate-scored object, with alignment confined to a labeled report-only production replay
  • adjudicates deterministic threshold-prefix selection for discrete weighted events and a disclosed ratio only for continuous covered wages
  • binds a vintage-pinned 2026 Trustees Alternative II target manifest while leaving unavailable granular targets, horizon extension, and other unresolved inputs explicit
  • requires per-margin, per-year canonical path comparison and incremental person/unit/represented-weight displacement disclosure

Review posture

This is a docs-only design for adversarial review. It changes no code, gate, floor, threshold, or run artifact. It does not resolve the separate re-drawn-T* null-with-reason surface.

The design incorporates independent contract and citation reviews, including exact-key accounting across roster-changing paths, homogeneous-unit calls to the existing function, seam-time threshold-prefix resolution floors, fresh replay state, chained earnings updates, stable synthetic IDs/ordinals/RNG coupling, and joint-vector deferral for atomic multi-cell migration units.

Checks

  • git diff --check
  • quarto pandoc docs/design/alignment_machinery.md --from gfm --to html
  • external citations checked against public sources; local file/line anchors checked against the repository and merged amendment 3h

MaxGhenis and others added 8 commits July 15, 2026 17:27
Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
@vercel

vercel Bot commented Jul 15, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
social-security-model Ready Ready Preview, Comment Jul 15, 2026 11:27pm

Request Review

@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Adversarial design referee — alignment machinery (draft)

Scope: full doc + PR body, cross-checked against a completed primary-sourced research pass (Li & O'Donoghue 2014 taxonomy; Stephensen 2016 logit scaling; DYNASIM/MINT/POLISIM practices; determinism recipes; Trustees intermediates), against docs/design/m6_projection_engine.md (phase-5 rule, decision 9, family B, §2.8.9, amendments 3g/3h), and against every cited engine/harness pin at the PR head. External citations were verified against the primary sources themselves (PDFs downloaded and read; SSA 2026 pages loaded live).

Verified sound

  • Decision-9 fidelity. The decision-9 and phase-5 quotes are exact against the branch tree. §1.2's four invariants (gate_scored = false, alignment_applied = true, certification_transfer = false, verdict only on the unaligned series), the §5.3 artifact-schema failure on a missing label/sidecar, the §7.1 no-network rule, and §9's demanded unaligned-gate-isolation check pin "certification stays unaligned forever" as law, with the storage/certification distinction handled correctly. Failure of the aligned producer explicitly cannot convert into an M6 gate failure (§5.3).
  • The existing function. Every §1.3 claim about build_alignment_displacement is code-exact (src/populace_dynamics/harness/m6_reporting.py:321-415): value-column parameterization, exact-key merge including year, one-sided rejection, per-year field maxima, frame-global n_intervened_rows, not_computed with alignment_layer_output_not_collected, reasons at :37-38. The doc's refusal to relabel n_intervened_rows as a per-year person count, and the §5.1 homogeneous-measure rule (no cross-unit rowwise max), are the correct honest readings. The moved call site (m6_runner.py:1120-1128, actual call at 1124) is correctly disclosed.
  • Engine pins. All verified at the PR head: eight-step order (loop.py:268-321), length-dependent newborn ordinals (loop.py:323-331) vs ordinal-keyed person streams (rng.py:158-169), scheduled-entry pre-registration (loop.py:192-237), parity-internal positive-only fertility (marital.py:288-352), continuous-only earnings interface and even/odd chain behavior (steps.py:169-287), the three-field even-year read (forward_earnings.py:1046-1070), allocator + length-dependent child sex draw (steps.py:381-437), the 2022 horizon pin (panel_builders.py:120-132), and the assembly-cache freshness argument (assembly.py:195-250). One subtle asymmetry checks out under scrutiny: mortality's per-person (p, u) is reconstructible non-invasively (pure model.probabilities + re-derivable keyed person streams, steps.py:113-134), while fertility's uniforms are positionally consumed inside a parity-evolving loop — so flagging fertility/covered-work as seam-blocked while leaving mortality mechanism-ready is right, not an oversight.
  • Citations: 15/15 verified, zero fabrications. Urban 2024 p. 2 (five processes aligned to OCACT), Johnson et al. 2017 appendix p. 47 (selection criterion to Trustees sex-by-age employment; wage alignment), Favreault & Smith 2004 pp. 7-8 (five fertility groups, 12 age-sex mortality groups — and the doc correctly files this under "DYNASIM3 lineage"), Smith & Favreault 2013 pp. 7-8 + fn 24 (donor swap, unequal-weight overshoot) and p. 2 fn 7, Smith et al. 2010 VII-2/VII-4 and I-5–I-6 (re-estimation recommendation), Toder et al. 2000 pp. 52-53, Li & O'Donoghue 2014 §§2.3/5.3-5.8, Harding et al. 2010 (alignment off for the module under test), McKay 2003 pp. 2/5-6 with the masking warning verbatim on p. 6, and Cumpston 2010 §2 — the DYNACAN mortality-coefficient masking sentence is confirmed in the primary source. SSA 2026 bindings verified live: release date June 9, 2026; Alternative II "best estimate ... as of ... February 2026" verbatim; IV.B4 title + fn a ("paid at some time during the year"); V.A1 = TFR + age-sex-adjusted death rates only; V.B1 = percentage change only, nominal and real, no level base; VI.G1 = AWI level + taxable payroll (not the covered-wage level), so the doc's V.B1/AWI/payroll distinctness claims are exact; V.B2 average-week CPS concepts confirmed via the V.C methods page; the XLSX and both PDFs resolve with the pinned filenames. The doc asserts no numeric Trustees values anywhere — prudent, since the 2026 intermediate ultimate TFR moved from the 2025 report's value, which would have invalidated any copied number.
  • Honesty of the mechanism's limits. Non-exactness under unequal contributions is admitted (§4.2), floored (§6.1, including the non-prefix-subset disclaimer and the anti-gaming pooling rule), and evaluated against the floor (§6.2, non-vacuous because the evaluator recomputes from the materialized panel). Adds/removals never netted; mother/child roles separated; the 3h summary is faithful to §2.8.2h at the pinned blob (L1183-1264 at 0e27be2 verified, including roster_absent_births staying upstream).
  • Mechanics. Docs-only (one new 763-line file), git diff --check clean, gfm render passes, doc-bound tests bind only m6_projection_engine.md + the gate YAML (untouched), CI green.

Findings

SHOULD-FIX 1 — the strongest method in the literature is never adjudicated (logit scaling, Stephensen 2016). §2.1 adjudicates only the weak member of the probability-adjustment family — annual linear probability-ratio scaling (the DYNASIM3 precedent) — and rejects it on three grounds: expected-not-realized margins, clipping at one, no unique displaced set. Logit scaling (IJM 9(3) 2016: multiply odds, not probabilities, by a common factor solved against the weighted margin) is exact on weighted expected margins, KL-minimal, deterministic in its factor, has a weighted variant, and cannot clip — it defeats the second objection entirely, and with a common-uniform reuse rule it also yields a unique deterministic flip set, weakening the third. The doc's grep-clean absence of "Stephensen"/"logit" means the selection-vs-scaling adjudication knocks down a strawman. The chosen mechanism plausibly still wins — logit scaling controls the expected weighted count, not the realized one, and realizing a count from scaled probabilities requires a selection step anyway — but that decisive ground must be argued against the strong method, not the weak one. Bound up with this: the registered flip ordering (p − u, additive draw-distance) is a decided-but-unadjudicated methodological choice. The odds-scaling family induces a different marginal-flip order (smallest odds ratio [u/(1−u)]/[p/(1−p)] flips first). For mortality the two coincide within the doc's strata (single year of age × sex under an age-sex mortality model makes within-stratum p constant), but for fertility (parity/cohort heterogeneity within mother-age cells) and especially the year-total covered-work margin (highly heterogeneous p in one stratum), additive ordering systematically prefers flipping low-p near-misses in absolute units where odds ordering prefers relative near-misses — a real compositional difference in who gets displaced. Fix: add the Stephensen adjudication to §2/§2.1 (realized-count control, displaced-set disclosure, and raw-draw coupling as the decisive grounds), and either register the additive ordering with that rationale or list the ordering rule as a ninth open decision.

SHOULD-FIX 2 — the V.A2 claim "no unique gross addition/removal decomposition" is false against the pinned table. Verified on the live 2026 single-year table: projected years publish the full column set for both categories — LPR inflow / outflow / adjustments-of-status / net, and temporary-or-unlawfully-present inflow / outflow / adjustments-of-status / net (2030 intermediate: 600 / 263 / +450 / 788 and 970 / 329 / −450 / 191; total net 978, arithmetic exact), with footnote e making AoS a pure cross-category transfer. Gross additions and removals are uniquely decomposed at year × legal-status granularity; what the table genuinely lacks is age-sex cells, the Social-Security-area-to-roster bridge, and within-category flow semantics (fn c mixes citizen emigration into LPR outflow; fn b's residual method conflates temporary-stock outflow components). The error is conservative in direction (it defers a margin the table partially binds), but a design whose §3.3 discipline is "definitions must not drift" cannot misdescribe its own pinned source, and decision 2's premise ("implement V.A2 net change without an arbitrary split") overstates the gap — the split is published; the age-sex disaggregation and bridge are what is missing. Rewrite the §3.2 row-1 binding and the decision-2 text accordingly.

SHOULD-FIX 3 — every m6-doc line anchor is stale against the merge target. The PR branch base (75d30dd) predates merged #216, and master's m6_projection_engine.md is +327/−8 lines against the branch copy. Verified drift on master: the phase-5 rule cited as :1534-1536 sits at 1838-1840; decision 9 cited as :2661-2663 sits at 2977-2979; everything at or after §2.8.2h's insertion shifts ~300 lines, including :1221-1300 and :2347-2362. The doc itself demonstrates the correct discipline twice — the 3h citation is pinned to a blob URL at 0e27be2, and the moved m6_runner.py call site is explicitly disclosed — but does not apply it to its own m6-doc anchors, and the PR body's "anchors checked against the repository and merged amendment 3h" holds only for the pre-3h branch tree. Fix: re-anchor the m6-doc references to post-3h master line numbers, or pin them to a commit SHA the way the 3h citation already is.

SHOULD-FIX 4 — seam ownership and the concrete gate-isolation mechanism are unnamed (a ninth open decision in substance). §4.1 says the wrapper "inserts reconciliation hooks at the named seams" of a loop that today has no seam surface (loop.py:268-321 invokes the eight modules back-to-back), and §4.4's "the aligned loop must accept that frozen ordinal registry" / "child materialization must accept the reserved identity" require engine-code changes. Two architectures satisfy the text — (a) the certified loop gains optional hook/registry parameters that the gate path never passes, or (b) a parallel production loop reimplements orchestration around the same certified module callables — and they carry different isolation-enforcement stories and different §9 "does not edit the eight-step engine loop" accounting. The doc names the right tests (zero-flip identity, add/suppress isolation, unaligned-gate-isolation) but not the structural mechanism: nothing pins, for example, that the gate runner's engine invocation signature can never accept an alignment specification, or that certification CI must prove byte-identical gate artifacts with the alignment machinery present versus absent. Name the seam-ownership choice (or add it to §8) and state the enforcement in structural terms, not only as test obligations.

NOTE 1 — publish the implied per-cell alignment factor as a first-class field. §5.2 publishes target, raw, and aligned totals plus raw residual per (margin, year, stratum), so the misspecification diagnostic is derivable — but the literature the doc itself cites states the diagnostic in factor terms (Toder's large-factor warning; DYNASIM's factors-as-diagnostic practice), and §6.3's corridors use shares and absolute displacements only. An explicit implied_alignment_factor = target / raw (with a defined zero-raw behavior) per applied cell would make cross-model and cross-vintage comparison direct.

NOTE 2 — key tie values on the unit, not on frame position. §4.4 derives tie streams from (draw index, spec hash, year, margin, stratum, revision) but does not say the per-unit tie value is keyed on alignment_unit_id. If tie values are assigned positionally over the sorted candidate frame, determinism for a fixed input holds, but tie values are not invariant under candidate-set perturbation — mildly at odds with the add/suppress isolation posture. One sentence pinning counter-based per-unit tie derivation closes it.

NOTE 3 — the emigration seam placement is decided without rationale while decision 2 is open. §3.2 fixes "a separately registered emigration branch runs after addition and before mortality" although the entire migration mechanism (units, donor pool, joint vector objective) is deferred to decision 2. Either state the ordering rationale (single-transition-per-year discreteness would do) or fold the placement into decision 2 so the future ratification is not half-pre-committed.

Verdict

REVISE. No blocking findings: the certification law is faithfully pinned, the existing accounting function is characterized code-exactly, the determinism recipes match the strongest published practice, all fifteen external citations verify against primary sources, and the doc's honesty about its own residuals (floors, non-exactness, unavailable margins) is exemplary. Revision is required because the central mechanism adjudication omits the strongest method in the literature (SHOULD-FIX 1), one pinned-source factual claim is wrong (SHOULD-FIX 2), the m6-doc anchors break on the merge target (SHOULD-FIX 3), and the gate-isolation enforcement needs a named structural mechanism (SHOULD-FIX 4). 0 BLOCKING / 4 SHOULD-FIX / 3 NOTE.

Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
Publish the implied per-cell alignment factor and key tie values on stable alignment units.

Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
Bind the published gross flow columns, preserve their unresolved roster bridge, and explain the emigration seam ordering.

Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
MaxGhenis and others added 2 commits July 15, 2026 19:22
Record the two orchestration architectures and require a closed unaligned gate facade with structural and byte-identity enforcement.

Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
Update every mutable M6 document line pin to the merged post-3h master tree while retaining the immutable amendment blob link.

Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
Distinguish adjustment-of-status state changes from roster additions and removals throughout the migration design.

Co-Authored-By: Codex gpt-5.6-sol <noreply@openai.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Fixes applied

Applied the alignment-design referee findings from comment 4986136518.

Finding → commit map

  • Base refresh → e090d59: merged current post-3h origin/master before deriving any M6 document anchors.
  • SHOULD-FIX 1 + NOTES 1–2 → bd836af: adjudicates Stephensen 2016 logit scaling on its strongest terms (exact expected weighted margins, KL minimum, no clipping, common-uniform determinism), explains why realized weighted-count control still requires selection, names and defends additive p - u versus odds-distance ordering, publishes the implied factor with zero-raw behavior, and keys tie values on alignment_unit_id.
  • SHOULD-FIX 2 + NOTE 3 → 9fd3485, clarified by 0c88c4f: binds V.A2's published LPR and temporary-or-unlawfully-present inflow/outflow/adjustment/net columns; leaves only the age-sex, roster, and within-category semantic bridges open; removes the repeated arbitrary-gross-split premise; and explains the emigration-before-mortality seam.
  • SHOULD-FIX 3 → 5a09426: re-anchors all six mutable m6_projection_engine.md references to the merged post-3h tree (1838-1840, 2100-2103, 2979-2981, 2665-2680, 1525-1604, 1009-1152). The immutable 0e27be2#L1183-L1264 amendment pin remains unchanged and exact.
  • SHOULD-FIX 4 → 0adc6ec: names hooks-in-loop versus a parallel production loop as open decision 9 and requires structural gate isolation through a closed unaligned-only facade, import/dependency and signature checks, alignment-bearing metadata rejection, and byte-identical present-versus-import-blocked gate artifacts.

Validation

  • PR diff against current origin/master: only docs/design/alignment_machinery.md.
  • git diff --check: pass.
  • Pandoc GFM render: pass.
  • PYTHONPATH=src /opt/homebrew/bin/python3.12 -m pytest -q tests/test_gate_m6_floors.py: 23 passed.
  • Independent final audit: all 4 SHOULD-FIXes and 3 NOTEs resolved; no stale pin, numbering collision, factual contradiction, or Markdown defect remains.

No gates.yaml or runs/ content was touched. PR #217 remains draft for the verification round.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant