Skip to content

docs(mmq): conclude llama.cpp#29953 -- the head is correct, measured, and a guard for upstream - #454

Merged
glennneuber merged 1 commit into
mainfrom
docs/mmq-29953-conclusion
Oct 5, 2026
Merged

glennneuber merged 1 commit into
mainfrom
docs/mmq-29953-conclusion

Conversation

@glennneuber

Copy link
Copy Markdown

Summary

Concludes the llama.cpp#29953 work. #29953's head, 3070d927f, is correct: it picks J_best once on the host, launches exactly that tile, and pads src1 by that config's padded y tile and ids_dst by J_best - 1. Nothing about its correctness is owed upstream.

Measured on the exact head, not an equivalent. A new head29953 variant swaps in the head's mmq.cu and mmq.cuh, the only two files its commit touches, after checking that the checkout is the head's parent; the stock build is byte-identical to 3070d927f.

head 3070d927f the published #29953, same harness
sm_120, 13 cases x stock, three guard-page modes and exact-size memcheck 190/190 pass, no abort, 0 memcheck errors fails in every mode it ran in: ids16 aborts every guarded run; memcheck reports 12,885 errors on one113 and 298 on ids16
sm_75, ids16 and dense321 x stock and two guard-page modes 18/18 aborts every run that can fault

With hclsys's GB10 re-test (ggml-cuda is byte-identical between the tree they tested and the head) and the ROCm host's gfx1151 run (#449), the head is measured clean on four architectures.

What upstream lacks is a guard. The head computes the y tile's size twice, in mmq_get_nbytes_shared() and in its new padding line, and those two drifting apart is what this bug was. tasks/mmq-29953-y-tile-guard.patch (+70/-9 on the head) routes both through one helper and checks it against the kernel's load loop at compile time, for all ten config tables. Verified with nvcc: the patch compiles; three mutations fail by the guard and only the guard, one of them naming a single table; it costs about 3 s per build. This is material for the maintainer to weigh. Nothing is posted to ggml-org.

Changes

  • docs/maxusai/upstream-mmq-29953-material.md: a Conclusion section, "The exact head, measured", "A guard upstream could take", and "For the fork" rewritten with the decision and a retirement gate. Corrections below.
  • docs/maxusai/tasks/:
    • mmq-rules-check.cu models the head's rule and carries its own get_J_max(), which #29953 deletes.
    • mmq-variant.py gains head29953, mmq-rules-gpu.sh gains BUILD_KINDS, and mmq-rules-table.py gains --summary.
    • New: the run scripts, the guard patch, its verification and timing scripts, and their outputs.
  • llama/compat/README.md, 903's entry: nvcc did build the guard. At a pin that contains #29953, retire 903's padding and keep a guard.

Nothing outside docs/maxusai/ and that README. No code or build change.

Corrections to the record

  • A mislabeled row. The ten-architecture table said "#29953 at 3070d927f: short in 3,520,756 shapes, worst 12 blocks". The checker modelled the published rule, 5bd8b0013, so read literally the row said upstream's current fix is short, which is false. Relabeled, and the head's rule added: covered.
  • "Where it stands" still described 5bd8b0013 as the head, and "For the fork" still said gfx1151 was waiting on Help wanted (ROCm): does hipcc agree that compat 903 is short on gfx1151 for q2_K/q3_K? #449 and the 903 re-cut was pending. Both were done.
  • The amendment section still offered the retracted NVFP4 src1_scale change as live. Marked retracted.
  • A section date mixed the local date with a UTC time: the head was pushed 2026-10-04 18:54Z.
  • The compat README said nvcc had not built 903's guard. It had, before refactor(compat): Pad through one helper in 903 and guard it at compile time #451 merged.

For the fork

Keep compat 903, as merged in #451, until the pin includes #29953. Not a backport: b11081 predates the prec_src1 refactor, so a backport would be an adaptation rather than upstream's code; #29953 is not merged; and 903 is measured and guarded. At that pin move, retire 903's padding by the gate in the doc (the series applies, the checker is covered, and the ids16/dense321 guard runs pass with a control), and keep a guard-only re-cut unless upstream takes one.

Test plan

  • sm_120 (RTX PRO 6000 Blackwell, CUDA 12.8): the exact head passes 190/190 across 13 cases and five modes; the published #29953 fails in every mode in the same harness (tasks/mmq-successor-results/matrix-head29953-sm120.md).
  • sm_75 (RTX 2080 Ti): the exact head passes 18/18 on ids16 and dense321; the control aborts every run that can fault (matrix-head29953-sm75.md).
  • No test process left on either GPU after the runs.
  • CPU checker with the head's rule: covered on ten architectures, 9,431,816 MoE and 203,200 dense shapes. Built against dd266785c and against 3070d927f it prints byte-identical output, the same as the build that used the library's get_J_max().
  • Guard patch on 3070d927f: six nvcc builds as intended (verify-29953-guard.txt); cost a median of 11.1 s against 8.2 s over five interleaved repeats (guard-timing.txt).
  • check_source_paths.py --changed-since origin/main (every referenced file resolves), name scan.
  • Not run: hipcc and MUSA builds of the guard patch. There is no ROCm or MUSA toolchain on this host.

ai-server/mlx-cuda

🤖 Generated with Claude Code

… and a guard for upstream

#29953's head (3070d927f) pads src1 by the launched config's padded y tile and
ids_dst by J_best - 1. Measured on the exact head, not an equivalent -- the
head29953 variant swaps in its two files and checks they are the head's:

- sm_120: 190/190 runs pass over 13 cases x stock, three guard-page modes and
  exact-size memcheck, no abort, 0 memcheck errors. In the same harness the
  published #29953 fails in every mode it ran in: ids16 aborts every guarded
  run, and memcheck reports 12,885 errors on one113 and 298 on ids16.
- sm_75: ids16 and dense321, 18/18; the control aborts every run that can fault.
- hclsys's GB10 re-test of 20c31408 is the head's result: ggml-cuda is
  byte-identical between the two.

The CPU checker now models the head's rule, covered on all ten architectures by
construction, and carries its own get_J_max(), which #29953 deletes: built
against dd266785c and against 3070d927f it prints byte-identical output, the
same as with the library's copy.

What upstream lacks is a guard: the head computes the y tile's size twice. A
patch on the head (+70/-9) routes both through one helper and checks it against
the load loop at compile time for all ten config tables. Six nvcc builds behave
as intended: the patch compiles; three mutations fail by the guard and only the
guard, one of them naming a single table; compiled out, it compiles. It costs
about 3 s per build (median of five).

Record fixes: the ten-architecture table labelled the published row "#29953 at
3070d927f", which read as the head being short in 3.5M shapes; stale "where it
stands" and "for the fork" sections; the retracted y-scale change still offered
as live in the amendment section; a section dated with the local date and a UTC
time; and in the compat README, "nvcc has not built it yet" and re-cut advice
for a pin that contains #29953.

The fork keeps 903 until its pin includes #29953, then retires the padding by
a written gate and keeps a guard.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber
glennneuber merged commit 803e979 into main Oct 5, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant