Repository navigation
docs(mmq): conclude llama.cpp#29953 -- the head is correct, measured, and a guard for upstream - #454
Merged
Merged
Conversation
… and a guard for upstream #29953's head (3070d927f) pads src1 by the launched config's padded y tile and ids_dst by J_best - 1. Measured on the exact head, not an equivalent -- the head29953 variant swaps in its two files and checks they are the head's: - sm_120: 190/190 runs pass over 13 cases x stock, three guard-page modes and exact-size memcheck, no abort, 0 memcheck errors. In the same harness the published #29953 fails in every mode it ran in: ids16 aborts every guarded run, and memcheck reports 12,885 errors on one113 and 298 on ids16. - sm_75: ids16 and dense321, 18/18; the control aborts every run that can fault. - hclsys's GB10 re-test of 20c31408 is the head's result: ggml-cuda is byte-identical between the two. The CPU checker now models the head's rule, covered on all ten architectures by construction, and carries its own get_J_max(), which #29953 deletes: built against dd266785c and against 3070d927f it prints byte-identical output, the same as with the library's copy. What upstream lacks is a guard: the head computes the y tile's size twice. A patch on the head (+70/-9) routes both through one helper and checks it against the load loop at compile time for all ten config tables. Six nvcc builds behave as intended: the patch compiles; three mutations fail by the guard and only the guard, one of them naming a single table; compiled out, it compiles. It costs about 3 s per build (median of five). Record fixes: the ten-architecture table labelled the published row "#29953 at 3070d927f", which read as the head being short in 3.5M shapes; stale "where it stands" and "for the fork" sections; the retracted y-scale change still offered as live in the amendment section; a section dated with the local date and a UTC time; and in the compat README, "nvcc has not built it yet" and re-cut advice for a pin that contains #29953. The fork keeps 903 until its pin includes #29953, then retires the padding by a written gate and keeps a guard. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Concludes the llama.cpp#29953 work. #29953's head,
3070d927f, is correct: it picksJ_bestonce on the host, launches exactly that tile, and pads src1 by that config's padded y tile andids_dstbyJ_best - 1. Nothing about its correctness is owed upstream.Measured on the exact head, not an equivalent. A new
head29953variant swaps in the head'smmq.cuandmmq.cuh, the only two files its commit touches, after checking that the checkout is the head's parent; the stock build is byte-identical to3070d927f.3070d927fids16aborts every guarded run; memcheck reports 12,885 errors onone113and 298 onids16ids16anddense321x stock and two guard-page modesWith hclsys's GB10 re-test (
ggml-cudais byte-identical between the tree they tested and the head) and the ROCm host's gfx1151 run (#449), the head is measured clean on four architectures.What upstream lacks is a guard. The head computes the y tile's size twice, in
mmq_get_nbytes_shared()and in its new padding line, and those two drifting apart is what this bug was.tasks/mmq-29953-y-tile-guard.patch(+70/-9 on the head) routes both through one helper and checks it against the kernel's load loop at compile time, for all ten config tables. Verified with nvcc: the patch compiles; three mutations fail by the guard and only the guard, one of them naming a single table; it costs about 3 s per build. This is material for the maintainer to weigh. Nothing is posted to ggml-org.Changes
docs/maxusai/upstream-mmq-29953-material.md: a Conclusion section, "The exact head, measured", "A guard upstream could take", and "For the fork" rewritten with the decision and a retirement gate. Corrections below.docs/maxusai/tasks/:mmq-rules-check.cumodels the head's rule and carries its ownget_J_max(), which #29953 deletes.mmq-variant.pygainshead29953,mmq-rules-gpu.shgainsBUILD_KINDS, andmmq-rules-table.pygains--summary.llama/compat/README.md, 903's entry: nvcc did build the guard. At a pin that contains #29953, retire 903's padding and keep a guard.Nothing outside
docs/maxusai/and that README. No code or build change.Corrections to the record
3070d927f: short in 3,520,756 shapes, worst 12 blocks". The checker modelled the published rule,5bd8b0013, so read literally the row said upstream's current fix is short, which is false. Relabeled, and the head's rule added: covered.5bd8b0013as the head, and "For the fork" still said gfx1151 was waiting on Help wanted (ROCm): does hipcc agree that compat 903 is short on gfx1151 for q2_K/q3_K? #449 and the 903 re-cut was pending. Both were done.src1_scalechange as live. Marked retracted.For the fork
Keep compat 903, as merged in #451, until the pin includes #29953. Not a backport: b11081 predates the
prec_src1refactor, so a backport would be an adaptation rather than upstream's code; #29953 is not merged; and 903 is measured and guarded. At that pin move, retire 903's padding by the gate in the doc (the series applies, the checker is covered, and theids16/dense321guard runs pass with a control), and keep a guard-only re-cut unless upstream takes one.Test plan
tasks/mmq-successor-results/matrix-head29953-sm120.md).ids16anddense321; the control aborts every run that can fault (matrix-head29953-sm75.md).dd266785cand against3070d927fit prints byte-identical output, the same as the build that used the library'sget_J_max().3070d927f: six nvcc builds as intended (verify-29953-guard.txt); cost a median of 11.1 s against 8.2 s over five interleaved repeats (guard-timing.txt).check_source_paths.py --changed-since origin/main(every referenced file resolves), name scan.ai-server/mlx-cuda🤖 Generated with Claude Code