Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
bebc258
docs(mmq): material for a successor to llama.cpp#27044 — #29941 is no…
glennneuber Oct 4, 2026
57f9fbc
compat: 903 pads every tile-read MMQ buffer for the widest tile
glennneuber Oct 4, 2026
ea754c4
docs(mmq): real-model reproduction of the over-read on qwen3.6:35b-a3b
glennneuber Oct 4, 2026
2edbc84
docs(mmq): #29941 covers the large-image src1 read; gaps are ids_dst …
glennneuber Oct 4, 2026
0387b5e
docs(mmq): llama.cpp#29953 reproduced — two reads it does not cover
glennneuber Oct 4, 2026
5656b15
docs(mmq): offer the mmq_get_nbytes_y_tile dedup as an untested alter…
glennneuber Oct 4, 2026
4115968
docs(mmq): the table generator dropped unnamed variant columns silently
glennneuber Oct 4, 2026
18476d8
compat: 903 pads the widest PADDED tile, not the widest tile
glennneuber Oct 4, 2026
5de3e4c
compat: the 903 amendment costs 864 bytes, on one arch/type combination
glennneuber Oct 4, 2026
f7e8e2e
docs(mmq): the seven-rule matrix, and a correction on which cases can…
glennneuber Oct 4, 2026
bfc3fe4
docs(mmq): the J=112 cells get a rate and a control; matrix complete …
glennneuber Oct 4, 2026
647b7f5
docs(mmq): flag the gfx1151 column as modelled, and link the ROCm ask
glennneuber Oct 4, 2026
ffd6dd9
docs(mmq): #29953 now carries both amendments; only the NVFP4 y scale…
glennneuber Oct 4, 2026
1ecdb17
docs(mmq): retract the NVFP4 y-scale claim; add the sm_75 dense result
glennneuber Oct 4, 2026
fa5ad16
compat: gfx1151 measured — amendment confirmed, four corrections from…
glennneuber Oct 4, 2026
6fcc318
docs(mmq): ten architectures — the widest-tile rule is short on CDNA3…
glennneuber Oct 4, 2026
05736ba
docs(mmq): gfx1151 measurements for #449 -- tools, results, and a HIP…
Oct 4, 2026
07b6272
docs(mmq): a per-config padding check and a compile-time guard; CDNA3…
Oct 4, 2026
2cfc99f
compat: 903 pads through one helper, guarded at compile time
Oct 5, 2026
0bf3e8e
Merge pull request #450 from MaxusAI/docs/mmq-rocm-gfx1151
glennneuber Oct 5, 2026
4dbe35c
Merge pull request #451 from MaxusAI/compat/903-padding-guard
glennneuber Oct 5, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 5 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,7 +102,7 @@ decision and its measurements live (`docs/maxusai/`). The
| **nemotron-3 vision** | fixed 512×512 canvas — **256 tokens per image**, whatever the aspect ratio. Upstream's MLX `nemotron_h` has been native-aspect since v0.34.3 | native-aspect dynamic resolution, **256–3,328 tokens**, position embeddings interpolated to the patch grid in-graph | patch `002`, ADR 0001 |
| **gemma4 image budget** | default limits **70–1,120 tokens** (40–280 before b10864); an under-budget image keeps its natural rounded grid and is letterbox-padded | every image scaled to *fill* the requested budget and snapped to gemma4's supported ladder (70/140/280/560/1120), never padded — off-ladder grids measurably break `box_2d` vertical grounding. The budget is a per-request option (`image_min_tokens`/`image_max_tokens`, defaults 70/1120) and the scheduler reloads when the resolved flags change. Upstream has since adopted the same default limits; the fill is still fork-only | patch `004`, ADR 0003/0008/0016 |
| **qwen2.5-vl on CUDA** | f16 vision matmuls accumulate in fp16; on some ordinary images a few elements of millions reach `inf` at `v.blk.31.ffn_down` and the caption collapses into one repeated glyph | fp32 accumulation forced for every Qwen-VL runner (`qwen2vl` and `qwen25vl`, one family under two converter spellings), keyed on the GGUF architecture. Offered upstream as [ollama#18070](https://github.com/ollama/ollama/pull/18070), still open | `llm/llama_server.go` |
| **MoE + MMQ on CUDA** | ids-path tail padding sized from `ne11`; under broadcast `ne11 == 1`, so the buffer gets no padding and the kernel overruns by up to a 512-row tile | padding sized from the flattened row count. Reported as [llama.cpp#27044](https://github.com/ggml-org/llama.cpp/issues/27044), still open | patch `903` |
| **MoE + MMQ on CUDA** | MMQ reads its buffers in whole tiles, but pads them from the batch (`get_J_max(ne12)` since [#29941](https://github.com/ggml-org/llama.cpp/pull/29941), `get_J_max(ne11)` before), which rounds down while the launched tile rounds up. Before #29941, a broadcast `ne11 == 1` gave no padding at all: an illegal memory access. Since, batches under 128 tokens can still read past the buffer | every tile-read buffer (`src1`, `ids_dst`, NVFP4 scales) padded for the widest tile that has a config, as before llama.cpp #24127. Offered as [llama.cpp#27044](https://github.com/ggml-org/llama.cpp/pull/27044), which upstream closed for #29941 | patch `903` |
| **gemma4 flash attention on CUDA** | `b11081` retunes the MMA configs and tile sizes for head dims 256/512 (`ce8caa6e6`). Its accuracy is unchanged in llama.cpp's `test-backend-ops`, but the rounding moves. With ollama's Q4_K_M, gemma4:26b leaves 6 of 27 think-on cases in loops that never end, against 1 without it; with ggml-org's Q4_0, a case loops the other way | the device half of `ce8caa6e6` is reverted to `b10969`'s tiling, which keeps the numerics production was gated on; the host half, the decode selection, stays upstream's. The revert shows no speed cost: in the paired runs on CUDA, gemma4:31b decoded 57 tok/s with either tiling, and gemma4:26b 176 tok/s against the new tiling's 137. On gfx1151 the patch changes no kernel | patch `908`, the [v0.34.4 fold record](docs/maxusai/tasks/upstream-sync-0.34.4.md) |

**Structured output and generation control**
Expand Down Expand Up @@ -223,9 +223,10 @@ Use GGML/llama-server for anything where throughput or comparability matters.
Every row in the tables above is a delta we would rather not have. Each is
offered upstream where it is upstream's to take, and deleted from here once it
lands there. The Qwen-VL accumulation gate
([ollama#18070](https://github.com/ollama/ollama/pull/18070)) and the MMQ
padding fix ([llama.cpp#27044](https://github.com/ggml-org/llama.cpp/issues/27044))
are filed and pending. The gemma4 tiling revert (`908`) is not upstream's to take:
([ollama#18070](https://github.com/ollama/ollama/pull/18070)) is filed and pending.
Upstream closed the MMQ padding fix ([llama.cpp#27044](https://github.com/ggml-org/llama.cpp/pull/27044))
for its narrower [#29941](https://github.com/ggml-org/llama.cpp/pull/29941), which still reads past the
buffer below 128 tokens. `903` now carries the widest-tile padding that a successor would offer. The gemma4 tiling revert (`908`) is not upstream's to take:
`ce8caa6e6` loses no precision in llama.cpp's `test-backend-ops`, so the revert is the fork's own
choice to keep production's gated numerics.

Expand Down
9 changes: 9 additions & 0 deletions docs/maxusai/qwen35moe-mmq-investigation.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,15 @@ flattened row count `ne12 * n_expert_used`. In the MoE broadcast case `ne11 == 1
`ggml_cuda_mmq_get_J_max()` returns 0, so the quantised buffer gets no tail padding while the
kernel overruns it by up to a 512-row tile. Fixed in `llama/compat/903-fix-mmq-ids-padding.patch`.

> **Update 2026-10-05.** 903 no longer carries the `ne12 * n_expert_used` line shown below.
> - The tile is at most 128 columns wide, not 512.
> - The line was short with one expert per token.
> - Upstream's own fix, #29941 (`ne12`), is short below 128 tokens.
>
> 903 now pads every buffer that MMQ reads in whole tiles for the widest tile. See
> [upstream-mmq-successor-material.md](upstream-mmq-successor-material.md). This document is the
> record of how the original fault was found, and stays as it was.

## Symptom

```
Expand Down
2 changes: 1 addition & 1 deletion docs/maxusai/retirement-register.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ _Nothing pending. `x/structured` moved to "Already retired" below on 2026-09-17.
| **think + format, two-pass** (ADR 0004's routes-layer flow, behind `OLLAMA_FORMAT_TWO_PASS=1` since v0.34.4): the marker-stop flow (pass one stops at the think-close tag and continues textually), pass-one metrics reconstruction, the second pass pinned to pass one's truncation window, and the lower-layer hooks the v0.34.4 fold restored | ADR 0002, 0004, 0010, [0045](adr/0045-think-format-single-pass-by-default-two-pass-in-production.md) | a single pass drafts its thinking on MLX without the qwen3.5-family retention, so production can drop the switch (ADR 0045, proposed) | `leak-repro5.sh` flat on the qwen3.5 family; the drafting probe with the single pass matching P0; the MLX think-on protocol with the single pass matching two-pass's loop counts, on Metal; then the two-pass route tests go with the flow | 2026-09-28: upstream's flow is a single pass since v0.34.4, and ours is production's flow (ADR 0045) |
| **gemma4 vision on MLX through our path**, upstream's tower excluded | ADR 0021, D1-A; the D1-B spike | upstream's tower gains a per-request image budget, or the budget seam is re-grafted onto it (D1-B) | vision goldens 12b/26b/31b; the MLX think-off campaign | 1,523/−2,271 lines against upstream; D1-B not started |
| **fp32 cuBLAS accumulation for the shared Qwen-VL clip graph** (arch `qwen2vl` and `qwen25vl` — one family, two converter spellings; widened 2026-09-19) | #214, ollama#18070 | #18070 or a per-op `PREC_F32` fix lands upstream | `poison_probe`; the fp16 canary | #18070 open |
| **`903` MMQ ids-path padding** | llama.cpp#27044 | fixed upstream | the MoE + MMQ functional gate | #27044 open |
| **`903` MMQ tile padding**: every buffer MMQ reads in whole tiles padded for the widest tile (since 2026-10-05; before, `get_J_max(ne12*n_expert_used)`) | llama.cpp#27044 (closed 2026-10-04 for #29941); [successor material](upstream-mmq-successor-material.md) | upstream pads for every tile it launches, src1, `ids_dst` and NVFP4 scales included, in both branches | `tasks/mmq-successor-gpu.sh` on the new pin: its `master` build at 0 memcheck errors with the exact allocation in every case, and `tasks/mmq-padding-check.cu` showing the pin's own rule short in no shape; then the MoE + MMQ functional gate | 2026-10-05: upstream's #29941 (`get_J_max(ne12)`, `dd266785c`) is short below 128 tokens, and master aborts on `test_mul_mat_id(q4_K, 256, 10, b, 576, 120, 1536)`. Switched to the widest-tile rule on the maintainer's word; at the first pin past `dd266785c`, re-cut the ids hunk against its `ne12` line |
| **`908` gemma4 FA tiling revert**: the device half of llama.cpp `ce8caa6e6` (MMA configs and tile sizes for head dims 256/512), back to b10969's | #375; the v0.34.4 fold record, "Gates 4–6 on CUDA" | a fold's CUDA GGUF loop-rate run on the pin's own tiling matches 908's. No upstream fix is due: `ce8caa6e6` loses no precision in `test-backend-ops` (2026-09-28) | the CUDA GGUF loop-rate run: gemma4:26b-a4b full think-on suite on the full ladder, the pin as shipped against the pin with 908, NOT CONVERGED counts (1 of 27 with 908, 6 without, at b11081); then the think-off cells of gemma4:31b, 26b and e4b against production (e2b moves with both halves) | carried from the v0.34.4 fold (2026-09-26) on the maintainer's word. Not an upstream bug: on 2026-09-28 the precision was unchanged, and on the public Q4_0 the loop runs the other way. Nothing is filed |
| **`801` clip node-stats meter** | #214 diagnostic | with the f32 gate above; it exists to localise that overflow | the fp16 canary under `OLLAMA_CLIP_NODE_STATS` | keep while the gate is carried |
| **`004` gemma4 budget fill** (`image_budget_fill`, `PAD_NONE`) | ADR 0008 | upstream fills to the ladder and stops letterbox-padding. Its default *limits* converged on ours (70/1120) in b10864 | `pinned_image_token_budget`, `token_ladder`, bbox conformance | limits converged; the fill is still ours |
Expand Down
22 changes: 22 additions & 0 deletions docs/maxusai/tasks/build-check.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
#!/usr/bin/env bash
# Build a GPU-free MMQ check against a llama.cpp checkout.
#
# build-check.sh <llama.cpp checkout> <source.cu> <output>
#
# The gencode list is load-bearing: ggml_cuda_highest_compiled_arch() reads __CUDA_ARCH_LIST__, and
# ggml_cuda_mmq_get_config() picks a different config table when the arch it is asked about was not compiled for.
# Without these flags every NVIDIA arch resolves to nvcc's default and the configs come out wrong (J capped at 64
# on sm_120 instead of 128), so the whole check silently models a machine that does not exist.
set -eu
SRC=$(cd "${1:?usage: $0 <llama.cpp checkout> <source.cu> <output>}" && pwd)
CU=$(cd "$(dirname "${2:?}")" && pwd)/$(basename "$2")
OUT=$3
CUDA_HOME=${CUDA_HOME:-/usr/local/cuda-12.8}
unset CPATH C_INCLUDE_PATH CPLUS_INCLUDE_PATH
cd "$SRC"
"$CUDA_HOME/bin/nvcc" -Wno-deprecated-gpu-targets -std=c++17 \
-gencode arch=compute_70,code=sm_70 -gencode arch=compute_75,code=sm_75 -gencode arch=compute_80,code=sm_80 -gencode arch=compute_90,code=sm_90 \
-gencode arch=compute_86,code=sm_86 \
-gencode arch=compute_89,code=sm_89 \
-gencode arch=compute_120,code=sm_120 \
-I. -Iggml/include -Iggml/src -Iggml/src/ggml-cuda "$CU" -o "$OUT" -lcublas -lcuda
22 changes: 22 additions & 0 deletions docs/maxusai/tasks/mmq-amend-29953-yscale.patch
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
diff --git a/ggml/src/ggml-cuda/mmq.cu b/ggml/src/ggml-cuda/mmq.cu
index 027a7e60..7f761ba4 100644
--- a/ggml/src/ggml-cuda/mmq.cu
+++ b/ggml/src/ggml-cuda/mmq.cu
@@ -236,7 +236,7 @@ void ggml_cuda_mul_mat_q(
ggml_cuda_pool_alloc<char> src1_q8_1(ctx.pool(), nbytes_src1_q8_1);
ggml_cuda_pool_alloc<float> src1_scale(ctx.pool());
if (src0->type == GGML_TYPE_NVFP4 && use_native_fp4) {
- src1_scale.alloc(ne13*ne12*ne11);
+ src1_scale.alloc(ne13*ne12*ne11 + J_best-1); // Needs to be padded for unconditional memory access.
}

{
@@ -304,7 +304,7 @@ void ggml_cuda_mul_mat_q(
ggml_cuda_pool_alloc<char> src1_q8_1(ctx.pool(), nbytes_src1_q8_1);
ggml_cuda_pool_alloc<float> src1_scale(ctx.pool());
if (src0->type == GGML_TYPE_NVFP4 && use_native_fp4) {
- src1_scale.alloc(ne12*n_expert_used);
+ src1_scale.alloc(ne12*n_expert_used + J_best-1); // Needs to be padded for unconditional memory access.
}

const int64_t ne11_flat = ne12*n_expert_used;
68 changes: 68 additions & 0 deletions docs/maxusai/tasks/mmq-amend-29953.patch
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
# Amendment to apply on top of llama.cpp#29953 (5bd8b0013). Two gaps that #29953 leaves open:
# 1. src1: the y tile load copies the whole padded shared-memory y tile from global memory,
# GGML_PAD(J*sizeof(block_q8_1_mmq), nthreads*sizeof(int)) bytes from the tile first column -- the same
# expression mmq_get_nbytes_shared() uses -- so J_best blocks is short whenever nthreads*4 does not divide
# J_best*144. Measured short in 1,608,722 of 4,717,824 shapes.
# 2. ids_dst (and, for native NVFP4, the y scales): a tile loads J entries with no bound, so the last expert's
# last tile reads up to J - 1 entries past ne12*n_expert_used. Short in all 4,717,824 shapes.
# Generated by mmq-variant.py (variant p29953fix); apply with git apply -p1 after #29953.
--- a/ggml/src/ggml-cuda/mmq.cu
+++ b/ggml/src/ggml-cuda/mmq.cu
@@ -187,7 +187,8 @@
const size_t y_block_size = use_native_fp4 ? sizeof(block_fp4_mmq) : sizeof(block_q8_1_mmq);
const size_t y_values_per_block = use_native_fp4 ? QK_FP4_MMQ : QK8_1_MMQ;

- int J_best = 0;
+ int J_best = 0;
+ size_t nbytes_pad_y = 0;
{
int64_t ncols_opt = ne11;
if (ids) {
@@ -216,8 +217,9 @@
const int ntiles_x = (ncols_opt + config.J - 1) / config.J;

if (ntiles_x < ntiles_J_best) {
- J_best = J;
+ J_best = J;
ntiles_J_best = ntiles_x;
+ nbytes_pad_y = GGML_PAD(config.J*sizeof(block_q8_1_mmq), config.nthreads*sizeof(int));
}
}
}
@@ -225,11 +227,11 @@

if (!ids) {
const size_t nbytes_src1_q8_1 = ne13*ne12 * ne11*ne10_padded * y_block_size/y_values_per_block +
- J_best * sizeof(block_q8_1_mmq);
+ nbytes_pad_y;
ggml_cuda_pool_alloc<char> src1_q8_1(ctx.pool(), nbytes_src1_q8_1);
ggml_cuda_pool_alloc<float> src1_scale(ctx.pool());
if (src0->type == GGML_TYPE_NVFP4 && use_native_fp4) {
- src1_scale.alloc(ne13*ne12*ne11);
+ src1_scale.alloc(ne13*ne12*ne11 + J_best);
}

{
@@ -276,7 +278,7 @@
GGML_ASSERT(ne1 == n_expert_used);

ggml_cuda_pool_alloc<int32_t> ids_src1(ctx.pool(), ne_get_rows);
- ggml_cuda_pool_alloc<int32_t> ids_dst(ctx.pool(), ne_get_rows);
+ ggml_cuda_pool_alloc<int32_t> ids_dst(ctx.pool(), ne_get_rows + J_best);
ggml_cuda_pool_alloc<int32_t> expert_bounds(ctx.pool(), ne02 + 1);

// gate/up activations are broadcast across experts (ne11 == 1): quantize each token once and
@@ -294,11 +296,11 @@
}

const size_t nbytes_src1_q8_1 = ne12*n_expert_used*ne10_padded * y_block_size/y_values_per_block +
- J_best * sizeof(block_q8_1_mmq);
+ nbytes_pad_y;
ggml_cuda_pool_alloc<char> src1_q8_1(ctx.pool(), nbytes_src1_q8_1);
ggml_cuda_pool_alloc<float> src1_scale(ctx.pool());
if (src0->type == GGML_TYPE_NVFP4 && use_native_fp4) {
- src1_scale.alloc(ne12*n_expert_used);
+ src1_scale.alloc(ne12*n_expert_used + J_best);
}

const int64_t ne11_flat = ne12*n_expert_used;
Loading
Loading