Skip to content

cuda: always use MMVQ for MUL_MAT_ID on sm_60 - #27828

Open
dfriehs wants to merge 1 commit into
ggml-org:masterfrom
dfriehs:p100-moe-mmvq
Open

cuda: always use MMVQ for MUL_MAT_ID on sm_60#27828
dfriehs wants to merge 1 commit into
ggml-org:masterfrom
dfriehs:p100-moe-mmvq

Conversation

@dfriehs

@dfriehs dfriehs commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Overview

On the P100, MMVQ is faster than the alternative for MUL_MAT_ID on all batch sizes up to 8, so include it in the list of architectures that always use MMVQ.

I believe some of the quantization types will also run faster with fewer warps active (calc_nwarps), but not all of them benefited from halving the number of warps. I'll follow up on that, either in this PR or a next one. See below, there are some speedups to be had from tuning the number of warps and enabling small_k for more quantization types. Tuning seems benefitial, but my first attempts ended in regressions on larger models and will have to wait.

Additional information

All tests were done on a single P100, powerlimited to 150 W, CUDA 12.9.2, driver 575.64.05.

scripts used for testing
_quants=(
  Q1_0 Q2_0 Q4_0 Q4_1 Q5_0 Q5_1 Q8_0
  IQ1_S IQ1_M IQ2_XXS IQ2_XS IQ2_S IQ3_XXS IQ3_S IQ4_XS IQ4_NL
  Q2_K Q3_K Q4_K Q5_K Q6_K
)

for q in "${_quants[@]}"
do
  llama-quantize --tensor-type "exps=$q" --tensor-type "embd|blk|out=q8_0" \
    --imatrix "Ling-3.0-Tiny-128x1.0B-imatrix.gguf" "Ling-3.0-Tiny-128x1.0B-BF16.gguf" "bailingmoe3-$q.gguf" "$q"
done

for q in MXFP4 NVFP4
do
  llama-quantize --tensor-type "exps=$q" --tensor-type "embd|blk|out=q8_0" \
    --imatrix "Ling-3.0-Tiny-128x1.0B-imatrix.gguf" "Ling-3.0-Tiny-128x1.0B-BF16.gguf" "bailingmoe3-$q.gguf" "Q8_0"
done
export CUDA_VISIBLE_DEVICES=0

_quants=(
  Q1_0 Q2_0 Q4_0 Q4_1 Q5_0 Q5_1 Q8_0
  IQ1_S IQ1_M IQ2_XXS IQ2_XS IQ2_S IQ3_XXS IQ3_S IQ4_XS IQ4_NL
  Q2_K Q3_K Q4_K Q5_K Q6_K
  MXFP4 NVFP4
)

for q in "${_quants[@]}"
do
  echo "### $q"
  llama-batched-bench -lm dio -fa on -m "bailingmoe3-$q.gguf" -ub 2048 -npp 2048 -ntg 128 -npl 1,2,3,4,5,6,7,8
  sleep 2
done
results for Ling 3.0 Tiny

Bolded values previously exceeded get_mmvq_mmid_max_batch_pascal_older.

Q1_0

1 2 3 4 5 6 7 8
master 93.57 160.74 205.60 232.63 251.32 272.34 289.34 302.14
PR 93.47 160.68 205.57 232.42 251.22 272.34 289.34 302.13

Q2_0

1 2 3 4 5 6 7 8
master 93.92 160.94 205.42 232.12 250.87 271.14 287.89 300.47
PR 93.68 160.64 205.16 231.97 250.55 270.75 287.41 300.15

Q4_0

1 2 3 4 5 6 7 8
master 94.92 166.06 215.03 243.58 264.47 287.69 232.84 248.23
PR 95.09 166.43 215.32 243.88 264.73 287.98 307.24 321.34

Q4_1

1 2 3 4 5 6 7 8
master 95.93 168.95 219.08 248.75 270.51 294.04 232.38 247.79
PR 95.68 168.41 218.47 248.47 270.28 293.73 314.21 328.45

Q5_0

1 2 3 4 5 6 7 8
master 90.78 157.85 200.10 225.35 242.93 263.05 212.67 228.17
PR 90.93 158.10 200.44 225.47 243.01 263.30 278.20 288.57

Q5_1

1 2 3 4 5 6 7 8
master 92.18 159.46 203.87 230.44 248.07 268.15 218.11 233.08
PR 91.98 159.10 203.36 230.05 247.78 267.89 285.19 295.42

Q8_0

1 2 3 4 5 6 7 8
master 90.51 154.42 196.01 220.32 183.60 203.84 223.53 238.01
PR 90.45 154.47 196.11 220.17 237.81 257.10 271.11 282.74

IQ1_S

1 2 3 4 5 6 7 8
master 87.03 157.85 202.83 228.34 246.51 265.94 229.18 245.03
PR 87.04 157.80 202.56 228.13 246.63 265.99 282.86 294.16

IQ1_M

1 2 3 4 5 6 7 8
master 83.14 152.28 193.60 217.04 233.22 250.47 161.82 170.68
PR 83.23 152.25 193.68 217.51 233.93 251.70 265.30 272.28

IQ2_XXS

1 2 3 4 5 6 7 8
master 82.22 147.15 183.75 205.72 220.32 206.62 225.94 241.61
PR 82.12 147.01 183.53 205.70 220.26 235.35 248.34 257.06

IQ2_XS

1 2 3 4 5 6 7 8
master 81.14 146.00 182.00 203.15 216.85 207.52 226.43 241.97
PR 81.02 145.92 181.93 203.03 216.98 230.94 244.01 251.22

IQ2_S

1 2 3 4 5 6 7 8
master 81.62 145.31 181.90 203.43 186.42 206.96 225.19 240.86
PR 81.83 145.49 181.95 203.48 217.35 231.38 244.35 251.41

IQ3_XXS

1 2 3 4 5 6 7 8
master 81.56 143.36 178.54 199.32 181.87 202.11 220.87 235.85
PR 81.45 143.22 178.26 199.16 212.32 227.38 237.38 246.68

IQ3_S

1 2 3 4 5 6 7 8
master 80.29 139.91 174.38 193.65 179.12 199.16 217.60 232.86
PR 80.39 140.03 174.45 193.41 206.19 218.42 230.00 234.84

IQ4_XS

1 2 3 4 5 6 7 8
master 84.74 156.83 200.13 225.09 242.13 205.94 225.42 240.94
PR 84.97 157.07 200.41 225.24 242.51 262.11 277.13 288.76

IQ4_NL

1 2 3 4 5 6 7 8
master 94.36 164.40 212.20 239.51 259.43 281.13 225.51 241.07
PR 94.15 164.16 212.03 239.31 258.94 281.01 298.99 310.75

Q2_K

1 2 3 4 5 6 7 8
master 83.55 160.36 205.09 231.39 183.61 203.56 223.00 237.68
PR 83.53 160.24 204.99 231.61 249.85 269.22 286.38 296.63

Q3_K

1 2 3 4 5 6 7 8
master 78.32 146.33 182.99 203.44 178.61 198.68 217.05 232.70
PR 78.32 145.94 182.51 202.75 215.48 230.24 237.35 247.43

Q4_K

1 2 3 4 5 6 7 8
master 91.66 155.20 197.12 222.25 239.32 209.13 228.94 243.89
PR 91.59 155.01 196.93 222.07 239.18 257.11 272.16 282.56

Q5_K

1 2 3 4 5 6 7 8
master 89.61 155.65 197.70 223.57 241.00 205.08 223.57 239.11
PR 89.51 154.62 196.01 221.78 238.49 257.36 271.31 281.66

Q6_K

1 2 3 4 5 6 7 8
master 84.31 145.79 182.96 204.37 179.61 199.24 217.53 232.61
PR 84.41 145.90 183.04 204.49 219.04 235.56 246.90 256.60

MXFP4

1 2 3 4 5 6 7 8
master 92.89 159.02 202.61 228.59 182.89 202.55 221.60 237.72
PR 92.77 159.09 202.85 228.46 246.78 267.12 282.25 294.00

NVFP4

1 2 3 4 5 6 7 8
master 84.30 147.62 185.73 206.83 183.66 204.28 222.62 238.58
PR 83.99 147.16 185.28 206.37 221.12 238.18 250.95 260.36

Requirements

@dfriehs
dfriehs requested a review from a team as a code owner August 27, 2026 20:28
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Aug 27, 2026
@dfriehs

dfriehs commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

For some types, halving warps improves performance for B=1, especially when re-enabling small_k as well. B=2+ is unaffected.

I really want to test a dense model too before this merges, but I would appreciate a review and possibly tests on other cards (especially sm_61, sm_70) regarding the type filter in mul_mat_vec_q_switch_ncols_dst::should_use_small_k and halving warps.

results for Ling 3.0 Tiny

IQ1_S

1 2 3 4 5 6 7 8
old 87.04 157.80 202.56 228.13 246.63 265.99 282.86 294.16
new 93.00 158.66 202.91 230.40 246.75 266.55 283.28 295.06

IQ1_M

1 2 3 4 5 6 7 8
old 83.23 152.25 193.68 217.51 233.93 251.70 265.30 272.28
new 90.71 151.91 193.86 219.27 234.12 251.60 265.24 272.16

IQ2_XXS

1 2 3 4 5 6 7 8
old 82.12 147.01 183.53 205.70 220.26 235.35 248.34 257.06
new 87.43 146.48 182.93 206.12 218.90 235.05 247.84 256.82

IQ2_XS

1 2 3 4 5 6 7 8
old 81.02 145.92 181.93 203.03 216.98 230.94 244.01 251.22
new 87.94 145.86 182.11 204.58 216.51 231.04 244.01 250.85

IQ2_S

1 2 3 4 5 6 7 8
old 81.83 145.49 181.95 203.48 217.35 231.38 244.35 251.41
new 87.64 145.29 181.84 204.69 216.99 231.68 244.62 251.06

IQ3_XXS

1 2 3 4 5 6 7 8
old 81.45 143.22 178.26 199.16 212.32 227.38 237.38 246.68
new 86.99 143.57 178.85 200.45 211.99 227.41 237.43 246.55

IQ3_S

1 2 3 4 5 6 7 8
old 80.39 140.03 174.45 193.41 206.19 218.42 230.00 234.84
new 85.46 140.17 174.48 194.96 206.24 218.47 230.13 234.88

IQ4_XS

1 2 3 4 5 6 7 8
old 84.97 157.07 200.41 225.24 242.51 262.11 277.13 288.76
new 91.51 156.54 199.19 226.83 241.72 261.79 276.85 288.20

Q2_K

1 2 3 4 5 6 7 8
old 83.53 160.24 204.99 231.61 249.85 269.22 286.38 296.63
new 91.86 160.83 205.00 233.47 249.38 269.37 285.53 296.43

Q3_K

1 2 3 4 5 6 7 8
old 78.32 145.94 182.51 202.75 215.48 230.24 237.35 247.43
new 85.05 146.02 182.60 204.14 215.70 230.60 237.48 246.94

NVFP4

1 2 3 4 5 6 7 8
old 83.99 147.16 185.28 206.37 221.12 238.18 250.95 260.36
new 87.37 147.66 185.24 208.48 221.91 238.77 250.87 261.10

@dfriehs

dfriehs commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

Tested Qwen 3.5 9B and had to slightly change the number of warps to not regress. No change for MoE, although I see 1-2 tok/s more at B=8.

Most types stayed relatively flat, but there's some improvements on B=2-4 and on IQ quants.

script used for quantization
_quants=(
  Q1_0 Q2_0 Q4_0 Q4_1 Q5_0 Q5_1 Q8_0
  IQ1_S IQ1_M IQ2_XXS IQ2_XS IQ2_S IQ3_XXS IQ3_S IQ4_XS IQ4_NL
  Q2_K Q3_K Q4_K Q5_K Q6_K
)

for q in "${_quants[@]}"
do
  llama-quantize --tensor-type "ssm_[ab]=f32" --tensor-type "blk=$q" --tensor-type "embd|out=q4_0" \
    --imatrix "Qwen3.5-9B-Base-imatrix.gguf" "Qwen3.5-9B-Base-BF16.gguf" "qwen35-$q.gguf" "$q"
done

for q in MXFP4 NVFP4
do
  llama-quantize --tensor-type "ssm_[ab]=f32" --tensor-type "blk=$q" --tensor-type "embd|out=q4_0" \
    --imatrix "Qwen3.5-9B-Base-imatrix.gguf" "Qwen3.5-9B-Base-BF16.gguf" "qwen35-$q.gguf" "Q8_0"
done
results for Qwen 3.5 9B

Q1_0

1 2 3 4 5 6 7 8
master 51.14 93.92 109.35 114.10 112.21 116.75 121.58 121.59
PR 51.14 99.80 114.44 121.10 112.17 117.56 121.10 123.20

Q2_0

1 2 3 4 5 6 7 8
master 51.31 91.99 107.29 111.78 110.29 113.72 118.51 115.69
PR 51.39 97.92 114.26 118.58 111.59 115.19 118.16 120.72

Q4_0

1 2 3 4 5 6 7 8
master 46.34 83.79 95.63 98.08 97.00 101.30 104.04 104.99
PR 46.37 88.65 98.95 102.62 98.99 102.42 105.65 105.78

Q4_1

1 2 3 4 5 6 7 8
master 50.24 91.64 97.61 105.02 104.37 109.44 113.40 112.71
PR 50.22 94.90 102.53 109.95 106.81 111.35 111.67 114.66

Q5_0

1 2 3 4 5 6 7 8
master 39.62 71.80 82.12 84.06 86.45 92.30 91.24 97.04
PR 39.64 73.18 85.93 90.13 90.32 94.12 94.31 97.87

Q5_1

1 2 3 4 5 6 7 8
master 41.04 72.11 83.94 88.52 89.13 94.21 97.33 98.57
PR 41.02 75.30 84.17 89.69 87.79 94.61 96.70 99.71

Q8_0

1 2 3 4 5 6 7 8
master 35.01 64.15 79.01 85.82 85.69 90.47 94.96 96.81
PR 35.01 64.02 79.61 86.33 85.96 92.27 95.24 96.98

IQ1_S

1 2 3 4 5 6 7 8
master 44.93 77.51 89.46 95.66 99.06 100.48 106.18 108.98
PR 44.90 80.23 97.09 102.77 98.97 104.70 105.90 109.62

IQ1_M

1 2 3 4 5 6 7 8
master 40.79 69.41 83.66 89.90 90.84 94.37 97.06 100.14
PR 41.30 75.32 89.94 91.85 91.39 95.20 98.25 100.51

IQ2_XXS

1 2 3 4 5 6 7 8
master 38.64 72.54 85.21 88.41 90.13 95.76 98.14 94.26
PR 41.12 77.50 89.39 94.20 91.90 97.18 100.90 96.84

IQ2_XS

1 2 3 4 5 6 7 8
master 36.81 69.08 78.75 85.06 85.17 89.79 91.17 86.75
PR 39.37 73.53 85.76 89.67 86.33 91.40 94.39 89.21

IQ2_S

1 2 3 4 5 6 7 8
master 37.89 67.75 77.72 84.17 85.10 89.83 91.12 88.75
PR 39.30 69.92 84.19 86.27 85.31 92.53 93.36 89.29

IQ3_XXS

1 2 3 4 5 6 7 8
master 35.54 63.56 76.36 83.64 83.79 90.61 90.44 89.95
PR 36.33 65.02 79.11 85.66 85.44 90.77 94.52 91.61

IQ3_S

1 2 3 4 5 6 7 8
master 32.49 56.64 68.55 79.63 80.93 86.10 88.23 88.63
PR 32.71 58.72 73.03 79.53 80.84 86.23 90.95 89.51

IQ4_XS

1 2 3 4 5 6 7 8
master 44.10 82.08 93.32 97.90 96.73 103.46 103.81 107.25
PR 48.19 86.32 102.51 103.33 99.51 105.67 108.86 111.61

IQ4_NL

1 2 3 4 5 6 7 8
master 43.53 81.39 94.41 98.98 96.58 104.22 108.31 109.31
PR 43.56 85.06 97.76 100.94 98.80 106.18 108.25 110.69

Q2_K

1 2 3 4 5 6 7 8
master 36.86 62.50 68.05 69.30 68.66 70.93 70.70 71.08
PR 36.85 65.36 69.74 72.39 70.15 71.23 71.78 71.58

Q3_K

1 2 3 4 5 6 7 8
master 29.31 52.25 60.96 65.41 67.63 70.94 73.48 74.95
PR 31.27 56.28 67.15 68.91 68.77 73.72 74.06 76.31

Q4_K

1 2 3 4 5 6 7 8
master 39.49 63.24 70.34 69.68 71.41 72.60 72.28 72.85
PR 39.47 66.07 72.91 72.47 71.30 72.80 73.06 72.54

Q5_K

1 2 3 4 5 6 7 8
master 36.83 57.98 65.23 67.60 66.42 68.26 69.89 70.39
PR 36.79 60.82 67.15 68.97 68.06 69.48 69.83 70.26

Q6_K

1 2 3 4 5 6 7 8
master 28.31 52.86 61.98 66.88 69.00 72.32 74.18 76.49
PR 28.32 51.73 63.09 66.90 68.84 72.11 74.89 75.86

MXFP4

1 2 3 4 5 6 7 8
master 39.98 73.58 86.12 92.50 94.36 97.23 104.07 103.89
PR 40.00 73.68 89.34 95.39 94.34 99.71 103.05 105.25

NVFP4

1 2 3 4 5 6 7 8
master 37.79 52.45 54.15 57.04 54.50 57.50 58.43 59.11
PR 39.15 53.94 54.91 56.73 54.90 58.08 59.42 59.76

@dfriehs dfriehs changed the title cuda: always use MMVQ for MUL_MAT_ID on sm_60 cuda: tune MMVQ for sm_60 Aug 28, 2026
@dfriehs

dfriehs commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

Very sorry for the spam, but testing another model at Q8_0 I think I get regressions again.

I think I'll have to leave the tuning for the future, and will drop everything except the MMVQ_MAX_BATCH_SIZE change which gave me good results on other models as well.

@dfriehs dfriehs changed the title cuda: tune MMVQ for sm_60 cuda: always use MMVQ for MUL_MAT_ID on sm_60 Aug 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant