Skip to content

fix(export): Split unquantized 3-D dense weights in quant-aware reverse conversion (GLM-5.3-Flash) - #2732

Draft
cjluo-nv wants to merge 1 commit into
mainfrom
chenjiel/fix-glm5next-conv1d-reverse
Draft

cjluo-nv wants to merge 1 commit into
mainfrom
chenjiel/fix-glm5next-conv1d-reverse

Conversation

@cjluo-nv

Copy link
Copy Markdown
Collaborator

What does this PR do?

Type of change: Bug fix

Unified HF export of GLM-5.3-Flash (glm5_next) writes transformers' in-memory tensor
names instead of the hub layout, and vLLM cannot load the checkpoint:

KeyError: 'layers.0.self_attn.conv1d.weight'

Root cause. transformers fuses GLM-5.3's KDA depthwise convs q_conv1d / k_conv1d /
v_conv1d (each [C, 1, K]) into one 3-D conv1d weight. The quant-aware reverse conversion
gets a dense SplitRule for that fusion. _apply_split_rule treats any 3-D .weight as a
stacked expert tensor and raises QuantConversionUnsupportedError. The reverse is atomic, so
every tensor then keeps its in-memory name (conv1d, forget_gate.A_log, ...). The only
trace is a UserWarning ("Quant-aware reverse weight conversion skipped").

Fix. The guard now rejects a 3-D weight only when it really can be a stacked expert:

  • it sits under .experts, or
  • it carries quantization state (weight_scale, ...).

Unquantized dense 3-D weights chunk cleanly along the output dim, the same way the 2-D case
already does.

Usage

No API change. For example, this now exports with hub tensor names:

python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path zai-org/GLM-5.3-Flash-BF16 \
    --qformat nvfp4_experts_only --kv_cache_qformat fp8_cast --export_path <out>

Testing

  • New unit tests in tests/unit/torch/export/test_quant_aware_conversion.py:
    • An unquantized fused [3C, 1, K] conv weight splits back to exactly the original q/k/v
      tensors.
    • A quantized 3-D weight outside .experts still falls back.
    • The existing test_stacked_3d_expert_raises_unsupported is unchanged and passes.
  • tests/unit/torch/export/test_quant_aware_conversion.py: 59 passed with transformers 5.14.1
    (15 passed, 44 skipped with 4.57.6, which lacks core_model_loading).
  • tests/unit/torch/export/ without test_export_diffusers.py: 349 passed.
    test_export_diffusers_models_non_quantized[get_tiny_dit] fails in my local env on main as
    well, unrelated to this change.
  • End to end (oci-jhb GB300, transformers 5.16.1). Four GLM-5.3-Flash
    nvfp4_experts_only + fp8_cast exports from zai-org/GLM-5.3-Flash-BF16:
    • Without the fix, vLLM failed with the KeyError above.
    • With it, the tensor-name set matches the published nvidia/GLM-5.3-Flash-NVFP4, apart from
      that checkpoint's extra dense-MLP quantization in layers 0-2.
    • The exports serve in vLLM and were evaluated on GPQA-D, AA-LCR, SciCode and Terminal-Bench
      2.1.

Before your PR is "Ready for review"

  • Is this change backward compatible?: ✅ (the stacked-expert fallback is unchanged)
  • If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: N/A
  • Did you write any new necessary tests?: ✅
  • Did you update Changelog?: ✅ (the bug is in 0.46 and 0.47)
  • Did you get Claude approval on this PR?: ❌ (draft)

Additional Information

Found while running an NVFP4 calibration-data study on GLM-5.3-Flash. Follows #1833 (quant-aware
reverse weight conversion).

🤖 Generated with Claude Code

…se conversion

The quant-aware reverse weight conversion rejected every 3-D tensor
matched by a dense split rule as a stacked expert. GLM-5.3-Flash
(glm5_next) fuses its KDA q/k/v depthwise conv1d weights ([C, 1, K]) into
one 3-D `conv1d`, so the whole reverse was skipped and the export kept
transformers' in-memory names (`conv1d`, `forget_gate.*`). vLLM then fails
with KeyError: 'layers.0.self_attn.conv1d.weight'.

Only reject 3-D weights under `.experts` or carrying quantization state;
unquantized dense 3-D weights chunk cleanly along the output dim.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 11, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Oct 11, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

QR code for preview link

🚀 View preview at
https://NVIDIA.github.io/Model-Optimizer/pr-preview/pr-2732/

Built to branch gh-pages at 2026-10-11 02:57 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

@codecov

codecov Bot commented Oct 11, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 69.46%. Comparing base (54d4416) to head (5374a5f).

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #2732      +/-   ##
==========================================
- Coverage   69.47%   69.46%   -0.01%     
==========================================
  Files         646      646              
  Lines       71672    71673       +1     
==========================================
- Hits        49796    49791       -5     
- Misses      21876    21882       +6     
Flag Coverage Δ
unit 59.88% <100.00%> (-0.01%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant