Repository navigation
fix(export): Split unquantized 3-D dense weights in quant-aware reverse conversion (GLM-5.3-Flash) - #2732
Draft
cjluo-nv wants to merge 1 commit into
Draft
fix(export): Split unquantized 3-D dense weights in quant-aware reverse conversion (GLM-5.3-Flash)#2732cjluo-nv wants to merge 1 commit into
cjluo-nv wants to merge 1 commit into
Conversation
…se conversion The quant-aware reverse weight conversion rejected every 3-D tensor matched by a dense split rule as a stacked expert. GLM-5.3-Flash (glm5_next) fuses its KDA q/k/v depthwise conv1d weights ([C, 1, K]) into one 3-D `conv1d`, so the whole reverse was skipped and the export kept transformers' in-memory names (`conv1d`, `forget_gate.*`). vLLM then fails with KeyError: 'layers.0.self_attn.conv1d.weight'. Only reject 3-D weights under `.experts` or carrying quantization state; unquantized dense 3-D weights chunk cleanly along the output dim. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
Contributor
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #2732 +/- ##
==========================================
- Coverage 69.47% 69.46% -0.01%
==========================================
Files 646 646
Lines 71672 71673 +1
==========================================
- Hits 49796 49791 -5
- Misses 21876 21882 +6
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Type of change: Bug fix
Unified HF export of GLM-5.3-Flash (
glm5_next) writes transformers' in-memory tensornames instead of the hub layout, and vLLM cannot load the checkpoint:
Root cause. transformers fuses GLM-5.3's KDA depthwise convs
q_conv1d/k_conv1d/v_conv1d(each[C, 1, K]) into one 3-Dconv1dweight. The quant-aware reverse conversiongets a dense
SplitRulefor that fusion._apply_split_ruletreats any 3-D.weightas astacked expert tensor and raises
QuantConversionUnsupportedError. The reverse is atomic, soevery tensor then keeps its in-memory name (
conv1d,forget_gate.A_log, ...). The onlytrace is a
UserWarning("Quant-aware reverse weight conversion skipped").Fix. The guard now rejects a 3-D weight only when it really can be a stacked expert:
.experts, orweight_scale, ...).Unquantized dense 3-D weights chunk cleanly along the output dim, the same way the 2-D case
already does.
Usage
No API change. For example, this now exports with hub tensor names:
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path zai-org/GLM-5.3-Flash-BF16 \ --qformat nvfp4_experts_only --kv_cache_qformat fp8_cast --export_path <out>Testing
tests/unit/torch/export/test_quant_aware_conversion.py:[3C, 1, K]conv weight splits back to exactly the original q/k/vtensors.
.expertsstill falls back.test_stacked_3d_expert_raises_unsupportedis unchanged and passes.tests/unit/torch/export/test_quant_aware_conversion.py: 59 passed with transformers 5.14.1(15 passed, 44 skipped with 4.57.6, which lacks
core_model_loading).tests/unit/torch/export/withouttest_export_diffusers.py: 349 passed.test_export_diffusers_models_non_quantized[get_tiny_dit]fails in my local env onmainaswell, unrelated to this change.
nvfp4_experts_only+fp8_castexports fromzai-org/GLM-5.3-Flash-BF16:KeyErrorabove.nvidia/GLM-5.3-Flash-NVFP4, apart fromthat checkpoint's extra dense-MLP quantization in layers 0-2.
2.1.
Before your PR is "Ready for review"
CONTRIBUTING.md: N/AAdditional Information
Found while running an NVFP4 calibration-data study on GLM-5.3-Flash. Follows #1833 (quant-aware
reverse weight conversion).
🤖 Generated with Claude Code