Skip to content

Don't carry source tensors a dequantizing load already folded into the weights - #2733

Draft
shengliangxu wants to merge 2 commits into
mainfrom
shengliangx/fix-carry-load-time-conversions
Draft

shengliangxu wants to merge 2 commits into
mainfrom
shengliangx/fix-carry-load-time-conversions

Conversation

@shengliangxu

@shengliangxu shengliangxu commented Oct 11, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

Type of change: Bug fix

The export copies every source tensor unplaced_source_keys reports into the checkpoint verbatim. That pass replayed only the model's stock conversion mapping, so it missed whatever loading added on top, and those additions consume tensors the stock mapping never mentions:

  • Dequantize-on-load quantizers. Mxfp4Config(dequantize=True), which examples/hf_ptq uses for gpt-oss, folds each expert's *_blocks/*_scales into a bf16 gate_up_proj/down_proj. FineGrainedFP8Config(dequantize=True) folds every FP8 weight_scale_inv into its weight.
  • A patched mapping. An on-the-fly compressed-tensors decompressor feeds weight_packed/weight_scale experts in.

Each such tensor read as unplaced, so the export carried it in beside the weight it had become: stale pre-quantization copies under the source names, plus matching exclude_modules entries.

  • Where it was found: an expert-parallel Kimi-K3 export through such a decompressor carried 494,592 packed expert tensors. The checkpoint grew from 1.5 TB to 2.8 TB, and rank 0 spent 21 of the run's 37 minutes writing them.
  • Where it hits on main today: gpt-oss. A tiny gpt-oss checkpoint loaded the way hf_ptq loads it reports all four expert block/scale tensors as unplaced.

Fix: from_pretrained records the conversions it actually ran on the model (model._weight_conversions, the list transformers' own save path replays, and the one our quant-aware reverse conversion in quant_aware_conversion.py already prefers), with quantizer and patch additions included. A key now counts as placed when either that record or the stock mapping resolves it to a parameter. The record keeps only conversions some key used, so the stock mapping still answers for weights one of our own loaders placed (the FSDP2 meta-device load).

Usage

N/A: internal to the export's carry-over of unplaced weights.

Testing

  • New tests/unit/torch/utils/test_unplaced_source_keys.py. It runs a real from_pretrained on CPU of a tiny FP8-block Llama (FineGrainedFP8Config(dequantize=True)) and a tiny gpt-oss (Mxfp4Config(dequantize=True)). Each checkpoint also holds one genuinely stray tensor.
    • Before the fix: every scale (7) and every expert block/scale tensor (4) is reported unplaced, alongside the stray.
    • After the fix: only the stray.
  • Decompressor case, simulated with an on-the-fly decompressor's converter rewrite on a tiny Qwen3-MoE: 24 of 24 packed expert keys were unplaced before the fix, none after.
  • Existing suites: tests/unit/torch/utils/test_model_load_utils.py, tests/examples/hf_ptq/test_carry_over_layouts.py, tests/examples/hf_ptq/test_example_utils.py and tests/unit/torch/export/{test_unified_export_hf,test_vllm_fakequant_hf,test_get_quantization}.py: 161 passed.
  • mypy and the rest of pre-commit: clean.
  • Kimi-K3 expert-parallel export (16 nodes, EP64) through such a decompressor: with the fix, no "Carrying N source weight(s)" line appears, and the export has the expected 995,715 keys (1.5 TB, down from 2.8 TB; the difference is exactly the 494,592 packed expert tensors), matching a reference export key for key. In the unfixed run the carry pass alone took 20:53 of the 37:16 wall time, per the log timestamps; the fixed run took 14:57. The two runs also differed in branch changes unrelated to this PR, so the time saved is read from the log rather than from the run-to-run difference.

Before your PR is "Ready for review"

  • Is this change backward compatible?: ✅ Exports stop carrying tensors that were already folded into weights.
  • If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: N/A
  • Did you write any new necessary tests?: ✅
  • Did you update Changelog?: N/A. The carry-over of unplaced weights is new in the unreleased 0.48.0 cycle.
  • Did you get Claude approval on this PR?: ❌ Not yet.

Additional Information

Prerequisite of the HF expert-parallel PTQ (hfep) stack, which is rebased on top of it.

The export carries every source tensor unplaced_source_keys reports, and
that pass replayed only the model's stock conversion mapping. Loading can add
to the mapping, and what it adds consumes tensors the stock mapping never
mentions:

- A dequantizing quantizer. Mxfp4Config(dequantize=True), which
  examples/hf_ptq uses for gpt-oss, folds each expert's *_blocks and *_scales
  into a bf16 gate_up_proj/down_proj; FineGrainedFP8Config(dequantize=True)
  folds every FP8 weight_scale_inv into its weight.
- A patched mapping, such as an on-the-fly decompressor that feeds
  compressed-tensors weight_packed/weight_scale experts in.

Each such tensor read as unplaced, so the export copied it in beside the
weight it had become. A tiny gpt-oss checkpoint loaded the way hf_ptq loads
it reports all four expert block/scale tensors; a tiny FP8-block Llama, all
seven scales. Found on an expert-parallel Kimi-K3 export through such a
decompressor: 494,592 packed expert tensors were carried, the checkpoint grew
from 1.5 TB to 2.8 TB, and rank 0 spent 21 of the run's 37 minutes writing
them.

from_pretrained records the conversions it actually ran on the model
(_weight_conversions, the list transformers' own save path replays), the
additions included. A key now counts as placed when either that record or
the stock mapping resolves it to a parameter: the record keeps only the
conversions some key used, so the stock mapping still answers for weights
one of our own loaders placed.

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Reassigning the filtered list kept the type of its first assignment,
list[dict | None], so handing an element to _resolve_target failed type
checking. Filter a tuple of the two candidate plans into a fresh name instead.

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 11, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Oct 11, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

QR code for preview link

🚀 View preview at
https://NVIDIA.github.io/Model-Optimizer/pr-preview/pr-2733/

Built to branch gh-pages at 2026-10-11 07:26 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

@codecov

codecov Bot commented Oct 11, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.85714% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 69.47%. Comparing base (54d4416) to head (2a49c89).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
modelopt/torch/utils/plugins/model_load_utils.py 92.85% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #2733      +/-   ##
==========================================
- Coverage   69.47%   69.47%   -0.01%     
==========================================
  Files         646      646              
  Lines       71672    71681       +9     
==========================================
+ Hits        49796    49798       +2     
- Misses      21876    21883       +7     
Flag Coverage Δ
unit 59.89% <92.85%> (-0.01%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant