Repository navigation
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
kaix-nv
added this pull request to stack #2563
October 5, 2026 06:11
kaix-nv
removed this pull request from stack #2563
October 5, 2026 06:15
kaix-nv
added this pull request to stack #2658
October 5, 2026 06:15
This was referenced Oct 5, 2026
kaix-nv
force-pushed
the
kaix/linear-attention-qat-example
branch
from
October 6, 2026 17:27
2829885 to
3837d28
Compare
kaix-nv
force-pushed
the
kaix/linear-attention-qat-example
branch
from
October 6, 2026 19:55
3837d28 to
b7401f7
Compare
kaix-nv
force-pushed
the
kaix/linear-attention-qat-example
branch
from
October 6, 2026 21:26
b7401f7 to
0357915
Compare
kaix-nv
force-pushed
the
kaix/linear-attention-qat-example
branch
from
October 8, 2026 05:35
0d898ac to
7717ea4
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## kaix/linear-attention-decode-first #2657 +/- ##
===================================================================
Coverage 78.11% 78.11%
===================================================================
Files 659 659
Lines 72634 72634
===================================================================
Hits 56740 56740
Misses 15894 15894
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
kaix-nv
removed this pull request from stack #2658
October 8, 2026 18:41
kaix-nv
added this pull request to stack #2714
October 8, 2026 18:41
kaix-nv
force-pushed
the
kaix/linear-attention-qat-example
branch
from
October 8, 2026 22:40
7717ea4 to
942115a
Compare
kaix-nv
force-pushed
the
kaix/linear-attention-qat-example
branch
from
October 10, 2026 17:16
db921e4 to
1dacb59
Compare
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Update the linear-attention example for the flat execution config, native state-only training, and checkpoint migration. Remove the retired CPU reference script while retaining its historical results and source link. Record the focused CPU, native GPU, Megatron checkpoint, and config migration validation. Documentation pre-commit hooks passed. Signed-off-by: Kai Xu <kaix@nvidia.com>
Document the QDQ scheduling and arithmetic mismatch, the native prefix/handoff/suffix solution, gradient flow, and the limits of current validation. Align configuration guidance with the current state QAT API. Signed-off-by: Kai Xu <kaix@nvidia.com>
Remove raw experiment JSON from the example and shorten the usage and state-alignment documents. Retain the mismatch derivation, supported policies, concise numerical results, and validation limits. Preserve development evidence in an untracked local archive. Training code and minimal QAT/QAD tests are unchanged. Validation: scoped pre-commit, documentation links and anchors, Python snippet syntax, archive integrity, and git diff --check. Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
kaix-nv
force-pushed
the
kaix/linear-attention-qat-example
branch
from
October 11, 2026 06:22
1dacb59 to
5731740
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Linear-attention series — 6 PRs
mainThe five open PRs form one native GitHub stack in the order shown. #2497 has landed, so #2519 targets
main. #2541 now targets #2657. Rebase each remaining descendant after its immediate parent merges.#2541 applies TensorQuantizer before native vLLM prefill/decode calls. Serving-time prefill-GEMM quantization remains deferred until an optimized fused kernel is available.
What does this PR do?
Type of change: New example.
Add a Megatron Bridge example for recurrent-state QAT and QAD. The training entry point selects a native chunked prefix and recurrent suffix, applies loss to the suffix, and keeps the phase context active through backward. Bridge owns optimization, distributed scheduling, and checkpoints. A frozen unquantized teacher enables QAD.
The example contains the training script, dependency files, launcher, two minimal training tests, and concise usage/alignment documentation. Raw experiment records and development links are kept outside the release PR. State formats, execution policies, and kernels are supplied by the parent PR.
Usage
Add
--teacher-model /path/to/unquantized-modelfor QAD. The ordinary INT8 recipe uses public vLLM kernels. INT8 + Hadamard and replay require a compatible native ReplaySSM fork. A KDA trainer additionally requires a compatible Bridge provider. TP/PP/EP options follow Bridge; context parallelism stays at one, and each local pipeline chunk must contain linear attention.Testing
Rebased onto #2519 at
03f6076dd8; this PR's head is57317404d0.git diff --checkpassed.git range-diffconfirms all nine example commits replayed without patch changes. The parent-relative diff is 8 files, 987 additions, with no runtime/kernel or workflow changes._forward_computehook now required by [2/6] Torch GDN/KDA decode QAT with INT8 recurrent state #2519. Coreac100f773f9dprovides the required hooks, but collection fails with the localnvidia_resiliency_extpackage; compatible published wheels require a newer system runtime. No passing QAT/QAD workflow result is claimed for this rebased head.140735b962,pytest tests/examples/megatron_bridge/test_linear_attention.pypassed 2 tests in 57.84 seconds. Those tests cover QAT/QAD student updates, a frozen QAD teacher, checkpoint save/resume, HF export/load, and Megatron reimport with restored state-quantization metadata and matching logits/gradients.Previous workflow checks exercised the real GDN Bridge training entry point with tiny random weights and mock data, including Hadamard token mode and replay QAT/QAD. KDA checks exercised Megatron layers rather than a Bridge KDA trainer. These do not establish pretrained-model quality recovery, full serving-engine equivalence, or performance.
Before your PR is "Ready for review"
CONTRIBUTING.md?: The optional public vLLM dependency and its license are listed separately. INT8/Hadamard and replay require the documented native fork.Additional Information
Step 3/6. Merge #2519 first, then this example, followed by #2541. This example trains state quantization; prefill GEMM quantization and approximate inverse remain follow-up work. Rebased onto #2519 at
03f6076dd8; the review diff contains only this example (8 files, 987 added lines).