Conversation
Signed-off-by: wdykas <wdykas@nvidia.com>
wdykas
marked this pull request as ready for review
September 30, 2026 15:19
Contributor
Author
|
/ok to test ab7f0e1 |
kvareddy
approved these changes
Sep 30, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Hybrid models can skip Mamba recurrent state updates when a continuation prefill chunk matches cached attention KV blocks, producing predictions from a state that represents the wrong token history.
Attention KV blocks and Mamba snapshots have independent retention. If a Mamba snapshot is evicted while its KV blocks remain, the first prompt chunk correctly recomputes tokens until a usable state boundary. On subsequent block-aligned chunks,
_compute_prefix_matchfalls through to the attention-only skip calculation: its hybrid fallback is guarded byfinished == 0.However,
add_requestrestores Mamba state only on the first chunk. A continuation carries the live state from the previous chunk, so matching KV blocks alone cannot justify advancing past additional recurrent transitions. Restoring an earlier state on the first chunk does not make later skips safe.Fix
Apply the hybrid zero-skip fallback to continuation chunks as well:
First-chunk skipping still uses a matching Mamba snapshot. Matched attention KV blocks remain shared, and attention-only continuation skipping is unchanged. Hybrid continuations recompute the transitions necessary to keep the live recurrent state aligned with the logical prompt position.
Controlled reproduction
A four-GPU test on the affected deployment revision (
d37db1077cb77ec08fb846485accd92ca39cd639) held model weights and the same 49,326-token prompt fixed. No weight update occurred between conditions.The failing trace skipped four 8,192-token chunks and one 6,912-token region without restoring Mamba state. The most likely next token changed: the cold top token's probability fell from 99.95% to 5.72%, while another token received 79.03%. The guard restored the entire cold next-token distribution exactly in this test.
This demonstrates an inference correctness defect. Its contribution to the RL reasoning-length divergence that prompted the investigation is still unproven.
Validation
DynamicInferenceContextimplementation with CPU tensors, through a one-processtorch.distributed.runinvocation; GPUs are hidden and unrelated distributed suite fixtures are omitted. The full GPU unit suite has not been run.tools/autoformat.shcompleted in check mode. Black, isort, Pylint, and Ruff passed. Its non-gating mypy step reported missing dependency stubs and type errors; this was not a clean mypy run.Use fresh caches when adopting the fix: an earlier faulty request may already have stored inconsistent Mamba state or suffix KV.
Pre-checks
No linked issue. This PR does not add or change a GPU kernel or public API.