Repository navigation
Conversation
|
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closed after checking the complete, unmodified latest stable pipeline (PaddleX 3.7.2 / PaddleOCR 3.7.0 / PaddleOCR-VL 1.6) on an 11 GiB RTX 2080 Ti. The synthetic blank A4 succeeds repeatedly with stable post-cleanup memory and no eager vision-attention calls. The retained older 3.3.11 runtime failure and inspection of the current kernel alone do not establish the same defect in the current default pipeline. No upstream action is requested.
The original proposal below described a compatibility workaround validated on an older offline runtime; it is not evidence of a current-default regression.
PaddleOCR-VL's eager vision attention allocates full
[batch, heads, queries, keys]score and FP32 softmax tensors. On GPUs that cannot use its SDPA path (for example T4), even a generated blank A4 PDF can exhaust a 16 GiB device. In a PaddleX 3.3.11 / PaddleOCR 3.3.2 / Paddle 3.2.1 native serving reproduction, the failing query/key shape was[1, 16, 10080, 72]; each FP32 score matrix is about 6.1 GiB. The current default branch retains the same dense fallback.This change processes at most 256 query positions at a time during no-grad inference, attending to all keys in every chunk. It preserves image resolution, model weights, token budgets and attention semantics. Query-dependent masks are sliced along the query axis; broadcast masks remain broadcast. The caller already rejects
output_attentions=Trueand discards weights, so chunks never rebuild the full weight matrix. Gradient-enabled execution and active training dropout retain the existing dense path. Zero-dropout generation is handled even when an older runtime leavesmodule.training=True.Validation:
Minimal synthetic serving reproducer (PyMuPDF + requests; start the native PaddleOCR-VL service on a T4):