Skip to content

Fix PaddleOCR-VL eager attention OOM with bounded query chunks - #5209

Closed
boin wants to merge 4 commits into
PaddlePaddle:developfrom
boin:fix/bounded-siglip-eager-attention
Closed

boin wants to merge 4 commits into
PaddlePaddle:developfrom
boin:fix/bounded-siglip-eager-attention

Conversation

@boin

@boin boin commented Sep 17, 2026 •

Copy link
Copy Markdown

Closed after checking the complete, unmodified latest stable pipeline (PaddleX 3.7.2 / PaddleOCR 3.7.0 / PaddleOCR-VL 1.6) on an 11 GiB RTX 2080 Ti. The synthetic blank A4 succeeds repeatedly with stable post-cleanup memory and no eager vision-attention calls. The retained older 3.3.11 runtime failure and inspection of the current kernel alone do not establish the same defect in the current default pipeline. No upstream action is requested.

The original proposal below described a compatibility workaround validated on an older offline runtime; it is not evidence of a current-default regression.


PaddleOCR-VL's eager vision attention allocates full [batch, heads, queries, keys] score and FP32 softmax tensors. On GPUs that cannot use its SDPA path (for example T4), even a generated blank A4 PDF can exhaust a 16 GiB device. In a PaddleX 3.3.11 / PaddleOCR 3.3.2 / Paddle 3.2.1 native serving reproduction, the failing query/key shape was [1, 16, 10080, 72]; each FP32 score matrix is about 6.1 GiB. The current default branch retains the same dense fallback.

This change processes at most 256 query positions at a time during no-grad inference, attending to all keys in every chunk. It preserves image resolution, model weights, token budgets and attention semantics. Query-dependent masks are sliced along the query axis; broadcast masks remain broadcast. The caller already rejects output_attentions=True and discards weights, so chunks never rebuild the full weight matrix. Gradient-enabled execution and active training dropout retain the existing dense path. Zero-dropout generation is handled even when an older runtime leaves module.training=True.

Validation:

  • Five tests (plus eight subtests) pass against this branch with Paddle 3.2.1, including CPU FP32, CUDA FP16, broadcast/query masks, tail chunks, bounded softmax shapes, autograd/dropout preservation and the no-grad/training-flag case.
  • A compatibility backport around the same eager kernel was exercised on T4 with the retained 3.3.11 offline model image. A generated blank A4 PDF changed from OOM after about 13 seconds to HTTP 200 after 10.3 seconds, with readiness restored after request cleanup. A private nine-page regression also completed; no private inputs are included here.
  • The full latest PaddleX model/server stack has not been deployed for this check; the unit tests use this branch, while the real serving measurement uses the compatibility backport.

Minimal synthetic serving reproducer (PyMuPDF + requests; start the native PaddleOCR-VL service on a T4):

import base64
import fitz
import requests

with fitz.open() as pdf:
    pdf.new_page(width=595, height=842)
    data = pdf.tobytes()

response = requests.post(
    "http://127.0.0.1:8080/layout-parsing",
    json={
        "file": base64.b64encode(data).decode(),
        "fileType": 0,
        "restructurePages": False,
        "returnMarkdownImages": False,
        "visualize": False,
    },
    timeout=120,
)
print(response.status_code)

@boin boin closed this Sep 17, 2026
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
3 out of 4 committers have signed the CLA.

✅ zhangyubo0722
✅ Bobholamovic
✅ zhang-prog
❌ boin
You have signed the CLA already but the status is still pending? Let us recheck it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants