Skip to content

Add opt-in width_scale / width_limit to text recognition resize - #5215

Open
Ahmed-Sinkeat wants to merge 1 commit into
PaddlePaddle:developfrom
Ahmed-Sinkeat:feat/rec-width-option
Open

Ahmed-Sinkeat wants to merge 1 commit into
PaddlePaddle:developfrom
Ahmed-Sinkeat:feat/rec-width-option

Conversation

@Ahmed-Sinkeat

Copy link
Copy Markdown

Refs PaddlePaddle/PaddleOCR#18349 (maintainers invited this PR there).

What

Two optional parameters on text recognition, both defaulting to the current behavior:

  • width_scale (float > 0, default 1.0): multiplier on the aspect-ratio-derived resize width.
  • width_limit (int, default None = the existing 3200): maximum resized width.

They are set under TextRecognition in the pipeline config (OCRReisizeNormImg → Paddle predictor → OCR pipeline). Word-box placement uses the capped width so batched crops stay consistent. Fixed input_shape ignores them; the Transformers recognition engine raises if either is set, instead of silently doing nothing. Fractional width_limit and non-finite width_scale are rejected.

Why

For dense connected scripts the stock resize leaves the CTC head too few time steps per character (CTC emits input_width / 8 steps). Details and the sweep are in the issue.

Evidence

Defaults unchanged. Against installed PaddleX 3.7.2, with width_scale=1.0, width_limit=None: preprocessed tensors are identical on 215/215 inputs (204 real Arabic line crops plus edge cases: padded, squashed, 1×1, 64 000 px wide), and text, scores and word boxes are identical through the full predictor on 20/20 crops for each of arabic_PP-OCRv5_mobile_rec, arabic_PP-OCRv3_mobile_rec, en_PP-OCRv5_mobile_rec and PP-OCRv5_server_rec.

Arabic gain (width_scale=2.0, width_limit=1280, two real scanned books, 10 hand-verified pages each, nothing tuned on them):

model book stock CER with option
arabic_PP-OCRv5_mobile_rec 1 41.4% 9.2%
arabic_PP-OCRv5_mobile_rec 2 36.8% 12.8%
arabic_PP-OCRv3_mobile_rec 1 37.1% 23.3%
arabic_PP-OCRv3_mobile_rec 2 31.1% 23.8% (95% interval −15.3 to +0.5 points, crosses zero)

Latin, not a general win. 200 IIIT5K test words (chosen by SHA-256 order before running; word-level, not book lines), lowercase alphanumeric exact match:

model stock with option
en_PP-OCRv5_mobile_rec 94.5% 93.0% (4 better / 7 worse; interval −5.0 to +1.5)
PP-OCRv5_server_rec 94.5% 94.5% (6 / 6)

So this should stay opt-in and documented as script/model-specific; the docs added here say to measure first.

Limits. arabic_PP-OCRv5_server_rec named in the issue is not published, so arabic_PP-OCRv3_mobile_rec and PP-OCRv5_server_rec stand in. The benchmark placeholder in the issue was unanswered, so I used IIIT5K. The end-to-end page numbers cover two books; the cross-model runs reuse their cached boxes. Scripts and raw numbers: https://github.com/Ahmed-Sinkeat/ocrarabic/tree/master/experiments/phase5_upstream_pr

Tests

python -m unittest tests.test_text_rec_width (5 tests; this is the first Python test under tests/, happy to move or drop it). Pre-commit checks pass locally: black, isort, flake8, license header, import check, py3.8 compatibility.

Docs: new section 4.3 in OCR.md / OCR.en.md. Not included: the PaddleOCR-side text_rec_width_scale / text_rec_width_limit arguments — I'll follow up there once the shape is agreed here.

Dense connected scripts such as Arabic get too few CTC time steps per
character when the recognition input keeps only the aspect-ratio width.
OCRReisizeNormImg gains width_scale (multiplier, default 1.0) and
width_limit (integer cap, default None = the existing 3200), threaded
through the Paddle predictor and the OCR pipeline's TextRecognition
config. With the defaults, preprocessed tensors, text, scores and word
boxes are identical to before. Word-box placement uses the capped width.
The Transformers engine raises if either option is set, and static
input_shape ignores them.

Refs PaddlePaddle/PaddleOCR#18349
@Ahmed-Sinkeat

Copy link
Copy Markdown
Author

Additional check on a standard Arabic benchmark: KITAB-Bench patsocr (ahmedheakl/arocrbench_patsocr, all 500 printed line images), official arabic_PP-OCRv5_mobile_rec, recognition only, this branch. The setting is the same width_scale=2.0, width_limit=1280 as above, not tuned on this data; the protocol was committed before running.

500 lines stock width_scale=2.0, width_limit=1280
CER 61.6% 3.4%
empty outputs 22 0
output length / reference length 0.39 0.99
lines better / worse 379 / 10

95% paired bootstrap interval for the CER change: −61.6 to −54.7 points. The stock errors are almost all deletions (24,841 deletions vs 94 substitutions): these lines have a median aspect ratio of 16:1, so the default resize gives about 1.2 CTC steps per character and characters collapse. The option restores the full output length.

CER is on whitespace-normalised text with tashkeel/tatweel removed (the references have no tashkeel). These are pre-cropped lines, so the number is not comparable with full-page results. Details and script: https://github.com/Ahmed-Sinkeat/ocrarabic/tree/master/experiments/phase6_kitab

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant