Repository navigation
Add opt-in width_scale / width_limit to text recognition resize - #5215
Ahmed-Sinkeat wants to merge 1 commit into
Conversation
Dense connected scripts such as Arabic get too few CTC time steps per character when the recognition input keeps only the aspect-ratio width. OCRReisizeNormImg gains width_scale (multiplier, default 1.0) and width_limit (integer cap, default None = the existing 3200), threaded through the Paddle predictor and the OCR pipeline's TextRecognition config. With the defaults, preprocessed tensors, text, scores and word boxes are identical to before. Word-box placement uses the capped width. The Transformers engine raises if either option is set, and static input_shape ignores them. Refs PaddlePaddle/PaddleOCR#18349
|
Additional check on a standard Arabic benchmark: KITAB-Bench
95% paired bootstrap interval for the CER change: −61.6 to −54.7 points. The stock errors are almost all deletions (24,841 deletions vs 94 substitutions): these lines have a median aspect ratio of 16:1, so the default resize gives about 1.2 CTC steps per character and characters collapse. The option restores the full output length. CER is on whitespace-normalised text with tashkeel/tatweel removed (the references have no tashkeel). These are pre-cropped lines, so the number is not comparable with full-page results. Details and script: https://github.com/Ahmed-Sinkeat/ocrarabic/tree/master/experiments/phase6_kitab |
Refs PaddlePaddle/PaddleOCR#18349 (maintainers invited this PR there).
What
Two optional parameters on text recognition, both defaulting to the current behavior:
width_scale(float > 0, default1.0): multiplier on the aspect-ratio-derived resize width.width_limit(int, defaultNone= the existing 3200): maximum resized width.They are set under
TextRecognitionin the pipeline config (OCRReisizeNormImg→ Paddle predictor → OCR pipeline). Word-box placement uses the capped width so batched crops stay consistent. Fixedinput_shapeignores them; the Transformers recognition engine raises if either is set, instead of silently doing nothing. Fractionalwidth_limitand non-finitewidth_scaleare rejected.Why
For dense connected scripts the stock resize leaves the CTC head too few time steps per character (CTC emits
input_width / 8steps). Details and the sweep are in the issue.Evidence
Defaults unchanged. Against installed PaddleX 3.7.2, with
width_scale=1.0, width_limit=None: preprocessed tensors are identical on 215/215 inputs (204 real Arabic line crops plus edge cases: padded, squashed, 1×1, 64 000 px wide), and text, scores and word boxes are identical through the full predictor on 20/20 crops for each ofarabic_PP-OCRv5_mobile_rec,arabic_PP-OCRv3_mobile_rec,en_PP-OCRv5_mobile_recandPP-OCRv5_server_rec.Arabic gain (
width_scale=2.0, width_limit=1280, two real scanned books, 10 hand-verified pages each, nothing tuned on them):arabic_PP-OCRv5_mobile_recarabic_PP-OCRv5_mobile_recarabic_PP-OCRv3_mobile_recarabic_PP-OCRv3_mobile_recLatin, not a general win. 200 IIIT5K test words (chosen by SHA-256 order before running; word-level, not book lines), lowercase alphanumeric exact match:
en_PP-OCRv5_mobile_recPP-OCRv5_server_recSo this should stay opt-in and documented as script/model-specific; the docs added here say to measure first.
Limits.
arabic_PP-OCRv5_server_recnamed in the issue is not published, soarabic_PP-OCRv3_mobile_recandPP-OCRv5_server_recstand in. The benchmark placeholder in the issue was unanswered, so I used IIIT5K. The end-to-end page numbers cover two books; the cross-model runs reuse their cached boxes. Scripts and raw numbers: https://github.com/Ahmed-Sinkeat/ocrarabic/tree/master/experiments/phase5_upstream_prTests
python -m unittest tests.test_text_rec_width(5 tests; this is the first Python test undertests/, happy to move or drop it). Pre-commit checks pass locally: black, isort, flake8, license header, import check, py3.8 compatibility.Docs: new section 4.3 in
OCR.md/OCR.en.md. Not included: the PaddleOCR-sidetext_rec_width_scale/text_rec_width_limitarguments — I'll follow up there once the shape is agreed here.