DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks

Jiaxin ZhangDezhi PengChongyu LiuPeirong ZhangLianwen Jin

article2024CVPR46 citations

Introduces DocRes, a unified generalist model that uses dynamic task-specific visual prompts to perform five distinct document restoration tasks—including dewarping, deshadowing, and binarization—while matching or exceeding the performance of specialized single-task systems.

Listen

Document processing systems frequently struggle with low-quality captured images affected by geometric distortions, shadows, uneven lighting, blurriness, and poor contrast. Conventional solutions deploy separate, dedicated machine learning models for each individual defect. This fragmented setup introduces substantial operational complexity, inflates maintenance costs, and misses the synergistic performance benefits of multi-task learning. Developing a unified restoration system has historically been difficult because disparate restoration tasks often share identical input images but require completely different visual outputs.

The article introduces and evaluates DocRes, a unified multi-task model capable of handling five distinct document restoration tasks: dewarping, deshadowing, appearance enhancement, deblurring, and binarization. The objective is to demonstrate that a single generalist model can replace multiple specialized systems and achieve competitive or superior performance without complex architectural changes.

The researchers developed a mechanism called Dynamic Task-Specific Prompts to instruct a single underlying neural network on which restoration task to execute. Unlike prior prompt techniques that rely on static tokens or computationally heavy visual pairs, this approach extracts visual cues directly from the input image, such as document masks, background approximations, Sauvola thresholds, and gradient maps. The researchers integrated these prompts into an off-the-shelf restoration network and trained the unified model on standard synthetic and real-world benchmark datasets across all five tasks, comparing its outputs directly against leading specialized systems.

The evaluation produced several key findings. First, the single DocRes model matched or outperformed specialized, state-of-the-art models across standard benchmarks, achieving top scores in deblurring and competitive metrics in dewarping, shadow removal, and contrast enhancement. Second, the model demonstrated strong multi-task synergy, particularly in binarization, where training alongside other tasks substantially improved accuracy on historical text benchmarks compared to baseline models. Third, DocRes exhibited superior real-world generalization; although trained on scanned historical documents, it successfully cleaned camera-captured modern smartphone documents containing unmodeled lighting and blur artifacts. Finally, the unified model proved highly efficient, requiring only 15.2 million parameters—roughly half to one-quarter the size of many single-task models.

These findings suggest that organizations can dramatically simplify automated document processing pipelines. Replacing a suite of independent models with a single multi-task system reduces infrastructure footprint, eases deployment overhead, and lowers overall maintenance costs. The shared visual representations also provide natural robustness against real-world noise without requiring distinct models for every anticipated distortion.

Stakeholders looking to streamline document processing pipelines should consider piloting unified restoration frameworks. Moving forward, engineering efforts should investigate trainable prompt generators with shared parameters to refine feature extraction, as well as single-pass architectures that can execute compound restorations in one step to prevent sequential error accumulation.

While the results demonstrate high confidence across standard benchmarks, certain trade-offs remain. For tasks like deblurring and shadow removal, the extracted prompt features currently act more as task indicators than performance boosters due to the use of basic heuristic filters. Additionally, applying the model iteratively to handle multiple compounding defects on a single image can lead to error accumulation across stages. Caution is advised when processing heavily degraded documents through multi-step pipelines until single-pass multi-defect models are developed.

  • Paper: Pre-Trained Image Processing Transformer, Hanting Chen et al. (2020). Read IPT first to see an earlier shared-network approach to multiple image-restoration tasks, the generalist restoration problem DocRes specializes for documents.

No sufficiently relevant recommendations were found.

Cover for DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks

Abstract

Document image restoration is a crucial aspect of Document AI systems, as the quality of document images significantly influences the overall performance. Prevailing methods address distinct restoration tasks independently, leading to intricate systems and the incapability to harness the potential synergies of multi-task learning. To overcome this challenge, we propose DocRes, a generalist model that unifies five document image restoration tasks including dewarping, deshadowing, appearance enhancement, deblurring, and binarization. To instruct DocRes to perform various restoration tasks, we propose a novel visual prompt approach called Dynamic Task-Specific Prompt (DTSPrompt). The DTSPrompt for different tasks comprises distinct prior features, which are additional characteristics extracted from the input image. Beyond its role as a cue for task-specific execution, DTSPrompt can also serve as supplementary information to enhance the model's performance. Moreover, DTSPrompt is more flexible than prior visual prompt approaches as it can be seamlessly applied and adapted to inputs with high and variable resolutions. Experimental results demonstrate that DocRes achieves competitive or superior performance compared to existing state-of-the-art task-specific models. This underscores the potential of DocRes across a broader spectrum of document image restoration tasks. The source code is publicly available at this https URL

Table of Contents

  • 1 Introduction
  • 2 Related works
  • 2.1 Document Image Restoration
  • 2.2 All-in-one image restoration
  • 3 Methodology
  • 3.1 Dynamic task-specific prompt
  • 3.2 Prompt fusion and restoration network
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Evaluation metrics
  • 4.3 Implementation details
  • 4.4 Results
  • 5 Discussions and conclusions
  • References
  • 1 Efficiency
  • 2 Additional ablation study
  • 3 Comparison with more SOTA
  • 4 More visualized results
  • 5 Further discussions about DTSPrompt

Citation

MLA
Zhang, J., et al. “DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks”. arXiv, 2024, http://arxiv.org/abs/2405.04408v1.
APA
Zhang, J., Peng, D., Liu, C., Zhang, P., & Jin, L. (2024). DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks. arXiv. http://arxiv.org/abs/2405.04408v1
Chicago
Zhang, J., D. Peng, C. Liu, P. Zhang, and L. Jin. 2024. “DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks”. arXiv. http://arxiv.org/abs/2405.04408v1.
Harvard
Zhang, J. et al. (2024) “DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2405.04408v1.
Vancouver
1. Zhang J, Peng D, Liu C, Zhang P, Jin L (2024) DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks. arXiv

BibTeX

@article{zhang2024docres,
  title = {DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks},
  author = {Zhang, Jiaxin and Peng, Dezhi and Liu, Chongyu and Zhang, Peirong and Jin, Lianwen},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2405.04408v1},
  eprint = {2405.04408}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/