MonkeyOCRv2: A Visual-Text Foundation Model
for Document AI

Yuliang Liu1, Zhang Li1*, Ziyang Zhang1*, Shuo Zhang1*, Qiang Liu2, Jiajun Song1, Zidun Guo1, Xinhan Wang1, Handong Zheng1, Yang Liu1, Dongliang Luo1, Zhiyin Ma1, Jiarui Zhang2, Xiang Bai1
1Huazhong University of Science and Technology    2Kingsoft Office  ·  *Equal contribution
MonkeyOCRv2 Overview

MonkeyOCRv2: a document-native vision encoder, pretraining corpus, and downstream document AI models.

Abstract

Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks — text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation — where replacing the original encoders with MonkeyOCRv2 consistently improves performance. Finally, kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, surpassing the previous best 3B model by 2.8% absolute with a vision encoder roughly 11× smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings.

113M
Pretraining Images
17
Languages
0.7B
Parsing Model
#1
MDPBench Open-source

News

2026.08.17 ⚡ Released training and evaluation instructions for Recognition and Overlapping Text Segmentation.
2026.07.24 ⚡ Released MonkeyOCRv2-B-Parsing-DFlash, enabling vLLM serving with DFlash for up to 2× faster inference.
2026.07.22 🏆 MonkeyOCRv2-B-Parsing ranks #1 among open-source models on the official MDPBench Leaderboard (83.3 overall across 17 languages).
2026.07.21 📦 Released MonkeyDoc v2, an open multilingual corpus for document-oriented pretraining.
2026.07.14 🚀 Released MonkeyOCRv2: vision encoder, MonkeyOCRv2-Parsing for multilingual document parsing, and MonkeyOCRv2-Und for efficient document understanding.

Multilingual Parsing Examples

Robust document parsing on real-world photographed documents — under challenging conditions such as underexposure, blur, wrinkles, bends and shadows — across 17 languages. Select an example to compare the captured photo (left) with the MonkeyOCRv2-B-Parsing output (right); detected image blocks are rendered inline in the parsing result.

Chinese (Simplified) Financial Report

Input photo

ZH Chinese (Simplified)
Results (.zip)
Chinese (Traditional) Exam Paper

Input photo

ZH-T Chinese (Traditional)
Results (.zip)
English Book

Input photo

EN English
Results (.zip)
Arabic Academic Paper

Input photo

AR Arabic
Results (.zip)
German Academic Paper

Input photo

DE German
Results (.zip)
Spanish Financial Report

Input photo

ES Spanish
Results (.zip)
French Academic Paper

Input photo

FR French
Results (.zip)
Hindi Magazine

Input photo

HI Hindi
Results (.zip)
Indonesian Slide

Input photo

ID Indonesian
Results (.zip)
Italian Academic Paper

Input photo

IT Italian
Results (.zip)
Japanese Academic Paper

Input photo

JP Japanese
Results (.zip)
Korean Book

Input photo

KO Korean
Results (.zip)
Portuguese Academic Paper

Input photo

PT Portuguese
Results (.zip)
Russian Academic Paper

Input photo

RU Russian
Results (.zip)
Thai Book

Input photo

TH Thai
Results (.zip)
Vietnamese Exam Paper

Input photo

VI Vietnamese
Results (.zip)

Text Detection Examples

Arbitrary-shaped scene text detection with the MonkeyOCRv2-AS encoder — accurate polygon localization of curved, vertical and artistic text in the wild.

Overlapping Text Segmentation Examples

Overlapping text segmentation with the MonkeyOCRv2-AS encoder on the MOT dataset — disentangling severely overlapping text instances (distinct colors denote different text layers and overlapping regions).

Model Zoo

1. Vision Encoder

ModelBackboneParamsPretraining
Resolution
Applicable TasksCheckpoint Link
MonkeyOCRv2-SViT-S28M1280*28*28Recognition / Parsing / Understanding 🤗HuggingFace
🤖ModelScope
MonkeyOCRv2-BViT-B113M1280*28*28Recognition / Parsing / Understanding 🤗HuggingFace
🤖ModelScope
MonkeyOCRv2-ASViTAEv2-S21M1760*32*32Detection / Segmentation 🤗HuggingFace
🤖ModelScope

2. Document Parsing Model (results on MDPBench)

ModelLinkTotal ParamsViTLLMAllDigit.Photo. Latin Avg.DEENESFRIDITNLPTVI Non-Latin Avg.ARHIJPKORUTHZHZH-T
MonkeyOCRv2-S-Parsing HuggingFace ModelScope 0.6B0.03B0.6B 82.587.980.783.287.383.676.873.685.487.285.587.481.981.791.287.169.988.778.079.884.474.7
MonkeyOCRv2-B-Parsing HuggingFace ModelScope 0.7B0.1B0.6B 83.388.181.784.287.784.575.278.486.588.686.187.983.282.190.787.271.987.680.180.883.675.3

DFlash-accelerated checkpoint for 2× faster inference: MonkeyOCRv2-B-Parsing-DFlash.

3. Document Understanding Model

ModelLinkTotal ParamsOverallDocVQAInfoVQADFKLCWTQChartQADT-VQAOCRBench
MonkeyOCRv2-S-Und HuggingFace ModelScope 1.7B55.979.344.565.137.643.062.063.152.2
MonkeyOCRv2-B-Und HuggingFace ModelScope 1.8B57.279.346.365.838.243.262.064.358.1

Evaluation Results

Document parsing results on MDPBench, a comprehensive multilingual benchmark for real-world document parsing. Bold: best among open-source; underline: second best.

ModelTotal ParamsViTLLMAllDigit.Photo. Latin Avg.DEENESFRIDITNLPTVI Non-Latin Avg.ARHIJPKORUTHZHZH-T
Closed-source VLMs
ChatGPT-5.2-2025-12-11---68.685.663.075.270.879.471.460.077.778.571.685.082.161.164.963.455.865.460.763.856.358.7
Claude-Sonnet-4.6---73.185.069.379.279.880.672.866.582.383.376.788.083.166.267.871.763.464.370.865.261.365.1
Doubao-2.0-pro---74.278.972.875.782.874.469.070.073.382.069.983.476.572.581.375.765.874.763.371.971.975.2
Gemini-3-pro---86.490.485.188.491.290.683.482.791.591.687.791.485.984.189.490.474.885.584.980.685.182.1
Open-source VLMs
InternVL-3.5-8B8.3B0.3B8B42.759.737.053.439.864.247.542.753.860.652.263.257.030.68.29.045.630.326.110.855.359.3
MinerU-2.51.2B0.7B0.5B46.361.940.863.068.878.454.757.367.575.260.458.846.027.41.39.039.114.78.611.372.962.2
DeepSeek-OCR3.4B0.4B3B51.880.742.254.555.058.344.143.260.969.352.453.054.148.956.952.249.128.236.249.459.759.2
MonkeyOCR-pro-3B3.7B0.7B3B52.268.047.065.171.777.955.962.166.274.566.371.140.237.64.64.255.260.542.69.172.252.4
Nanonets-OCR-s4.7B0.7B4B63.778.858.771.375.178.561.262.570.381.069.675.967.555.059.561.855.951.243.539.567.461.5
Nanonets-OCR2-3B3.7B0.7B3B64.279.259.371.476.776.461.866.168.478.574.174.266.056.260.259.252.154.745.544.668.365.1
Qwen3.5-Instruct-9B9.7B0.7B9B65.774.862.772.572.872.072.064.466.277.674.579.174.058.253.456.255.760.354.756.760.867.5
GLM-OCR0.9B0.4B0.5B67.377.963.778.782.784.575.876.279.782.880.277.469.254.321.739.665.561.264.227.478.576.7
Qwen3-VL-Instruct-8B8.3B0.3B8B68.378.465.073.673.771.469.366.268.579.178.382.273.462.563.158.459.961.957.962.062.673.8
HunyuanOCR1B0.4B0.6B68.380.264.372.475.073.163.066.169.980.361.481.980.663.768.373.155.668.952.260.766.864.2
PaddleOCR-VL0.9B0.6B0.3B69.687.663.672.178.279.362.966.077.478.467.972.066.666.765.868.459.977.856.957.878.268.5
olmOCR27.7B0.7B7B70.479.967.276.775.777.372.568.970.681.072.088.084.063.359.060.859.470.665.859.268.663.4
MinerU-2.5-Pro1.2B0.7B0.5B71.086.266.174.678.379.563.467.478.079.772.178.674.267.056.672.259.177.662.661.876.569.7
PaddleOCR-VL-1.60.9B0.6B0.3B75.082.872.678.084.179.769.274.881.682.074.776.479.371.669.465.668.782.570.762.378.075.7
HunyuanOCR-1.51B0.4B0.6B76.886.273.679.779.680.474.270.081.584.578.486.482.473.571.871.665.575.767.477.780.877.2
Kimi-K2.51T0.4B1T77.585.075.081.685.986.272.771.080.686.677.487.686.272.975.874.572.570.961.867.081.778.6
PaddleOCR-VL-1.50.9B0.6B0.3B78.387.475.281.284.883.075.778.183.985.280.680.278.974.971.367.769.586.076.068.484.875.7
chandra-ocr-25.3B0.5B4.8B79.787.877.182.786.686.569.770.384.687.482.790.785.676.478.281.168.880.374.078.573.876.3
dots.mocr3B1.2B1.8B80.590.577.281.782.687.471.370.184.589.383.286.879.979.283.383.675.078.771.277.984.679.6
MonkeyOCRv2-S-Parsing 🤗0.6B0.03B0.6B82.587.980.783.287.383.676.873.685.487.285.587.481.981.791.287.169.988.778.079.884.474.7
MonkeyOCRv2-B-Parsing 🤗0.7B0.1B0.6B83.388.181.784.287.784.575.278.486.588.686.187.983.282.190.787.271.987.680.180.883.675.3

Document understanding performance comparison across different vision foundation models. The evaluation benchmarks are selected following TextMonkey and DT-VQA.

ModelParamsOverallDocVQAInfoVQADFKLCWTQChartQADT-VQAOCRBench
CLIP-B86M16.020.124.22.313.812.822.222.310.6
SigLIP2-B93M24.927.023.53.116.717.435.041.535.1
RADIOv2.5-B98M37.560.331.229.930.429.751.144.223.1
OpenVision-B87M44.063.330.719.833.131.158.362.652.9
DINOv3-B86M16.126.520.85.613.214.028.915.83.9
SAM-B90M25.237.822.24.717.517.646.533.321.9
SAM2-B69M22.332.521.92.715.816.640.230.318.4
oCLIP24M12.414.819.51.47.411.417.919.27.4
DiT86M8.911.320.90.95.29.912.09.21.9
MonkeyOCRv2-S* 🤗Link28M55.979.344.565.137.643.062.063.152.2
MonkeyOCRv2-B* 🤗Link113M57.279.346.365.838.243.262.064.358.1

1. Text Recognition

Text recognition results on Common Benchmarks, Union14M-Benchmark, OST, and Chinese Benchmarks. We follow the training and evaluation protocols of OpenOCR.

ModelOverall Union14M-Benchmark Chinese Benchmarks Occlusion SceneText
AvgArtisticContextlessCurveGeneralMulti-OrientedMulti-WordsSaliency AvgSceneWebDocumentHandwriting
ABINet73.775.771.774.780.479.869.076.877.670.366.663.298.253.175.0
MAERec81.685.279.084.289.184.687.185.986.383.184.483.099.565.676.4
CPPD80.481.976.582.986.283.578.781.983.581.782.782.499.462.379.6
IGTR-AR81.084.977.082.490.484.491.284.084.781.782.081.799.563.876.3
SMTR80.485.076.883.989.183.787.789.384.682.783.483.099.365.173.5
SVTRv283.186.179.386.190.685.189.086.786.283.383.583.399.567.080.0
 
CRNN (ResNet)58.749.251.262.348.168.213.060.441.468.863.868.297.046.158.0
CRNN (MonkeyOCRv2-S)67.365.263.773.071.174.528.672.173.474.273.074.996.951.862.4
PARSeq (ViT)82.284.376.583.487.684.988.884.384.482.484.282.899.563.079.9
PARSeq (MonkeyOCRv2-S)84.387.678.686.492.185.493.988.787.783.784.683.299.567.381.5

2. Formula Recognition

Formula recognition results on OmniDocBench 1.6, MathWriting, and UniMER-Test.

ModelParams OverallOmniDocBench 1.6MathWriting SPECPEHWESCE
CDMExpRateCDMExpRateCDMExpRate CDMExpRateCDMExpRateCDMExpRateCDMExpRate
Pix2tex25.5M53.823.369.427.00.40.096.272.464.97.124.50.667.632.8
Texify312M67.340.476.546.426.62.098.591.070.428.252.723.679.351.3
UniMERNet-B325M89.564.590.459.563.812.399.193.396.080.594.064.393.777.0
UniMERNet-S202M89.864.090.159.165.912.799.193.495.977.793.763.994.176.9
 
UniMERNet-T (Swin)107M89.461.889.957.265.612.999.192.394.969.993.361.993.876.6
UniMERNet-T (MonkeyOCRv2-S)110M90.966.490.861.170.816.299.293.896.179.294.369.594.078.6

1. Text Detection

Text detection results on Total-Text, CTW1500, ICDAR2015 and ArT. We follow the training and evaluation protocols of MMOCR and DPText-DETR.

Text Detection Results

2. Document Tampering Detection

Document tampering detection results on the DocTamper benchmark.

MethodParams OverallDocTamper-TestDocTamper-FCDDocTamper-SCD
IoUFIoUPRFIoUPRFIoUPRF
PSCC-Net5M13.731.317.025.083.039.013.019.082.030.011.015.083.025.0
UperNet67M49.354.070.066.060.062.030.057.035.043.048.057.058.057.0
CAT-Net114M67.371.078.075.069.072.066.085.070.076.058.065.065.065.0
Swin-UPer81M66.771.779.075.072.073.064.080.070.075.057.066.068.067.0
SegFormer85M70.374.081.077.074.075.069.082.074.078.061.068.070.069.0
Mask2Former69M69.778.084.082.083.082.066.081.075.078.059.070.079.074.0
ConvNext122M69.775.384.081.078.079.062.076.071.074.063.071.074.073.0
ConvNextV2121M72.777.786.082.079.081.065.079.075.077.067.074.076.075.0
InternImage128M73.377.784.081.077.079.072.083.079.081.064.073.074.073.0
ASC-Former80M68.280.881.591.887.889.861.374.977.176.061.978.075.076.5
DTD66M77.079.784.081.077.079.079.088.082.085.068.075.076.075.0
 
FFDN* (ViTAEv2)69M70.782.769.476.288.782.079.092.584.488.363.679.176.577.8
FFDN (MonkeyOCRv2-AS)71M78.287.587.494.891.893.379.990.487.488.967.281.079.880.4

* denotes models trained with the ViTAEv2 pretrained by DeepSolo.

3. Overlapping Text Segmentation

Overlapping text segmentation results on the MOT dataset.

ModelmIoUTextIoUOccIoUOccdIoUOv
Unet62.280.265.740.7
Deeplab v367.983.271.249.3
OCRNet65.881.068.547.8
Segformer69.083.674.149.3
MaskFormer68.483.570.351.4
TexRNet68.984.273.249.3
EAFormer69.183.874.250.5
WASNet70.884.874.453.1
 
Mask2Former (ResNet)70.384.773.352.8
Mask2Former (MonkeyOCRv2-AS)76.688.683.457.7
MOTS (ResNet)72.685.277.554.9
MOTS (MonkeyOCRv2-AS)76.988.682.659.4

MonkeyDoc v2

MonkeyDoc v2 is currently the largest document image pre-training image-text pair dataset, comprising 113 million document images across 17 languages. So far, 52M synthetic samples and 41M real-world samples have been released, with more being organized progressively. Annotations are produced by a multi-expert labeling toolchain: structure detection & reading order by dots.mocr, content recognition by three expert models (dots.mocr, PaddleOCR-VL, Qwen3-VL), block-level agreement filtering, and page-level quality control verified by Qwen3 / Qwen3-VL.

Citation

If you use any part of this release — the encoders, MonkeyOCRv2-Parsing, MonkeyOCRv2-Und, MDPBench, or MonkeyDoc v2 — please cite:

@article{liu2026monkeyocrv2,
  title   = {MonkeyOCRv2: A Visual-Text Foundation Model for Document AI},
  author  = {Liu, Yuliang and Li, Zhang and Zhang, Ziyang and Zhang, Shuo and
             Liu, Qiang and Song, Jiajun and Guo, Zidun and Wang, Xinhan and
             Zheng, Handong and Liu, Yang and Luo, Dongliang and Ma, Zhiyin and
             Zhang, Jiarui and Bai, Xiang},
  journal = {arXiv preprint arXiv:2607.11562},
  year    = {2026}
}