About Me

I am a Principal Architect at Baidu, leading a pre-training data team for the ERNIE large language model, with a focus on long-context data curation and strategy for pre-training and mid-training. Previously, I led the large model team at PaddlePaddle, working across infrastructure, development tools, algorithms, and applications for the ERNIE model series.

I have led over 10 open-source projects at Baidu that have collectively earned more than 200,000 GitHub stars, including PaddleOCR, ERNIE, PaddleFormers, and PaddleX. These projects power critical AI applications worldwide, spanning multilingual document understanding, large language model training, and large-scale deployment.

My technical expertise spans large language model pre-training and data strategy, particularly for long-context modeling, as well as computer vision, vision-language models, and autonomous driving. I hold a Ph.D. from Nanyang Technological University, Singapore (2017) and a B.Eng. from Harbin Institute of Technology, China (2013). Before joining Baidu in 2018, I was a Data Scientist at HP Labs Singapore. At Baidu, I also collaborated extensively with the Apollo team on autonomous driving technologies, contributing to Robotaxi and AD 2.0.

🔥

We're Hiring!

My team is looking for talented interns and full-time engineers interested in large language model pre-training, especially long-context data curation, data strategy, data quality, and large-scale pre-training data pipelines. Candidates with experience in large-scale data processing, machine learning, natural language processing, or distributed systems are especially welcome. Please feel free to send your resume to my email.

News

  • Jul 2026 Our three works, RISTER, Real5-OmniDocBench, and RT-DocLayout, have been accepted by ECCV 2026!
  • Jun 2026 We open-sourced PP-OCRv6, a three-tier OCR model family (1.5M/7.7M/34.5M) that surpasses mainstream VLMs with major accuracy gains, unified support for 50 languages, and faster inference across CPU, Apple M4, and GPU deployments.
  • May 2026 We open-sourced PaddleOCR-VL-1.6, achieving SOTA on OmniDocBench v1.6 with over 96.3% accuracy, significantly improving table, ancient document, rare character, and seal recognition while maintaining full compatibility with v1.5.
  • May 2026 A collaboration with Renmin University of China on SUDER has been accepted by KDD 2026!
  • May 2026 I gave a talk at VALSE 2026 and CCIG 2026 on "PaddleOCR Multimodal Document Intelligent Parsing".
  • Mar 2026 Our two works, PP-OCRv5 and PaddleOCR-VL, have been accepted by CVPR 2026!

Key Open-Source Projects

PaddleOCR

A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages with ultra-lightweight system, powering popular projects like Umi-OCR, OmniParser, MinerU, and RAGFlow.

PaddleFormers

An easy-to-use library of pre-trained large language models based on PaddlePaddle. Implements 4D parallel strategies through unified Trainer API, supporting SFT/DPO paradigms and integrating PEFT, MergeKit, and Quantization APIs for efficient LLM development.

ERNIE

Official repository for ERNIE 4.5, featuring both MoE and Dense models across LLMs and multimodal architectures. Provides end-to-end development pipeline for training, compression, and inference, supporting full-cycle industrial deployment.

PaddleX

All-in-One low-code development tool for AI models built on PaddlePaddle. Integrates over 200 ready-to-use pre-trained models covering OCR, object detection, and time series forecasting, supporting complete workflow from training to deployment.

PaddleDetection

End-to-end object detection toolkit providing 30+ algorithms and 250+ pre-trained models. Supports object detection, instance segmentation, keypoint detection, and multiple object tracking with complete pipeline from development to deployment.

PaddleNLP

Easy-to-use NLP library with 45+ architectures and 500+ pretrained models. Supports wide-range of tasks from research to industrial applications including Neural Search, Question Answering, Information Extraction, and Sentiment Analysis.

PaddleSeg

Easy-to-use image segmentation library providing 45+ models and 150+ pre-trained models. Supports Semantic Segmentation, Interactive Segmentation, Panoptic Segmentation, Image Matting, and 3D Segmentation with complete flow from labeling to deployment.

PaddleClas

Comprehensive toolkit for image classification and recognition. Encompasses advanced algorithms including PP-HGNet, PP-LCNetv2, and PP-LCNet, providing 35 series with 164 ImageNet pre-trained models for industrial and academic applications.

Publications

Conference Papers

[C26] Embedding Rotation Invariance for Provable Multi-Oriented Scene Text Recognition ECCV 2026

Zhibin Ma, Pengwen Dai, Yi Liu, Xugong Qin, Chenyun Yu, Xiaochun Cao

[C25] Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild ECCV 2026

Changda Zhou, Ziyue Gao, Xueqing Wang, Tingquan Gao, Cheng Cui, Jing Tang, Yi Liu

[C24] RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild ECCV 2026

Cheng Cui, Tingquan Gao, Xueqing Wang, Changda Zhou, Hongen Liu, Ting Sun, Yubo Zhang, Zelun Zhang, Jiaxuan Liu, Manhui Lin, Yue Zhang, Suyin Liang, Yiqing Xiang, Yi Liu

[C23] SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards KDD 2026

Jixiang Hong, Yiran Zhang, Guanzhong Wang, Yi Liu, Ji-Rong Wen, Rui Yan

[C22] Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing CVPR 2026

Cheng Cui, Ting Sun, Suyin Liang, Tingquan Gao, Zelun Zhang, Jiaxuan Liu, Xueqing Wang, Changda Zhou, Hongen Liu, Manhui Lin, Yue Zhang, Yubo Zhang, Jing Zhang, Jun Zhang, Xing Wei, Yi Liu, Dianhai Yu, Yanjun Ma

[C21] PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks CVPR 2026

Cheng Cui, Yubo Zhang, Ting Sun, Xueqing Wang, Hongen Liu, Manhui Lin, Yue Zhang, Tingquan Gao, Changda Zhou, Jiaxuan Liu, Zelun Zhang, Jing Zhang, Jun Zhang, Yi Liu

[C20] Sortblock: Similarity-Aware Feature Reuse for Diffusion Model AAAI 2026

Hanqi Chen, Xu Zhang, Xiaoliu Guan, Lielin Jiang, Guanzhong Wang, Zeyu Chen, Yi Liu

[C19] SUTrack: Towards Simple and Unified Single Object Tracking AAAI 2025

Xin Chen, Ben Kang, Wanting Geng, Jiawen Zhu, Yi Liu, Dong Wang, Huchuan Lu

[C18] Exploring Enhanced Contextual Information for Video-Level Object Tracking AAAI 2025

Ben Kang, Xin Chen, Simiao Lai, Yang Liu, Yi Liu, Dong Wang

[C17] Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation ICRA 2024

Jiayi Ni, Senqiao Yang, Ran Xu, Jiaming Liu, Xiaoqi Li, Wenyu Jiao, Zehui Chen, Yi Liu, Shanghang Zhang

[C16] DETRs Beat YOLOs on Real-time Object Detection CVPR 2024

Yian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei, Guanzhong Wang, Qingqing Dang, Yi Liu, Jie Chen

[C15] Video4MRI: An Empirical Study on Brain Magnetic Resonance Image Analytics with CNN-Based Video Classification Frameworks ISBI 2023

Yuxuan Zhang, Qingzhong Wang, Jiang Bian, Yi Liu, Yanwu Xu, Dejing Dou, Haoyi Xiong

[C14] Context Matters: Cross-Domain Cell Detection in Histopathology Images via Contextual Regularization MIUA 2023

Ziqi Wen, Qingzhong Wang, Jiang Bian, Xuhong Li, Yi Liu, Haoyi Xiong

[C13] Lightweight Image Super-Resolution with Superpixel Token Interaction ICCV 2023

Aiping Zhang, Wenqi Ren, Yi Liu, Xiaochun Cao

[C12] Towards Efficient 3D Human Motion Prediction using Deformable Transformer-based Adversarial Network ICRA 2022

Hua Yu, Xuanzhe Fan, Yaqing Hou, Yi Liu, Cai Kang, Dongsheng Zhou, Qiang Zhang

[C11] MUSCLE: Multi-task Self-supervised Continual Learning to Pre-train Deep Models for X-Ray Images of Multiple Body Parts MICCAI 2022

Weibin Liao, Haoyi Xiong, Qingzhong Wang, Yan Mo, Xuhong Li, Yi Liu, Zeyu Chen, Siyu Huang, Dejing Dou

[C10] PP-HumanSeg: Connectivity-Aware Portrait Segmentation With a Large-Scale Teleconferencing Video Dataset WACV 2022

Lutao Chu, Yi Liu, Zewu Wu, Shiyu Tang, Guowei Chen, Yuying Hao, Juncai Peng, Zhiliang Yu, Zeyu Chen, Baohua Lai, Haoyi Xiong

[C9] EdgeFlow: Achieving Practical Interactive Segmentation with Edge-Guided Flow ICCV 2021

Yuying Hao, Yi Liu, Zewu Wu, Lin Han, Yizhou Chen, Guowei Chen, Lutao Chu, Shiyu Tang, Zhiliang Yu, Zeyu Chen, Baohua Lai

[C8] A New Reconstruction Method in Gaze Estimation with Natural Head Movement MVA 2017

Yi Liu, Bu-Sung Lee, Andrzej Sluzek, Deepu Rajan, Martin J. McKeown

[C7] Feasibility Analysis of Eye Typing with a Standard Webcam ECCV Workshop 2016

Yi Liu, Bu-Sung Lee, Andrzej Sluzek, Deepu Rajan, Martin J. McKeown

[C6] GazeTry: Swipe Text Typing Using Gaze OzCHI 2015

Yi Liu, Chi Zhang, Chonho Lee, Bu-Sung Lee, Alex Q. Chen

[C5] A Robust Recognition Approach in Eye-Based Dwell-Free Typing IEEE PIC 2015

Yi Liu, Bu-Sung Lee, Martin J. McKeown, Chonho Lee

[C4] Feasibility Analysis and Adaptive Thresholding for Mobile Applications Controlled by EEG Signals EUSIPCO 2015

Chonho Lee, Jiawei Chin, Yi Liu, Bu-Sung Lee, Martin J. McKeown

[C3] A Wavelet Entropy-Based Change Point Detection on Network Traffic: A Case Study of Heartbleed Vulnerability IEEE CCTA 2014

Chonho Lee, Yi Liu, Lim Hui Tan, Wei Goh, Bu-Sung Lee, Chai Kiat Yeo

[C2] A Motion Accuracy Evaluator Based on Body Parts Movement by MapReduce Video Processing IEEE BHI 2014

Chonho Lee, Yoshihiro Terada, Yi Liu, Bu-Sung Lee

[C1] Analysis of Visually Guided Tracking Performance in Parkinson's Disease IEEE e-Health 2014

Yi Liu, Chonho Lee, Bu-Sung Lee, John Keith Robert Stevenson, Martin J. McKeown

Journal Papers

[J7] DSDC-GCN: Decoupled Static-Dynamic Co-Occurrence Graph Convolutional Networks for Skeleton-Based Action Recognition IEEE Transactions on Circuits and Systems for Video Technology 2025

Tianming Zhuang, Zhen Qin, Yi Ding, Zhiguang Qin, Ji Geng, Yi Liu, Kim-Kwang Raymond Choo

[J6] EGAvatar: Efficient GAN Inversion for Generalizable Head Avatar From Few-Shot Images IEEE Transactions on Visualization and Computer Graphics 2025

Hao-Pan Ren, Wei Duan, Wan-Yu Li, Yi Liu, Yu-Dong Guo, Shi-Sheng Huang, Ju-Yong Zhang, Hua Huang

[J5] Forecasting When to Forecast: Accelerating Diffusion Models with Confidence-Gated Taylor Knowledge-Based Systems 2025

Xiaoliu Guan, Lielin Jiang, Hanqi Chen, Xu Zhang, Jiaxing Yan, Guanzhong Wang, Yi Liu, Zetao Zhang, Yu Wu

[J4] MTPret: Improving X-Ray Image Analytics With Multitask Pretraining IEEE Transactions on Artificial Intelligence 2024

Weibin Liao, Haoyi Xiong, Qingzhong Wang, Yi Liu, Zeyu Chen, Qinghua Zheng, Dejing Dou

[J3] Distilling Ensemble of Explanations for Weakly-Supervised Pre-Training of Image Segmentation Models Machine Learning 2023

Xuhong Li, Haoyi Xiong, Yi Liu, Dingfu Zhou, Zeyu Chen, Yaqing Wang, Dejing Dou

[J2] CamType: Assistive Text Entry Using Gaze with an Off-the-Shelf Webcam Machine Vision and Applications 2019

Yi Liu, Bu-Sung Lee, Deepu Rajan, Andrzej Sluzek, Martin J. McKeown

[J1] Robust Eye-Based Dwell-Free Typing International Journal of Human–Computer Interaction 2016

Yi Liu, Bu-Sung Lee, Martin J. McKeown