I am a Principal Architect at Baidu, leading a pre-training data team for the ERNIE large language model, with a focus on long-context data curation and strategy for pre-training and mid-training. Previously, I led the large model team at PaddlePaddle, working across infrastructure, development tools, algorithms, and applications for the ERNIE model series.
I have led over 10 open-source projects at Baidu that have collectively earned more than 200,000 GitHub stars, including PaddleOCR, ERNIE, PaddleFormers, and PaddleX. These projects power critical AI applications worldwide, spanning multilingual document understanding, large language model training, and large-scale deployment.
My technical expertise spans large language model pre-training and data strategy, particularly for long-context modeling, as well as computer vision, vision-language models, and autonomous driving. I hold a Ph.D. from Nanyang Technological University, Singapore (2017) and a B.Eng. from Harbin Institute of Technology, China (2013). Before joining Baidu in 2018, I was a Data Scientist at HP Labs Singapore. At Baidu, I also collaborated extensively with the Apollo team on autonomous driving technologies, contributing to Robotaxi and AD 2.0.
We're Hiring!
My team is looking for talented interns and full-time engineers interested in large language model pre-training, especially long-context data curation, data strategy, data quality, and large-scale pre-training data pipelines. Candidates with experience in large-scale data processing, machine learning, natural language processing, or distributed systems are especially welcome. Please feel free to send your resume to my email.