Biography
- I am a Senior Researcher at Tencent Hunyuan. My research focuses on CUA/Code/Search and long horizon agentic models, reinforcement learning, and synthetic data.
- Previously, I created WizardLM at Microsoft, which contributed the state-of-the-art LLM family WizardLM, WizardCoder, and WizardMath. I also created the widely adopted Evol-Instruct method.
- I received my M.S. degree from the Institute of Computational Linguistics at Peking University, advised by Prof. Houfeng Wang.
- We are hiring research interns. If you have strong experience in LLMs, RL, agents and are excited to work with Hunyuan, feel free to contact me.
News
- [Jul 2026] We introduce TurnOPD, a turn-aware on-policy distillation methodology for efficient long-horizon agent RL training.
- [Jun 2026] We introduce VeriEvol, a verifiable Evol-Instruct methodology for scaling multimodal mathematical reasoning.
- [Mar 2026] Our RL papers RubricBench and Beyond Length Scaling are featured on Hugging Face Daily Papers.
- [Jan 2026] AgentMath is accepted by ICLR 2026.
- [Jan 2026] We release OffSeeker, studying online reinforcement learning for deep research agents.
- [Aug 2025] We release Hunyuan-Large-Vision, which ranks #5 globally on LMArena-Vision and is the #1 VLM of China.
- [May 2025] We release Hunyuan-TurboS, which ranks #7 globally on LMArena-Text and is the #2 LLM of China.
- [Oct 2024] Wizard models achieve 3M+ HF downloads, 9K+ GitHub stars, and #4 globally (#1 open source) on LMSYS Arena.
- [April 2024] Release WizardLM-2, which outperforms GPT-4 on MT-Bench, GPT4-Turbo on AlpacaEval 2.0 and Claude 3 Sonnet on Arena-Hard.
- [June 2023] WizardLM achieves the 1st-rank of the opensource models on Standford AlpacaEval Leaderboard.
- [April 2023] Release WizardLM expertized in following complex instructions. [Github] (Over 9K Stars) [WizardLM Pages] [Huggingface] [Beeboom: 12 Best LLMs in 2024] [华尔街见闻]
- 1 long paper accepted by ICLR 2024!
- 1 long paper about Adversarial Knowledge Stimulated Contrastive Prompting for Few-shot Language Learners accepted by ACL 2023!
- 1 long paper about Self-Supervised Multi-Modal Sequential Recommendation arxiv!
- 1 long paper about Knowledge Stimulated Contrastive Prompting for Low-Resource Stance Detection accepted by EMNLP 2022!
- 1 long paper about Multimodal Dialogue Response Generation accepted by ACL 2022!
Selected Publications [Google Scholar]
(*: Equal contribution, #: The intern I mentored)
Agentic Model
AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
Haipeng Luo, Huawen Feng, Qingfeng Sun, Can Xu, Kai Zheng, Yufei Wang, Tao Yang, Han Hu, Yansong Tang, Di Wang
ICLR 2026TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Yuhang Zhou#, Kai Zheng, Haoling Li, Dengyun Peng, Can Xu, Jingjing Chen
arXiv preprint arXiv:2607.05804VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct
Haoling Li#*, Kai Zheng*, Jie Wu, Can Xu, Qingfeng Sun, Han Hu, Yujiu Yang
arXiv preprint arXiv:2606.23543OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents
Yuhang Zhou#, Kai Zheng, Qiguang Chen, Mengkang Hu, Qingfeng Sun, Can Xu, Jingjing Chen
arXiv preprint arXiv:2601.18467
RLVR & RLHF
RubricBench: Aligning Model-Generated Rubrics with Human Standards
Qiyuan Zhang, Junyi Zhou, Yufei Wang, Fuyuan Lyu, Yidong Ming, Can Xu, Qingfeng Sun, Kai Zheng, Peng Kang, Xue Liu, Chen Ma
ACL 2026Beyond Length Scaling: Synergizing Breadth and Depth for Generative Reward Models
Qiyuan Zhang, Yufei Wang, Tianhe Wu, Can Xu, Qingfeng Sun, Kai Zheng, Xue Liu, Chen Ma
ACL 2026 FindingsWizardLM: Empowering Large Language Models to Follow Complex Instructions
Can Xu*, Qingfeng Sun*, Kai Zheng*, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Daxin Jiang
ICLR 2024
General LLM&VLM
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought
Tencent Hunyuan Team
arXiv preprint arXiv:2505.15431Adversarial Knowledge Stimulated Contrastive Prompting for Few-shot Language Learners
Kai Zheng, Qingfeng Sun, Yaming Yang, Tengchao Lv, Yeyong Pi, Changlin Zhao, Fei Xu, Qi Zhang
ACL 2023 FindingsSelf-Supervised Multi-Modal Sequential Recommendation
Kunzhe Song, Qingfeng Sun, Can Xu, Kai Zheng, Yaming Yang
arXiv preprint arXiv:2304.13277Knowledge Stimulated Contrastive Prompting for Low-Resource Stance Detection
Kai Zheng, Qingfeng Sun, Yaming Yang, Fei Xu
EMNLP 2022 FindingsMultimodal Dialogue Response Generation
Qingfeng Sun, Yujing Wang, Can Xu, Kai Zheng, Yaming Yang, Huang Hu, Fei Xu, Jessica Zhang, Xiubo Geng, Daxin Jiang
ACL 2022
Open-source Projects
- WizardLM: WizardLM, WizardCoder, WizardMath, and Evol-Instruct.
- WizardLM-2: strong open instruction-following models from the WizardLM family.
- Evol-Instruct: automatic instruction evolution for building complex instruction-following data.
- Llama-X: open academic research to achieve sota LLMs.
- OffSeeker: online reinforcement learning for deep research agents.
- VeriEvol: verifiable Evol-Instruct method vlm&&llm.
Experiences
- Now, Senior Researcher, Tencent Hunyuan X.
- Previously, Research Scientist, Microsoft AI.
- ERNIE-LLM team Baidu(百度文心), Research Intern.
- M.S., Institute of Computational Linguistics, Peking University.
Academic Services
Program Committee for
- ICLR 2025
- NeurIPS 2024
- NAACL 2024
- COLM 2024
- EMNLP 2022, 2023, 2024
- ACL 2022, 2023
- KDD 2022, 2023
