About Me
I am Zihan Wang (王子涵), an incoming M.S. student in Artificial Intelligence at
Tsinghua University (Shenzhen), advised by Prof. Yujiu Yang. I received my B.Eng. in Computer Science and Technology from Xi’an Jiaotong University.
My research centers on NLP / Large Language Models (LLMs), with a focus on building intrinsically motivated, self-evolving, and reasoning-capable agents — particularly around agentic reinforcement learning, self-evolving agents, coding / GUI agents, and environment scaling.
🐈 I am always open to discussion and collaboration, and I am currently looking for internship opportunities. Feel free to reach out via email or WeChat: wishme25
🔥 News
- 2026.09🎓 Incoming M.S. student in AI at Tsinghua University (Shenzhen).
- 2026.08🚀 Released AgentOPSD, a recursive self-distillation method for agentic RL, featured as 🤗 HF Daily Paper #1!
- 2026.07🚀 Released SkillRise on cross-task skill evolution for agentic RL.
- 2026.05🎉 CreativeBench was accepted to ACL 2026!
- 2026.05🔥🔥 Our new work SDAR was released, featured as 🤗 HF Daily Paper #2!
- 2025.05🎉 Alignment for Efficient Tool Calling was accepted to EMNLP 2025 (Main)!
- 2025.05🎉 Reducing Tool Hallucination via Reliability Alignment was accepted to ICML 2025!
📝 Selected Publications
* Equal contribution. † Corresponding author. See my full list on Google Scholar.
🤖 Agentic RL & Skill Evolution

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Jie Wu, Zhengzhou Cai, Yueqing Sun, Ziang Ye, Linji Hao, Qi Gu, Xunliang Cai, Yongliang Shen, Yujiu Yang
Paper HF Code BIB
🤗 HuggingFace Daily Paper #1
- A critic-free, recursive self-distillation method that provides turn-level credit assignment for long-horizon, multi-turn agentic reinforcement learning.

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
Zhiyuan Yao, Yuxin Chen, Zhengxi Lu, Zishan Xu, Yueqing Sun, Yifu Guo, Yuquan Lu, Zhengzhou Cai, Kangning Zhang, Zhuowen Han, Zi-Han Wang, Ziang Ye, Qi Gu, Xunliang Cai, Weiwen Liu, Yongliang Shen
- A unified RL framework that learns reusable skills across related tasks by organizing instances into progressively challenging sequences.

SDAR: Self-Distilled Agentic Reinforcement Learning
Zhengxi Lu, Zhiyuan Yao, Zhuowen Han, Zi-Han Wang, Jinyang Wu, Qi Gu, Xunliang Cai, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
Paper HF Code BIB
🤗 HuggingFace Daily Paper #2
- Extends on-policy self-distillation to multi-turn agents, delivering dense token-level guidance on top of coarse trajectory-level rewards.

Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
Huichi Zhou*, Yihang Chen*, Siyuan Guo, Xue Yan, Kin Hei Lee, Zihan Wang, Ka Yiu Lee, Guchun Zhang, Kun Shao, Linyi Yang†, Jun Wang†
- Memory-based online reinforcement learning that lets LLM agents continually adapt from experience without updating the underlying model weights.
🎨 Machine Creativity & Tool Learning

CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges
Zi-Han Wang, Lam Nguyen, Zhengyang Zhao, Mengyue Yang, Chengwei Qin, Yujiu Yang, Linyi Yang
- A benchmark for machine creativity in code generation — covering combinatorial and exploratory creativity with executable evaluation and the inference-time enhancement strategy EvoRePE.

Alignment for Efficient Tool Calling of Large Language Models
Hongshen Xu*, Zihan Wang*, Zichen Zhu, Lei Pan, Xingyu Chen, Lu Chen, Kai Yu
- A multi-objective alignment framework that combines probabilistic knowledge-boundary estimation with dynamic decision making to reduce unnecessary tool calls while preserving performance.

Reducing Tool Hallucination via Reliability Alignment
Hongshen Xu, Zichen Zhu, Lei Pan, Zihan Wang, Su Zhu, Da Ma, Ruisheng Cao, Lu Chen, Kai Yu
- Defines and categorizes tool hallucinations (tool selection vs. tool usage) and introduces reliability-oriented alignment for more robust and efficient tool interaction.
💼 Internships
- Mar. 2026 – Present: Research Intern at
LongCat Team, Meituan, advised by Qi Gu, Chengcheng Han, and Yueqing Sun. Agentic RL, multimodal productivity agents, and environment scaling (LongCat 2.0). - Aug. 2025 – Feb. 2026: Research Intern at Southern University of Science and Technology (SUSTech), advised by Linyi Yang. AutoResearch / AI Scientist.
- Feb. 2025 – May 2025: Research Intern, Peking University. LLM pre-training, data selection and mixing.
- Aug. 2024 – Feb. 2025: Research Intern, Shanghai Jiao Tong University. LLM alignment, tool learning, and tool-use agents.
🎖 Honors and Awards
- 2024–2025: Baidu Artificial Intelligence and Big Data Elite Class.
- 2023–2024: Outstanding Award (Meritorious Winner), Mathematical Contest in Modeling (MCM/ICM).
- 2023–2024: Provincial First Prize, National College Students’ Mathematics Competition.
- 2023–2024: Outstanding Volunteer Award.
- 2022–2023: National Scholarship.
📖 Educations
- 2026.09 - now (incoming): M.S. in Artificial Intelligence, Tsinghua University (Shenzhen).
- 2022.09 - 2026.07: B.Eng. in Computer Science and Technology, Xi’an Jiaotong University. GPA 3.97/4.30, Rank 5/193.
🧑⚖️ Academic Services
Conference Reviewer
- AAAI 2026.
- ACL Rolling Review (ARR), 2025–2026.
