About Me

I am Zihan Wang (王子涵), an incoming M.S. student in Artificial Intelligence at Tsinghua University (Shenzhen), advised by Prof. Yujiu Yang. I received my B.Eng. in Computer Science and Technology from Xi’an Jiaotong University.

My research centers on NLP / Large Language Models (LLMs), with a focus on building intrinsically motivated, self-evolving, and reasoning-capable agents — particularly around agentic reinforcement learning, self-evolving agents, coding / GUI agents, and environment scaling.

🐈 I am always open to discussion and collaboration, and I am currently looking for internship opportunities. Feel free to reach out via email or WeChat: wishme25

🔥 News

📝 Selected Publications

* Equal contribution. † Corresponding author. See my full list on Google Scholar.

🤖 Agentic RL & Skill Evolution

Preprint
AgentOPSD

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Jie Wu, Zhengzhou Cai, Yueqing Sun, Ziang Ye, Linji Hao, Qi Gu, Xunliang Cai, Yongliang Shen, Yujiu Yang

Paper HF Code BIB 🤗 HuggingFace Daily Paper #1

  • A critic-free, recursive self-distillation method that provides turn-level credit assignment for long-horizon, multi-turn agentic reinforcement learning.
Preprint
SkillRise

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Zhiyuan Yao, Yuxin Chen, Zhengxi Lu, Zishan Xu, Yueqing Sun, Yifu Guo, Yuquan Lu, Zhengzhou Cai, Kangning Zhang, Zhuowen Han, Zi-Han Wang, Ziang Ye, Qi Gu, Xunliang Cai, Weiwen Liu, Yongliang Shen

Paper HF BIB

  • A unified RL framework that learns reusable skills across related tasks by organizing instances into progressively challenging sequences.
Preprint
SDAR

SDAR: Self-Distilled Agentic Reinforcement Learning

Zhengxi Lu, Zhiyuan Yao, Zhuowen Han, Zi-Han Wang, Jinyang Wu, Qi Gu, Xunliang Cai, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen

Paper HF Code BIB 🤗 HuggingFace Daily Paper #2

  • Extends on-policy self-distillation to multi-turn agents, delivering dense token-level guidance on top of coarse trajectory-level rewards.
Preprint
Memento

Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

Huichi Zhou*, Yihang Chen*, Siyuan Guo, Xue Yan, Kin Hei Lee, Zihan Wang, Ka Yiu Lee, Guchun Zhang, Kun Shao, Linyi Yang†, Jun Wang†

Paper HF Code BIB

  • Memory-based online reinforcement learning that lets LLM agents continually adapt from experience without updating the underlying model weights.

🎨 Machine Creativity & Tool Learning

ACL 2026
CreativeBench

CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges

Zi-Han Wang, Lam Nguyen, Zhengyang Zhao, Mengyue Yang, Chengwei Qin, Yujiu Yang, Linyi Yang

Paper Homepage Code BIB

  • A benchmark for machine creativity in code generation — covering combinatorial and exploratory creativity with executable evaluation and the inference-time enhancement strategy EvoRePE.
EMNLP 2025
Alignment for Efficient Tool Calling

Alignment for Efficient Tool Calling of Large Language Models

Hongshen Xu*, Zihan Wang*, Zichen Zhu, Lei Pan, Xingyu Chen, Lu Chen, Kai Yu

Paper BIB

  • A multi-objective alignment framework that combines probabilistic knowledge-boundary estimation with dynamic decision making to reduce unnecessary tool calls while preserving performance.
ICML 2025
Reducing Tool Hallucination

Reducing Tool Hallucination via Reliability Alignment

Hongshen Xu, Zichen Zhu, Lei Pan, Zihan Wang, Su Zhu, Da Ma, Ruisheng Cao, Lu Chen, Kai Yu

Paper BIB

  • Defines and categorizes tool hallucinations (tool selection vs. tool usage) and introduces reliability-oriented alignment for more robust and efficient tool interaction.

💼 Internships

  • Mar. 2026 – Present: Research Intern at LongCat Team, Meituan, advised by Qi Gu, Chengcheng Han, and Yueqing Sun. Agentic RL, multimodal productivity agents, and environment scaling (LongCat 2.0).
  • Aug. 2025 – Feb. 2026: Research Intern at Southern University of Science and Technology (SUSTech), advised by Linyi Yang. AutoResearch / AI Scientist.
  • Feb. 2025 – May 2025: Research Intern, Peking University. LLM pre-training, data selection and mixing.
  • Aug. 2024 – Feb. 2025: Research Intern, Shanghai Jiao Tong University. LLM alignment, tool learning, and tool-use agents.

🎖 Honors and Awards

  • 2024–2025: Baidu Artificial Intelligence and Big Data Elite Class.
  • 2023–2024: Outstanding Award (Meritorious Winner), Mathematical Contest in Modeling (MCM/ICM).
  • 2023–2024: Provincial First Prize, National College Students’ Mathematics Competition.
  • 2023–2024: Outstanding Volunteer Award.
  • 2022–2023: National Scholarship.

📖 Educations

  • 2026.09 - now (incoming): M.S. in Artificial Intelligence, Tsinghua University (Shenzhen).
  • 2022.09 - 2026.07: B.Eng. in Computer Science and Technology, Xi’an Jiaotong University. GPA 3.97/4.30, Rank 5/193.

🧑‍⚖️ Academic Services

Conference Reviewer

  • AAAI 2026.
  • ACL Rolling Review (ARR), 2025–2026.