Peking University
Sep 2023 – Jun 2027 (expected)Dual degree in Computer Science
School of Mathematical Sciences
I am an undergraduate at Peking University, pursuing a dual degree in Mathematics and Computer Science. I expect to graduate in June 2027.
My research focuses on large language models, including agentic reasoning, reinforcement learning and post-training, safety alignment, and decoding for language diffusion models.
Outside research, I share course notes and small projects on my blog, and enjoy listening to JJ Lin.
Dual degree in Computer Science
School of Mathematical Sciences
High school
* Equal contribution.
Submitted to EMNLP 2026; under review.
A multi-agent framework that uses sub-agents for active context management in long-horizon deep research, with supervised fine-tuning data for learning delegation.
EMNLP 2026.
A search-based decoding method that combines local and global information for discrete optimization, with an accelerated variant that concentrates search on selected decoding steps.
ICML 2025 Workshop: Models of Human Feedback for AI Alignment.
A safety-alignment method that ranks candidate responses using intermediate Transformer representations and a lightweight similarity scorer.
School of Mathematical Sciences, Peking University
School of Mathematical Sciences, Peking University
An open-source multi-agent framework for long-horizon deep research. Trains research agents to delegate subtasks and use compact, evidence-grounded reports for active context management.
LLM agents · Deep research · Supervised fine-tuning
Built an agent trained through behavior cloning and self-play PPO, with an asymmetric actor-critic model and IMPALA-style asynchronous training. Ranked 5th out of 28 teams in the final round of the Botzone competition.
Reinforcement learning · Self-play · Multi-process training
Maintenance and updates to a campus venue-booking tool.
Python · Campus tools