UCLA · COMPUTER SCIENCE

Yihang Chen 陈奕行

Exploring learning & generative intelligence.

I am a PhD student in computer science at UCLA , advised by Prof. Cho-Jui Hsieh . I received my bachelor’s degree in computational mathematics at Peking University in July 2021, advised by Prof. Liwei Wang ; and my master’s degree in data science at EPFL in February 2024, advised by Prof. Volkan Cevher .

Yihang Chen
PhD student · Los Angeles, CA
01

Research

My research focuses on LLM post-training, generative models, and AI safety:

02

What's new

03

Publications

* Equal contribution
Negative self-certainty supplies the intrinsic reward for image-generation RL.
arXiv IMAGE GENERATION · INTRINSIC REWARDS

IRIS: Intrinsic Reward Image Synthesis

Yihang Chen* , Yuanhao Ban*, Yunqi Hong, Cho-Jui Hsieh
arXiv preprint: September 2025.
[arXiv]

We propose IRIS, the first framework to improve autoregressive T2I models with reinforcement learning using only an intrinsic reward.

Compare policies at each turn and update them using multi-step preferences.
TMLR LLM ALIGNMENT · MARKOV GAMES

Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees

Yongtao Wu*, Luca Viano*, Kimon Antonakopoulos, Yihang Chen , Zhenyu Zhu, Quanquan Gu, Volkan Cevher
Transactions on Machine Learning Research, 2025. Abridged in Language Gamification Workshop @ NeurIPS, Spotlight, 2024
[arXiv] / [workshop] / [workshop poster]

We model the multi-step preference alignment problem as a two-player constant-sum Markov game. We propose the MPO, a natural actor-critic algorithm; and the OMPO, an optimistic online gradient descent algorithm.

Importance weighting changes the bias–variance trade-off under covariate shift.
ICML LEARNING THEORY · DISTRIBUTION SHIFT

High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization

Yihang Chen , Fanghui Liu, Taiji Suzuki, Volkan Cevher
41st International Conference on Machine Learning ( ICML ) , 2024
[arXiv] / [poster]

We study kernel ridge regression in high dimensions under covariate shifts and analyzes the role of importance re-weighting. We also provide asymptotic expansion of kernel functions/vectors under covariate shift.

Randomly retain edges using layerwise ratios, then train the sparse network.
NeurIPS NEURAL NETWORKS · PRUNING

Sanity-Checking Pruning Methods: Random Tickets can Win the Jackpot

Jingtong Su*, Yihang Chen* , Tianle Cai*, Tianhao Wu, Ruiqi Gao, Liwei Wang, Jason D. Lee.
34th Annual Conference on Neural Information Processing Systems ( NeurIPS ) , 2020
[arXiv] / [code] / [slides]

We sanity-check prune-at-init methods, and find them hardly exploits any information from the training data. We propose "zero-shot" pruning, which only relies on simple data-independent pruning ratios for each layer.

Honors and Awards

ICLR 2025 Notable Reviewers, 2025.
NeurIPS 2024 Scholar Award & Top Reviewers, 2024.
Research Scholars MSc Program, EPFL, 2021-2022.
The Elite Undergraduate Training Program of Applied Math, 2019-2021.
Excellent Graduate of Peking University, 2021.
National Scholarship, People's Republic of China (Top 1%), 2020.
Shing Tung Yau Mathematics Awards, Chia Chiao Lin Medals, Bronze, 2020.

Invited Talks

Mila GFlowNet meeting, Order-Preserving GFlowNets, 2023.10.04.

Services

Conference Reviewers: ICLR 2025, 2026; ICML 2025; NeurIPS 2024, 2025, AISTATS 2025, CVPR 2026, ECCV 2026.

Fun Facts

During my stay in Switzerland, I finished the Via Alpina hiking.