About Me
I am a third-year PhD student in the Department of Electrical and Computer Engineering at Carnegie Mellon University, advised by Prof. Rashmi Vinayak. I am a member of the TheSys Group, the Catalyst Group, and the Parallel Data Lab. My research interests span large-scale systems for LLMs and the learning algorithms underlying modern language models.
I received my bachelor’s degree from the Yao Class (Special Pilot CS Class) at Tsinghua University. During my undergraduate studies, I had the privilege of working with Prof. Yi Wu and Prof. Yang Gao on reinforcement learning and its applications to systems.
Research
My research spans large language models and large-scale distributed systems. My systems work focuses on improving the reliability and efficiency of large-scale LLM workloads. Current topics include silent data corruption in hyperscale LLM training clusters, in collaboration with Meta, and efficient LLM serving on heterogeneous hardware.
Building on this systems background, I am expanding my research toward the learning aspects of LLM training, with an initial focus on reinforcement-learning-based post-training. More broadly, I am interested in fundamental questions about how language models learn and behave, including how their learning dynamics and behavior evolve at scale.
Publications
-
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
Yixuan Mei, Zikun Li, Zixuan Chen, Shiqi Pan, Mengdi Wu, Xupeng Miao, Zhihao Jia, KV Rashmi
Preprint Paper
-
SEVI: Silent Data Corruption of Vector Instructions in Hyper-Scale Datacenters
Yixuan Mei, Shreya Varshini, Harish Dixit, Sriram Sankar, KV Rashmi
-
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
Yixuan Mei, Yonghao Zhuang, Xupeng Miao, Juncheng Yang, Zhihao Jia, Rashmi Vinayak
-
Quarl: A Learning-Based Quantum Circuit Optimizer
Zikun Li, Jinjun Peng, Yixuan Mei, Sina Lin, Yi Wu, Oded Padon, Zhihao Jia
-
SpeedyZero: Mastering Atari with Limited Data and Time
Yixuan Mei*, Jiaxuan Gao*, Weirui Ye, Shaohuai Liu, Yang Gao, Yi Wu
* Equal contribution.
Experience
Google LLC — Software Engineering Intern
Research on training stability and convergence in asynchronous agentic RL post-training, using MaxText, Tunix, and TPU Inference.
Google LLC — Software Engineering Intern / Student Researcher
Research on scheduling algorithms for improving resource utilization in long-context LLM serving systems.
Teaching
- Teaching Assistant — 15-750: Algorithms in the Real World · Fall 2024
2026
OOPSLA 2024