I build machine learning systems for large language models, focusing on efficient inference and agentic serving, speculative decoding, GPU kernel optimization, and post-training infrastructure.
I am an ML systems researcher at Carnegie Mellon's Catalyst Lab and Parallel Data Lab, advised by Zhihao Jia. I spent two years at Snowflake AI Research, working with Samyam Rajbhandari and Aurick Qiao. I am completing my Ph.D. in Computer Science at Carnegie Mellon University.
Highlights include: SpecInfer, whose tree-based verification is used by EAGLE-3, Medusa, and REST in vLLM and SGLang; SuffixDecoding, productionized at Snowflake for up to 5.3× agentic serving speedup; and FlexLLM, for up to 6.8× higher fine-tuning throughput.