Deep Learning Systems · LLM Inference
Senior Deep Learning Algorithm Engineer at NVIDIA, working on SGLang & TensorRT-LLM — delivering fast, stable, and correct LLM inference across GB200 / GB300 and Rubin.
Previously CMU INI & Tongji CS. I care about the intersection of engineering, optimization, and research with real-world impact.
01 — About
I’m a Senior Deep Learning Algorithm Engineer at NVIDIA on the TensorRT-LLM / Inference team, where I contribute to large-scale LLM inference and serving, from Day-0 model enablement and distributed-runtime integration to GPU performance and correctness. A lot of my day-to-day lives inside SGLang: Qwen3.8 and Kimi K3 bring-up, PP / DCP / TP / EP serving, KV and Mamba-state transfer, FP8 / NVFP4 MoE kernels, and agentic workload optimization across GB200 / GB300 NVL72 and Rubin.
On the side, I collaborate with LMSYS and UC Berkeley’s Sky Computing Lab on multi-LLM serving (Prism, kvcached), and with CMU’s AI Common Sense Lab under Deepak Pathak on physics-grounded reasoning (Sim2Reason) — turning physics simulators into infinite training environments for LLMs.
Before NVIDIA full-time, I was at CMU INI (M.S., 2023–2024), and Tongji University for my B.E. in CS. I’ve been lucky to publish first-author papers at CIKM and IEEE TAFFC along the way.
02 — Now
Updated 2026
Contributing to Qwen3.8 Day-0 support and Rubin optimization in SGLang — DCP GQA KV-head mapping, PP Mamba-pool sizing, Mooncake staging buffers, prefill-memory fixes, and FP8 / NVFP4, MTP, PP / DCP, and PD recipes for 8K / 1K serving.
Delivering the Kimi K3 Day-0 and PP-prefill roadmap — deep chunked-PP prefill, KDA state transfer, PP×DCP layer mapping, overlapped P2P scheduling, and integrated PP-prefill with TP / DCP decode for disaggregated serving.
Contributing to Qwen3.5 release stabilization and PD disaggregation — Mamba / GQA correctness, heterogeneous-TP KV and state transfer, GPU staging and ring allocation, and NIXL / Mooncake transport improvements on GB200.
Working on agentic workload optimization across SGLang AgentX and AA-AgentPerf — pacing-correct clients, cache and state capacity, TP / DEP decode, PD serving, and TensorRT-LLM’s NVFP4 stack for large-scale agent workloads.
03 — Selected Work
Contributed to DCP GQA KV-head mapping, PP Mamba-pool sizing, Mooncake staging buffers, and prefill-memory fixes; delivered FP8 / NVFP4, MTP, PP / DCP, and PD recipes reaching 5,126 tok/s/GPU at 8K / 1K. Rubin bring-up work delivered up to 46.7% higher TPS/GPU than GB300 at matched concurrency.
Contributed to deep chunked-PP prefill, KDA state transfer, PP×DCP layer mapping, and overlapped P2P scheduling. PP8 reached 5,958 input tok/s/GPU while hiding 91% of handoff latency; integrated PP-prefill with TP / DCP decode for disaggregated serving.
Contributed to Mamba / GQA correctness and heterogeneous-TP KV / state transfer. GPU staging and ring allocation reduced RDMA operations by about 1,000×, delivering 4.9× Mooncake throughput and up to 64% higher NIXL throughput on GB200.
Contributed to Qwen3.5 AgentX across pacing-correct clients, cache and state capacity, TP / DEP decode, and PD serving; expanded the strict GB300 Pareto set from 9 to 14 points. Delivered TensorRT-LLM’s NVFP4 stack for AA-AgentPerf at up to 1,840 agents.
Two-tier SLO-aware scheduler (global KV-pressure placement + GPU-local priority admission). 50+ LLMs concurrently on shared infra, 2× cost reduction & 3.3× SLO attainment.
OS-style virtual memory for LLM KV caches. CUDA virtual-memory APIs decouple physical from virtual GPU memory, enabling elastic allocation and cross-model cache sharing. Integrated with SGLang & vLLM.
Procedurally generate MuJoCo scenes & QA pairs through a DSL, then RL-finetune LLMs on synthetic supervision. +10–15pp on IPhO Mechanics across model scales (Qwen2.5-32B: 19.8%→25.2%). With Mihir Prabhudesai, Aryan Satpathy, Deepak Pathak, Katerina Fragkiadaki.
Semi-supervised model with modality-correlation contrastive learning for audio / image / text fusion. Strong results on CMU-MOSEI. First-author, IEEE Transactions on Affective Computing 2025.
04 — Publications
ICML 2026
CMU AI Common Sense Lab · Core contributor
Sim2Reason turns physics simulators into scalable QA generators through a domain-specific language, then uses the synthetic interactions for label-free reinforcement learning.
Additional publications
IEEE TAFFC · 2025 · First author
ACM CIKM · 2022 · First author
05 — Experience
Jan 2025 — Present
TensorRT-LLM · Deep Learning Inference · Santa Clara, CA
Contributing to SGLang and TensorRT-LLM: Qwen3.8 / Kimi K3 Day-0 support, PP / DCP serving, Rubin bring-up, and AgentX optimization.
May 2024 — Aug 2024
Autonomous Vehicle ML & Driveworks Platform · Santa Clara, CA
RoadCast 2.1 — distributed data logging with dwProto serialization, batched I/O, NVSci/TCP/DFS.
Feb 2023 — May 2023
Cloud Architecture & MLOps · Beijing
TB-scale data lake (Kinesis → S3 → Glue/Athena), SageMaker MLOps with CodePipeline + EKS/ECS.
Sep 2022 — Dec 2022
Distributed Video Quality DL Evaluation · Shanghai
Multimodal DL/ML video QoE evaluation (D-VQA, multimodal Transformer), Spark on HDFS.
2023 — 2024
M.S. Information Technology · Information Networking Institute · GPA 3.84/4.00
2019 — 2023
B.E. Computer Science & Technology · GPA 90.8/100 · First-class Scholarship
06 — Writing
A few longer reflections on projects I’ve been part of — currently lives on LinkedIn.
SGLang · GB300 · 2025
A few rewarding months making SGLang fast, stable, and correct at scale on GB300 with the LMSYS & NVIDIA folks.
linkedin →
Sim2Reason · CMU · 2026
Why simulators may be where language models finally learn how the world works.
linkedin →
Prism · LMSYS × Berkeley · 2025
From AWS & ByteDance to NVIDIA & SGLang — and how Prism became my favorite collaboration so far.
linkedin →