Deep Learning Systems · LLM Inference

Yangmin Li a.k.a. YAMY

Senior Deep Learning Algorithm Engineer at NVIDIA, working on SGLang & TensorRT-LLM — delivering fast, stable, and correct LLM inference across GB200 / GB300 and Rubin.

Previously CMU INI & Tongji CS. I care about the intersection of engineering, optimization, and research with real-world impact.

Yangmin Li

01 — About

Building the systems that serve LLMs.

I’m a Senior Deep Learning Algorithm Engineer at NVIDIA on the TensorRT-LLM / Inference team, where I contribute to large-scale LLM inference and serving, from Day-0 model enablement and distributed-runtime integration to GPU performance and correctness. A lot of my day-to-day lives inside SGLang: Qwen3.8 and Kimi K3 bring-up, PP / DCP / TP / EP serving, KV and Mamba-state transfer, FP8 / NVFP4 MoE kernels, and agentic workload optimization across GB200 / GB300 NVL72 and Rubin.

On the side, I collaborate with LMSYS and UC Berkeley’s Sky Computing Lab on multi-LLM serving (Prism, kvcached), and with CMU’s AI Common Sense Lab under Deepak Pathak on physics-grounded reasoning (Sim2Reason) — turning physics simulators into infinite training environments for LLMs.

Before NVIDIA full-time, I was at CMU INI (M.S., 2023–2024), and Tongji University for my B.E. in CS. I’ve been lucky to publish first-author papers at CIKM and IEEE TAFFC along the way.

02 — Now

What I’m working on

Updated 2026

Contributing to Qwen3.8 Day-0 support and Rubin optimization in SGLang — DCP GQA KV-head mapping, PP Mamba-pool sizing, Mooncake staging buffers, prefill-memory fixes, and FP8 / NVFP4, MTP, PP / DCP, and PD recipes for 8K / 1K serving.

Delivering the Kimi K3 Day-0 and PP-prefill roadmap — deep chunked-PP prefill, KDA state transfer, PP×DCP layer mapping, overlapped P2P scheduling, and integrated PP-prefill with TP / DCP decode for disaggregated serving.

Contributing to Qwen3.5 release stabilization and PD disaggregation — Mamba / GQA correctness, heterogeneous-TP KV and state transfer, GPU staging and ring allocation, and NIXL / Mooncake transport improvements on GB200.

Working on agentic workload optimization across SGLang AgentX and AA-AgentPerf — pacing-correct clients, cache and state capacity, TP / DEP decode, PD serving, and TensorRT-LLM’s NVFP4 stack for large-scale agent workloads.

03 — Selected Work

Projects & contributions

SGLang · NVIDIA · 2026

Qwen3.8 Day-0 support & Rubin optimization

Contributed to DCP GQA KV-head mapping, PP Mamba-pool sizing, Mooncake staging buffers, and prefill-memory fixes; delivered FP8 / NVFP4, MTP, PP / DCP, and PD recipes reaching 5,126 tok/s/GPU at 8K / 1K. Rubin bring-up work delivered up to 46.7% higher TPS/GPU than GB300 at matched concurrency.

LMSYS Day-0 blog →
SGLang · NVIDIA · 2026

Kimi K3 Day-0 & PP-prefill roadmap

Contributed to deep chunked-PP prefill, KDA state transfer, PP×DCP layer mapping, and overlapped P2P scheduling. PP8 reached 5,958 input tok/s/GPU while hiding 91% of handoff latency; integrated PP-prefill with TP / DCP decode for disaggregated serving.

LMSYS Day-0 blog →
SGLang · NVIDIA · 2026

Qwen3.5 release stabilization & PD disaggregation

Contributed to Mamba / GQA correctness and heterogeneous-TP KV / state transfer. GPU staging and ring allocation reduced RDMA operations by about 1,000×, delivering 4.9× Mooncake throughput and up to 64% higher NIXL throughput on GB200.

SGLang v0.5.10 release →
SGLang · TensorRT-LLM · 2026

Agentic workload optimization

Contributed to Qwen3.5 AgentX across pacing-correct clients, cache and state capacity, TP / DEP decode, and PD serving; expanded the strict GB300 Pareto set from 9 to 14 points. Delivered TensorRT-LLM’s NVFP4 stack for AA-AgentPerf at up to 1,840 agents.

AA-AgentPerf article →
LMSYS × Berkeley Sky Lab · 2025–

Prism — Multi-LLM serving on shared GPUs

Two-tier SLO-aware scheduler (global KV-pressure placement + GPU-local priority admission). 50+ LLMs concurrently on shared infra, 2× cost reduction & 3.3× SLO attainment.

arXiv 2505.04021 · Multi-LLM/prism-research →
LMSYS × Berkeley Sky Lab · 2025–

kvcached — GPU virtual memory for KV cache

OS-style virtual memory for LLM KV caches. CUDA virtual-memory APIs decouple physical from virtual GPU memory, enabling elastic allocation and cross-model cache sharing. Integrated with SGLang & vLLM.

Apache 2.0 · ovg-project/kvcached →
CMU AI Common Sense Lab · 2024–

Sim2Reason — Solving Physics Olympiad via RL on simulators

Procedurally generate MuJoCo scenes & QA pairs through a DSL, then RL-finetune LLMs on synthetic supervision. +10–15pp on IPhO Mechanics across model scales (Qwen2.5-32B: 19.8%→25.2%). With Mihir Prabhudesai, Aryan Satpathy, Deepak Pathak, Katerina Fragkiadaki.

sim2reason.github.io →
CMU ADMIS Lab · TAFFC’25

CorMulT — Modality correlation-aware Transformer

Semi-supervised model with modality-correlation contrastive learning for audio / image / text fusion. Strong results on CMU-MOSEI. First-author, IEEE Transactions on Affective Computing 2025.

04 — Publications

Selected papers

Prism architecture for coordinated multi-LLM serving with GPU memory ballooning

OSDI 2026

Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning

LMSYS / UC Berkeley Sky Computing Lab · Co-author

Prism treats GPU memory as an elastic resource, combining memory ballooning with coordinated model placement and request scheduling for multi-LLM serving.

Sim2Reason physics simulator environments and zero-shot reasoning evaluation

ICML 2026

Sim2Reason: Solving Physics Olympiad via RL on Physics Simulators

CMU AI Common Sense Lab · Core contributor

Sim2Reason turns physics simulators into scalable QA generators through a domain-specific language, then uses the synthetic interactions for label-free reinforcement learning.

05 — Experience

Where I’ve been

  1. Jan 2025 — Present

    NVIDIA — Senior Deep Learning Algorithm Engineer

    TensorRT-LLM · Deep Learning Inference · Santa Clara, CA

    Contributing to SGLang and TensorRT-LLM: Qwen3.8 / Kimi K3 Day-0 support, PP / DCP serving, Rubin bring-up, and AgentX optimization.

  2. May 2024 — Aug 2024

    NVIDIA — System Software Engineer Intern

    Autonomous Vehicle ML & Driveworks Platform · Santa Clara, CA

    RoadCast 2.1 — distributed data logging with dwProto serialization, batched I/O, NVSci/TCP/DFS.

  3. Feb 2023 — May 2023

    Amazon Web Services — Cloud Development Associate Intern

    Cloud Architecture & MLOps · Beijing

    TB-scale data lake (Kinesis → S3 → Glue/Athena), SageMaker MLOps with CodePipeline + EKS/ECS.

  4. Sep 2022 — Dec 2022

    ByteDance / TikTok — R&D Engineer Intern

    Distributed Video Quality DL Evaluation · Shanghai

    Multimodal DL/ML video QoE evaluation (D-VQA, multimodal Transformer), Spark on HDFS.

  5. 2023 — 2024

    Carnegie Mellon University

    M.S. Information Technology · Information Networking Institute · GPA 3.84/4.00

  6. 2019 — 2023

    Tongji University

    B.E. Computer Science & Technology · GPA 90.8/100 · First-class Scholarship

06 — Writing

Notes from the road

A few longer reflections on projects I’ve been part of — currently lives on LinkedIn.