I am a graduate student at Xi'an Jiaotong University (XJTU), focusing on AI infrastructure, LLM serving, and RL post-training systems.
I enjoy turning systems ideas into practical open-source implementations: efficient rollout execution, distributed training workflows, weight synchronization, and cross-platform GPU optimization for GRPO-style workloads.
- π Currently contributing to: vLLM-Omni, a framework for efficient omni-modality model inference and serving.
- π Main work: Leading Vime framework research, fork-roadmap planning, and PR delivery across CUDA and ROCm; integrating Vime with RL-Kernel for reproducible RL training and rollout.
- π¬ Research interests: Efficient inference, distributed attention, GRPO/RLHF systems, and cross-platform GPU performance.
| Project | Focus | Status |
|---|---|---|
| vLLM-Omni | Efficient omni-modality model inference and serving in the vLLM ecosystem | π₯ Contributor |
| Vime | RL framework integration, roadmap planning, and end-to-end training/rollout validation | β‘ Contributor |
| RL-Kernel | GPU kernels and strict runtime validation consumed by Vime on CUDA and ROCm | π€ Maintainer |
| Area | Selected Work |
|---|---|
| Vime framework and delivery | Framework investigation, fork-version roadmap planning, upstream contribution delivery, and reproducible experiment documentation |
| CUDA + ROCm integration | Led the Vime provider boundary across both GPU stacks, preserving Vime's loss semantics and native fallback |
| Distributed Attention | Developed and validated paged/CP attention paths, including FlashInfer RoPE-fused attention, CP drift checks, and bitwise ROCm schedules |
| Deterministic runtime and performance | CUDA Graph capture, tensor-parallel all-reduce optimization, strict runtime modes, and cross-configuration kernel validation |
| End-to-end RL validation | Ran matched native/provider train-rollout consistency experiments, TP/CP ablations, bitwise checks, throughput profiling, and performance tuning on Qwen3 workloads |
| Provider integration | Designed structured provider contracts, TP vocabulary partition handling, autograd checks, and technical write-ups for portable log-probability execution |


