Yaqi Xia (夏亚奇)

Wuhan University HongYi Postdoc Research Fellow, School of Computer Science, Wuhan University

I am Yaqi Xia, a HongYi Postdoc Research Fellow at Wuhan University, working with Prof. Dazhao Cheng. My research focuses on distributed and high-performance systems for AI/ML, with an emphasis on efficient long-context LLM inference, scalable graph learning systems, and ML systems optimization. My work has appeared in top systems venues and journals including SC, PPoPP, ATC, IEEE TPDS, and HPDC, with Best Paper Runner-up and Best Paper Award Nomination honors. I received my Ph.D. from Wuhan University in 2025. Prior to that, I obtained both my Bachelor's and Master's degrees from Xidian University under the supervision of Prof. Rui Song.


Education
  • Wuhan University

    Wuhan University

    Ph.D. in Artificial Intelligence Sep. 2021 - Dec. 2025

  • Xidian University

    Xidian University

    M.S. in Electronics and Communication Engineering Sep. 2018 - Jul. 2021

  • Xidian University

    Xidian University

    B.S. in Communication Engineering Sep. 2014 - Jul. 2018

Honors & Awards
  • Best Paper Award Nomination, SC 26 2026
  • Best Paper Runner-up, IEEE TPDS24 2024
  • Best Paper Award Nomination, ACM HPDC23 2023
  • First Prize of the 'DiDi' Outstanding PhD Student Academic Forum 2025
  • The Best Presentation Award at the Graduate Future Innovation Academic Forum-Doctoral Forum, Wuhan University 2025
  • Second Prize of First 'Tianzhi Cup' Artificial Intelligence Challenge 2019
Experience
  • Research Center for Graph Computing, Zhejiang Lab

    Research Center for Graph Computing, Zhejiang Lab

    Research Intern Aug. 2023 - Dec. 2023

Service
  • Associate Editor, The Journal of Supercomputing
  • Reviewer, IEEE TPDS/TDSC
News
2026
Our work 'DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference' has been nominated for the Best Paper Award at SC 26.
Jul 24
Our work 'DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference' is accepted by SC 2026.
Jul 02
Our work 'QDiTx: Arbitrary Bit-Width Quantization Framework for Diffusion Transformer' is accepted by IEEE TCAD 2026.
Jul 02
2025
Our work 'JanusQuant: Accurate and Efficient 2-bit KV Cache Quantization for Long-context Inference' is accepted by PPoPP 2026.
Nov 11
Our work 'MPMoE: Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism' has been selected as the Runner-Up of the 2024 Best Paper Award from IEEE TPDS.
Aug 12
Our work 'MXBLAS: Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library' is accepted by SC 2025.
Jun 27
Our work 'Voltrix: Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel Optimization' is accepted by ATC 2025.
Apr 27
I am looking for highly self-motivated Bachelor, Master and PhD students. Feel free to get in touch with me via email If you're interested in collaborating.
Feb 14
2024
Our work 'Harnessing Inter-GPU Shared Memory for Seamless MoE Communication-Computation Fusion' is accepted by PPoPP 2025.
Dec 12
Our work 'Redundancy-Free and Load-Balanced TGNN Training With Hierarchical Pipeline Parallelism' is accepted by TPDS 2024.
Nov 11
Our work 'Scaling New Heights :Transformative Cross-GPU Sampling for Training Billion-Edge Graphs' is accepted by SC 2024.
May 16
Our work 'Accelerating Distributed DLRM Training with Optimized TT Decomposition and Micro-Batching' is accepted by SC 2024.
Apr 16
Our work 'Raptor-T :A Fused and Memory-Efficient Sparse Transformer for Long and Variable-Length Sequences' is accepted by TC 2024.
Apr 16
Our work 'MPMoE :Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism' is accepted by TPDS 2024.
Apr 08
2023
Our work 'Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism' is accepted by HPDC 2023 and be selected as Best Paper Nomination (only two nominations)!
Aug 07
Our work 'MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism' is accepted by IPDPS 2023.
Jul 18
Selected Publications (view all )
DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference
DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference

Chengyu Sun, Yaqi Xia†, Ruirui Pan, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

2026 International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2026 ConferenceCCF-ABest Paper Award Nomination

We present DySpin, a plug-and-play library that advances dynamic sparse long-context inference.

DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference
DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference

Chengyu Sun, Yaqi Xia†, Ruirui Pan, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

2026 International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2026 ConferenceCCF-ABest Paper Award Nomination

We present DySpin, a plug-and-play library that advances dynamic sparse long-context inference.

MXBLAS:Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library.
MXBLAS:Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library.

Weihu Wang*, Yaqi Xia*, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(* equal contribution; † corresponding author)

2025 International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2025 ConferenceCCF-A

We present MXBLAS, a high-performance MX-GEMM library that unifies support across the full spectrum of MX-format variations.

MXBLAS:Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library.
MXBLAS:Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library.

Weihu Wang*, Yaqi Xia*, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(* equal contribution; † corresponding author)

2025 International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2025 ConferenceCCF-A

We present MXBLAS, a high-performance MX-GEMM library that unifies support across the full spectrum of MX-format variations.

Voltrix:Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel Optimization
Voltrix:Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel Optimization

Yaqi Xia*, Weihu Wang*, Donglin Yang, Xiaobo Zhou†, Dazhao Cheng†(* equal contribution; † corresponding author)

2025 USENIX Annual Technical Conference (ATC) 2025 ConferenceCCF-A

We introduce Voltrix-SpMM, a revolutionary GPU kernel design for sparse matrix-matrix multiplication.

Voltrix:Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel Optimization
Voltrix:Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel Optimization

Yaqi Xia*, Weihu Wang*, Donglin Yang, Xiaobo Zhou†, Dazhao Cheng†(* equal contribution; † corresponding author)

2025 USENIX Annual Technical Conference (ATC) 2025 ConferenceCCF-A

We introduce Voltrix-SpMM, a revolutionary GPU kernel design for sparse matrix-matrix multiplication.

Redundancy-free and load-balanced TGNN training with hierarchical pipeline parallelism
Redundancy-free and load-balanced TGNN training with hierarchical pipeline parallelism

Yaqi Xia, Zheng Zhang, Donglin Yang, Chuang Hu, Xiaobo Zhou, Hongyang Chen, Qianlong Sang, Dazhao Cheng†(† corresponding author)

IEEE Transactions on Parallel and Distributed (TPDS) 2024 JournalCCF-A

This work introduces Sven, a co-designed algorithm-system library aimed at accelerating TGNN training on a multi-GPU platform.

Redundancy-free and load-balanced TGNN training with hierarchical pipeline parallelism
Redundancy-free and load-balanced TGNN training with hierarchical pipeline parallelism

Yaqi Xia, Zheng Zhang, Donglin Yang, Chuang Hu, Xiaobo Zhou, Hongyang Chen, Qianlong Sang, Dazhao Cheng†(† corresponding author)

IEEE Transactions on Parallel and Distributed (TPDS) 2024 JournalCCF-A

This work introduces Sven, a co-designed algorithm-system library aimed at accelerating TGNN training on a multi-GPU platform.

Scaling New Heights :Transformative Cross-GPU Sampling for Training Billion-Edge Graphs
Scaling New Heights :Transformative Cross-GPU Sampling for Training Billion-Edge Graphs

Yaqi Xia, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2024 ConferenceCCF-A

In this paper, we introduced HyDRA, a pioneering framework for sampling-based GNN training on large-scale graphs.

Scaling New Heights :Transformative Cross-GPU Sampling for Training Billion-Edge Graphs
Scaling New Heights :Transformative Cross-GPU Sampling for Training Billion-Edge Graphs

Yaqi Xia, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2024 ConferenceCCF-A

In this paper, we introduced HyDRA, a pioneering framework for sampling-based GNN training on large-scale graphs.

MPMoE :Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism
MPMoE :Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism

Zheng Zhang, Yaqi Xia, Hulin Wang, Donglin Yang, Chuang Hu, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

IEEE Transactions on Parallel and Distributed (TPDS) 2024 JournalCCF-ABest Paper Runner-up

In this paper, we present the design and implementation of MPMoE, a high-performance library that accelerates MoE training with adaptive and memory-efficient pipeline parallelism.

MPMoE :Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism
MPMoE :Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism

Zheng Zhang, Yaqi Xia, Hulin Wang, Donglin Yang, Chuang Hu, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

IEEE Transactions on Parallel and Distributed (TPDS) 2024 JournalCCF-ABest Paper Runner-up

In this paper, we present the design and implementation of MPMoE, a high-performance library that accelerates MoE training with adaptive and memory-efficient pipeline parallelism.

Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism
Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism

Yaqi Xia, Zheng Zhang, Hulin Wang, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

The 32nd International Symposium on High-Performance Parallel and Distributed Computing (ACM HPDC) 2023 ConferenceCCF-ABest Paper Nomination

This paper presents Sven, an algorithm and system co-designed TGNN training library for the end-to-end performance optimization on multi-node multi-GPU systems.

Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism
Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism

Yaqi Xia, Zheng Zhang, Hulin Wang, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

The 32nd International Symposium on High-Performance Parallel and Distributed Computing (ACM HPDC) 2023 ConferenceCCF-ABest Paper Nomination

This paper presents Sven, an algorithm and system co-designed TGNN training library for the end-to-end performance optimization on multi-node multi-GPU systems.

ASFM-Net :Asymmetrical Siamese Feature Matching Network for Point Completion
ASFM-Net :Asymmetrical Siamese Feature Matching Network for Point Completion

Yaqi Xia*, Yan Xia*, Wei Li, Rui Song, Kailang Cao, Uwe Stilla†(* equal contribution; † corresponding author)

Proceedings of the 29th ACM international conference on multimedia (ACM MM) 2021 ConferenceCCF-A

We tackle the problem of object completion from point clouds and propose a novel point cloud completion network employing an Asymmetrical Siamese Feature Matching strategy, termed as ASFM-Net.

ASFM-Net :Asymmetrical Siamese Feature Matching Network for Point Completion
ASFM-Net :Asymmetrical Siamese Feature Matching Network for Point Completion

Yaqi Xia*, Yan Xia*, Wei Li, Rui Song, Kailang Cao, Uwe Stilla†(* equal contribution; † corresponding author)

Proceedings of the 29th ACM international conference on multimedia (ACM MM) 2021 ConferenceCCF-A

We tackle the problem of object completion from point clouds and propose a novel point cloud completion network employing an Asymmetrical Siamese Feature Matching strategy, termed as ASFM-Net.

All publications