夏亚奇 Yaqi Xia

武汉大学 武汉大学计算机学院弘毅博士后
[个人简历]

我是夏亚奇,现为武汉大学弘毅博士后,合作导师为程大钊教授。我的研究聚焦于面向人工智能与机器学习的分布式系统和高性能计算,重点关注长上下文大模型高效推理、可扩展图学习系统及机器学习系统优化。

我的研究成果发表于 SC、PPoPP、USENIX ATC、IEEE TPDS 和 HPDC 等系统领域重要会议与期刊,并曾获最佳论文亚军和最佳论文提名。

Portrait of 夏亚奇

学生招募

欢迎对分布式系统、高性能计算以及人工智能与机器学习系统感兴趣,且具有较强自驱力的本科生、硕士生和博士生与我联系。如有合作意向,请通过邮件联系我。

荣誉与奖励
  • CCF博士论文激励计划(CCF Computility 2026)
  • 最佳论文提名,SC 26 2026
  • 最佳论文亚军,IEEE TPDS 24 2024
  • 最佳论文提名,ACM HPDC 23 2023
  • 滴滴优秀博士生学术论坛一等奖 2025
  • 武汉大学研究生未来创新学术论坛博士生论坛最佳汇报奖 2025
  • 首届“天智杯”人工智能挑战赛二等奖 2019
学术服务
  • 副主编 The Journal of Supercomputing
  • IEEE TPDS / IEEE TDSC 审稿人
科研项目
  • 云端协同计算项目

    云端协同计算项目

    华为 2012 实验室

    项目负责人 在研

  • 中国博士后科学基金面上资助

    中国博士后科学基金面上资助

    主持人

  • 中国博士后国家资助计划

    中国博士后国家资助计划

    主持人

  • 之江实验室科研项目

    之江实验室科研项目

    之江实验室图计算研究中心

    项目负责人

News
2026
My doctoral dissertation has been selected for the 2026 CCF Doctoral Dissertation Incentive Program (CCF Computility).
Jul 24
Our work 'DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference' has been nominated for the Best Paper Award at SC 26.
Jul 24
Our work 'DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference' is accepted by SC 2026.
Jul 02
Our work 'QDiTx: Arbitrary Bit-Width Quantization Framework for Diffusion Transformer' is accepted by IEEE TCAD 2026.
Jul 02
2025
Our work 'JanusQuant: Accurate and Efficient 2-bit KV Cache Quantization for Long-context Inference' is accepted by PPoPP 2026.
Nov 11
Our work 'MPMoE: Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism' has been selected as the Runner-Up of the 2024 Best Paper Award from IEEE TPDS.
Aug 12
Our work 'MXBLAS: Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library' is accepted by SC 2025.
Jun 27
Our work 'Voltrix: Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel Optimization' is accepted by ATC 2025.
Apr 27
2024
Our work 'Harnessing Inter-GPU Shared Memory for Seamless MoE Communication-Computation Fusion' is accepted by PPoPP 2025.
Dec 12
Our work 'Redundancy-Free and Load-Balanced TGNN Training With Hierarchical Pipeline Parallelism' is accepted by TPDS 2024.
Nov 11
Our work 'Scaling New Heights :Transformative Cross-GPU Sampling for Training Billion-Edge Graphs' is accepted by SC 2024.
May 16
Our work 'Accelerating Distributed DLRM Training with Optimized TT Decomposition and Micro-Batching' is accepted by SC 2024.
Apr 16
Our work 'Raptor-T :A Fused and Memory-Efficient Sparse Transformer for Long and Variable-Length Sequences' is accepted by TC 2024.
Apr 16
Our work 'MPMoE :Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism' is accepted by TPDS 2024.
Apr 08
2023
Our work 'Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism' is accepted by HPDC 2023 and be selected as Best Paper Nomination (only two nominations)!
Aug 07
Our work 'MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism' is accepted by IPDPS 2023.
Jul 18
代表性论文 (查看全部 )
DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference
DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference

Chengyu Sun, Yaqi Xia†, Ruirui Pan, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

2026 International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2026 ConferenceCCF-ABest Paper Award Nomination

We present DySpin, a plug-and-play library that advances dynamic sparse long-context inference.

DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference
DySpin: A Plug-and-Play Library Advancing Dynamic Sparse Long-Context Inference

Chengyu Sun, Yaqi Xia†, Ruirui Pan, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

2026 International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2026 ConferenceCCF-ABest Paper Award Nomination

We present DySpin, a plug-and-play library that advances dynamic sparse long-context inference.

MXBLAS:Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library.
MXBLAS:Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library.

Weihu Wang*, Yaqi Xia*, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(* equal contribution; † corresponding author)

2025 International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2025 ConferenceCCF-A

We present MXBLAS, a high-performance MX-GEMM library that unifies support across the full spectrum of MX-format variations.

MXBLAS:Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library.
MXBLAS:Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library.

Weihu Wang*, Yaqi Xia*, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(* equal contribution; † corresponding author)

2025 International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2025 ConferenceCCF-A

We present MXBLAS, a high-performance MX-GEMM library that unifies support across the full spectrum of MX-format variations.

Voltrix:Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel Optimization
Voltrix:Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel Optimization

Yaqi Xia*, Weihu Wang*, Donglin Yang, Xiaobo Zhou†, Dazhao Cheng†(* equal contribution; † corresponding author)

2025 USENIX Annual Technical Conference (ATC) 2025 ConferenceCCF-A

We introduce Voltrix-SpMM, a revolutionary GPU kernel design for sparse matrix-matrix multiplication.

Voltrix:Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel Optimization
Voltrix:Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel Optimization

Yaqi Xia*, Weihu Wang*, Donglin Yang, Xiaobo Zhou†, Dazhao Cheng†(* equal contribution; † corresponding author)

2025 USENIX Annual Technical Conference (ATC) 2025 ConferenceCCF-A

We introduce Voltrix-SpMM, a revolutionary GPU kernel design for sparse matrix-matrix multiplication.

Redundancy-free and load-balanced TGNN training with hierarchical pipeline parallelism
Redundancy-free and load-balanced TGNN training with hierarchical pipeline parallelism

Yaqi Xia, Zheng Zhang, Donglin Yang, Chuang Hu, Xiaobo Zhou, Hongyang Chen, Qianlong Sang, Dazhao Cheng†(† corresponding author)

IEEE Transactions on Parallel and Distributed (TPDS) 2024 JournalCCF-A

This work introduces Sven, a co-designed algorithm-system library aimed at accelerating TGNN training on a multi-GPU platform.

Redundancy-free and load-balanced TGNN training with hierarchical pipeline parallelism
Redundancy-free and load-balanced TGNN training with hierarchical pipeline parallelism

Yaqi Xia, Zheng Zhang, Donglin Yang, Chuang Hu, Xiaobo Zhou, Hongyang Chen, Qianlong Sang, Dazhao Cheng†(† corresponding author)

IEEE Transactions on Parallel and Distributed (TPDS) 2024 JournalCCF-A

This work introduces Sven, a co-designed algorithm-system library aimed at accelerating TGNN training on a multi-GPU platform.

Scaling New Heights :Transformative Cross-GPU Sampling for Training Billion-Edge Graphs
Scaling New Heights :Transformative Cross-GPU Sampling for Training Billion-Edge Graphs

Yaqi Xia, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2024 ConferenceCCF-A

In this paper, we introduced HyDRA, a pioneering framework for sampling-based GNN training on large-scale graphs.

Scaling New Heights :Transformative Cross-GPU Sampling for Training Billion-Edge Graphs
Scaling New Heights :Transformative Cross-GPU Sampling for Training Billion-Edge Graphs

Yaqi Xia, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) 2024 ConferenceCCF-A

In this paper, we introduced HyDRA, a pioneering framework for sampling-based GNN training on large-scale graphs.

MPMoE :Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism
MPMoE :Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism

Zheng Zhang, Yaqi Xia, Hulin Wang, Donglin Yang, Chuang Hu, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

IEEE Transactions on Parallel and Distributed (TPDS) 2024 JournalCCF-ABest Paper Runner-up

In this paper, we present the design and implementation of MPMoE, a high-performance library that accelerates MoE training with adaptive and memory-efficient pipeline parallelism.

MPMoE :Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism
MPMoE :Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism

Zheng Zhang, Yaqi Xia, Hulin Wang, Donglin Yang, Chuang Hu, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

IEEE Transactions on Parallel and Distributed (TPDS) 2024 JournalCCF-ABest Paper Runner-up

In this paper, we present the design and implementation of MPMoE, a high-performance library that accelerates MoE training with adaptive and memory-efficient pipeline parallelism.

Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism
Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism

Yaqi Xia, Zheng Zhang, Hulin Wang, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

The 32nd International Symposium on High-Performance Parallel and Distributed Computing (ACM HPDC) 2023 ConferenceCCF-ABest Paper Nomination

This paper presents Sven, an algorithm and system co-designed TGNN training library for the end-to-end performance optimization on multi-node multi-GPU systems.

Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism
Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism

Yaqi Xia, Zheng Zhang, Hulin Wang, Donglin Yang, Xiaobo Zhou, Dazhao Cheng†(† corresponding author)

The 32nd International Symposium on High-Performance Parallel and Distributed Computing (ACM HPDC) 2023 ConferenceCCF-ABest Paper Nomination

This paper presents Sven, an algorithm and system co-designed TGNN training library for the end-to-end performance optimization on multi-node multi-GPU systems.

ASFM-Net :Asymmetrical Siamese Feature Matching Network for Point Completion
ASFM-Net :Asymmetrical Siamese Feature Matching Network for Point Completion

Yaqi Xia*, Yan Xia*, Wei Li, Rui Song, Kailang Cao, Uwe Stilla†(* equal contribution; † corresponding author)

Proceedings of the 29th ACM international conference on multimedia (ACM MM) 2021 ConferenceCCF-A

We tackle the problem of object completion from point clouds and propose a novel point cloud completion network employing an Asymmetrical Siamese Feature Matching strategy, termed as ASFM-Net.

ASFM-Net :Asymmetrical Siamese Feature Matching Network for Point Completion
ASFM-Net :Asymmetrical Siamese Feature Matching Network for Point Completion

Yaqi Xia*, Yan Xia*, Wei Li, Rui Song, Kailang Cao, Uwe Stilla†(* equal contribution; † corresponding author)

Proceedings of the 29th ACM international conference on multimedia (ACM MM) 2021 ConferenceCCF-A

We tackle the problem of object completion from point clouds and propose a novel point cloud completion network employing an Asymmetrical Siamese Feature Matching strategy, termed as ASFM-Net.

全部论文