Ruixiang Xue薛瑞翔

研究内容

项目

智能点云压缩

面向 MPEG AI-PCC 的深度学习点云压缩算法研究与开发。

  • 针对点云数据表征设计 Transformer 网络,提高空间特征提取效率。
  • 利用点云降采样过程中的 密度变化 指导重建模式选择。
  • 在公开测试条件下,相较基线实现 39% 几何压缩增益34% 属性压缩增益
标准化提案
  1. MPEG m70061 · 2024.11 AI-PCC CfP Response Proposal from Nanjing University and OPPO (Track1) T. Chen, R. Xue, J. Zhang, J. Wang, Z. Ma, S. Xia, R. Xue, Z. Sun, C. Ma, Y. Yu, H. Yu, D. Wang
  2. MPEG m70062 · 2024.11 AI-PCC CfP Response Proposal from Nanjing University and OPPO (Track2) T. Chen, R. Xue, J. Zhang, J. Wang, Z. Ma, S. Xia, R. Xue, Z. Sun, C. Ma, Y. Yu, H. Yu, D. Wang
  3. MPEG m70395 · 2024.11 [AI-GC][CfP-related] Improved Cross-Platform Reproducibility for AI-Based Point Cloud Compression J. Zhang, T. Chen, R. Xue, Z. Ma, S. Xia, Z. Sun, C. Ma, Y. Yu, H. Yu, D. Wang
  4. MPEG m70396 · 2024.11 [AI-GC][CfP-related] On Model Quantization for AI-Based Point Cloud Compression J. Zhang, T. Chen, R. Xue, Z. Ma, S. Xia, Z. Sun, C. Ma, Y. Yu, H. Yu, D. Wang
  5. MPEG m64417 · 2023.07 [AI-3DGC EE 5.4] Update On the Training Datasets for Attribute Compression J. Wang, R. Xue, J. Li, Z. Ma, H. Wei, Y. Yu, V. Zakharchenko, D. Wang
  6. MPEG m64418 · 2023.07 [AI-3DGC EE 5.1] Performance Comparison of Dynamic Point Cloud Geometry Compression for LiDAR Point Clouds R. Xue, J. Wang, J. Li, Z. Ma, H. Wei, Y. Yu, V. Zakharchenko, D. Wang
  7. MPEG m62176 · 2023.01 [AI-3DGC] On the Training Datasets for Attribute Compression J. Wang, R. Xue, J. Li, Z. Ma, H. Wei, Y. Yu, V. Zakharchenko, D. Wang
  8. MPEG m62177 · 2023.01 [AI-3DGC][EE5.1-related][EE5.3-related] Dynamic Point Cloud Geometry Compression for LiDAR Point Cloud with Ego-Motion Compensation R. Xue, J. Wang, J. Li, Z. Ma, H. Wei, Y. Yu, V. Zakharchenko, D. Wang
  9. MPEG m60353 · 2022.07 SparsePCGCv2: Multihead Neighborhood Point Attention for Sparse Point Cloud Ruixiang Xue, Jianqiang Wang, Zhan Ma, Honglian Wei, Yue Yu, Vladyslav Zakharchenko, Dong Wang
  10. MPEG m59552 · 2022.04 SparsePCGCv2: Improved SparsePCGC with attention mechanism Jianqiang Wang, Ruixiang Xue, Zhan Ma, Honglian Wei, Yue Yu, Vladyslav Zakharchenko, Dong Wang
专利申请
  1. 2024 编解码方法、编码器、解码器以及存储介质 马展,薛瑞翔,魏红莲
  2. 2023 基于隐式神经网络和深度图投影的激光雷达点云序列表征方法 薛瑞翔,李嘉欣,马展
  3. 2023 基于神经网络的动态激光雷达点云多尺度几何无损压缩方法 马展,薛瑞翔
  4. 2022 基于神经网络的点云几何压缩后处理方法 马展,薛瑞翔
  5. 2022 基于注意力机制和稀疏卷积的点云几何无损压缩方法 薛瑞翔,王剑强,马展 对应申请名:点云几何信息的压缩、解压缩及点云视频编解码方法、装置
三维高斯泼溅编码

面向 3D Gaussian Splatting 场景的紧凑化、压缩、渐进式传输与区域自适应质量控制。

  • 提出表示重组方法,将预训练 3DGS 重构为区域感知的分层表示,实现灵活的区域自适应质量控制。
  • 提出前馈式分层压缩方法,利用跨层依赖关系建模,支持可截断码流、渐进式压缩与流式传输。
  • 在紧凑化/压缩性能上相较基线实现 70% / 35% 总体增益,并支持基于 ROI 的 3DGS 渐进式压缩。
自动驾驶世界模型

参与自动驾驶世界模型方向国家重点研发项目及相关技术路线调研,聚焦街景大幅度新视角外推。

  • 探索 UniSplat 前馈重建与 VISTA 生成模型的双向增强:用几何约束提升生成一致性,用生成先验改进前馈渲染质量。
  • 调研三维/四维重建、重建-生成耦合与自动驾驶世界模型等技术路线。
智能辅助摄影
Virtual Studio Pipeline
Unreal Engine 组合导入
参考场景图

四张参考场景图像作为虚拟摄影棚世界创建输入。

场景图像 1
场景图像 1
场景图像 2
场景图像 2
场景图像 3
场景图像 3
场景图像 4
场景图像 4
Marble 场景 3DGS

四个场景 3DGS 视频预览;视频不会自动播放。

Marble 预览 1
Marble 预览 2
Marble 预览 3
Marble 预览 4
参考人物图

七张人物图像作为独立主体资产输入。

人物图像 1
人物图像 1
人物图像 2
人物图像 2
人物图像 3
人物图像 3
人物图像 4
人物图像 4
人物图像 5
人物图像 5
人物图像 6
人物图像 6
人物图像 7
人物图像 7
SAM 人物 mask

SAM 输出七张人物 mask,为后续人物 3DGS 重建提供干净主体。

SAM mask 1
SAM mask 1
SAM mask 2
SAM mask 2
SAM mask 3
SAM mask 3
SAM mask 4
SAM mask 4
SAM mask 5
SAM mask 5
SAM mask 6
SAM mask 6
SAM mask 7
SAM mask 7
SHARP 人物 3DGS

SHARP 将分割后的人物重建为七个可组合的人物 3DGS 资产。

人物 3DGS 1
人物 3DGS 1
人物 3DGS 2
人物 3DGS 2
人物 3DGS 3
人物 3DGS 3
人物 3DGS 4
人物 3DGS 4
人物 3DGS 5
人物 3DGS 5
人物 3DGS 6
人物 3DGS 6
人物 3DGS 7
人物 3DGS 7
虚拟 Shooting

在虚拟摄影棚中调整视角并生成拍摄候选结果。

虚拟拍摄结果
虚拟拍摄结果
神经渲染照片

通过神经渲染将原始虚拟拍摄结果提升为更接近照片的结果。

神经渲染输出
神经渲染输出
空间重构流程
相机姿态 视角变换
输入照片

用户提供单张参考图像。

参考图像
参考图像
图像外扩

这里只展示外扩后的第三张结果。

外扩结果
外扩结果
SHARP 3DGS 预览

录屏展示前景-背景 SHARP 3DGS 组合与受限视角重构效果。

SHARP/3DGS Reframe 录屏
新视角渲染

显示作为后续修复输入的新视角渲染结果。

修复输入
修复输入
重构照片

输出最终空间重构结果。

最终结果
最终结果

教育经历

南京大学

信息与通信工程博士研究生

2021.09 - 2027.06

  • NJU Vision Lab,导师:马展教授、陈彤研究员
  • 研究方向:智能点云压缩三维高斯泼溅压缩隐式神经表达智能辅助摄影
  • 博士中期考核优秀

杭州电子科技大学

电子信息工程工学学士

2017.09 - 2021.06

工作经历

吉利

人工智能中心 · 算法实习生

2026.05 - 至今

围绕街景新视角外推开展研究,探索前馈重建模型(UniSplat)与生成模型(VISTA)之间的双向增强。

  • 利用前馈街景重建模型提升生成模型的几何一致性。
  • 利用生成先验改进前馈街景重建方法的渲染质量损失,以提升大幅度新视角外推效果。

OPPO

研究院 · 算法实习生

2024.02 - 2024.11

围绕智能点云压缩算法研究与 MPEG AI-PCC 标准化开展工作。

  • 研发和评估面向 MPEG AI-PCC 标准化的智能点云压缩算法。
  • 开发点云编解码器软件,并在不同测试序列上评估压缩性能。
  • 多次参加 MPEG 国际会议,提交 10 份标准化提案,并参与 5 项发明专利申请。

论文

  1. 3D Gaussian Splatting Compression with Object Scalability
    ECCV2026 · European Conference on Computer Vision
    Ruixiang Xue, Tong Chen, Zhan Ma
    CCF-B三维高斯泼溅压缩
    PaperCode
    Abstract

    We introduce a framework towards scalable, finer-grained object-level 3DGS compression. First, a post-training method named RecastGS is proposed to reorganize pretrained 3DGS into a layered representation and progressively distills cumulative submodels to improve rate-distortion efficiency. Leveraging multi-view SAM predictions from user click prompts, Gaussians are further partitioned into user-defined regions of interest (ROI), enabling region-adaptive quality control without retraining. Second, built upon this reorganized region-aware layered hierarchy, a feed-forward 3DGS compression method named LayeredCGS is proposed to compress position using a lightweight point cloud codec and attributes with a layer-wise context model to exploit cross-layer correlations. Extensive experiments show that LayeredCGS achieves 35% BD-Rate gain over the existing feed-forward method FCGS. With progressive distillation in RecastGS enabled, our method further outperforms most per-scene optimization methods. Moreover, the proposed method supports ROI-aware compression and flexible bitstream truncation, achieving up to 2 dB higher ROI PSNR at comparable bitrates compared with the uniform quality allocation baseline while enabling low-latency preview and progressive quality refinement. The code will be released at https://github.com/RuixiangXue/ScalableGSC.

    BibTeX
    @inproceedings{xue2026objectscalable3dgs,
      title={3D Gaussian Splatting Compression with Object Scalability},
      author={Xue, Ruixiang and Chen, Tong and Ma, Zhan},
      booktitle={European Conference on Computer Vision (ECCV)},
      year={2026}
    }
  2. A Versatile Point Cloud Compressor Using Universal Multiscale Conditional Coding – Part I: Geometry
    TPAMI · IEEE Transactions on Pattern Analysis and Machine Intelligence
    Vol. 47, No. 1, pp. 269-287, Jan. 2025 · DOI: 10.1109/TPAMI.2024.3462938
    Jianqiang Wang, Ruixiang Xue, Jiaxin Li, Dandan Ding, Yi Lin, Zhan Ma
    SCI 一区CCF-AIF 18.6共同一作智能点云压缩
    Abstract

    A universal multiscale conditional coding framework, Unicorn, is proposed to compress the geometry and attribute of any given point cloud. Geometry compression is addressed in Part I of this paper, while attribute compression is discussed in Part II. We construct the multiscale sparse tensors of each voxelized point cloud frame and properly leverage lower-scale priors in the current and (previously processed) temporal reference frames to improve the conditional probability approximation or content-aware predictive reconstruction of geometry occupancy in compression. Unicorn is a versatile, learning-based solution capable of compressing static and dynamic point clouds with diverse source characteristics in both lossy and lossless modes. Following the same evaluation criteria, Unicorn significantly outperforms standard-compliant approaches like MPEG G-PCC, V-PCC, and other learning-based solutions, yielding state-of-the-art compression efficiency while presenting affordable complexity for practical implementations.

    BibTeX
    @article{wang2025unicorngeometry,
      title={A Versatile Point Cloud Compressor Using Universal Multiscale Conditional Coding -- Part I: Geometry},
      author={Wang, Jianqiang and Xue, Ruixiang and Li, Jiaxin and Ding, Dandan and Yi, Lin and Ma, Zhan},
      journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
      volume={47},
      number={1},
      pages={269--287},
      year={2025},
      doi={10.1109/TPAMI.2024.3462938}
    }
  3. NeRI: Implicit Neural Representation of LiDAR Point Cloud Using Range Image Sequence
    ICASSP 2024 · IEEE International Conference on Acoustics, Speech, and Signal Processing
    ICASSP 2024, pp. 8020-8024 · DOI: 10.1109/ICASSP48485.2024.10446596
    Ruixiang Xue, Jiaxin Li, Tong Chen, Dandan Ding, Xun Cao, Zhan Ma
    CCF-B第一作者隐式神经表达智能点云压缩
    PaperCode
    Abstract

    This paper proposes the NeRI, an implicit neural representation (INR) based LiDAR point cloud compressor. In NeRI, we first transform a sequence of 3D LiDAR frames into a 2D range image sequence through range image projection over time. Then, we employ a neural network conditioned on the temporal frame index and associated LiDAR sensor pose to fit input range images as closely as possible. The optimized network parameters, which implicitly represent the input LiDAR data, are later lossily compressed. NeRI decoder is then initialized using decoded parameters to generate range images for reconstructing the 3D LiDAR sequence accordingly. Extensive experimental results demonstrate the significant superiority of NeRI regarding the compression efficiency and decoding speed compared to state-of-the-art 2D and 3D compressors for LiDAR point cloud.

    BibTeX
    @inproceedings{xue2024neri,
      title={NeRI: Implicit Neural Representation of LiDAR Point Cloud Using Range Image Sequence},
      author={Xue, Ruixiang and Li, Jiaxin and Chen, Tong and Ding, Dandan and Cao, Xun and Ma, Zhan},
      booktitle={ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
      pages={8020--8024},
      year={2024},
      doi={10.1109/ICASSP48485.2024.10446596}
    }
    Poster
    NeRI: Implicit Neural Representation of LiDAR Point Cloud Using Range Image Sequence poster
  4. A Versatile Point Cloud Compressor Using Universal Multiscale Conditional Coding – Part II: Attribute
    TPAMI · IEEE Transactions on Pattern Analysis and Machine Intelligence
    Vol. 47, No. 1, pp. 252-268, Jan. 2025 · DOI: 10.1109/TPAMI.2024.3462945
    Jianqiang Wang, Ruixiang Xue, Jiaxin Li, Dandan Ding, Yi Lin, Zhan Ma
    SCI 一区CCF-AIF 18.6第二作者智能点云压缩
    Abstract

    A universal multiscale conditional coding framework, Unicorn, is proposed to code the geometry and attribute of any given point cloud. Attribute compression is discussed in Part II of this paper, while geometry compression is given in Part I of this paper. We first construct the multiscale sparse tensors of each voxelized point cloud attribute frame. Since attribute components exhibit very different intrinsic characteristics from the geometry element, e.g., 8-bit RGB color versus 1-bit occupancy, we process the attribute residual between lower-scale reconstruction and current-scale data. Similarly, we leverage spatially lower-scale priors in the current frame and (previously processed) temporal reference frame to improve the probability estimation of attribute intensity through conditional residual prediction in lossless mode or enhance the attribute reconstruction through progressive residual refinement in lossy mode for better performance. The proposed Unicorn is a versatile, learning-based solution capable of compressing a great variety of static and dynamic point clouds in both lossy and lossless modes. Following the same evaluation criteria, Unicorn significantly outperforms standard-compliant approaches like MPEG G-PCC, V-PCC, and other learning-based solutions, yielding state-of-the-art compression efficiency with affordable encoding/decoding runtime.

    BibTeX
    @article{wang2025unicornattribute,
      title={A Versatile Point Cloud Compressor Using Universal Multiscale Conditional Coding -- Part II: Attribute},
      author={Wang, Jianqiang and Xue, Ruixiang and Li, Jiaxin and Ding, Dandan and Yi, Lin and Ma, Zhan},
      journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
      volume={47},
      number={1},
      pages={252--268},
      year={2025},
      doi={10.1109/TPAMI.2024.3462945}
    }
  5. GRNet: Geometry Restoration for G-PCC Compressed Point Clouds Using Auxiliary Density Signaling
    TVCG · IEEE Transactions on Visualization and Computer Graphics
    Vol. 30, No. 10, pp. 6740-6753, Oct. 2024 · DOI: 10.1109/TVCG.2023.3336936
    Gexin Liu, Ruixiang Xue, Jiaxin Li, Dandan Ding, Zhan Ma
    SCI 一区CCF-AIF 6.5第二作者智能点云压缩
    Paper
    Abstract

    The lossy Geometry-based Point Cloud Compression (G-PCC) inevitably impairs the geometry information of point clouds, which deteriorates the quality of experience (QoE) in reconstruction and/or misleads decisions in tasks such as classification. To tackle it, this work proposes GRNet for the geometry restoration of G-PCC compressed large-scale point clouds. By analyzing the content characteristics of original and G-PCC compressed point clouds, we attribute the G-PCC distortion to two key factors: point vanishing and point displacement. Visible impairments on a point cloud are usually dominated by an individual factor or superimposed by both factors, which are determined by the density of the original point cloud. To this end, we employ two different models for coordinate reconstruction, termed Coordinate Expansion and Coordinate Refinement, to attack the point vanishing and displacement, respectively. In addition, 4-byte auxiliary density information is signaled in the bitstream to assist the selection of Coordinate Expansion, Coordinate Refinement, or their combination. Before being fed into the coordinate reconstruction module, the G-PCC compressed point cloud is first processed by a Feature Analysis Module for multiscale information fusion, in which kNN-based Transformer is leveraged at each scale to adaptively characterize neighborhood geometric dynamics for effective restoration. Following the common test conditions recommended in the MPEG standardization committee, GRNet significantly improves the G-PCC anchor and remarkably outperforms state-of-the-art methods on a great variety of point clouds (e.g., solid, dense, and sparse samples) both quantitatively and qualitatively. Meanwhile, GRNet runs fairly fast and uses a smaller-size model when compared with existing learning-based approaches, making it attractive to industry practitioners.

    BibTeX
    @article{liu2024grnet,
      title={GRNet: Geometry Restoration for G-PCC Compressed Point Clouds Using Auxiliary Density Signaling},
      author={Liu, Gexin and Xue, Ruixiang and Li, Jiaxin and Ding, Dandan and Ma, Zhan},
      journal={IEEE Transactions on Visualization and Computer Graphics},
      volume={30},
      number={10},
      pages={6740--6753},
      year={2024},
      doi={10.1109/TVCG.2023.3336936}
    }

荣誉

*
南京大学学业一等奖学金
2021 - 2025
*
浙江省第十二届大学生创业计划竞赛特等奖
2020