Xidian University logo
CVIA · Industry & Aerospace Vision Group
Computer Vision for Industry and Aerospace · XDU
Faculty Profile

宋锐 · Rui Song

Professor and PhD advisor. His research and applications center on 3D spatial intelligence, 3D localization and measurement, 6D pose estimation, and engineering vision systems.

宋锐 · Rui Song
Professor · PhD Advisor

Xidian University · School of Telecommunications Engineering · ISN State Key Laboratory

Institute of Image Transmission and Processing

Email: rsong at xidian.edu.cn

Office: Room B-201, Science Building · Tel: +86(29)88202607

Mail: Mailbox 103, No. 2 Taibai South Road, Yanta District, Xi’an, China

3D spatial intelligence3D localization and measurement6D pose estimationIndustrial machine vision
Academic Profile

He received the B.S., M.S., and Ph.D. degrees from Xidian University in 2003, 2006, and 2009, respectively, and was a visiting researcher at the University of Southern California from 2012 to 2013.

He is a member of the ISN State Key Laboratory and the Institute of Image Transmission and Processing, a member of IEEE, a council member of the Shaanxi Society of Image and Graphics, and a committee member of both the CCF-CV and CSIG-3DV technical committees. He also serves as a reviewer for IEEE TPAMI, TNNLS, TMM, TCSVT, TGRS, TIP, Pattern Recognition, and Remote Sensing.

He leads the Shaanxi University Youth Innovation Team and serves as chief scientist of the Shaanxi “Scientist + Engineer” team.

His honors include individual commendation for Chang’e-5, a first-class Ministry of Education Science and Technology Progress Award, a second-class Shaanxi Science and Technology Award, a first-class Surveying and Mapping Science and Technology Award, a second-class AVIC Science and Technology Progress Award, the Shaanxi Rising Science Star award, and Xidian University Huashan Scholar recognition.

His main research interests include 3D spatial intelligence, 3D localization and 6D pose estimation in computer vision, as well as intelligent processing of remote sensing imagery.

His research follows the principle that problems should be distilled from engineering practice and results should be validated in real applications, with continuous closed-loop advancement from algorithm design to hardware implementation and system deployment.

EducationB.S., M.S., and Ph.D. from Xidian University · Visiting scholar at USC
Recent publicationsMore than 60 SCI papers in the past five years, including over 40 as first or corresponding author
Engineering focusMachine vision algorithms for aerospace, aviation, deep-sea, and industrial intelligence applications
Research areas3D spatial intelligence, 3D localization, and 6D pose estimation and tracking
News
  • [2026-03-06] PhD openings are still available for 2026. Students interested in computer vision, image analysis, 3D spatial measurement, and 6D pose estimation are welcome to email him.
  • [2026-02-27] One additional academic master’s opening is available for 2026. Interested students are welcome to contact him by email.
  • [2026-02-21] Congratulations to PhD student Zhiqiang Liu. His paper “Exploring 6D Object Pose Estimation with Deformation” has been accepted by CVPR 2026.
  • [2026-02-09] Congratulations to PhD student Rong Qi. Her paper “Leveraging Spatiotemporal Cues for Self-Supervised Stereo Depth Estimation in Endoscopic Videos” has been accepted by IEEE Transactions on Medical Imaging.
  • [2025-06-16] Congratulations to master’s student Luyuan Liu. His paper “ASMF: A Self-Supervised Atmospheric Scatter Model-Based Fusion Network for Infrared Image Enhancement” has been accepted by IEEE Transactions on Geoscience and Remote Sensing.
  • [2025-02-27] Congratulations to master’s student Qingyuan Wang. His paper “SCFlow2: Plug-and-Play Object Pose Refiner with Shape-Constraint Scene Flow” has been accepted by CVPR 2025.

Research & Output

Research Interests

His major research directions are 3D spatial intelligence, 3D localization, and 6D pose estimation for computer vision. In recent years he has built a continuous line of work on geometric perception of complex space targets, target pose estimation, structure-constrained modeling, point-cloud understanding, and deployment of engineering vision systems, while also covering key algorithms and architectures for image/video compression and transmission.

Over the past five years, he has led multiple government-funded research projects and more than ten industry-commissioned R&D tasks from institutes affiliated with CASC, AVIC, CAS, NORINCO, and CETC. During the same period, he has published more than 60 SCI papers, including more than 40 as first or corresponding author, with over 50 papers in CAS Q2 or above, together with multiple national invention patents.

Research Results & Applications

Research Results & Applications

Chang’e-5 compression engine

Core Compression Engine for Chang’e-5

For high-fidelity downlink of lunar-surface operation videos captured by the lander robotic arm, the group developed a real-time multi-channel, multi-mode encoding unit for the ascender sampling and separation-monitoring cameras. The system supports five camera streams with different resolutions, operating modes, compression formats, and analysis requirements.

Tiangong/Tianzhou docking system

Real-Time Monitoring and Guidance System for Tiangong / Tianzhou Docking

To support real-time video monitoring and guidance during rendezvous and docking between the space station and cargo spacecraft, the group developed a multi-channel HD video communication and image-processing system for stable mission support and ground analysis.

Deep-sea ultra-HD compression system

Ultra-HD Video Compression System for 10,000-Meter Deep-Sea Missions

For live broadcasting during Mariana Trench deep-sea trials, the group developed a multi-channel cinema-grade 4K encoding system together with ship-side decoding equipment, achieving stable performance in a harsh deep-sea environment.

SSA pose perception unit

Target Pose Perception Unit for Space Situational Awareness

For attitude perception and early warning of complex space targets, the group built a real-time intelligent perception unit based on target detection, pose estimation, and pose tracking, enabling more autonomous space-situation warning capability.

Aerial refueling localization

Precision Localization System for Aerial Refueling

To meet the need for precise localization in aerial refueling, the group developed an integrated system for visual perception, target measurement, and motion guidance, supporting high-precision relative pose estimation under complex flight conditions.

2D-3D registration and DSM projection

2D–3D Registration and Real-Time DSM Projection for UAV Monitoring Imagery

The group studied key techniques for 2D–3D registration and real-time DSM projection to fuse UAV surveillance imagery with geospatial data. Pixels in the UAV field of view can be projected directly onto a 3D DSM, improving real-time monitoring, localization, and situational understanding.

Algorithm Competitions

Awards and Challenge Results

2023 ICCV BOP Challenge champion

2023 · ICCV BOP Challenge — Single-Model Track Champion

A joint team from Xidian University, EPFL, and Magic Leap won first place in the single-model track of the BOP Challenge and was invited to present at the R6D workshop. The winning method followed a detection–estimation–refinement pipeline and led clearly in both RGB and RGB-D settings.

2022 ECCV BOP Challenge champion

2022 · ECCV BOP Challenge — Single-Model Track Champion

A joint team from Xidian University, EPFL, and Magic Leap won first place in the single-model track of the BOP Challenge and was invited to present at the corresponding international workshop.

2021 Completion3D benchmark first place

2021 · Completion3D Benchmark — 1st Place

ASFM-Net ranked first on Stanford’s Completion3D leaderboard and became one of the group’s representative achievements in point-cloud completion.

2019 Tianzhi Cup second place

2019 · “Tianzhi Cup” AI Challenge — 2nd Place in Subject I

The team reconstructed 3D structures of target regions from multi-view satellite imagery, demonstrating strong algorithmic capability for demanding engineering tasks.

Projects

Projects
  • 2025–2026: 2D–3D matching algorithm research for an institute under China North Industries Group.
  • 2023–2026: Target pose-estimation algorithms, datasets, and management software for an institute under CASC.
  • 2023–2026: Vision-based localization system for aerial refueling under AVIC-related tasks.
  • 2021–2023: Prototype machine-vision system for a specific aircraft design task.
  • 2021–2023: Target pose-estimation system for a space-station related task at a CAS institute.
  • 2020–2021: 3D reconstruction of complex space targets.
  • 2019–2021: Real-time onboard image-processing algorithms and architecture under a key defense laboratory project.
Systems
  • 2023–2025: TDI camera compression system for a satellite mission.
  • 2024–2026: Image processor for a satellite monitoring camera.
  • 2022–2025: Highly integrated onboard compression system for a space camera.
  • 2022–2023: Prototype IP for DSC video compression/decompression.
  • 2022–2024: PNG image decoding IP design.
  • 2021–2023: Real-time panoramic scene-generation system for a defense platform.
  • 2020–2022: On-orbit compression project at a CAS institute.
  • 2019–2021: Deep-sea ultra-HD (4096×2160@60fps) video codec system.
  • 2019–2020: Compression software for an ultra-high-resolution (5000×5000@5fps) multi-mode camera.
  • 2020–2022: Hybrid-mode compression software for an ultra-HD satellite payload (8000×6000@5fps).
  • 2012–2017: Video encoding/decoding system for a space camera mission (including Tianzhou-1).
  • 2014–2017: Ground test equipment for spaceborne cameras.
  • 2015–2017: Test circuit development for Chang’e-5 related tasks.
  • 2014–2016: Multi-channel HD video communication system for space applications.
  • 2011–2015: H.264/AVC codec IP development under AVIC-related industrial tasks.

Publications

Selected Publications & Output
Exploring 6D Object Pose Estimation with Deformation

Exploring 6D Object Pose Estimation with Deformation

Zhiqiang Liu, Rui Song, Jiaojiao Li, Yinlin Hu
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

This paper introduces deformation modeling for 6D pose estimation under non-rigid deformation and heavy occlusion, improving robustness and estimation accuracy while preserving geometric consistency.

SCFlow2: Plug-and-Play Object Pose Refiner with Shape-Constraint Scene Flow

SCFlow2: Plug-and-Play Object Pose Refiner with Shape-Constraint Scene Flow

Qingyuan Wang, Rui Song, Jiaojiao Li, Kerui Cheng, David Ferstl, Yinlin Hu
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

We propose a plug-and-play pose refinement algorithm that injects shape-constrained scene flow into the optimization loop, improving both estimation accuracy and scene adaptability from an existing initial pose.

Pseudo Flow Consistency for Self-Supervised 6D Object Pose Estimation

Pseudo Flow Consistency for Self-Supervised 6D Object Pose Estimation

Yang Hai, Rui Song, Jiaojiao Li, David Ferstl, Yinlin Hu
IEEE/CVF International Conference on Computer Vision (ICCV), 2023

This work presents a self-supervised 6D pose-estimation framework without depth or additional annotations, and improves real-scene generalization and refinement via pseudo-flow supervision derived from multi-view geometric consistency.

Shape-Constraint Recurrent Flow for 6D Object Pose Estimation

Shape-Constraint Recurrent Flow for 6D Object Pose Estimation

Yang Hai, Rui Song, Jiaojiao Li, Yinlin Hu
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023

The paper proposes a shape-constrained recurrent-flow framework that implicitly embeds 3D object shape into iterative matching and pose solving for more efficient and accurate 6D pose refinement.

Rigidity-Aware Detection for 6D Object Pose Estimation

Rigidity-Aware Detection for 6D Object Pose Estimation

Yang Hai, Rui Song, Jiaojiao Li, Mathieu Salzmann, Yinlin Hu
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023

For 6D pose estimation under complex occlusion, the paper proposes a rigidity-aware detection strategy that uses visibility-guided sampling to obtain more stable detection initialization.

Structure-Aware Graph Convolution Network for Point Cloud Parsing

Structure-Aware Graph Convolution Network for Point Cloud Parsing

Fengda Hao, Jiaojiao Li, Rui Song, Yunsong Li, Kailang Cao
IEEE Transactions on Multimedia, 2023

This work proposes SA-GCN, a structure-aware graph convolution network with adaptive dilated KNN, learnable graph filters, and structure-aware feature transformation for point-cloud classification and segmentation.

HE2LM-AD: Hierarchical and Efficient Attitude Determination Framework with Adaptive Error Compensation Module Based on ELM Network

HE2LM-AD: Hierarchical and Efficient Attitude Determination Framework with Adaptive Error Compensation Module Based on ELM Network

Kailang Cao, Jiaojiao Li, Rui Song, Yunsong Li
ISPRS Journal of Photogrammetry and Remote Sensing, 2023

For remote-sensing satellite attitude determination, the paper proposes an efficient hierarchical framework combining adaptive EKF, ELM-based error compensation, and weighted smoothing to improve high-precision attitude estimation.

Lightweight Centroid Locating Method for the Satellite Target

Lightweight Centroid Locating Method for the Satellite Target

Luyuan Liu, Luyao Han, Jiaojiao Li, Hui Xia, Peng Rao, Rui Song
Journal of Xidian University, 2023

For onboard processors with limited computational capability, the paper proposes an ultra-lightweight satellite centroid localization method based on line and contour features, enabling accurate and real-time localization at low cost.

Mixed Feature Prediction on Boundary Learning for Point Cloud Semantic Segmentation

Mixed Feature Prediction on Boundary Learning for Point Cloud Semantic Segmentation

Fengda Hao, Jiaojiao Li, Rui Song, Yunsong Li, Kailang Cao
Remote Sensing, 2022

This work proposes a boundary-aware self-supervised pretraining framework for point-cloud semantic segmentation, improving boundary-region and fine-structure segmentation quality.

Research on the On-orbit Real-time Space Target Detection Algorithm

Research on the On-orbit Real-time Space Target Detection Algorithm

Luyao Han, Chan Tan, Yunmeng Liu, Rui Song
Spacecraft Recovery & Remote Sensing, 2021

For real-time on-orbit detection of space targets against deep-space backgrounds, the paper proposes a target-detection algorithm based on morphological processing and inter-frame matching, validated on FPGA and DSP platforms.

ASFM-Net: Asymmetrical Siamese Feature Matching Network for Point Completion

ASFM-Net: Asymmetrical Siamese Feature Matching Network for Point Completion

Yaqi Xia, Yan Xia, Wei Li, Rui Song, Kailang Cao, Uwe Stilla
ACM International Conference on Multimedia (ACM MM), 2021

ASFM-Net improves point-cloud completion by learning prior knowledge in feature space through an asymmetrical Siamese feature-matching autoencoder.

Robust Interpolation of Correspondences for Large Displacement Optical Flow

Robust Interpolation of Correspondences for Large Displacement Optical Flow

Peng Zhang, Hui Xu, Rui Song, etc.
Pattern Recognition, 2017

A robust interpolation framework for large-displacement optical flow improves dense flow recovery from sparse but reliable correspondences.

Efficient Coarse-to-Fine PatchMatch for Large Displacement Optical Flow

Efficient Coarse-to-Fine PatchMatch for Large Displacement Optical Flow

Peng Zhang, Hui Xu, Rui Song, etc.
IEEE Transactions on Image Processing, 2016

This work develops an efficient coarse-to-fine PatchMatch strategy for large-displacement optical flow, balancing accuracy and efficiency.

Full Publication List

All Publications

The entries below are uniformly formatted in an IEEE Transactions-style reference format. Some papers are still in early access or online-ahead-of-print status, so formal volume, issue, and page information may not yet be available and is therefore kept as early-publication metadata.

Computer Vision, 3D Vision, and 6D Pose Estimation

  1. R. Wang, R. Song, J. Zhang, and Y. Nie, "Leveraging Spatiotemporal Cues for Self-Supervised Stereo Depth Estimation in Endoscopic Videos," IEEE Trans. Med. Imaging, early access, Jan. 2026, doi: 10.1109/TMI.2026.3659145.
  2. Y. Sun, Y. Ma, Z. Chen, Z. Liu, B. Chen, and R. Song, "A Sliding Window for Data Reuse in Deep Convolution Operations to Reduce Bandwidth Requirements and Resource Utilization," Electronics, vol. 14, no. 3, art. 582, 2025.
  3. F. Hao, J. Li, R. Song, Y. Li, and K. Cao, "Structure-Aware Graph Convolution Network for Point Cloud Parsing," IEEE Trans. Multimed., vol. 25, pp. 7025–7036, 2023, doi: 10.1109/TMM.2022.3216951.
  4. Y. Hai, R. Song, J. Li, D. Ferstl, and Y. Hu, "Pseudo Flow Consistency for Self-Supervised 6D Object Pose Estimation," in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 14075–14085, doi: 10.1109/ICCV51070.2023.01294.
  5. Y. Hai, R. Song, J. Li, M. Salzmann, and Y. Hu, "Rigidity-Aware Detection for 6D Object Pose Estimation," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Vancouver, BC, Canada, 2023, pp. 8927–8936.
  6. Y. Hai, R. Song, J. Li, and Y. Hu, "Shape-Constraint Recurrent Flow for 6D Object Pose Estimation," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Vancouver, BC, Canada, 2023, pp. 4831–4840, doi: 10.1109/CVPR52729.2023.00468.
  7. F. Hao, J. Li, R. Song, Y. Li, and K. Cao, "Mixed Feature Prediction on Boundary Learning for Point Cloud Semantic Segmentation," Remote Sens., vol. 14, no. 19, p. 4757, Sep. 2022, doi: 10.3390/rs14194757.
  8. F. Hao, R. Song, J. Li, K. Cao, and Y. Li, "Cascaded Geometric Feature Modulation Network for Point Cloud Processing," Neurocomputing, vol. 492, pp. 474–487, Jul. 2022, doi: 10.1016/j.neucom.2022.04.007.
  9. Y. Xia, Y. Xia, W. Li, R. Song, K. Cao, and U. Stilla, "ASFM-Net: Asymmetrical Siamese Feature Matching Network for Point Completion," in Proc. 29th ACM Int. Conf. Multimedia, Oct. 2021, pp. 1938–1947, doi: 10.1145/3474085.3475348.
  10. B. Han, X. Jia, R. Song, F. Ran, and P. Rao, "Auto Complementary Exposure Control for High Dynamic Range Video Capturing," IEEE Access, vol. 9, pp. 144285–144299, 2021, doi: 10.1109/ACCESS.2021.3118416.
  11. 韩璐瑶, 谭婵, 刘云猛, and 宋锐, "在轨实时空间目标检测算法研究," 航天返回与遥感, vol. 42, no. 6, pp. 122–131, 2021, doi: 10.3969/j.issn.1009-8518.2021.06.012.
  12. S. Li and R. Song, "Bilateral Adaptive Quantization in HEVC," Multimed. Tools Appl., vol. 78, no. 2, pp. 2385–2399, 2019, doi: 10.1007/s11042-018-6312-y.
  13. R. Song, Y. Li, Y. Jia, Y. Wang, and P. Rao, "Efficient, Robust and Divisible Paired Comparison for Subjective Quality Assessment," Multimed. Tools Appl., vol. 77, no. 11, pp. 13597–13613, 2018, doi: 10.1007/s11042-017-4977-2.
  14. Y. Li, Y. Hu, R. Song, P. Rao, and Y. Wang, "Coarse-to-Fine PatchMatch for Dense Correspondence," IEEE Trans. Circuits Syst. Video Technol., vol. 28, no. 9, pp. 2233–2245, 2018, doi: 10.1109/TCSVT.2017.2720175.
  15. R. Song, Y. Yuan, Y. Li, and Y. Wang, "Extra Sign Bit Hiding Algorithm Based on Recovery of Transform Coefficients," Circuits Syst. Signal Process., vol. 37, no. 9, pp. 4128–4135, 2018, doi: 10.1007/s00034-017-0740-1.
  16. Y. Hu, Y. Li, and R. Song, "Robust Interpolation of Correspondences for Large Displacement Optical Flow," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Jul. 2017, pp. 4791–4799, doi: 10.1109/CVPR.2017.509.
  17. Y. Hu, R. Song, and Y. Li, "Efficient Coarse-to-Fine Patch Match for Large Displacement Optical Flow," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2016, pp. 5704–5712, doi: 10.1109/CVPR.2016.615.
  18. Y. Hu, R. Song, Y. Li, P. Rao, and Y. Wang, "Highly Accurate Optical Flow Estimation on Superpixel Tree," Image Vis. Comput., vol. 52, pp. 167–177, Aug. 2016, doi: 10.1016/j.imavis.2016.06.004.
  19. Y. Tian, Y. Wang, R. Song, and H. Song, "Accurate Vehicle Detection and Counting Algorithm for Traffic Data Collection," in 2015 Int. Conf. Connected Vehicles Expo (ICCVE), 2016, pp. 285–290, doi: 10.1109/ICCVE.2015.60.
  20. Y. Jia, Y. Wang, R. Song, and J. Li, "Decoder Side Information Generation Techniques in Wyner-Ziv Video Coding: A Review," Multimed. Tools Appl., vol. 74, no. 6, pp. 1777–1803, 2015, doi: 10.1007/s11042-013-1718-z.
  21. J. Y. Lin, R. Song, C.-H. Wu, T. Liu, H. Wang, and C.-C. J. Kuo, "MCL-V: A Streaming Video Quality Assessment Database," J. Vis. Commun. Image Represent., vol. 30, pp. 1–9, 2015, doi: 10.1016/j.jvcir.2015.02.012.

Intelligent Processing of Remote Sensing Imagery

  1. J. Li, H. Wu, R. Song, et al., "DF-PEM: Dual-Flow Prompt-Expert Mamba for Multimodal Remote Sensing Incremental Classification," IEEE Trans. Geosci. Remote Sens., early access, Jan. 2026, doi: 10.1109/TGRS.2026.3674175.
  2. J. Li, H. Wu, R. Song, H. Xu, Y. Li, and Q. Du, "Physics-Guided Time-Interactive-Frequency Network for Cross-Domain Few-Shot Hyperspectral Image Classification," IEEE Trans. Neural Netw. Learn. Syst., vol. 37, pp. 438–452, 2026, doi: 10.1109/TNNLS.2025.3608294.
  3. J. Li, S. Duan, H. Xu, R. Song, et al., "NukesFormers: Unpaired Hyperspectral Image Generation with Non-Uniform Domain Alignment," IEEE Trans. Geosci. Remote Sens., early access, 2025, doi: 10.1109/TGRS.2025.3634312.
  4. J. Li, Y. Ji, H. Xu, R. Song, et al., "UAT: Exploring Latent Uncertainty for Semi-Supervised Object Detection in Remote-Sensing Imagery," IEEE Trans. Geosci. Remote Sens., vol. 63, pp. 1–12, 2025, Art no. 5633212, doi: 10.1109/TGRS.2025.3586701.
  5. L. Liu, R. Song, J. Li, W. Wang, and Q. Wen, "ASMF: A Self-Supervised Atmospheric Scatter Model-Based Fusion Network for Infrared Image Enhancement," IEEE Trans. Geosci. Remote Sens., vol. 63, pp. 1–13, 2025, Art no. 5004613, doi: 10.1109/TGRS.2025.3581064.
  6. J. Li, Z. Zhang, R. Song, H. Xu, Y. Li, and Q. Du, "Contrastive MLP Network Based on Adjacent Coordinates for Cross-Domain Zero-Shot Hyperspectral Image Classification," IEEE Trans. Circuits Syst. Video Technol., vol. 35, pp. 8377–8390, 2025, doi: 10.1109/TCSVT.2025.3549365.
  7. J. Li, D. Zhu, R. Song, H. Xu, Y. Li, and Q. Du, "Multi-Feature Interaction and Degradation Estimation Transformer for Spectral Compressive Imaging," IEEE Trans. Circuits Syst. Video Technol., early access, 2025, doi: 10.1109/TCSVT.2025.3543569.
  8. Y. Leng, J. Li, R. Song, Y. Li, and Q. Du, "Uncertainty-Guided Discriminative Priors Mining for Flexible Unsupervised Spectral Reconstruction," IEEE Trans. Neural Netw. Learn. Syst., early access, 2025, doi: 10.1109/TNNLS.2025.3526159.
  9. J. Li, H. Li, H. Xu, R. Song, Y. Li, and Q. Du, "Background Suppression Network With Attention Collapse Inhibited Transformer for Optical Remote Sensing Object Detection," IEEE Trans. Geosci. Remote Sens., vol. 63, pp. 1–13, 2025, Art no. 5602413, doi: 10.1109/TGRS.2024.3520299.
  10. J. Li, S. Du, R. Song, Y. Li, and Q. Du, "Progressive Spatial Information-Guided Deep Aggregation Convolutional Network for Hyperspectral Spectral Super-Resolution," IEEE Trans. Neural Netw. Learn. Syst., vol. 36, no. 1, pp. 1677–1691, Jan. 2025, doi: 10.1109/TNNLS.2023.3325682.
  11. J. Li, Z. Zhang, Y. Liu, R. Song, Y. Li, and Q. Du, "SWFormer: Stochastic Windows Convolutional Transformer for Hybrid Modality Hyperspectral Classification," IEEE Trans. Image Process., vol. 33, pp. 5482–5495, 2024, doi: 10.1109/TIP.2024.3465038.
  12. J. Li, S. Duan, Y. Leng, R. Song, Y. Li, and Q. Du, "Residual Mask in Cascaded Convolutional Transformer for Spectral Reconstruction," IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–15, 2024, Art no. 5523615, doi: 10.1109/TGRS.2024.3427633.
  13. J. Li, Z. Zhang, R. Song, Y. Li, and Q. Du, "SCFormer: Spectral Coordinate Transformer for Cross-Domain Few-Shot Hyperspectral Image Classification," IEEE Trans. Image Process., vol. 33, pp. 840–855, 2024, doi: 10.1109/TIP.2024.3351443.
  14. J. Li, P. Tian, R. Song, H. Xu, Y. Li, and Q. Du, "PCViT: A Pyramid Convolutional Vision Transformer Detector for Object Detection in Remote-Sensing Imagery," IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–15, 2024, Art no. 5608115, doi: 10.1109/TGRS.2024.3360456.
  15. J. Li, Y. Liu, R. Song, W. Liu, Y. Li, and Q. Du, "HyperMLP: Superpixel Prior and Feature Aggregated Perceptron Networks for Hyperspectral and LiDAR Hybrid Classification," IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–14, 2024, Art no. 5505614, doi: 10.1109/TGRS.2024.3355037.
  16. K. Cao, J. Li, R. Song, Z. Liu, and Y. Li, "Model-Driven Deep Pipeline With Uncertainty-Aware Bundle Adjustment for Satellite Photogrammetry," IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–13, 2024, Art no. 5605313, doi: 10.1109/TGRS.2024.3352072.
  17. C. Wu, J. Li, R. Song, Y. Li, and Q. Du, "HPRN: Holistic Prior-Embedded Relation Network for Spectral Super-Resolution," IEEE Trans. Neural Netw. Learn. Syst., vol. 35, no. 8, pp. 11409–11423, Aug. 2024, doi: 10.1109/TNNLS.2023.3260828.
  18. J. Li, Y. Leng, R. Song, W. Liu, Y. Li, and Q. Du, "MFormer: Taming Masked Transformer for Unsupervised Spectral Reconstruction," IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1–12, 2023, Art no. 5508412, doi: 10.1109/TGRS.2023.3264976.
  19. J. Li, Y. Diao, R. Song, B. Xi, Y. Li, and Q. Du, "Class-Specific Autoaugment Architecture Based on Schmidt Mathematical Theory for Imbalanced Hyperspectral Classification," IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1–15, 2023, Art no. 5525315, doi: 10.1109/TGRS.2023.3317885.
  20. S. Duan, J. Li, R. Song, Y. Li, and Q. Du, "Unmixing-Guided Convolutional Transformer for Spectral Reconstruction," Remote Sens., vol. 15, no. 10, p. 2619, 2023, doi: 10.3390/rs15102619.
  21. K. Cao, J. Li, R. Song, and Y. Li, "HE²LM-AD: Hierarchical and Efficient Attitude Determination Framework With Adaptive Error Compensation Module Based on ELM Network," ISPRS J. Photogramm. Remote Sens., vol. 195, pp. 418–431, Jan. 2023, doi: 10.1016/j.isprsjprs.2022.12.010.
  22. C. Wu, J. Li, R. Song, Y. Li, and Q. Du, "RepCPSI: Coordinate-Preserving Proximity Spectral Interaction Network With Reparameterization for Lightweight Spectral Super-Resolution," IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1–13, 2023, Art no. 5508313, doi: 10.1109/TGRS.2023.3264675.
  23. J. Li, Y. Liu, R. Song, Y. Li, K. Han, and Q. Du, "Sal²RN: A Spatial–Spectral Salient Reinforcement Network for Hyperspectral and LiDAR Data Fusion Classification," IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1–14, 2023, Art no. 5500114, doi: 10.1109/TGRS.2022.3231930.
  24. J. Li et al., "Deep Hybrid 2-D–3-D CNN Based on Dual Second-Order Attention With Camera Spectral Sensitivity Prior for Spectral Super-Resolution," IEEE Trans. Neural Netw. Learn. Syst., vol. 34, no. 2, pp. 623–634, Feb. 2023, doi: 10.1109/TNNLS.2021.3098767.
  25. Z. Liu, J. Li, R. Song, C. Wu, W. Liu, Z. Li, and Y. Li, "Edge Guided Context Aggregation Network for Semantic Segmentation of Remote Sensing Imagery," Remote Sens., vol. 14, no. 6, p. 1353, Mar. 2022, doi: 10.3390/rs14061353.
  26. J. Li, S. Du, R. Song, C. Wu, Y. Li, and Q. Du, "HASIC-Net: Hybrid Attentional Convolutional Neural Network With Structure Information Consistency for Spectral Super-Resolution of RGB Images," IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–15, 2022, Art no. 5522515, doi: 10.1109/TGRS.2022.3142258.
  27. J. Li, Y. Ma, R. Song, B. Xi, D. Hong, and Q. Du, "A Triplet Semisupervised Deep Network for Fusion Classification of Hyperspectral and LiDAR Data," IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–13, 2022, Art no. 5540513, doi: 10.1109/TGRS.2022.3213513.
  28. J. Li, S. Zi, R. Song, Y. Li, Y. Hu, and Q. Du, "A Stepwise Domain Adaptive Segmentation Network With Covariate Shift Alleviation for Remote Sensing Imagery," IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–15, 2022, Art no. 5618515, doi: 10.1109/TGRS.2022.3152587.
  29. J. Li et al., "Feature Guide Network With Context Aggregation Pyramid for Remote Sensing Image Segmentation," IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., vol. 15, pp. 9900–9912, 2022, doi: 10.1109/JSTARS.2022.3221860.
  30. B. Xi, J. Li, Y. Li, R. Song, D. Hong, and J. Chanussot, "Few-Shot Learning With Class-Covariance Metric for Hyperspectral Image Classification," IEEE Trans. Image Process., vol. 31, pp. 5079–5092, 2022, doi: 10.1109/TIP.2022.3192712.
  31. Y. Li, Y. Zheng, J. Li, R. Song, and J. Chanussot, "Hyperspectral Pansharpening With Adaptive Feature Modulation-Based Detail Injection Network," IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–17, 2022, Art no. 5538117, doi: 10.1109/TGRS.2022.3206880.
  32. J. Li, H. Zhang, R. Song, W. Xie, Y. Li, and Q. Du, "Structure-Guided Feature Transform Hybrid Residual Network for Remote Sensing Object Detection," IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–13, 2022, Art no. 5610713, doi: 10.1109/TGRS.2021.3103964.
  33. B. Xi et al., "Multi-Direction Networks With Attentional Spectral Prior for Hyperspectral Image Classification," IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–15, 2022, Art no. 5500915, doi: 10.1109/TGRS.2020.3047682.
  34. B. Xi, J. Li, Y. Li, R. Song, W. Sun, and Q. Du, "Multiscale Context-Aware Ensemble Deep KELM for Efficient Hyperspectral Image Classification," IEEE Trans. Geosci. Remote Sens., vol. 59, no. 6, pp. 5114–5130, Jun. 2021, doi: 10.1109/TGRS.2020.3022029.
  35. J. Li, C. Wu, R. Song, Y. Li, and W. Xie, "Residual Augmented Attentional U-Shaped Network for Spectral Reconstruction From RGB Images," Remote Sens., vol. 13, no. 1, p. 115, Dec. 2020, doi: 10.3390/rs13010115.
  36. J. Li, C. Wu, R. Song, Y. Li, and F. Liu, "Adaptive Weighted Attention Network With Camera Spectral Sensitivity Prior for Spectral Reconstruction From RGB Images," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), 2020, pp. 462–463.
  37. J. Li et al., "Hyperspectral Image Super-Resolution by Band Attention Through Adversarial Learning," IEEE Trans. Geosci. Remote Sens., vol. 58, no. 6, pp. 4304–4318, Jun. 2020, doi: 10.1109/TGRS.2019.2962713.
  38. B. Xi et al., "Deep Prototypical Networks With Hybrid Residual Attention for Hyperspectral Image Classification," IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., vol. 13, pp. 3683–3700, 2020, doi: 10.1109/JSTARS.2020.3004973.
  39. J. Li, R. Cui, B. Li, R. Song, Y. Li, and Q. Du, "Hyperspectral Image Super-Resolution With 1D–2D Attentional Convolutional Neural Network," Remote Sens., vol. 11, no. 23, p. 2859, Dec. 2019, doi: 10.3390/rs11232859.
  40. J. Li, Y. Li, R. Song, S. Mei, and Q. Du, "Local Spectral Similarity Preserving Regularized Robust Sparse Hyperspectral Unmixing," IEEE Trans. Geosci. Remote Sens., vol. 57, no. 10, pp. 7756–7769, Oct. 2019, doi: 10.1109/TGRS.2019.2916296.

Teaching

Courses

Engineering Optimization

Audience: 2nd year, spring semesterType: ElectiveHours: 32

Goal: The course introduces how higher mathematics, matrices, and related knowledge are applied in real engineering problems, and trains students to translate engineering problems into mathematical formulations and solve them approximately with optimization methods. It also serves as a foundation for later study in machine learning, data mining, and deep learning.

Slides:Download linkCode: und5

Computer Communication Networks

Audience: 3rd year, spring semesterType: ElectiveHours: 80

Goal: The course covers basic concepts and architectures of computer communication networks, protocol mechanisms and key protocols, basic analytical methods for communication problems, and recent advances in networking technology, with emphasis on system understanding and engineering capability.

Slides:Download linkCode: mb6l
Writing Resources

LaTeX Template for Master’s Thesis

CVIA recommends LaTeX for master’s theses to reduce time wasted on formatting and separate content from layout. The Graduate School of Xidian University provides an official LaTeX template, and alumni from the institute have also fixed several issues in that template.

Official template:Graduate School link

LaTeX Template for Undergraduate Thesis

An undergraduate student from the 2020 class shared an undergraduate thesis template on GitHub. Students working on their bachelor thesis are encouraged to start from that template.

Template link:GitHub repository

Openings

Latest Updates

news iconOne additional academic master’s opening for 2026: Interested students are welcome to contact us by email. Replies are usually fast during working hours, and an online meeting can be arranged.

news iconPhD openings are still available for 2026: Students interested in computer vision, image analysis, 3D spatial measurement, and 6D pose estimation are welcome to get in touch.

What We Look For

  • Backgrounds in communications, electronics, computer science, surveying and mapping, or related fields.
  • Honesty, responsibility, solid work ethic, and a pursuit of excellence.
  • Strong initiative and independent thinking.
  • Good mathematical foundation, logical thinking, and English reading/writing ability.

What the Lab Offers

  • Entry training: topic introduction, research fundamentals, literature reading, research methodology, and paper-writing practice.
  • Research environment: essential equipment and materials for scientific work.
  • Academic guidance: at least one group meeting per week plus one-to-one guidance based on interest and task allocation.
  • Training style: the group does not rigidly separate academic and professional master’s students during training. Greater emphasis is placed on individual interest, capability, and strengths, with support for both algorithm research and engineering-system research.

Application Contact and Materials

  • Before applying, please email your officially recognized transcript covering all undergraduate years and mark the three courses that interest you most.
  • Attach a detailed CV, especially competition, practice, and project experience.
  • Attach English scores such as CET-6, TOEFL, GRE, or IELTS if available.
  • Please highlight mathematics and programming abilities, including courses such as calculus and linear algebra, and grades related to C/C++, Verilog, Python, and similar topics.
  • It is recommended to request a read receipt and confirm within one week that the email has been received.

Training and Career Paths

  • PhD graduates mainly go to universities, research institutes, and technology companies. The group currently has 2 PhD students in training and 5 PhD alumni.
  • Master’s graduates mainly go to Intel, Huawei, ZTE, HONOR, HiSilicon, ByteDance, Baidu, and institutes under CETC, CASC, AVIC, and the Chinese Academy of Sciences. The group currently has 23 master’s students in training and 103 master’s alumni.
  • If you need a recommendation letter, please make sure the faculty member knows you well enough.