Advancing deep learning denoising methods with self-supervised training and multi-channel image priors
Supervised deep learning methods have saturated nearly every step of medical imaging data acquisition and reconstruction pipelines, complementing older numerical methods and physics-informed prior information with a new generation of data-specific prior information. However, supervised learning faces a fundamental limitation: the quality of data-specific prior information learned by a network is proportional to the quality of available training data. Said another way, deep learning methods excel at reproducing the current state of the art (human-level segmentation performance, denoising to match standard-dose reference CT data), but they are ill conditioned to advance it beyond our ability to quantify network performance. We believe multi-channel imaging data and prior knowledge of the relationships between image channels provide a unique opportunity to advance deep learning performance more broadly. We illustrate this concept using in vivo, preclinical micro-CT data acquired in the mouse with the objective of reconstructing 10 phases of the cardiac cycle (~15 ms temporal resolution). Balancing the ionizing radiation dose against physical constraints on the data acquisition and resolution requirements, individual cardiac phases exhibit a high degree of image noise (280 HU). In this study, we summarize a data pre-processing method and self-supervised network training strategy which enables robust denoising of individual cardiac phases without compromising spatial or temporal resolution (noise level: 150 HU). We then propose further modifications to the network structure and training procedure that allow the trained network to extrapolate to lower levels of image noise in the domain of the network training data (55 HU) and to generalize to single and multi-channel CT data from related imaging domains (photon-counting CT data; time-resolved perfusion imaging, higher spatial resolution).