Enhanced Lung Cancer Survival Prediction Using Semi-Supervised Pseudo-Labeling and Learning from Diverse PET/CT Datasets

Mohammad R Salmanpour; Arman Gorji; Amin Mousavi; Ali Fathi Jouzdani; Nima Sanati; Mehdi Maghsudi; Bonnie Leung; Cheryl Ho; Ren Yuan; Arman Rahmim

doi:10.3390/cancers17020285

Enhanced Lung Cancer Survival Prediction Using Semi-Supervised Pseudo-Labeling and Learning from Diverse PET/CT Datasets

Cancers (Basel). 2025 Jan 17;17(2):285. doi: 10.3390/cancers17020285.

Authors

Mohammad R Salmanpour^{1

2

3}, Arman Gorji^{3

4}, Amin Mousavi³, Ali Fathi Jouzdani^{3

4}, Nima Sanati^{3

4}, Mehdi Maghsudi³, Bonnie Leung⁵, Cheryl Ho¹, Ren Yuan^{2

5}, Arman Rahmim^{1

2

6}

Affiliations

¹ BC Cancer Research Institute, Vancouver, BC V5Z 1L3, Canada.
² Department of Radiology, University of British Columbia, Vancouver, BC V6T 1Z4, Canada.
³ Technological Virtual Collaboration (TECVICO Corp.), Vancouver, BC V6L 1L7, Canada.
⁴ Neuroscience and Artificial Intelligence Research Group (NAIRG), Department of Neuroscience, Hamadan University of Medical Sciences, Hamadan 6517838736, Iran.
⁵ BC Cancer, Vancouver Center, Vancouver, BC V5Z 1L3, Canada.
⁶ Department of Physics & Astronomy, University of British Columbia, Vancouver, BC V6T 1Z4, Canada.

PMID: 39858067
DOI: 10.3390/cancers17020285

Abstract

Objective: This study explores a semi-supervised learning (SSL), pseudo-labeled strategy using diverse datasets such as head and neck cancer (HNCa) to enhance lung cancer (LCa) survival outcome predictions, analyzing handcrafted and deep radiomic features (HRF/DRF) from PET/CT scans with hybrid machine learning systems (HMLSs).

Methods: We collected 199 LCa patients with both PET and CT images, obtained from TCIA and our local database, alongside 408 HNCa PET/CT images from TCIA. We extracted 215 HRFs and 1024 DRFs by PySERA and a 3D autoencoder, respectively, within the ViSERA 1.0.0 software, from segmented primary tumors. The supervised strategy (SL) employed an HMLS-PCA connected with six classifiers on both HRFs and DRFs. The SSL strategy expanded the datasets by adding 408 pseudo-labeled HNCa cases (labeled by the Random Forest algorithm) to 199 LCa cases, using the same HMLS techniques. Furthermore, principal component analysis (PCA) linked with four survival prediction algorithms were utilized in the survival hazard ratio analysis.

Results: The SSL strategy outperformed the SL method (p << 0.001), achieving an average accuracy of 0.85 ± 0.05 with DRFs from PET and PCA + Multi-Layer Perceptron (MLP), compared to 0.69 ± 0.06 for the SL strategy using DRFs from CT and PCA + Light Gradient Boosting (LGB). Additionally, PCA linked with Component-wise Gradient Boosting Survival Analysis on both HRFs and DRFs, as extracted from CT, had an average C-index of 0.80, with a log rank p-value << 0.001, confirmed by external testing.

Conclusions: Shifting from HRFs and SL to DRFs and SSL strategies, particularly in contexts with limited data points, enabling CT or PET alone, can significantly achieve high predictive performance.

Keywords: deep and handcrafted radiomic features; lung cancer; machine learning; supervised and semi-supervised strategy; survival prediction.