PGAGP: Predicting pathogenic genes based on adaptive network embedding algorithm

Front Genet. 2023 Jan 20:13:1087784. doi: 10.3389/fgene.2022.1087784. eCollection 2022.

Abstract

The study of disease-gene associations is an important topic in the field of computational biology. The accumulation of massive amounts of biomedical data provides new possibilities for exploring potential relations between diseases and genes through computational strategy, but how to extract valuable information from the data to predict pathogenic genes accurately and rapidly is currently a challenging and meaningful task. Therefore, we present a novel computational method called PGAGP for inferring potential pathogenic genes based on an adaptive network embedding algorithm. The PGAGP algorithm is to first extract initial features of nodes from a heterogeneous network of diseases and genes efficiently and effectively by Gaussian random projection and then optimize the features of nodes by an adaptive refining process. These low-dimensional features are used to improve the disease-gene heterogenous network, and we apply network propagation to the improved heterogenous network to predict pathogenic genes more effectively. By a series of experiments, we study the effect of PGAGP's parameters and integrated strategies on predictive performance and confirm that PGAGP is better than the state-of-the-art algorithms. Case studies show that many of the predicted candidate genes for specific diseases have been implied to be related to these diseases by literature verification and enrichment analysis, which further verifies the effectiveness of PGAGP. Overall, this work provides a useful solution for mining disease-gene heterogeneous network to predict pathogenic genes more effectively.

Keywords: biological network; disease-gene prediction; network embedding; network propagation; random projection.

Grants and funding

This work was supported by the Training Program for Excellent Young Innovators of Changsha (Grant No. kq2206056, kq2206058), the National Natural Science Foundation of China (Grant No. 81873780), The Foundation of Project of Hunan Health and Family Planning Commission (Grant No. 202202082739), The Foundation of the Education Department of Hunan Province (Grant No. 21A0586), and the Application Characteristic Discipline of Hunan Province.