Epistatic Net allows the sparse spectral regularization of deep neural networks for inferring fitness functions

Amirali Aghazadeh; Hunter Nisonoff; Orhan Ocal; David H Brookes; Yijie Huang; O Ozan Koyluoglu; Jennifer Listgarten; Kannan Ramchandran

doi:10.1038/s41467-021-25371-3

Epistatic Net allows the sparse spectral regularization of deep neural networks for inferring fitness functions

Nat Commun. 2021 Sep 1;12(1):5225. doi: 10.1038/s41467-021-25371-3.

Authors

Amirali Aghazadeh¹, Hunter Nisonoff², Orhan Ocal¹, David H Brookes³, Yijie Huang¹, O Ozan Koyluoglu¹, Jennifer Listgarten^{1

2}, Kannan Ramchandran⁴

Affiliations

¹ Department of Electrical Engineering and Computer Sciences, Berkeley, CA, USA.
² Center for Computational Biology, Berkeley, CA, USA.
³ Biophysics Graduate Group, University of California, Berkeley, CA, USA.
⁴ Department of Electrical Engineering and Computer Sciences, Berkeley, CA, USA. kannanr@eecs.berkeley.edu.

Abstract

Despite recent advances in high-throughput combinatorial mutagenesis assays, the number of labeled sequences available to predict molecular functions has remained small for the vastness of the sequence space combined with the ruggedness of many fitness functions. While deep neural networks (DNNs) can capture high-order epistatic interactions among the mutational sites, they tend to overfit to the small number of labeled sequences available for training. Here, we developed Epistatic Net (EN), a method for spectral regularization of DNNs that exploits evidence that epistatic interactions in many fitness functions are sparse. We built a scalable extension of EN, usable for larger sequences, which enables spectral regularization using fast sparse recovery algorithms informed by coding theory. Results on several biological landscapes show that EN consistently improves the prediction accuracy of DNNs and enables them to outperform competing models which assume other priors. EN estimates the higher-order epistatic interactions of DNNs trained on massive sequence spaces-a computational problem that otherwise takes years to solve.

Publication types

Research Support, N.I.H., Extramural
Research Support, U.S. Gov't, Non-P.H.S.

MeSH terms

Algorithms*
Bacteria
Green Fluorescent Proteins
Neural Networks, Computer*

Substances

Green Fluorescent Proteins

Abstract

Publication types

MeSH terms

Substances

Grants and funding