Scholars@Duke publication: LASSO and Elastic Net Tend to Over-Select Features

LASSO and Elastic Net Tend to Over-Select Features

Publication , Journal Article

Liu, L; Gao, J; Beasley, G; Jung, SH

Published in: Mathematics

September 1, 2023

Machine learning methods have been a standard approach to select features that are associated with an outcome and to build a prediction model when the number of candidate features is large. LASSO is one of the most popular approaches to this end. The LASSO approach selects features with large regression estimates, rather than based on statistical significance, that are associated with the outcome by imposing an (Formula presented.) -norm penalty to overcome the high dimensionality of the candidate features. As a result, LASSO may select insignificant features while possibly missing significant ones. Furthermore, from our experience, LASSO has been found to select too many features. By selecting features that are not associated with the outcome, we may have to spend more cost to collect and manage them in the future use of a fitted prediction model. Using the combination of (Formula presented.) - and (Formula presented.) -norm penalties, elastic net (EN) tends to select even more features than LASSO. The overly selected features that are not associated with the outcome act like white noise, so that the fitted prediction model may lose prediction accuracy. In this paper, we propose to use standard regression methods, without any penalizing approach, combined with a stepwise variable selection procedure to overcome these issues. Unlike LASSO and EN, this method selects features based on statistical significance. Through extensive simulations, we show that this maximum likelihood estimation-based method selects a very small number of features while maintaining a high prediction power, whereas LASSO and EN make a large number of false selections to result in loss of prediction accuracy. Contrary to LASSO and EN, the regression methods combined with a stepwise variable selection method is a standard statistical method, so that any biostatistician can use it to analyze high-dimensional data, even without advanced bioinformatics knowledge.

Duke Scholars

Author Georgia Marie Beasley Surgical Oncology

Author Sin-Ho Jung Biostatistics & Bioinformatics, Division of Integrative Geno ...

Published In

Mathematics

DOI

10.3390/math11173738

EISSN

2227-7390

Publication Date

September 1, 2023

Volume

Issue

Related Subject Headings

49 Mathematical sciences

Citation

APA

Chicago

ICMJE

MLA

NLM

Liu, L., Gao, J., Beasley, G., & Jung, S. H. (2023). LASSO and Elastic Net Tend to Over-Select Features. Mathematics, 11(17). https://doi.org/10.3390/math11173738

Liu, L., J. Gao, G. Beasley, and S. H. Jung. “LASSO and Elastic Net Tend to Over-Select Features.” Mathematics 11, no. 17 (September 1, 2023). https://doi.org/10.3390/math11173738.

Liu L, Gao J, Beasley G, Jung SH. LASSO and Elastic Net Tend to Over-Select Features. Mathematics. 2023 Sep 1;11(17).

Liu, L., et al. “LASSO and Elastic Net Tend to Over-Select Features.” Mathematics, vol. 11, no. 17, Sept. 2023. Scopus, doi:10.3390/math11173738.

Liu L, Gao J, Beasley G, Jung SH. LASSO and Elastic Net Tend to Over-Select Features. Mathematics. 2023 Sep 1;11(17).

Published In

Mathematics

DOI

10.3390/math11173738

EISSN

2227-7390

Publication Date

September 1, 2023

Volume

Issue

Related Subject Headings

49 Mathematical sciences