Your browser doesn't support javascript.
Causal Discovery in High-dimensional, Multicollinear Datasets.
Jia, Minxue; Yuan, Daniel Y; Lovelace, Tyler C; Hu, Mengying; Benos, Panayiotis V.
  • Jia M; Department of Computational and Systems Biology, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.
  • Yuan DY; Joint CMU-Pitt PhD Program in Computational Biology, Pittsburgh, PA, USA.
  • Lovelace TC; Department of Computational and Systems Biology, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.
  • Hu M; Joint CMU-Pitt PhD Program in Computational Biology, Pittsburgh, PA, USA.
  • Benos PV; Department of Computational and Systems Biology, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.
Front Epidemiol ; 22022.
Article in English | MEDLINE | ID: covidwho-2231452
ABSTRACT
As the cost of high-throughput genomic sequencing technology declines, its application in clinical research becomes increasingly popular. The collected datasets often contain tens or hundreds of thousands of biological features that need to be mined to extract meaningful information. One area of particular interest is discovering underlying causal mechanisms of disease outcomes. Over the past few decades, causal discovery algorithms have been developed and expanded to infer such relationships. However, these algorithms suffer from the curse of dimensionality and multicollinearity. A recently introduced, non-orthogonal, general empirical Bayes approach to matrix factorization has been demonstrated to successfully infer latent factors with interpretable structures from observed variables. We hypothesize that applying this strategy to causal discovery algorithms can solve both the high dimensionality and collinearity problems, inherent to most biomedical datasets. We evaluate this strategy on simulated data and apply it to two real-world datasets. In a breast cancer dataset, we identified important survival-associated latent factors and biologically meaningful enriched pathways within factors related to important clinical features. In a SARS-CoV-2 dataset, we were able to predict whether a patient (1) had Covid-19 and (2) would enter the ICU. Furthermore, we were able to associate factors with known Covid-19 related biological pathways.
Keywords

Full text: Available Collection: International databases Database: MEDLINE Type of study: Experimental Studies / Prognostic study Language: English Year: 2022 Document Type: Article Affiliation country: Fepid.2022.899655

Similar

MEDLINE

...
LILACS

LIS


Full text: Available Collection: International databases Database: MEDLINE Type of study: Experimental Studies / Prognostic study Language: English Year: 2022 Document Type: Article Affiliation country: Fepid.2022.899655