Complex diseases are defined as diseases that are affected by multiple genetic and environmental factors. Prediction of these diseases is a crucial step in personalized and preventive medicine. In recent years, high-throughput or "omics" data (genomics, epigenomics, proteomics, metabolomics, etc.) have become increasingly abundant and accessible, allowing for a deeper understanding of these complex diseases. Each type of omics data has associations with different diseases and traits, but each plays a non-overlapping role in complex disease prediction and can even be associated with each other. In this dissertation, I propose several methods that leverage these omics data to predict complex diseases.For the first project, I present functional HAUDI (fHAUDI), an extension of our previous work, GAUDI and HAUDI, that leverages functional annotations to improve polygenic risk scores (PRS) for admixed individuals. fHAUDI outperforms HAUDI, a PRS for admixed individuals without annotations, in simulations and real data. Moreover, it outperforms other comparison methods that can leverage functional annotations but are optimized for single-ancestry populations. In my second project, I move beyond the genetic component of PRS and present OMEGA (Omics Multi-modality Embedding via Graphical & Articulated data), which leverages a Mixture-of-Experts (MoE) framework to integrate the information from multiple omics data represented by multiple modalities, such as tabular data, image data, and text data, to predict phenotypes. We leverage proteomics and metabolomics to predict incident disease in the UK Biobank. OMEGA outperforms other comparator methods that only consider tabular data for disease prediction. Additionally, we see that including the different modalities does improve predictionFor my last project, I propose a method for feature importance in our OMEGA framework that leverages integrated gradients and gradient boosting methods. We validate this method in large-scale simulations using omics data from the UK Biobank used in the original OMEGA analysis. Under certain conditions, we are able to recover the true causal variants using this method.
Brian Chen (Fri,) studied this question.