Abstract Correlation among the observations in high-dimensional regression modeling can be a major source of confounding. We present a new open-source package, plmmr, to implement penalized linear mixed models in R. This R package estimates correlation among observations in high-dimensional data and uses those estimates to improve prediction with the best linear unbiased predictor. The package uses memory mapping so that genome-scale data can be analyzed on ordinary machines even if the size of data exceeds random-access memory. We present here the methods, workflow, and file-backing approach upon which plmmr is built, and we demonstrate its computational capabilities with two examples from real genome-wide association studies data.
Peter et al. (Thu,) studied this question.