Dataset resources for the publication: "Phenotypic reversion and target prioritization for cellular inflammation via representation learning with foundation models. " see GithHub: https: //github. com/pfizer-opensource/phenotypeᵣeversion imrufull. h5ad contains the transcript count matrix. the 'geneₜarget' column reflects the gene that was knocked down via CRISPRi, while the 'condition' column refers to whether or not the cells were treated with IL-1B & TNFa If the cytokine was introduced the treatment condition = 'Treated' else 'Untreated'. The controls in the 'geneₜarget' column are denoted as no target ('NO-TARGET') or safe target ('SAFETARGET'). . obs contains keys: 'condition', 'cellID', 'cellₜreat', 'ngenes', 'nGene', 'nUMI', 'log10GenesPerUMI', 'mitoRatio', 'geneₜarget', 'guide', 'welltag', 'flask' both treated and untreated condition have same controls, and roughly same sets of perturbations except: treated has unique perturbations: 'SH3PXD2A', 'EMCN', 'RAB11FIP3', 'CKAP5', 'POLR2K' untreated has unique perturbations: 'VWF', 'HNRNPA3', 'TRIM13', 'CORO1C', 'DDX27', 'PACSIN2' For all. h5ad files, adata. X will contain the embedding as np array, adata. obsm'umap' will contain the coordinates of UMAP applied to the embedding, e. g. adata. obsm'umap' = UMAP (adata. X) For convenience, the single cell embeddings as well as UMAP reductions of those embeddings for the different single cell foundation models have also been provided in this data repository. Place all the files in this repository in your local project directory: phenotypeᵣeversion/data/.
Daniel Wong (2026) studied this question.