Abstract The relative abundance of cellular and structural components in cancer tissues reflects disease biology and outcomes. High-throughput technologies, e.g., digital pathology and single-cell transcriptomics (scRNA-seq), profile tissue regions or cells, revealing the heterogeneous building blocks (elements) that comprise bulk phenotypes. Yet, translational applications increasingly demand that models not only infer sample-level distributions but also provide element-level characterizations (e.g., of specific histology regions, cells) that explain tumor macroscopic behavior. Achieving such granularity from sample-level compositional data without laborious annotations remains a challenge, requiring models capable of connecting local properties to global phenotypes. Existing methods, using attention to weight element contributions to sample-level predictions, lack probabilistic grounding and biological interpretability; thus, have limited utility in clinical settings. Here, we explicitly model sample-level compositional constraints with element-level assignments using Optimal Transport (OT). Particularly, we introduce Composer, a domain-agnostic dual-task machine learning model trained without element-level annotations to 1) estimate sample-level compositions, and, 2) importantly infer element labels as interpretable compositional allocations. The model's performance was evaluated across data modalities, cancer types, and tasks. We highlight 3 concepts: a) Tissue type classification on whole-slide images (WSIs): On 60 hematoxylin mean absolute error: 0.11), classifying correctly (AUROC 0.98) the tissue type (tumor, stroma, necrosis, adipose, other) of individual regions within WSIs without training on regional annotations. b) Tumor segmentation on WSIs: On 185 H MAE 0.15), distinguishing efficiently tumor from non-tumor regions (AUROC 0.91) within WSIs. c) Single-cell type annotation in scRNA-seq: Using 156 scRNA-seq HGSOC samples (Vazquez-Garcia et al. 2022), Composer not only inferred bulk cell type distributions (JSD: 0.15; MAE: 0.06), but also accurately classified (AUROC: 0.97) individual cells to their designated type (T cells, monocytes, fibroblasts, cancer cells, other). In summary, the proposed OT-based weakly supervised framework provides an effective approach for linking element-level representations to sample-level compositional profiles in cancer. Its applicability across data analyses empowers spatial and molecular characterization of tumor ecosystems, biomarker quantification, and potential applications to precision oncology. Citation Format: Georgios Asimomitis, Kevin M. Boehm, Konstantinos Liosis, Armaan Kohli, Tom Pollard, Andrew Aukerman, Arfath Pasha, Anika Begum, Lora H. Ellenson, Jinru Shia, Hong A. Zhang, Nikolaus Schultz, Sohrab P. Shah, Francisco Sanchez-Vega, . Inferring tissue element identities from sample-level compositional data in cancer abstract. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 5488.
Asimomitis et al. (Fri,) studied this question.