PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 6, 2026JMIR Medical Informatics0 citationsOpen Access

Scalable and Privacy-Conscious End-to-End Processing of Large-Scale Clinical Data for Precision Medicine: Empirical Evaluation Study

View Full Paper
JLJungwoo LeeSHSangwon HwangKLKyu Hee Lee

Key Points

  • This study aims to evaluate the effectiveness of a Parquet-based pipeline for large-scale clinical data analysis, focusing on efficiency and privacy.
  • Analyzed electronic health record data from a large medical center in Korea.
  • Compared Parquet, CSV, PostgreSQL, and DuckDB for storage and processing.
  • Applied multilabel classification models using Extreme Gradient Boosting to address class imbalance.
  • Conducted statistical equivalence testing and assessed privacy risks through membership inference attacks.
  • Parquet reduced disk access time significantly from 940.2 to 44.2 seconds.
  • End-to-end processing latency decreased notably across feature transformation and model training.
  • Classification performance was statistically equivalent across various metrics, adhering to prespecified clinical equivalence margins.
  • Membership inference attacks showed no measurable increase in privacy risk.

Abstract

Background In large-scale clinical data analysis, CSV and traditional relational database management system–based approaches are widely used but impose substantial storage and processing constraints that delay research preparation and hinder multicenter collaboration. Although column-oriented storage formats such as Apache Parquet have gained attention in data science, systematic end-to-end evaluations in clinical environments remain limited, particularly regarding efficiency and scalability. Objective This study aimed to empirically evaluate whether a Parquet-based end-to-end pipeline could improve computational efficiency and scalability in large-scale clinical data analysis while preserving predictive performance and protecting privacy. Methods Electronic health record data comprising 13.76 million rows from a large academic medical center in Korea were analyzed using Parquet, CSV, PostgreSQL, and DuckDB environments. Standardized SQL workloads and multilabel classification models—implemented using graphics processing unit–accelerated Extreme Gradient Boosting and classifier chain (CC) ensembles to address class imbalance—were applied to evaluate storage efficiency, time to analysis, and predictive performance. Statistical equivalence testing with prespecified clinical margins and bootstrap resampling ensured rigorous comparison, while privacy risks were assessed through advanced membership inference attacks (MIA), including shadow MIA and likelihood ratio attacks. Results Compared with CSV, Parquet demonstrated enhanced computational efficiency by lowering disk access from 940.2 to 44.2 seconds (95.3% reduction). End-to-end processing latency was substantially reduced across feature transformation (15.0 vs 9.3 s) and model training (8.1 vs 6.7 s). To address complex clinical correlations, we implemented CC and one-vs-rest architectures, which effectively captured interdependencies between disease labels. Classification performance remained statistically equivalent across area under the receiver operating characteristic curve, area under the precision-recall curve, accuracy, and F1-score, with all differences falling within prespecified clinical equivalence margins (P<.001). Notably, the CC ensemble demonstrated high technical rigor, minimizing Hamming loss (2.2×10–4) and ensuring robustness even in imbalanced cohorts. MIA performed at chance level (area under the curve=0.500), suggesting no measurable increase in privacy risk. Conclusions By significantly mitigating data processing bottlenecks, a Parquet-based pipeline enabled high-throughput, large-scale clinical evidence generation without compromising model integrity or patient privacy. This framework provides a scalable and robust infrastructure for precision medicine, facilitating agile multicenter collaborations and real-world data analysis in resource-constrained clinical environments.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lee et al. (2026) studied this question.

synapsesocial.com/papers/69aa70f8531e4c4a9ff5b456https://doi.org/10.2196/83487
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Smoking and Type 1 Versus Type 2 Myocardial Infarction Among People With HIV in the United States: Results from the Center for AIDS Research Network Integrated Clinical Systems Cohort2024 · 3 citations
  2. 2Clinical Decision Support to Increase Emergency Department Naloxone Coprescribing: Implementation Report2024 · 11 citations
  3. 3Mapping ICD-10 and ICD-10-CM Codes to Phecodes: Workflow Development and Initial Evaluation2019 · 571 citations
  4. 4Beyond TPC-DS, a benchmark for Big Data OLAP systems (BDOLAP-Bench)2022 · 16 citations
  5. 5Automatic de-identification of textual documents in the electronic health record: a review of recent research2010 · 338 citations