What type of study is this?

September 10, 2025

mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs

Key Points

mLoRA enables simultaneous fine-tuning of two Llama-2-13B models on four GPUs, increasing efficiency and accessibility.
Average fine-tuning task completion time can be reduced by 30% compared to existing methods like FSDP.
The novel LoRA-aware pipeline parallelism scheme enhances GPU utilization and reduces communication overhead.
This approach allows developers to adapt large language models to various tasks simultaneously, promoting cost-effective solutions.

Abstract

Transformer-based large language models (LLMs) have demonstrated outstanding performance across diverse domains, particularly in the emerging pretrain-then-finetune paradigm. LoRA, a parameter-efficient fine-tuning method, is commonly used to adapt a base LLM to multiple downstream tasks. Further, LLM platforms enable developers to fine-tune multiple models and develop various domain-specific applications simultaneously. However, existing model parallelism schemes suffer from high communication overhead and inefficient GPU utilization. In this paper, we present mLoRA, a parallelism-efficient fine-tuning system designed for training multiple LoRA across GPUs and machines. mLoRA introduces a novel LoRA-aware pipeline parallelism scheme that efficiently pipelines LoRA adapters and their distinct fine-tuning stages across GPUs and machines, along with a new LoRA-efficient operator to enhance GPU utilization. Our extensive evaluation shows that mLoRA can significantly reduce average fine-tuning task completion time, e.g., by 30%, compared to state-of-the-art methods like FSDP. More importantly, mLoRA enables simultaneous fine-tuning of larger models, e.g., two Llama-2-13B models on four NVIDIA RTX A6000 48GB GPUs, which is not feasible for FSDP due to high memory requirements. Hence, mLoRA not only increases fine-tuning efficiency but also makes it more accessible on cost-effective GPUs.

Connected Papers

Building similarity graph...

Analyzing shared references across papers

Discussion

Cite this study

Ye et al. (Sat,) studied this question.

www.synapsesocial.com/papers/68c1d97154b1d3bfb60fabb0 — DOI: https://doi.org/10.14778/3725688.3725718

Also consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

ArcheType: A Novel Framework for Open-Source Column Type Annotation Using Large Language Models· 2024 · 19 citations
Parameter-efficient fine-tuning of large-scale pre-trained language models· 2023 · 868 citations
DeepSpeed· 2020 · 698 citations
A Review of Current Trends, Techniques, and Challenges in Large Language Models (LLMs)· 2024 · 195 citations
Unifying Large Language Models and Knowledge Graphs: A Roadmap

Authors

Zhengmao Ye

Dengchun Li

Zhibin Hu

Journals

Proceedings of the VLDB Endowment

Actions

Institutions

Sichuan University

The University of Texas at Arlington

Academia Sinica

References and Citations

Connected Papers

Building similarity graph...

Analyzing shared references across papers

mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs

Key Points

Abstract

Citation Network

Connected Papers

Discussion

Cite this study

Also consider

Authors

Journals

Actions

Institutions

References and Citations

Citation Network

Connected Papers

Discussion