April 29, 2024Open Access

LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report

Key Points

Key points are not available for this paper at this time.

Abstract

Low Rank Adaptation (LoRA) has emerged as one of the most widely adopted methods for Parameter Efficient Fine-Tuning (PEFT) of Large Language Models (LLMs). LoRA reduces the number of trainable parameters and memory usage while achieving comparable performance to full fine-tuning. We aim to assess the viability of training and serving LLMs fine-tuned with LoRA in real-world applications. First, we measure the quality of LLMs fine-tuned with quantized low rank adapters across 10 base models and 31 tasks for a total of 310 models. We find that 4-bit LoRA fine-tuned models outperform base models by 34 points and GPT-4 by 10 points on average. Second, we investigate the most effective base models for fine-tuning and assess the correlative and predictive capacities of task complexity heuristics in forecasting the outcomes of fine-tuning. Finally, we evaluate the latency and concurrency capabilities of LoRAX, an open-source Multi-LoRA inference server that facilitates the deployment of multiple LoRA fine-tuned models on a single GPU using shared base model weights and dynamic adapter loading. LoRAX powers LoRA Land, a web application that hosts 25 LoRA fine-tuned Mistral-7B LLMs on a single NVIDIA A100 GPU with 80GB memory. LoRA Land highlights the quality and cost-effectiveness of employing multiple specialized LLMs over a single, general-purpose LLM.

Connected Papers

Building similarity graph...

Analyzing shared references across papers

Discussion

Cite this study

Zhao et al. (Mon,) studied this question.

www.synapsesocial.com/papers/68e6d055b6db64358764dff2 — DOI: https://doi.org/10.48550/arxiv.2405.00732

Also consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

LoRA Is Slower Than You Think· 2025
PeriodicLoRA: Breaking the Low-Rank Bottleneck in LoRA Optimization· 2024 · 1 citations
LoRA-GA: Low-Rank Adaptation with Gradient Approximation· 2024 · 4 citations
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs· 2025 · 1 citations
Bayesian-LoRA: LoRA based Parameter Efficient Fine-Tuning using Optimal Quantization levels and Rank Values trough Differentiable Bayesian Gates

Authors

Justin Zhao

Timothy C. Wang

Wael Abid

Actions

References and Citations

Connected Papers

Building similarity graph...

Analyzing shared references across papers

LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report

Key Points

Abstract

Citation Network

Connected Papers

Discussion

Cite this study

Also consider

Authors

Actions

References and Citations

Citation Network

Connected Papers

Discussion