LoRA vs QLoRA: choose a fine-tuning method
Understand LoRA and QLoRA memory tradeoffs, choose practical settings, and compare adapters fairly in Tensorant before deployment.
What you will build
Choose and validate an adapter training method for your GPU budget.
Steps in this guide
Choose LoRA when the model fits comfortably on your training GPU. Choose QLoRA when the base model's weight memory prevents a practical run or forces you onto a more expensive GPU. Then evaluate the finished adapter: a lower memory requirement is useful only if the model still performs your task reliably.
Both methods train an adapter instead of updating every base-model weight. In Tensorant, the choice appears as LoRA or QLoRA · 4-bit (less GPU memory) when you configure an experiment or training run. This guide explains the difference and uses the support-router ticket-classification task to make the decision concrete.
1. Understand what changes during training
LoRA freezes the original weight matrix and learns a low-rank update. Instead of training a full matrix with millions of entries, it trains two smaller matrices whose product modifies the original operation. See the original LoRA paper.
For an example square layer with 4,096 inputs and outputs:
Original matrix: 4,096 × 4,096 = 16,777,216 entries
Rank-16 update: (4,096 × 16) + (16 × 4,096) = 131,072 entriesThat calculation illustrates one layer, not your model's total adapter size. Real layers have different shapes, and adapter coverage varies. Tensorant targets the model's linear layers and trains assistant replies rather than the prompt text. Its fixed training settings explain the implementation.
QLoRA also learns a LoRA adapter, but loads the frozen base model in 4-bit form during training. The QLoRA paper introduces NF4 quantization, double quantization, and paged optimizers as memory-saving techniques. Tensorant's QLoRA choice uses 4-bit NF4 base weights; do not assume that every option from a research paper is a selectable platform setting.
The trained adapter is distinct from the base model. Keep the exact base-model identity and revision when you deploy it. Hugging Face's PEFT checkpoint documentation explains why adapter files are much smaller than complete model checkpoints and still depend on the original base model.
2. Compare the practical tradeoffs
| Decision | LoRA | QLoRA |
|---|---|---|
| Frozen base weights during training | Higher precision | 4-bit NF4 |
| Main benefit | Adapter training without quantizing the base weights | Less base-weight memory |
| GPU choice | Needs room for higher-precision base weights | Can make a smaller GPU feasible |
| Speed | Measure on the chosen GPU and workload | Quantization does not guarantee faster training |
| Output | Adapter for its base model | Adapter for its base model |
| Tensorant vision training | Supported with an image dataset and Pro plan | Not available |
Memory savings are not proportional to the complete training process. Activations, optimizer state for the adapter, attention buffers, and runtime overhead still take memory. A long sequence or larger micro batch can exhaust a GPU even with 4-bit weights. Hugging Face's quantized adapter training guide describes the distinction between quantized base weights and trainable adapter parameters.
For the 1.5B text-model size, Tensorant's initial readiness estimate is about 7.6 GiB for LoRA and 5.0 GiB for QLoRA. For an 8B size, it is about 23.4 GiB and 9.2 GiB respectively. These are platform launch estimates that omit sequence length and batch size, not observed peak-memory results. Read GPU memory before choosing hardware.
3. Start with the smallest useful task
For a ticket router, use Qwen/Qwen2.5-1.5B-Instruct with four labels: billing, account, technical, and other. The task requires consistent label selection, so a small instruct model is a reasonable starting candidate. Increasing parameter count before checking the base model adds cost without telling you which problem needs fixing.
Create a Classification project, import labeled conversations under Sources, review them, and freeze a dataset under Versions. Follow the Qwen fine-tuning tutorial for the exact steps and example file format.
Check that test tickets cover the difficult boundaries. For example:
| Ticket | Reference | Boundary being tested |
|---|---|---|
| “I cannot sign in to download my invoice.” | account |
A billing noun should not override the access problem |
| “The invoice export crashes after I choose a date range.” | technical |
A billing-related feature can have a technical failure |
| “I need a copy of last month's paid invoice.” | billing |
A billing request without a product failure |
| “Something went wrong. Can you help?” | other |
Avoid a confident label when the input is underspecified |
Keep near copies of these tickets in the same data partition. Tensorant's exact-prompt and source grouping helps, but does not replace reviewing semantic duplicates.
4. Configure a LoRA baseline experiment
Open Experiments, choose New experiment, and select the frozen version. Choose LoRA as the method and start with the default training settings:
| Parameter | Starting value | Purpose |
|---|---|---|
| Epochs | 2 | Two passes through the training examples |
| Learning rate | 0.0002 | Update size |
| Rank | 16 | Adapter capacity |
| Alpha | 32 | Adapter scaling relative to rank |
| Training context length | 2,048 | Maximum conversation length used in training |
| Micro batch size | 1 | Conversations processed at once |
| Gradient accumulation | 4 | Four micro batches contribute to one update |
Run Check dataset at the same sequence length before launch. If meaningful examples exceed it, shorten unnecessary material or raise the length with adequate GPU memory. Cutting off the answer is not a solution.
Choose an available Ampere-or-newer NVIDIA GPU that passes readiness with room to spare. Tensorant uses a single GPU per run; GH200 and GB200 machines are unsupported. For evaluation, choose Classification, use classification accuracy as the primary metric, and configure all four allowed labels.
Set runtime and hourly-price limits. Choose Check readiness, Save draft, Review and launch, and Approve and start. The experiment measures the base model and adapter with matching serving settings. See experiment setup.
5. Try QLoRA as a controlled change
Create a second experiment on the same frozen version. Match model revision, epochs, learning rate, rank, alpha, sequence length, scoring, and serving settings. Change the training method to QLoRA · 4-bit (less GPU memory).
For a method comparison, use the same training GPU if feasible. For a budget comparison, deliberately choose a smaller GPU and record that the comparison includes both method and hardware. Do not call a change in training time a method advantage when you also changed GPU type.
Tensorant pins a model revision when you save a draft. Copy the resolved commit from the first experiment when configuring the second if you need to ensure both use exactly the same weights. Reusing main at different times can resolve to different commits.
Training in 4-bit does not make production inference 4-bit automatically. Choose the adapter's deployment GPU based on serving readiness and context requirements. Tensorant does not currently allow FP8 weight quantization for adapter serving or published models. KV cache precision: FP8 is a separate setting that works with adapters; evaluate its effect on your answers before adopting it. See FP8 and adapters.
6. Decide from quality, spend, and failures
After both experiments complete, open Evaluations and compare the completed tuned evaluations on the same version. Check the selected metric and per-category changes. Read regressed answers rather than choosing whichever row has the larger overall number.
Keep a simple decision record:
Dataset version and model commit:
Method and training GPU:
Tuned classification accuracy:
Confusion between billing and technical:
Regressed answers requiring review:
Settled RunPod spend:
Reason for selecting this adapter:A small score difference with an uncertainty interval crossing zero is not strong evidence that one method wins. If both adapters satisfy the routing requirement, choose the configuration that offers suitable cost and operational simplicity. If both fail the same boundary cases, improve labels and examples before increasing rank.
For training on unfamiliar hardware, a small trial can check formatting, memory, and throughput. It requires a completed baseline, costs GPU time, counts as a training run, and produces an adapter that cannot be published. A trial is a systems check, not an evaluation of task quality.
7. Fix the right failure before repeating
If training runs out of memory, inspect the longest examples first. Reduce an unnecessarily large micro batch, use QLoRA for a text model, or choose a larger GPU. Reduce context length only when the examples still fit.
If training loss falls while validation loss rises, try fewer epochs on the next run. If the model returns prose instead of labels, inspect whether every assistant answer follows the same contract and whether your application supplies the training system instruction. If one category performs badly, add new independent examples for that category instead of duplicating existing rows.
Publish only the selected full-run adapter under API, after its GPU is released. Use the deployment guide to validate requests, latency behavior, and the serving contract. Track the complete fine-tuning cost, including comparisons and subsequent inference, when deciding whether the memory savings are useful for your application.