Estimate LLM fine-tuning cost before launching
Budget training, base-versus-tuned evaluation, storage, iteration, and inference with Tensorant's RunPod controls and transparent cost formulas.
What you will build
Set a realistic budget for a model experiment and its ongoing serving.
Steps in this guide
LLM fine-tuning cost includes more than the training GPU. Budget the baseline evaluation, adapter evaluation, persistent storage, repeated experiments, any data-generation or judge calls, and the inference service your application will use afterward.
Tensorant runs compute in your RunPod account. The console gives you hourly-price and runtime controls, dataset checks, launch estimates, and billing visibility. Use them to make each experiment a deliberate decision, then reconcile estimates with the provider's settled bill.
This guide builds a budget for the support-router example using Qwen/Qwen2.5-1.5B-Instruct. All dollar values below are illustrative planning assumptions, not RunPod prices or promised runtimes. Replace them with the GPU prices, observed durations, and storage charges for your organization.
1. Separate one experiment from the whole project
An experiment in Tensorant uses three GPU stages:
- Deploy the base model and evaluate it on the frozen test examples.
- Train an adapter on the training split and release the training GPU.
- Deploy the adapter, evaluate it on the same test examples, and clean up.
Those stages may use different training and inference GPUs. Their price limits and maximum hours are separate settings in Experiments.
Your total project budget also needs to cover items outside that workflow:
| Item | What drives it |
|---|---|
| Data preparation | Review effort, imports, generation calls, and quality-judge calls |
| Repeated experiments | Method, data, and hyperparameter changes |
| Network volume | Persisted data, cached models, checkpoints, and adapters |
| Deployment left running | GPU time between evaluation and shutdown |
| Published inference | Dedicated uptime or serverless worker runtime and warm time |
| Platform subscription | Your plan's features and usage limits |
Use RunPod's billing documentation and Tensorant's plans and billing to confirm which provider bills each item. A plan allowing a training run does not mean the GPU itself is free.
2. Calculate the launch estimate from limits
The experiment's Maximum GPU list-price estimate uses the maximum duration and hourly-price limit of each stage:
Maximum GPU list-price estimate =
training maximum hours × training hourly-price limit
+ 2 × hours per inference deployment × inference hourly-price limitFor an illustrative plan:
| Setting | Assumption |
|---|---|
| Training maximum hours | 2 |
| Training maximum $/hour | $1.20 |
| Hours per inference deployment | 1 |
| Inference maximum $/hour | $0.80 |
The estimate is 2 × $1.20 + 2 × 1 × $0.80 = $4.00.
This figure does not cover everything RunPod may bill. It leaves out network storage, CPU and disk fees, and GPU release time. A kept deployment or a model left running after the experiment also creates additional spend. Approved maximum compute estimate ($) records what you approved; it is not an account-level spending cutoff. Read experiment costs.
The runtime and price limits are the controls to configure carefully. Do not set a low runtime merely to produce a smaller estimate if the run is likely to hit it and require starting over.
3. Estimate duration from your actual workload
For training, model size alone is insufficient. Example count, token length, epochs, batch settings, hardware, initialization, and model download all influence elapsed time.
A useful rough workload calculation is:
Effective batch size = micro batch size × gradient accumulation
Approximate optimizer steps =
ceil(training examples / effective batch size) × epochsFor example, 800 training examples, micro batch size 1, accumulation 4, and 2 epochs imply roughly 400 optimizer steps. This is a planning approximation; partial batches and the trainer's configuration affect the exact step count. Two datasets with 800 rows can still have very different costs if their token lengths differ.
Before starting a GPU, open Versions, select the dataset, set the intended Training sequence length, and choose Check dataset. Inspect median, 95th-percentile, and maximum conversation lengths, plus rows exceeding the limit. This tokenizer-based check does not start a GPU. See Dataset readiness.
When using a new model, GPU, or sequence length for a larger workload, consider a small training trial. It requires a completed baseline and counts as a billed training run. Its throughput and extrapolated full-run cost can improve your estimate, but its small sample can miss long examples that later exhaust memory.
For evaluation, count test prompts, bound output tokens, and consider the serving GPU's throughput. A label-only classifier can use a much smaller output allowance than a long-answer assistant. With a judge in manual evaluations, allow one extra model call per test answer and check its separate provider bill. Automated experiments have no judge.
4. Make the first experiment informative
Build a reviewed dataset before buying more GPU time. For support-router, inconsistent billing and technical labels can waste every experiment, regardless of model size.
- Create a Classification project with the Qwen instruct model.
- Import representative labeled conversations under Sources.
- Correct labels and remove unnecessary identifiers in Review.
- Freeze a version in Versions and inspect category coverage and independent test groups.
- Choose a metric and acceptance threshold that reflect your application before launch.
The Qwen guide walks through this dataset and experiment. Keep a test set that represents real difficult requests; an easy test can make a cheap adapter look useful while leaving expensive operational mistakes unresolved.
Start with default training settings and a small supported model. Choose LoRA if it fits with room for your examples. Try QLoRA when memory pressure forces a costly GPU choice. A lower-memory method is not automatically the faster or cheaper configuration. Compare quality and settled spend, as described in LoRA vs QLoRA.
5. Budget iteration before approving launch
Reserve money for a small number of specific iterations. For example, plan three experiments:
| Experiment | Question it answers |
|---|---|
| First reviewed dataset with LoRA | Does adapting the small model improve routing? |
| A corrected dataset version | Do new independent boundary examples fix observed mistakes? |
| QLoRA or one parameter change | Can the same requirement be met with a better compute configuration? |
Using the $4.00 illustrative launch estimate, three experiments allocate $12.00 for their maximum GPU list-price estimates. Add a separate allowance for storage, non-GPU machine fees, generation, trials, cleanup, and serving. Do not describe $12.00 as the project's total bill.
Choose Check readiness, resolve failures, Save draft, then Review and launch. Read the estimate and warnings before Approve and start. Check RunPod balance and available training runs under Connections and your plan.
Use Continue automatically through evaluation, training and cleanup after approval when you do not need stage reviews. A paused experiment can leave a deployment billing while it waits for you. Keep Keep the tuned model running after the comparison off unless you have a planned Playground session, and stop that deployment afterward.
6. Reconcile estimates with billed spend
An illustrative completed experiment might have 45 minutes of training-machine runtime and 12 minutes for each evaluation deployment. If their hourly GPU rates were $1.20 and $0.80:
Training GPU component: 0.75 × $1.20 = $0.90
Two evaluation GPU components: 2 × 0.20 × $0.80 = $0.32
Total GPU component: $1.22Use total billable machine runtime in this calculation, not only the interval between the first and last optimizer step. Startup and model preparation consume time too. The illustrative $1.22 still excludes the other charges already identified.
Open the experiment's linked Runs and Deployments to review spend. Tensorant initially shows estimated spend, then actual spend when RunPod billing arrives. “May still change” means the bill has not settled. The network volume is shared, so it appears separately under Connections, rather than as a charge on each run. See estimated and actual spend.
If a run reaches its runtime limit, resolve Needs recovery promptly. A stopped machine can still incur storage charges, and Tensorant cannot resume training from its checkpoint. Check whether an adapter was completed; otherwise cancel and launch a properly sized new run.
7. Include production inference in the decision
Training is occasional; inference may run every day. At an illustrative dedicated GPU rate of $0.80 per hour, 30 days of continuous operation would have a GPU component of 720 × $0.80 = $576, before other charges. That can dominate a small training budget.
Under API, choose Wake on request for occasional traffic when callers can tolerate cold starts. RunPod bills worker startup, requests, and warm time. Always on suits steady traffic or applications requiring ready capacity, and bills while idle. Compare these choices in serverless vs dedicated inference.
Use separate production and development API keys, keep keys on your server, and configure appropriate key restrictions under Keys and usage. Track request volume and output length alongside reviewed model quality.
For a ticket router, a useful business metric is cost per correctly routed ticket, with a separate count of costly misroutes. Include serving, review, and correction costs when comparing a fine-tuned model with a larger prompted model. Select the model that meets the routing requirement at a sustainable total cost, then revisit the budget when traffic or ticket patterns change.