Train
Experiments
Measure the base model, train an adapter, compare the two and release the GPUs, in one guided workflow.
An experiment answers one question: does fine-tuning make this model better on your data? It measures the base model, trains an adapter, measures the adapter on the same test examples, tells you whether it meets your thresholds, and releases the GPUs. You set it up once and approve its cost, and it runs each step for you.
Use an experiment for your first adapter and whenever you want a fair, repeatable comparison. Everything it starts also appears where it would if you did it by hand: its deployments under Deployments, its run under Runs and its evaluations under Evaluations.
Before you start
- A frozen version with training, validation and test examples.
- RunPod compute and a RunPod network volume, connected by an owner in Connections.
- A free training slot and a free deployment slot. The Compute card in Connections shows your limits. Always-on published models share the deployment slots.
- The member or owner role. Viewers can follow experiments but not start them. See roles.
What an experiment does
- Baseline. Deploys the base model, evaluates it on the version's test examples, then stops the deployment.
- Train. Trains an adapter on the training examples, then releases the training GPU.
- Deploy. Deploys the adapter with the same GPU, context length and serving settings as the base model.
- Compare. Evaluates the adapter on the same test examples with the same settings.
- Cleanup. Stops the tuned deployment (unless you keep it), confirms nothing is still running, and records whether the adapter meets your thresholds.
Before each launch, the experiment checks GPU capacity and prices again, because they change.
Create an experiment
- Open Experiments under Tune and choose New experiment.
- Fill in Setup, Compute limits, Evaluation and Acceptance thresholds. The sections are described below.
- Choose Check readiness. Each check shows pass, warning or fail. See Readiness checks.
- When every check passes, choose Save draft.
Saving does not start anything. It fixes the model revision to one commit, so later changes on Hugging Face do not affect the experiment.
Setup
| Field | What to choose |
|---|---|
| Experiment name | A name you will recognise later. |
| Frozen dataset | The version to train and test on. Dataset coverage shows how its examples split by category. See Coverage. |
| Model revision | Leave at main unless you need a specific branch, tag or commit. |
| Successful launch preset | Appears when a saved preset fits this model. It copies the settings of an experiment that worked. See Launch presets. |
| Training method | LoRA, or QLoRA · 4-bit (less GPU memory) for a smaller GPU. Vision models use LoRA only. |
Compute limits
These are hard limits for the training GPU and the two deployments.
| Field | Default | What to change and why |
|---|---|---|
| Training GPU | None | Pick a GPU with enough memory for the model. See GPUs. |
| Training maximum hours | 2 | Raise it for large datasets or many epochs. Training that is not done by then stops and needs recovery. |
| Training maximum $/hour | 2 | The highest hourly list price the GPU may have. |
| Inference GPU | None | The GPU for both deployments. |
| Hours per inference deployment | 1 | Raise it if a large test set needs longer to evaluate. |
| Inference maximum $/hour | 2 | The highest hourly list price the inference GPU may have. |
| Inference context length | 4096 | Each test prompt plus Maximum answer tokens must fit. Raise it for long prompts. |
| Inference idle shutdown (minutes) | 15 | A deployment with no requests for this long stops. |
When a GPU costs more than its limit, choose Raise the limit to $price/h to match it.
Advanced serving settings apply to both deployments. FP8 quantization is not available here, because it is not yet verified with adapters. See Serving settings.
Advanced training parameters holds the training settings as JSON (epochs, learning rate, rank, alpha, context length, batch size, gradient accumulation and disk sizes). The defaults are a good start. See Hyperparameters for what each one does.
Evaluation
Both evaluations use these settings, so the comparison is fair. See Presets and metrics.
| Field | What it does |
|---|---|
| Task | Q&A, Classification, Structured JSON or Numeric. Starts from the project's use case. |
| Primary metric | The metric the decision is based on. |
| Maximum answer tokens | The longest answer the model may write. 512 by default. |
| Allowed labels, Response JSON schema, Field paths, tolerances | Shown for the task that needs them. |
| Evaluation system prompt (optional) | Leave blank to keep each example's own system prompt. |
An experiment's evaluations have no LLM judge. To score with a judge, run the evaluations by hand. See The LLM judge.
Acceptance thresholds
Set what "good enough" means for your use case.
| Field | Default | What it requires |
|---|---|---|
| Minimum tuned score (%) | 50 | The adapter's average score on the primary metric. |
| Minimum improvement (percentage points) | 1 | How much better than the base model the adapter must score. |
| Minimum independent test groups | 10 | How many independent groups of test examples the comparison must cover. |
| Require the entire 95% interval to be above zero | Off | The improvement must hold up when the test examples are re-sampled. Needs at least 10 independent groups. |
Two more choices:
- Continue automatically through evaluation, training and cleanup after approval. See Automatic or stage by stage.
- Keep the tuned model running after the comparison. See Keep the tuned model running.
Approve and start
- Select the draft and choose Review and launch.
- Check the GPUs, their hours and price limits, and the Maximum GPU list-price estimate.
- If warnings are listed, read them and tick I reviewed these launch and dataset warnings.
- Keep Approved maximum compute estimate ($) at the estimate, or raise it.
- Choose Approve and start.
The approved amount is recorded with the experiment. It does not stop spending: each stage's hour and price limits do.
Automatic or stage by stage
With Continue automatically… ticked, the experiment runs from start to finish without stopping.
Without it, the experiment pauses after each step and shows Paused. Review the step, then choose Review and continue and Approve and start to run the next one.
Warning
While a paused experiment has a deployment ready, RunPod keeps billing that GPU and its idle timer keeps counting. If the deployment stops before you continue, the experiment fails. Continue within the Inference idle shutdown time.
The experiment also pauses when a GPU has no capacity at launch time, or when another run or deployment is using the slot it needs. Wait, then choose Review and continue.
Follow an experiment
All experiments lists each experiment with its stage and status: Draft, Running, Paused, Cancelling, Cancelled, Failed or Completed. Select one to see:
- The stages, with the current one marked, and the Next action.
- This experiment needs attention, with the reason, when something stopped it.
- Training details, which opens the run with its loss chart and logs. See Follow a run.
- Compare answers, once it completed, which opens the comparison. See Compare two evaluations.
Read the result
When the experiment completes, it says Meets configured acceptance thresholds or Does not meet configured acceptance thresholds, with the reasons. It also shows:
- Tuned score: the adapter's average score on the primary metric.
- Change: the difference from the base model, in percentage points.
- Regressed answers: how many test examples scored lower with the adapter than with the base model.
An adapter misses the thresholds when its score is too low, its improvement is too small, the test set has too few independent groups, or (if you required it) the 95% range does not lie fully above zero.
Choose Compare answers to see which examples improved and which got worse. That tells you what data to add next. See Read a comparison.
Meeting the thresholds shows the adapter did better on this test set. It does not guarantee results in real use.
Keep the tuned model running
If you ticked Keep the tuned model running after the comparison, the tuned deployment stays up when the experiment completes. It becomes an ordinary deployment.
- Open playground lets you try it. See Playground.
- Stop it stops it now.
It stops by itself when idle or at its time limit. Until then it uses a deployment slot and RunPod bills it.
After an experiment
Follow-ups
A completed experiment shows Next dataset iteration, where you note what to fix in the next version.
- Under New follow-up, describe the Observed failure pattern and the New data or review action.
- Optionally set a category, priority and status.
- Optionally link new examples you collected for it. Only training examples created in this project after the experiment can be linked.
- Choose Record follow-up.
Collect new examples opens Sources, and Open Review opens Review.
Launch presets
Once an experiment completed and its GPUs are released, you can save its settings as a preset: enter a Launch preset name and choose Save successful settings. A preset keeps the model revision and the training and serving settings. It records settings that ran, not whether the adapter met your thresholds.
Presets are shared across your organization's projects. New experiments on the same base model offer them under Successful launch preset.
Cancel an experiment
- Choose Cancel and clean up.
- Choose Confirm cancellation.
The experiment stops its deployments and training run, then shows Cancelled. Your datasets and any finished adapter are kept. If its training run needs recovery, the experiment waits until you recover or cancel that run. See Needs recovery.
A failed experiment cleans up the same way. It cannot be resumed: fix the cause and create a new one. Its completed evaluations stay under Evaluations.
Costs
GPUs run in your RunPod account and RunPod bills them. The Maximum GPU list-price estimate assumes each of the three GPU stages runs for its full hours at its price limit:
training hours × training $/hour + 2 × hours per deployment × inference $/hour
With the defaults: 2 × $2 + 2 × 1 × $2 = $8. Most stages end sooner: each deployment stops when its evaluation is done, and training ends when the adapter is ready.
The estimate is not a billing cap. It leaves out the network volume, CPU and disk fees, and the time it takes to release a GPU. A paused deployment and a kept tuned model are billed while they run.
Tips
- Start with the default training settings. Change one thing at a time between experiments, so you know what made the difference.
- Use Continue automatically unless you want to inspect each stage. A paused deployment keeps costing money.
- Require the 95% interval only when your test set has at least 10 independent groups. With fewer, the adapter cannot meet the thresholds.
- After an experiment that ran well, save a launch preset so the next one starts from settings that worked.
- Use follow-ups to plan your next version from the answers that regressed.
Limits
| Limit | Value |
|---|---|
| Experiment name | 160 characters |
| Training maximum hours | 0.1 to 72, and at most the limit on the Compute card |
| Hours per inference deployment | 0.1 to 72, and at most the limit on the Compute card |
| Maximum $/hour, training and inference | $0.01 to $100 |
| Inference context length | 512 to 32,768 tokens |
| Inference idle shutdown | 5 to 120 minutes |
| Maximum answer tokens | 16 to 8,192, and below the inference context length |
| Evaluation system prompt | 10,000 characters |
| Minimum independent test groups | 1 to 100,000 |
| Approved maximum compute estimate | Up to $100,000, and at least the estimate |
| Follow-ups per experiment | 100 |
| Linked examples per follow-up | 100 |
| Experiments listed | The 100 newest |
When something goes wrong
| Problem | What to do |
|---|---|
| "Connect RunPod and storage in Connections before launching experiments." | An owner connects RunPod and a network volume in Connections. |
| The GPU list is empty or does not load | Choose Refresh GPUs. If it stays empty, an owner can use Test compute in Connections. |
| A test example does not fit the context length | Raise Inference context length, lower Maximum answer tokens, or shorten the system prompt. |
| "Incompatible reference answers: …" | Some test answers do not fit the task. See Reference answers are checked first. |
| The experiment is Paused because readiness changed | Usually the GPU has no capacity right now. Choose Review and continue later, or cancel and pick another GPU. |
| "An active run or unresolved cleanup already uses the compute slot" | Another training run is using the slot. Wait for it or cancel it, then choose Review and continue. |
| "Stop the existing inference deployment and resolve cleanup before deploying another" | A deployment or always-on published model is using the slot. Stop it or wait, then choose Review and continue. |
| "Image configuration changed…" | Tensorant was updated after you saved the draft. Create a new experiment. |
| A deployment, evaluation or training stage failed | Open the linked deployment, evaluation or run to read its error. Fix the cause and create a new experiment. For training, see Failure and recovery. |
| "Experiment stage unavailable…" | Often a deployment stopped while the experiment was paused. Cleanup continues by itself. Create a new experiment. |
| "Experiment is waiting for maintenance to finish…" | Nothing to do. It continues when maintenance ends. |
| "Your role in this organization cannot make changes" | You are a viewer. Ask an owner for the member role. |