Skip to content
Training runs

Train

Training runs

Train a LoRA or QLoRA adapter on a frozen version, on one GPU in your RunPod account.

A training run trains one adapter for your project's base model, on one GPU in your RunPod account. It learns from the training examples of a frozen version and checks its progress on the validation examples after every epoch. The test examples are never used for training: they are kept for evaluations.

An experiment creates and launches a run for you. Configure a run by hand when you want to control each step yourself. Either way, every run is listed under Runs.

Before you start

  • A frozen version with training and validation examples.
  • A completed baseline evaluation on that version. Your adapter is compared with it later.
  • RunPod compute and storage, connected by an owner in Connections.
  • The member or owner role. Viewers can open runs and download their files. See roles.

Configure and save a run

Saving a run never starts a GPU.

  1. Open Runs under Tune and choose Configure a run.
  2. Under Setup, choose the Dataset version and a Completed baseline of that version.
  3. Choose the Training GPU. See GPUs.
  4. Choose the Training method.
  5. Change the hyperparameters and limits if you need to. The defaults are a good start.
  6. Choose Save training run.

The run is saved as a draft. Choose the pencil next to its name to rename it. A draft's settings cannot change: to try other settings, configure another run.

Setup

Field What to choose
Dataset version The frozen version to train on.
Completed baseline The base model's evaluation on this version. The adapter is compared with it.
Training GPU The GPU RunPod starts.
Model revision Leave at main unless you need a specific branch, tag or commit. It is fixed to one commit when you save.
Training method LoRA is the standard choice. QLoRA · 4-bit (less GPU memory) loads the model in 4-bit, so it fits on a smaller, cheaper GPU. Vision models use LoRA only.

Hyperparameters

These control how the adapter learns. Start with the defaults, then change one at a time based on the loss chart and your evaluation results.

Field Default When to change it
Epochs 2 How many times training reads all training examples. Raise it if validation loss is still falling at the end. Lower it if validation loss starts rising, which means the adapter is memorising the training examples.
Learning rate 0.0002 How big each update is. Lower it if the loss jumps around or rises.
Adapter size (LoRA rank) 16 How much the adapter can learn. Raise it for harder tasks or large datasets. A larger rank makes a larger adapter.
Adapter scaling (LoRA alpha) 32 How strongly the adapter's changes count (alpha ÷ rank). If you change the rank, change alpha in proportion, for example rank 32 with alpha 64.
Context length 2048 The longest example, in tokens, the trainer uses. Set it to fit your longest examples. Longer needs more GPU memory.
Micro batch size 1 Examples processed at once. Raise it to train faster if the GPU has memory to spare. Lower it if the GPU runs out of memory.
Gradient accumulation 4 Each update learns from micro batch size × gradient accumulation examples. Raise it for steadier training without using more memory.

Every field shows its range. See Limits.

Limits and storage

Field Default What it does
Maximum runtime (hours) 2 The longest the run may take, from launch, including GPU start-up and model download. A run that hits it stops and needs recovery.
GPU list-price limit ($/hour) 2 The highest hourly list price the GPU may have. The launch is refused if the GPU costs more.
Container disk (GB) 50 The GPU machine's own disk.
Output volume (GB) 80 Shown only when your storage is not a RunPod network volume.

Settings every run uses

These are fixed. View configuration shows the full configuration of a run.

  • Only the assistant replies are trained, not the prompts.
  • The adapter covers every linear layer of the model. For a vision model, only its language part: see Vision models.
  • After each epoch, the run measures validation loss and saves a checkpoint. The adapter comes from the last checkpoint.

Vision models

A run of a vision model trains like any other, with these differences:

  • It needs images. A vision model can't be trained on a version without images, and a text model can't be trained on a version with images.
  • LoRA only. QLoRA is not available for vision models yet.
  • Images count toward Context length. Each image becomes hundreds or thousands of tokens. Examples are never shortened to fit. See Images and the context length.
  • Examples are checked on the GPU first. If an image is missing or an example is too long, the run fails before training, and Training logs say which rows. RunPod bills the minutes the GPU ran.

GPUs

Training GPU lists RunPod's Secure Cloud GPUs in the datacenter of your network volume, with each GPU's memory and hourly price. GPUs that can start now are listed first. The others are disabled with the reason.

If the GPU costs more than your GPU list-price limit, choose Raise the limit to $price/h to match it.

Some GPUs cannot train:

  • GPUs older than NVIDIA's Ampere generation. The run fails when training starts.
  • GPUs on arm64 machines, such as the GH200 and GB200. They are refused when you save.

Will the model fit?

At launch, Tensorant estimates the GPU memory training needs:

  • LoRA: about 2.6 bytes per parameter, plus 4 GiB.
  • QLoRA: about 0.7 bytes per parameter, plus 4 GiB.

For an 8-billion-parameter model that is about 23 GiB with LoRA and about 9 GiB with QLoRA. A smaller GPU is refused. The estimate ignores context length and batch size, so a GPU that passes can still run out of memory. If it does, lower Micro batch size or Context length, switch to QLoRA, or choose a larger GPU.

Readiness checks

When you launch, Tensorant checks the run before RunPod starts anything. If a check fails, the launch is refused and each failed check is listed with its reason.

Check What it verifies
Runpod Your RunPod API key works.
Storage Your network volume can be reached and is still in the same datacenter.
Model The base model is public, not gated and has a chat template. A model outside the supported families gets a warning, and an MXFP4 checkpoint (such as openai/gpt-oss-20b) gets one too: the trainer can't train MXFP4 weights. See Projects.
Training Gpu The GPU has capacity in your volume's datacenter, costs no more than your limit, and has enough memory.
Runtime Your maximum runtime is within the limit on the Compute card. Listed only when it fails.

If a check about the platform's images fails, contact support. A passing check is not a promise: capacity can disappear a minute later, and the memory estimate can be wrong.

Costs

GPUs run in your RunPod account and RunPod bills them. Two limits keep a run in check:

  • GPU list-price limit ($/hour): the launch is refused if the GPU costs more.
  • Maximum runtime (hours): the GPU is stopped when the run reaches it.

Your organization can run a limited number of runs at a time. The Compute card in Connections shows the limits. A run holds a slot from launch until its GPU is released.

Estimated and actual spend

Each run shows its spend in the run list and on the run:

  • Estimated, marked "billing pending": the GPU's hourly price times how long it ran. Tensorant shows this until RunPod's billing has the run.
  • Actual spend: what RunPod billed for the run's GPU machine, including its CPU and disk, with when Tensorant last checked. For example, "Actual spend $1.23 · synced 5 min ago".

Tensorant checks RunPod's billing about every 30 minutes, and again soon after the GPU is released. RunPod's billing can take a while to update, so the console shows an estimate until then. "May still change" means the amount hasn't settled yet. It usually settles a few hours after the GPU is released.

Cap (est.) is what the run would cost if it used all its time at the current price. Neither number is a billing cap.

Your network volume isn't part of any run's spend, because every run shares it. Connections shows what RunPod billed for it this month.

Launch a run

Try a bounded trial first

In Configure a run, set Run purpose to Small training trial. A trial uses a reproducible subset of the version's training and validation examples; it never includes held-out test examples. Both trial and full runs currently require a completed baseline evaluation for the version.

Defaults are 128 training rows, 16 validation rows, 20 optimizer steps, seed 42 and a 20-minute runtime limit. Trials support at most 512 training rows, 64 validation rows, 50 optimizer steps and 30 minutes. A smaller dataset uses its available rows. The full-run settings you enter are saved separately from these trial bounds.

Save the draft, then use the same paid launch confirmation as an ordinary run. The trial runs on your GPU and is billed by RunPod. Formatting checks run on the sampled rows; training records optimizer-step times and peak allocated/reserved GPU memory. Existing runtime and cleanup controls apply. A model download can consume the time limit before any training measurements are available.

The Training trial report shows the actual measurements and, after a successful trial, extrapolates full-run time and GPU cost using the saved dataset size and training settings. It explains missing measurements and suggests changes when memory is tight or the saved runtime is too short. Estimates omit full-data preparation, validation, checkpoint/export time and storage; a small sample can miss unusually long examples or memory failures. It is not a model-quality evaluation or a price quote.

After the trial completes and its GPU is released, Create full training draft makes a new draft using the saved full-run settings. It starts from the base model and requires a separate launch confirmation. Trial adapters are labeled diagnostic and cannot be deployed, evaluated as a tuned model, or published through Tensorant.

Launch the saved draft

  1. Select the draft.
  2. Choose Launch on RunPod.
  3. Read the confirmation, then choose Launch GPU.

Tensorant runs the readiness checks, then asks RunPod for a GPU in your volume's datacenter.

Follow a run

Select a run to open it. The page refreshes by itself.

Status What it means
Draft Saved. No GPU is running.
Launching, Checking launch The GPU machine is being requested from RunPod.
Running, Preparing The GPU is starting, downloading the model and checking the examples.
Training · step N / M Training.
Adapter ready The adapter is verified. Its GPU is being released, or was released.
Completed An evaluation of the adapter completed.
Cancelling, Cancelled The run is being cancelled, or was cancelled.
Failed The run failed. The reason is shown on it. See Failure and recovery.
Needs recovery The runtime limit was reached or the adapter could not be packaged. See Needs recovery.

A run shows:

  • Loss: training loss after every step, validation loss after every epoch. Both should go down. If validation loss goes up while training loss keeps falling, the adapter is memorising the training examples: use fewer epochs next time.
  • Spend: the actual spend once RunPod's billing is synced, or an estimate until then. See Estimated and actual spend.
  • Estimates: elapsed time, time remaining and cost cap.
  • Training logs: the end of the trainer's log. Copy logs copies it.
  • Open Pod in RunPod opens the GPU machine in RunPod's console.
  • Download job files: the training and validation examples and the training configuration, as a ZIP. It never holds test examples.

When training finishes

  1. The adapter is packaged from the last checkpoint and checked.
  2. The run shows Adapter ready and the GPU is released.
  3. Choose Manage inference to deploy the adapter.
  4. When the deployment is ready, choose Evaluate adapter to compare it with its baseline. See Evaluate an adapter.

When that evaluation completes, the run shows Completed.

The adapter

  • Download adapter downloads a ZIP with adapter_config.json, adapter_model.safetensors and the tokenizer files. Use it with any tool that loads PEFT LoRA adapters, on the same base model and revision. Every role can download it, and each download is recorded in the audit trail.
  • Publish to Hugging Face uploads the adapter and a model card to your organization's Hugging Face account. Nothing from your dataset is uploaded. See Publishing adapters to Hugging Face.
  • The run's files (checkpoints, logs and the adapter) stay on your network volume under runs/run ID/. Tensorant does not delete them. RunPod bills the volume for the space they use.

Cancel a run

  1. Choose Cancel run.
  2. Confirm with Cancel run.

The GPU machine is deleted. Saved output stays on your network volume. Training cannot be resumed: to train again, create a retry draft.

A draft costs nothing and has no Cancel run. Runs cannot be deleted.

Failure and recovery

Failed

The run shows the reason, for example "RuntimeError: training stopped; inspect logs". Read Training logs for the details. Common causes:

  • The GPU ran out of memory. Lower Micro batch size or Context length, switch to QLoRA, or choose a larger GPU.
  • The GPU is older than Ampere. Choose a newer GPU.
  • An assistant reply could not be found in the model's chat template.
  • For a vision model, an example is too long once images are counted. Raise Context length or remove the example.

Fix the cause, then configure a new run or create a retry draft.

Needs recovery

A run needs recovery when it reached its runtime limit, or when its adapter could not be packaged. Tensorant stops the GPU machine instead of deleting it, so nothing is lost. Storage charges continue until you resolve it.

  1. Choose Verify adapter on RunPod.
  2. If the adapter was written, the run becomes Adapter ready and the GPU machine is deleted.
  3. If not, the run says "No completed adapter manifest exists on the volume…". This is usual when the runtime limit stopped training. Choose Cancel run, then configure a new run with a longer Maximum runtime (hours).

Training cannot continue from a checkpoint, so a new run starts from the beginning and RunPod bills its GPU time again.

If your storage is not a RunPod network volume, the panel asks you to download /workspace/data/adapter.zip from the stopped machine in RunPod and upload it with Verify and recover adapter.

Create a retry draft

When a failed or cancelled run's GPU is released, Create retry draft saves a new draft with the same settings. Launch it like any draft. To change a setting, configure a new run instead.

GPU release failed

If RunPod does not confirm the GPU machine was deleted, the run shows GPU release failed. Tensorant keeps trying. Choose Retry cleanup to try now, and check the machine with Open Pod in RunPod. Until it is released, the run keeps its slot.

Unresolved launch

If RunPod's answer to a launch was lost, the run stays at Checking launch and Tensorant keeps looking for the machine. It never creates a second one.

If nothing appears, wait ten minutes after the launch and check your Pods in RunPod. If there is no Pod named tune-run ID, choose Confirm absent Pod. The run is marked cancelled, and you can create a retry draft. If RunPod shows several Pods with that name, delete the extra ones there.

Tips

  • Use QLoRA to train a larger model on a cheaper GPU.
  • Set Maximum runtime with room to spare. A run that hits the limit has to start over.
  • Watch the loss chart. Validation loss that rises in later epochs means fewer epochs will do better.
  • Compare runs with evaluations, not loss alone. Lower loss does not always mean better answers.

Limits

Limit Value
GPUs per run 1
Epochs 0.1 to 20
Learning rate 0.000001 to 0.01
Adapter size (LoRA rank) 4 to 256
Adapter scaling (LoRA alpha) 1 to 512
Context length 128 to 32,768 tokens, in steps of 128
Micro batch size 1 to 32
Gradient accumulation 1 to 128
Maximum runtime 0.1 to 72 hours, and at most the limit on the Compute card
GPU list-price limit $0.01 to $100 an hour
Container disk 30 to 500 GB
Output volume 30 to 1,000 GB
Runs at a time Shown on the Compute card
Recovered adapter upload Smaller than 5 GB
Run name 160 characters

When something goes wrong

Problem What to do
"Complete a baseline evaluation for this dataset version first" Run a baseline evaluation on this version and wait for it to complete.
"Connect RunPod compute in Connections first" An owner connects RunPod in Connections.
"Training must use the exact revision evaluated by the managed baseline" Set Model revision to main, or to the commit the baseline's deployment served.
"Evaluate the base model on a deployment without FP8 quantization, then train." FP8 cannot serve adapters. Run a new baseline on a deployment without FP8, then train against it.
"An active run or unresolved cleanup already uses the compute slot" Wait for the other run to finish, or cancel it. A run with GPU release failed holds the slot until Retry cleanup succeeds.
"GPU list price exceeds your hourly limit" The price went up. Configure a run with a higher limit or a cheaper GPU.
"Resolve launch readiness checks before training: …" Fix what each check says. See Readiness checks. A GPU with no capacity often has some again later.
"Storage unavailable…" An owner can use Test storage connection in Connections.
"RuntimeError: training stopped; inspect logs" Read Training logs. See Failed.
"Pod disappeared before a verified adapter was collected" The machine was deleted in RunPod. Create a retry draft.
"No completed adapter manifest exists on the volume…" No adapter was written. Cancel the run, then train again with a longer runtime.
"Your role in this organization cannot make changes" You are a viewer. Ask an owner for the member role.