Skip to content
Deployments

Serve

Deployments

Run the base model or a trained adapter on a GPU in your RunPod account, for evaluations, generation and the Playground.

A deployment runs one model, the project's base model or one of its trained adapters, on one GPU in your RunPod account. While it is Ready, you can use it in evaluations, in recipes and quality checks, and in the Playground.

A deployment is for work inside Tensorant: your applications cannot call it. To call a model from your own application, publish it.

Every deployment stops by itself at limits you set: a maximum runtime, an idle shutdown and a price limit. Experiments start and stop their own deployments; this page covers the ones you start by hand.

Before you start

  • RunPod compute and a network volume, connected by an owner in Connections. Until then, Deploy model is disabled.
  • The member or owner role. Viewers can only look. See roles.
  • To deploy an adapter: a training run that shows Adapter ready or Completed, and GPU released.
  • A free deployment slot.

Deploy a model

  1. Open Deployments under Use in the sidebar and choose Deploy model.
  2. Under Model to serve, choose Base model or an adapter (listed by its run's name).
  3. Choose the Inference GPU. See GPUs.
  4. Set the runtime, idle and price limits and the context length. See Settings.
  5. Optional: open Advanced serving settings. See Advanced serving settings.
  6. Choose Review deployment, check the cost limits, and choose Deploy inference GPU.

Tensorant first runs the readiness checks. If one fails, nothing starts and the dialog says why. Otherwise the deployment appears as Launching, then Starting, then Ready. The first start downloads the model to your network volume; later starts reuse it.

Choose Rename to change a deployment's name at any time.

Settings

Field Default Range What it does
Model revision main Branch, tag or commit Base model only. Fixed to one commit when you deploy.
Maximum runtime (hours) 1 0.1 to the limit shown under the field The longest it runs, counted from launch, used or not.
Idle shutdown (minutes) 15 5 to 120 Stops a ready deployment that gets no requests for this long.
GPU list-price limit ($/hour) 2 0.01 to 100 The launch is refused if the GPU costs more.
Serving context length 4,096 512 to 32,768 The most tokens one request may hold, prompt and answer together.

An adapter is always served on the base model and revision it was trained on. If its baseline was evaluated on a deployment, the adapter is served with that deployment's context length and advanced settings, read-only, so the comparison stays fair.

Vision models

A deployment of a vision model reads images as well as text, up to 4 per request. Nothing changes in the Deploy model dialog. Images count toward Serving context length. Images for evaluations are read from your network volume, so they never leave your RunPod account.

GPUs

Inference GPU lists RunPod Secure Cloud GPUs in your network volume's datacenter, with their list price per hour. GPUs that can start now come first; the others are disabled with the reason. If the GPU you want costs more than your limit, choose Raise the limit to $price/h. It works as for training runs.

Arm64 GPUs, such as the GH200 and GB200, are not supported.

Will the model fit?

Tensorant estimates the GPU memory a model needs from its parameter count: about 2.6 bytes per parameter plus 2 GiB, or about 1.6 bytes per parameter with FP8 quantization. An 8-billion-parameter model needs about 21 GiB. A GPU below the estimate is refused.

The estimate leaves out the context length, so a GPU that passes can still run out of memory. If it does, lower Serving context length or choose a larger GPU.

Readiness checks

Before anything starts, Tensorant checks the deployment. A failed check shows as "Resolve launch readiness checks before deploying:" followed by the check's name and reason.

Check What it checks Fix
runpod RunPod accepts your API key. Check the key in Connections.
storage Your network volume is reachable and in the expected datacenter. Check storage in Connections.
model The base model is public, ungated and has a chat template. A model outside the supported families gets a warning. A vision model also needs an image processor and no custom code. Choose another base model. See Projects.
serving_settings Your advanced serving settings are valid for this model. See Checks.
inference_gpu The GPU has capacity in your volume's datacenter, fits your price limit, and has enough memory. Choose another GPU, or raise the price limit.
runtime The maximum runtime is within the limit. Lower Maximum runtime (hours).
registry, inference_image, credentials The platform is ready to serve. Contact support.

Passing the checks does not guarantee the model will serve: capacity can disappear a minute later, and the memory estimate can be wrong.

Follow a deployment

Each row shows the model, its status, its GPU and price, and its limits, such as "$0.69/h estimated · 1h maximum · 15m idle limit". If it stopped or failed, the reason is under its name.

Spend shows what RunPod billed for the deployment's GPU machine, such as "Actual spend $0.72 · synced 10 min ago". Until RunPod's billing has it, it shows an estimate marked "billing pending". See Estimated and actual spend.

Status Meaning
Launching / Checking launch Tensorant is asking RunPod for the GPU.
Starting The GPU is loading the model. A ready deployment that stops answering also returns here.
Ready It answers, and you can choose it elsewhere.
Stopping / Stopped The GPU is being, or has been, deleted.
Failed RunPod refused the launch, or the GPU disappeared.

A second label shows the GPU: GPU in use, Releasing GPU, GPU released and so on. Open Pod in RunPod opens it in your RunPod console. On a ready adapter, Evaluate adapter opens an evaluation of its run.

When a deployment stops by itself

Reason shown Why
"Runtime limit reached" The maximum runtime passed.
"Idle timeout reached" It got no requests for the idle shutdown time.
"Model did not become ready within 30 minutes…" It took too long to start. Often the GPU is too small or the context length too long.
"Model stopped answering and did not recover within 30 minutes…" It stopped answering for 30 minutes.
"Inference Pod stopped unexpectedly…" RunPod reports the GPU is no longer running.
A reason about memory or serving settings See When the engine does not start.

The serving log is on your network volume, at inference/deployment ID/inference.log. The deployment ID is the part after base- or adapter- in its served model name.

Stop a deployment

  1. Choose Stop inference.
  2. Confirm with Stop inference. Requests in progress may fail. Model files and adapters stay on your network volume.

A stopped deployment cannot be restarted: deploy a new one.

A deployment tagged API serves a published model. If you stop it here, Tensorant starts another GPU for the model. Stop the model on the API screen instead.

Cleanup and unresolved launches

  • GPU release failed. Deleting the GPU failed. Tensorant keeps retrying. Choose Retry cleanup to retry now, and check in RunPod that the Pod is gone.
  • Launch outcome unknown. RunPod's answer to the launch was lost, and the deployment stays in Checking launch. Tensorant looks for a Pod named tune-infer-deployment ID. If none appears, wait ten minutes after the launch, check your Pods in RunPod, then choose Confirm absent Pod. This marks the deployment Stopped.
  • Multiple matching Pods. Delete the extra Pods with that name in RunPod.

Deployment slots

Your organization can run a limited number of deployments at a time, shown in the page header as "up to N at a time". Every deployment counts, from every project: the ones you start, the ones experiments start, and the GPUs of always-on published models. A deployment holds its slot until its GPU is released. When no slot is free, stop a deployment or wait for one to be released.

Costs

RunPod bills the GPU to your RunPod account from start until it is deleted, used or not. The most a deployment can cost is about its maximum runtime times its price limit. Your network volume is billed separately, and Connections shows it. Stopping a deployment keeps the model files on your volume for the next start.

Each deployment shows an estimate until Tensorant has synced RunPod's billing for it, then the actual spend. RunPod's billing can take a while to update, so an amount may still change for a few hours after the GPU is released.

Limits

Limit Value
Maximum runtime (hours) Up to the limit shown under the field
Idle shutdown (minutes) 5 to 120
GPU list-price limit ($/hour) 0.01 to 100
Serving context length 512 to 32,768 tokens
Time to become ready, or to recover 30 minutes
Deployments at a time Shown in the page header
GPUs per deployment 1

When something goes wrong

Message What to do
"Resolve launch readiness checks before deploying: …" Fix each check it names. See Readiness checks.
"Stop the existing inference deployment and resolve cleanup before deploying another" No slot is free. Stop a deployment or retry its cleanup.
"Inference runtime exceeds the server limit" Lower Maximum runtime (hours) to the limit shown.
"Selected GPU is unavailable on Secure Cloud" or "GPU list price is unavailable or exceeds your hourly limit" The GPU's capacity or price changed. Choose another GPU or raise the price limit.
"Choose a verified adapter from this project" Wait until the run shows Adapter ready.
"Complete training Pod cleanup before deploying its adapter" Wait for GPU released on the run, or retry its cleanup.
"Storage configuration changed…" The adapter is on a different volume. Connect the volume it was trained on.
"Model did not become ready within 30 minutes…" Read inference.log on your volume. Try a larger GPU or a shorter context length.
"GPU release failed" Choose Retry cleanup. See Cleanup and unresolved launches.
"Your role in this organization cannot make changes" You are a viewer. Ask an owner for the member role.