Skip to content
Publishing models

Serve

Publishing models

Give the base model or a trained adapter a name your applications call through the API, always on or waking on request.

A published model is a name, such as support, that your applications call through the API. Behind the name, a GPU in your RunPod account serves the project's base model or one of its trained adapters. You choose whether it stays always on or wakes on request.

Published models are managed on the API screen, under Use in the sidebar, in the Published models panel. It lists every published model of your organization, from every project.

Before you start

  • The owner role. Only owners publish, change, start, stop and delete published models. Everyone else can see them. See roles.
  • RunPod compute and a network volume, connected in Connections.
  • The project open whose model you want to publish.
  • To publish an adapter: a training run that shows Adapter ready or Completed, and GPU released.
  • For always on: a free deployment slot.

Always on or wake on request

Always on Wake on request
Runs on One dedicated RunPod GPU RunPod Serverless workers, each on its own GPU
While nobody calls it The GPU keeps running and billing Nothing runs, nothing is billed
First request after a quiet spell Answered at once Waits for a cold start, which can take a minute
You pay for The GPU, the whole time it runs Each worker while it runs, including its warm time
Runs for The days you choose, then it ends (you can extend it) Until you stop it
Uses a deployment slot Yes No

Choose wake on request when traffic is occasional or bursty, for development and testing, or when callers can wait for a cold start now and then. You pay only while workers run.

Choose always on when traffic is steady through the day, or when every answer must come quickly. You pay for the GPU all the time, so it costs more than wake on request when the model sits idle.

You can switch later with Change; the name stays the same.

Cold starts

A wake-on-request model with no worker running shows Asleep. When a request arrives, RunPod starts a worker; this is a cold start. For the caller:

  • The request is held while the worker starts, for up to about a minute, then answered as usual. Set your client's timeout to allow for this.
  • If the worker is still not ready after that minute, the API answers 503 with the code model_starting and a Retry-After: 15 header. Send the request again after 15 seconds. See Chat completions.

Once a worker answers, the model shows Warm and later requests are answered at once. The worker stops after Keep a worker warm (seconds) without requests. Raise it to have fewer cold starts; you pay for the time a worker waits. Model files stay on your network volume, so a worker does not download them again.

Publish a model

  1. Open the project, then API, and choose Publish a model.
  2. Enter a Name. See Names.
  3. Under Model to serve, choose Base model or an adapter.
  4. Under How it stays available, choose Wake on request or Always on.
  5. Fill in the GPU, price limit, context length and the fields of your choice. See Settings.
  6. Optional: open Advanced serving settings. See Advanced serving settings.
  7. Choose Review, check the cost, and choose Publish (wake on request) or Publish and start (always on).

The model shows Starting. An always-on model is ready when its GPU serves, usually after a few minutes. A wake-on-request model is ready when it shows Asleep.

Then create an API key and call the model by its name.

Settings

Field Default Range What it does
Model revision main Branch or commit Base model only. Fixed to one commit when you save.
Run for (days) 7 1 to the limit shown, at most 30 Always on. When it stops serving. You can extend it.
Starts allowed in 24 hours 24 1 to 96 Always on. Caps how often a GPU is started for this model, failed starts included.
Max workers 1 1 to the limit shown Wake on request. RunPod adds workers up to this number as requests grow.
Keep a worker warm (seconds) 60 5 to 3,600 Wake on request. How long a worker waits after the last request. Longer means fewer cold starts and more cost.
GPU None Always on: Secure Cloud GPUs in your volume's datacenter. Wake on request: Serverless GPU groups, priced at the most expensive GPU in the group.
GPU list-price limit ($/hour) 2 0.01 to 100 The highest price per hour, per worker for wake on request.
Context length (tokens) 4,096 512 to 32,768 The most tokens one request may hold, prompt and answer together.

Published models cannot use FP8 quantization. See FP8 and adapters.

Names

A name has up to 63 lowercase letters, digits, dots, underscores or hyphens, and starts with a letter or digit, such as support or support-v2. It is unique in your organization. Applications send it as model, so it cannot be changed: to rename, publish under the new name and delete the old one.

Vision models

A published vision model accepts up to 4 images per request, inline as base64 or as https links. Links are downloaded by the GPU in your RunPod account, so they must be reachable from the internet. See Images.

States

State Meaning Does the API answer?
Starting A GPU or Serverless endpoint is being set up. No: 503 model_starting, retry after 30 seconds
Running An always-on model's GPU is serving. Yes
Switching A replacement is starting; this one serves until it is ready. Yes
Asleep Wake on request, no worker running. Yes, after a cold start
Warm Wake on request, a worker is running. Yes
Cannot start Starting failed. See When a start fails. No: 503 model_unavailable
Stopped / Ended An owner stopped it, or an always-on model reached its end. No: 503 model_unavailable

Shared servers

Models of your organization that need the same setup share one server (one always-on GPU, or one Serverless endpoint): same base model and revision, same way of staying available, same GPU, context length and advanced serving settings. Each row shows its server, for example "Server: Qwen/Qwen2.5-1.5B-Instruct · NVIDIA GeForce RTX 4090 · 2 models".

  • A server serves its base model and every adapter on it, so publishing another adapter usually starts no new GPU.
  • Models on a server share its memory and capacity: heavy traffic to one slows the others.
  • A server never serves another organization's models. Each API key reaches only the models it may call. See Keys and usage.
  • An always-on server uses the lowest price limit and start limit of its models, and runs until the latest end. A Serverless server uses the highest Max workers and warm time.
  • Stopping one model keeps the server running for the others. When the last model leaves, the GPU or endpoint is released.

An always-on server appears in Deployments as Shared server · base model, tagged API. A Serverless server appears in RunPod as an endpoint named tune-serve- followed by its ID.

Change a model

Choose Change on the model's row, edit what you need, and choose Save changes. The name stays. To choose another adapter, open the model's project first.

The model keeps serving while the change takes effect:

  • Another adapter on the same base model is loaded into the running server. An always-on server takes adapters up to the LoRA rank it was started for (at least 16); a higher rank starts a replacement GPU.
  • Another GPU, context length, revision, advanced settings or way of staying available moves the model to a matching server, starting one if needed. The old one keeps answering until the new one is ready.
  • Price limit, start limit or end date changes in place.

A new always-on GPU needs a free deployment slot. If none is free, tick "Allow a pause of a few minutes…" to stop the running GPU first; the API answers 503 during the pause. After a switch, the old GPU keeps running for about 11 minutes so streaming answers can finish, and RunPod bills both.

Extend an always-on model

Choose Change, enter Run for (days from now), and save. The new end is that many days from now. A replacement GPU starts about 30 minutes before the old one's end, so RunPod bills both for a few minutes.

Start, stop and delete

  • Stop. Choose Stop, then Stop serving. Callers get an error until you start it again. Its GPU is released unless other models share its server.
  • Start. Choose Start on a stopped or ended model. For always on, enter Run for (days).
  • Delete. Stop the model first, then choose Delete. Callers are told the model does not exist. Its usage history stays.

Stopping a model's GPU in Deployments does not stop the model: Tensorant starts another GPU. Stop it on the API screen.

When a start fails

If a GPU cannot start, or a Serverless endpoint cannot be set up, the model shows Cannot start with the reason, and Tensorant tries again after 10 minutes. Every start counts toward Starts allowed in 24 hours, failed ones too.

If the serving engine fails on a new server for a reason that would repeat, such as running out of memory, Tensorant does not retry by itself, so you are not billed for repeated failures. Then either:

  • change the settings (often the advanced serving settings, context length or GPU), which starts it with the new ones; or
  • choose Try again to start it with the same settings.

See When the engine does not start for each reason.

Costs

  • Always on: RunPod bills the server's GPU from start until its end or until you stop it, used or not. Models on the same server share the GPU.
  • Wake on request: RunPod bills each worker while it starts, answers and stays warm. Nothing while asleep.
  • Your network volume is billed separately.
  • What you were billed: on the API page, a wake-on-request model shows its server's Serverless spend this month, as RunPod billed it. Models on the same server share one endpoint, so they share that amount. An always-on model's GPU is listed under Deployments, with its own spend. RunPod's billing can take a while to update: the line says when it was last synced.
  • Your plan limits how many models you may publish and how many requests and tokens you may use a month. See Plans and billing.

Limits

Limit Value
Published models per organization 20, or fewer if your plan says so
Name 63 characters
Run for (days) Up to the limit shown, at most 30
Starts allowed in 24 hours 1 to 96
Max workers Up to your organization's limit, shown under the field, counted across all wake-on-request models
Keep a worker warm (seconds) 5 to 3,600
GPU list-price limit ($/hour) 0.01 to 100
Context length (tokens) 512 to 32,768
Retry after a failed start 10 minutes
Cold start wait in the API Up to about 1 minute, then 503 model_starting

Limits on requests, such as answer length and time, are in Chat completions.

When something goes wrong

Message What to do
"Use up to 63 lowercase letters, digits…" or "This organization already publishes a model with that name" Choose another name. See Names.
"An organization can publish 1,000 models…" or "Your plan (plan) allows N published models…" Delete a model you no longer need, or change your plan.
"Only owners can do this" Ask an owner.
"Choose a verified adapter from this project" or "Complete training Pod cleanup before publishing its adapter" Wait until the run shows Adapter ready and GPU released.
"Serverless worker limit reached…" Your wake-on-request models together hit the worker limit. Lower Max workers, or stop another model.
"A Serverless worker on this GPU costs…" or "This Serverless GPU is unavailable or has no price" Raise the price limit or choose another GPU.
"No free deployment slot…" Stop a deployment, or for a change, tick "Allow a pause…". See Deployment slots.
"Start limit reached: N starts in 24 hours…" Find out why starts failed, then raise Starts allowed in 24 hours or wait.
"Resolve launch readiness checks before deploying: …" Fix each check. See Readiness checks.
A reason about memory or serving settings See When the engine does not start.
"Stopped because the organization was disabled or lost its connections." Reconnect RunPod and storage in Connections, then Start the model.
503 model_starting in your application The model is starting or waking. Retry after the Retry-After seconds.