Serverless vs dedicated GPU inference for fine-tuned LLMs
Choose RunPod Serverless or a dedicated GPU for LLM inference using cold starts, traffic patterns, worker costs, and Tensorant's publishing controls.
What you will build
Choose and configure an LLM inference mode that fits your application's response-time needs and compute budget.
Steps in this guide
LLM inference setup depends on when requests arrive and how long users can wait. An occasional support-routing job has different needs from a continuously used interactive tool.
Tensorant publishes LLM APIs through Wake on request, using RunPod Serverless workers, or Always on, using a dedicated GPU. Both serve base models or LoRA adapters. This guide uses support-router from the deployment guide.
Step 1: Write down your application's requirements
Before choosing a GPU or worker count, describe your workload:
| Question | Why it matters |
|---|---|
| Does a person wait for the answer? | A cold start can interrupt an interactive workflow. |
| When do tickets arrive? | Idle gaps give workers time to stop. |
| How many requests overlap? | Daily volume can hide capacity problems. |
| How large are requests? | Context and generation length affect memory and time. |
| What happens during failures? | A review queue gives you room to recover. |
| What is the monthly budget? | Hourly limits do not cap total spending. |
For a queued router, you might accept classification within several minutes. An interactive console might need answers within seconds. Choose these requirements for your application before measuring a GPU.
Separate model quality from serving performance. Select an adapter using held-out LLM evaluation, then compare inference modes with that same adapter, prompt, output length, and context settings.
Step 2: Understand what each mode runs
| Behavior | Wake on request | Always on |
|---|---|---|
| Compute | RunPod Serverless workers | One dedicated RunPod GPU |
| Idle state | Asleep when no worker is running | GPU continues serving |
| First request after inactivity | May wait for a cold start | Avoids waking a sleeping worker |
| Scaling control | Max workers caps worker count | Capacity of the selected server |
| Compute billing | Worker startup, answering, and warm time | The whole time the GPU runs |
| Deployment slot | Not required | Required |
| Lifetime | Until stopped | Ends after Run for (days) unless extended |
Start with wake on request for intermittent calls and queues with long quiet periods. Consider always on for steady requests or users who cannot wait for a worker. Its capacity remains finite, and overload or failures can still interrupt service.
An internal Deployment has maximum runtime and idle shutdown controls and serves Tensorant evaluations, recipes, and the Playground. Your application calls a published model, configured on API. Do not mistake the idle shutdown setting on a Deployment for a published model's warm-worker setting. See deployments and publishing models.
Step 3: Start with a controlled configuration
Open the model's project, then API, then Publish a model. Enter support-router as the name and select the trained adapter under Model to serve.
For an occasional queue, choose Wake on request under How it stays available. Start with Max workers of 1 and Keep a worker warm (seconds) of 60. Choose a Serverless GPU group with enough memory, set an acceptable GPU list-price limit ($/hour), and use Context length (tokens) of 4,096 for short tickets. Review the cost before choosing Publish.
Tensorant prices a Serverless GPU group using its most expensive GPU. The price limit is per worker, so keep the worker cap low until you have capacity measurements.
For an interactive tool, choose Always on, a suitable GPU, and Run for (days) that covers your intended period. Set Starts allowed in 24 hours to constrain repeated starts and choose Publish and start after reviewing. Check that the model reaches Running and plan for extending its end date.
Keep advanced settings at defaults initially. Conversations at once, context length, precision, and GPU affect memory and throughput. See advanced serving settings before adjusting them.
Step 4: Measure cold and warm requests separately
Create a key limited to support-router and store it as TUNE_API_KEY on your server. Store the copied API address, ending in /v1, as TUNE_API_BASE_URL. Use synthetic tickets for the first timing checks.
This request saves the answer separately and prints the HTTP status and total elapsed time:
curl --silent --show-error --max-time 150 \
--output support-probe.json \
--write-out 'status=%{http_code} total_seconds=%{time_total}\n' \
"$TUNE_API_BASE_URL/chat/completions" \
-H "Authorization: Bearer $TUNE_API_KEY" \
-H 'Content-Type: application/json' \
--data-binary @- <<'JSON'
{
"model": "support-router",
"messages": [
{
"role": "system",
"content": "Classify the support ticket as billing, account, technical, or other. Reply with exactly one label and no other text."
},
{"role": "user", "content": "My invoice was charged twice."}
],
"temperature": 0,
"max_tokens": 16
}
JSONFor a cold sample, wait for Asleep, then send the request. Include any retry waits in the time through successful completion. For warm samples, send requests while Warm. Repeat across representative prompt lengths and concurrency, within your key and plan limits.
Record completion times, errors, tokens, and concurrent requests. Keep cold and warm results separate and review the 95th percentile alongside the median. This curl request measures a whole answer. For longer conversations, measure streaming and time to first token separately.
Step 5: Give cold starts an explicit retry policy
When a request wakes a sleeping worker, Tensorant waits for up to about a minute. If the worker is still not ready, the API returns 503 model_starting and Retry-After: 15. A model showing Starting can return 503 model_starting immediately with Retry-After: 30.
Honor that header, allow time for startup, and bound the attempts. The adapter deployment guide includes a complete Python client.
Retry 429 rate_limit_exceeded after its header's delay. Fix monthly quota_exceeded errors, stopped models, bad keys, and wrong model names before retrying.
If a stream stops before data: [DONE], its answer is incomplete. Do not apply a partial routing label or repeat downstream actions automatically. See chat completions for limits and errors.
Step 6: Estimate cost using billed worker time
Serverless cost is not just generation time. Workers are billed while starting, answering, and waiting through their warm period. Longer warm time reduces some cold starts but adds paid idle time. Persistent storage is a separate charge. RunPod's Serverless billing documentation explains worker phases and storage costs.
Use these planning formulas, then replace estimates with observed billing:
Dedicated compute = GPU hourly price × running hours
Serverless compute = worker hourly price × total billed worker hoursHere is an illustrative calculation with assumed rates, not quoted RunPod prices. Suppose a dedicated GPU costs $0.50 per hour and the Serverless worker equivalent costs $0.80 per hour. For a 30-day month:
| Assumed usage | Compute calculation | Estimated compute |
|---|---|---|
| Dedicated GPU running continuously | 720 hours × $0.50 | $360 |
| Serverless, 60 total billed worker hours | 60 hours × $0.80 | $48 |
| Serverless, 500 total billed worker hours | 500 hours × $0.80 | $400 |
Under these assumptions the compute costs meet at 450 billed worker hours. That threshold includes startup and warm time, not only useful work. With multiple workers, add their hours together. Two workers running for one wall-clock hour contribute two worker hours.
Both configurations must meet your capacity and latency needs. The example excludes storage, the Tensorant plan, and overlapping GPUs during changes. A slower GPU can be worse value despite a lower hourly rate. See cost planning.
Step 7: Tune one control at a time
If isolated requests repeatedly wake workers and users wait too long, increase Keep a worker warm (seconds) and compare both latency and billed time. Tensorant permits 5 to 3,600 seconds. A longer setting is useful when requests tend to follow one another closely, but less useful when they remain hours apart.
If bursts exceed one worker's capacity, consider raising Max workers within the organization limit. Check the application's Requests at once limit too. Extra workers do not help when requests are refused before reaching the model because the key or plan limit is already exhausted.
For sustained demand, compare Always on using the same workload. Change the mode through Change on the published model. Its name stays stable, so callers can keep using support-router. Switching may start replacement compute, and an always-on replacement needs a free deployment slot. Review the cost confirmation and the switching behavior before saving.
Step 8: Check the decision against production behavior
Review Usage on API for errors and tokens. Allow for billing sync delays. Always-on spend appears under Deployments; Serverless spend appears on the published model's API row.
Compatible models can share a server. One busy model can slow another or keep workers warm. Account for shared load when interpreting measurements.
Keep a manual review path for the router, alert on persistent errors, and track response times from your application. Revisit the serving mode when traffic patterns change. For always on, extend the scheduled end before it expires. When the model is no longer needed, stop it on API so its published availability and compute lifecycle stay aligned.