Deploy a LoRA adapter as an LLM API on RunPod
Publish a fine-tuned support-routing model in Tensorant, choose its RunPod serving mode, and call an LLM API with secure keys and bounded retries.
What you will build
Publish a trained LoRA adapter and connect a support-routing application to its chat completions API.
Steps in this guide
Training produces an adapter. Your application still needs a running base model, the correct adapter, an address to call, and a reliable way to handle startup and errors. This guide takes a support-routing adapter from a completed Tensorant training run to an application-facing LLM API on GPUs in your own RunPod account.
The example uses Qwen/Qwen2.5-1.5B-Instruct and four labels: billing, account, technical, and other. A ticket about a duplicate charge should go to billing. A request to change a login email should go to account. Follow the Qwen fine-tuning guide to prepare and train that model first.
Step 1: Check that the adapter is ready to serve
Open the project's training run. It must be a full run showing Adapter ready or Completed, with GPU released. A trial run is useful for checking setup, but its adapter cannot be deployed or published.
Record the run name, dataset version, and its evaluation comparison before moving on. The adapter must be served with the same base model and revision it was trained on. Tensorant uses that identity when you select the run, so you do not need to rebuild or merge its weights manually.
Ask an organization owner to confirm that RunPod compute and a network volume are connected in Settings, then Connections. Publishing models and creating API keys require the owner role. See connections and training runs.
Step 2: Validate the adapter inside Tensorant
An internal Deployment lets you evaluate and try the model inside Tensorant. A published model is what your application calls. A ready Deployment alone does not create a public API model name.
- Open Deployments under Use and choose Deploy model.
- Under Model to serve, choose the adapter by its training run name.
- Choose an Inference GPU that can hold the base model, adapter, and request context.
- Set Maximum runtime (hours), Idle shutdown (minutes), GPU list-price limit ($/hour), and Serving context length. For a short classification smoke test, start with a one-hour runtime and 4,096 context tokens if the model fits.
- Choose Review deployment, then Deploy inference GPU.
Wait for Ready. Open Playground, select the adapter as Model A, set Model B to None, and set Temperature to 0. Use the same system prompt as the dataset, then try ordinary tickets and ambiguous ones. Check that the answer is exactly one permitted label, without an explanation or Markdown.
For an actual release decision, compare the adapter and base model on frozen test data. Review per-label errors as well as overall accuracy using the evaluation guide. Stop the internal deployment when finished if you no longer need it. It also stops at its configured limits.
Step 3: Choose how the API model stays available
Use Wake on request for an asynchronous ticket queue or occasional development traffic. RunPod Serverless workers start when needed and stop after their warm period. The first request after an idle period can wait for a cold start.
Use Always on when your support tool needs prompt responses throughout the day. One dedicated GPU stays running and billing until the model's scheduled end or until you stop it. A free deployment slot is required.
For this first integration, choose Wake on request, Max workers of 1, and Keep a worker warm (seconds) of 60. These are starting settings for a small queue, not a throughput guarantee. See serverless versus dedicated inference before committing to production settings.
Step 4: Publish the adapter
- Open the project, then API under Use.
- In Published models, choose Publish a model.
- Enter
support-routeras the Name. Applications send this name in themodelfield. - Under Model to serve, choose the evaluated adapter.
- Under How it stays available, choose Wake on request and set the worker settings above.
- Choose the GPU, set a GPU list-price limit ($/hour) you accept, and set Context length (tokens) to 4,096 for this short-ticket example.
- Keep Advanced serving settings at their defaults initially. Choose Review, check the costs, then Publish.
For Always on, set Run for (days) and use Publish and start instead. An always-on model is ready at Running. A wake-on-request model is ready at Asleep, which means it can accept a request and wake a worker.
Published models cannot use Quantization set to FP8, computed at start. That setting is different from FP8 KV cache precision. See advanced serving settings if you need to adjust memory later.
Step 5: Create an application-specific API key
On API, choose Copy address. The address ends in /v1. Store it as TUNE_API_BASE_URL in your application's environment.
Under API keys, choose Create key. Name it for the application, select Only the models I choose under May call, and select support-router. Start with Requests per minute of 60 and Requests at once of 4, then review these limits against your workload and plan.
Choose Create key, Copy key, and I copied it. Store the key as TUNE_API_KEY using your server's secret configuration. Keep it out of browser code, source control, and logs. The key is shown only once. See keys and usage for rotation and revocation.
Step 6: Call chat completions from your backend
Save this Python program as support_router.py. It uses the standard library, checks the output against the allowed labels, and retries only startup and temporary rate-limit responses. Set the two environment variables from Step 5 before running it.
import json
import os
import sys
import time
from urllib.error import HTTPError
from urllib.parse import urlsplit
from urllib.request import Request, urlopen
LABELS = {"billing", "account", "technical", "other"}
SYSTEM_PROMPT = (
"Classify the support ticket as billing, account, technical, or other. "
"Reply with exactly one label and no other text."
)
def classify(ticket):
if not ticket.strip() or len(ticket) > 4000:
raise ValueError("Provide a nonempty ticket of at most 4,000 characters.")
base = os.environ["TUNE_API_BASE_URL"].rstrip("/")
parsed = urlsplit(base)
if (
parsed.scheme != "https"
or not parsed.hostname
or parsed.path != "/v1"
or parsed.username
or parsed.password
or parsed.query
or parsed.fragment
):
raise ValueError("Use the HTTPS API address copied from Tensorant.")
payload = {
"model": "support-router",
"messages": [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": ticket},
],
"temperature": 0,
"max_tokens": 16,
}
request = Request(
base + "/chat/completions",
data=json.dumps(payload).encode("utf-8"),
headers={
"Authorization": "Bearer " + os.environ["TUNE_API_KEY"],
"Content-Type": "application/json",
},
method="POST",
)
for attempt in range(4):
try:
with urlopen(request, timeout=150) as response:
result = json.load(response)
except HTTPError as error:
try:
detail = json.loads(error.read()).get("error", {})
except (ValueError, AttributeError):
detail = {}
code = detail.get("code", "unknown")
retryable = (error.code, code) in {
(503, "model_starting"), (429, "rate_limit_exceeded")
}
if not retryable or attempt == 3:
raise RuntimeError(f"API error {error.code}: {code}") from error
try:
delay = int(error.headers.get("Retry-After", "15"))
except ValueError:
delay = 15
if not 0 <= delay <= 60:
raise RuntimeError("Retry delay exceeds this client's limit.")
time.sleep(delay)
continue
choice = result["choices"][0]
if choice["finish_reason"] != "stop":
raise RuntimeError("The model returned an incomplete answer.")
label = choice["message"]["content"].strip().lower()
if label not in LABELS:
raise RuntimeError("The answer needs manual review.")
return label
raise RuntimeError("Retry limit reached.")
if __name__ == "__main__":
if len(sys.argv) != 2:
raise SystemExit('Usage: python support_router.py "ticket text"')
print(classify(sys.argv[1]))Try a synthetic ticket:
python support_router.py "My invoice was charged twice."Use the dataset's exact system prompt if yours differs. The character limit is an application guard, not a token count. Prompts and answers still need to fit the published context length.
Tensorant accepts chat completions and model listing. It does not accept tool calling or response_format. Output validation belongs in your application. Do not let a routing label authorize refunds or account changes.
Step 7: Handle startup and failures deliberately
A sleeping model can hold the request for about a minute while it wakes. If still starting, the API returns 503 model_starting with Retry-After: 15. A model in Starting can return the same code immediately with a 30-second retry delay. The example allows time for startup and caps retries at four attempts.
Do not retry an invalid key, unknown model, malformed request, monthly quota error, or stopped model indefinitely. Put unclassified tickets in a review queue when the application cannot finish. Apply any downstream ticket update once, using the ticket's own identifier to prevent duplicate actions. The chat completions reference explains each error and streaming option.
Step 8: Release carefully and keep a rollback path
Keep a small acceptance set of synthetic and approved tickets. Check the published API with it before directing normal traffic to the model. Record the adapter run, serving settings, and application prompt together.
To update the adapter, open its project, choose Change on support-router, select the new adapter, and save. Tensorant keeps the name stable while the change takes effect. Retain the previous evaluated adapter so an owner can select it again if the new model fails your acceptance checks.
Review Usage for errors and token consumption and check RunPod spend. Compatible published models can share a server, so their traffic competes for its capacity. Stop published models on API, rather than stopping their GPU in Deployments. Tensorant can replace a published model's GPU when it is stopped elsewhere. See publishing models for switching costs, end dates, and shared servers.