Fine-tune Qwen on RunPod with Tensorant
Build a support-ticket classifier with Qwen, review labeled examples, compare a LoRA adapter against the base model, and publish an API.
What you will build
Train, evaluate, and publish a Qwen support-ticket router.
Steps in this guide
Fine-tuning is useful when a general model understands your input but needs to follow a specific decision rule consistently. A support-ticket router is a good first project: it reads a customer message and returns one label that your application can validate.
In this guide, use Qwen/Qwen2.5-1.5B-Instruct to classify tickets as billing, account, technical, or other. Tensorant prepares a frozen dataset, evaluates the base model, trains a LoRA adapter on your RunPod GPU, and compares both models on the same held-out examples. Publish the adapter as support-router once its answers meet your requirements.
The official Qwen model card describes a 1.54-billion-parameter instruction model with a chat template. It is also on Tensorant's supported model list. Start with this small text model to make the complete workflow easier to inspect before moving to a larger model.
1. Define the routing contract
Decide label boundaries before collecting data. Otherwise, fine-tuning teaches inconsistent decisions.
| Label | Route here when the main issue concerns |
|---|---|
billing |
Charges, invoices, refunds, or payment methods |
account |
Login, passwords, account access, or profile settings |
technical |
A product failure, error, broken integration, or unexpected behavior |
other |
Sales, general questions, or a request without enough information to route |
Choose a rule for mixed tickets. For this example, use the issue preventing the customer's next action. A customer who cannot log in to download an invoice goes to account; a customer disputing an invoice goes to billing. Include these boundary cases in your examples.
Use the same system instruction during training and in your application:
Classify the support ticket as billing, account, technical, or other. Reply with exactly one label and no other text.A label is a routing decision, not permission to change an account or issue a refund. Keep downstream actions behind your application's normal authorization checks.
2. Prepare representative examples
Collect resolved tickets with labels verified by your support team. Remove secrets and customer identifiers before import. Preserve the actual problem: replacing every ticket with a generic sentence makes the task artificially easy.
Save UTF-8 JSON Lines, one conversation per line. These four example rows show the format for support-tickets.jsonl:
{"messages":[{"role":"system","content":"Classify the support ticket as billing, account, technical, or other. Reply with exactly one label and no other text."},{"role":"user","content":"My card was charged twice for the same monthly subscription."},{"role":"assistant","content":"billing"}],"category":"billing"}
{"messages":[{"role":"system","content":"Classify the support ticket as billing, account, technical, or other. Reply with exactly one label and no other text."},{"role":"user","content":"The password reset email never arrives, so I cannot sign in."},{"role":"assistant","content":"account"}],"category":"account"}
{"messages":[{"role":"system","content":"Classify the support ticket as billing, account, technical, or other. Reply with exactly one label and no other text."},{"role":"user","content":"The CSV export fails with an error whenever I select last month."},{"role":"assistant","content":"technical"}],"category":"technical"}
{"messages":[{"role":"system","content":"Classify the support ticket as billing, account, technical, or other. Reply with exactly one label and no other text."},{"role":"user","content":"Can someone help me with this?"},{"role":"assistant","content":"other"}],"category":"other"}Build a pilot dataset from, for example, 200 distinct tickets covering all labels, mixed issues, short messages, and unfamiliar wording. That is a planning target, not a minimum that guarantees quality. Four format examples alone are insufficient for a useful comparison.
Keep related conversations and paraphrases together when planning the split. Tensorant groups identical prompts, shared source documents, and shared images, but it cannot recognize every related ticket. Do not distribute near copies of one ticket across training and test. See evaluation data leakage.
3. Connect RunPod compute and storage
An organization owner completes this once:
- Create a RunPod API key with read and write access.
- Open Settings, then Connections. On Compute, choose Set up compute, paste the key, and choose Connect RunPod.
- Create a RunPod network volume in a datacenter with both the GPUs you need and S3 API support. Check RunPod's supported datacenters.
- Create a separate S3 API key. Under Data storage, choose Set up storage and enter the volume ID, datacenter, access key, and secret.
- Choose Connect storage.
Tensorant starts GPUs in your RunPod account and keeps project data on your network volume. See Connections for connection tests and billing visibility.
4. Create the project and import the dataset
- Open the project switcher and choose New project.
- Choose Classification as the use case and name the project
Support router. - Enter
Qwen/Qwen2.5-1.5B-Instructas the Hugging Face base model. - Set the goal to consistent ticket routing using the four labels, then choose Create project.
- Open Sources, choose Add sources, and upload
support-tickets.jsonl. - Choose Training and validation, then Import files. Standard
messagesrecords are detected as conversations.
Read the import report. Skipped rows commonly mean invalid JSON, incorrect role order, or message fields other than role and content. Fix those rows before continuing. You do not need a generation endpoint or a recipe for already labeled examples. See Sources.
5. Review and freeze a version
Open Review and inspect each ticket and its label. Set Answer type for quality checks to a classification label where appropriate. Correct disagreements, save changes with Save & recheck, and approve only examples you want the model to learn or be scored against.
Use Review similar prompts to find word overlap, and inspect related tickets yourself. Imported examples do not have generated source quotes; they can still be approved.
Open Versions and choose Create version. The first automatic split needs at least 10 independent groups and every version needs training, validation, and test examples. Roughly a tenth of groups go to validation and a tenth to test. To reserve a separate test file, import it as Held-out evaluation instead. See group splitting.
Run Check dataset at a training sequence length of 2,048. Fix overlong conversations before training, then freeze a new version if necessary. Aim for at least 10 independent test groups so the comparison can show an uncertainty interval.
6. Configure and launch a controlled experiment
Open Experiments, choose New experiment, and select the frozen version. Use these starting settings:
| Setting | Starting choice |
|---|---|
| Training method | LoRA |
| Training parameters | Keep the defaults: 2 epochs, rank 16, alpha 32, learning rate 0.0002 |
| Training sequence length | 2,048, if dataset readiness confirms it fits |
| Inference context length | 4,096, if every prompt plus answer fits |
| Task and primary metric | Classification and classification accuracy |
| Maximum answer tokens | 32 |
| Allowed labels | billing, account, technical, other, one per line |
Choose available training and inference GPUs with memory headroom. Set hourly price and runtime limits you are willing to approve. Readiness uses an estimate, so longer inputs or larger batches can still exhaust memory.
Set acceptance thresholds before seeing the result. For example, require 90% tuned accuracy, at least a 2-percentage-point improvement, and 10 independent test groups. These are example product requirements; choose yours based on misrouting costs. If requiring the entire 95% interval above zero, keep the minimum test-group requirement at 10 or more.
Choose Check readiness, resolve failures, then Save draft. Select Continue automatically through evaluation, training and cleanup after approval unless you want to inspect each stage. Choose Review and launch, read the maximum GPU list-price estimate, then Approve and start. See Experiments.
7. Inspect the comparison before publishing
When complete, choose Compare answers. Read regressed answers and inspect accuracy by category. A high overall score can hide poor other handling or confusion between account and billing issues.
Review the confusion matrix in the evaluation results. If technical tickets containing the word “invoice” go to billing, collect new independent examples that teach the boundary. Record the failure under Next dataset iteration, add training examples, and create a new version. Do not turn individual test answers into memorization targets.
If the adapter does not clearly improve routing, keep the base model as your candidate and fix data or prompting first. See evaluating a fine-tuned LLM.
8. Publish the adapter and call it
Once the full run shows Adapter ready or Completed and its GPU is released, an owner can open API and choose Publish a model. Name it support-router, select the adapter, and choose Wake on request for occasional traffic or Always on for steady traffic. Review the GPU and cost settings before publishing.
Create a server-side API key and copy the API address from the console. With TUNE_API_BASE_URL holding that address ending in /v1 and TUNE_API_KEY supplied by your secrets manager:
import os
from openai import OpenAI
client = OpenAI(
base_url=os.environ["TUNE_API_BASE_URL"],
api_key=os.environ["TUNE_API_KEY"],
)
answer = client.chat.completions.create(
model="support-router",
temperature=0,
max_tokens=32,
messages=[
{"role": "system", "content": "Classify the support ticket as billing, account, technical, or other. Reply with exactly one label and no other text."},
{"role": "user", "content": "My export stops at 90% and reports a timeout."},
],
)
choice = answer.choices[0]
label = (choice.message.content or "").strip().lower()
if choice.finish_reason != "stop" or label not in {"billing", "account", "technical", "other"}:
label = "other"
print(label)Install the client with pip install openai. Send the other fallback queue to human review, including incomplete or invalid responses. A wake-on-request model may return 503 model_starting during a cold start; follow Retry-After with a bounded retry policy. Use API quickstart for key creation and adapter deployment for serving checks. Validate labels in your application and monitor reviewed routing mistakes as traffic changes.