Practical guides to fine-tune, evaluate and deploy LLMs.
Work through a useful example, understand the tradeoffs and follow the steps in Tensorant. From your first dataset to a model your application can call.
Start here
Build a support-ticket router with Qwen.
Prepare labeled tickets, fine-tune an adapter, compare it with the base model and publish a working API. One guide takes you through the whole workflow.
Prepare
Labeled tickets and a frozen dataset
Train and evaluate
Qwen, LoRA and a fair baseline
Publish
support-router and a chat completions API
Fine-tuning
Fine-tune Qwen on RunPod with Tensorant
Build a support-ticket classifier with Qwen, review labeled examples, compare a LoRA adapter against the base model, and publish an API.
Getting started7 min readLoRA vs QLoRA: choose a fine-tuning method
Understand LoRA and QLoRA memory tradeoffs, choose practical settings, and compare adapters fairly in Tensorant before deployment.
Intermediate7 min readFine-tuning vs RAG: build a grounded support assistant
Choose retrieval, fine-tuning, or both for a support assistant, with source-conditioned examples, held-out evaluation, and Tensorant API steps.
Intermediate8 min readEstimate LLM fine-tuning cost before launching
Budget training, base-versus-tuned evaluation, storage, iteration, and inference with Tensorant's RunPod controls and transparent cost formulas.
Getting started7 min read
Evaluation
Find out what improved, what regressed and whether the comparison is fair.
How to evaluate a fine-tuned LLM against its base model
Build useful LLM evals in Tensorant, score a support-ticket classifier, compare base and tuned models, and investigate regressions before deployment.
Intermediate8 min readHow to prevent data leakage in LLM evaluation
Prepare held-out LLM evals in Tensorant, separate related sources and prompts, review near-duplicates, and preserve trustworthy tests as your dataset grows.
Intermediate9 min read
Deployment
Publish your model and call it from the application you already have.
Inference
Choose how your model runs, with clear tradeoffs in cost and response time.
Ready to use your own data?
Start a project, connect RunPod and follow a guide at your own pace.