How to evaluate a fine-tuned LLM against its base model
Build useful LLM evals in Tensorant, score a support-ticket classifier, compare base and tuned models, and investigate regressions before deployment.
What you will build
Compare a fine-tuned adapter with its base model on held-out examples and make a clear deployment decision.
Steps in this guide
A fine-tuned LLM should earn its place in your application by handling useful tasks better than the base model. Training loss helps diagnose training, but it does not tell you whether a support ticket reaches the right team or whether an extracted field is correct.
This guide builds an evaluation for a support-ticket router using Qwen/Qwen2.5-1.5B-Instruct. The model returns one of four labels: billing, account, technical, or other. You can use the same process for structured extraction, question answering, and numeric tasks by choosing the scoring preset that matches the output.
If you need to create your first adapter, start with the Qwen fine-tuning guide. You need a frozen dataset version with test examples, member or owner access, and a ready deployment or configured model endpoint.
Step 1: Define the answer your application needs
Write the output contract before scoring anything. For this router, the contract is a single label with no explanation or Markdown.
| Label | Route this kind of request |
|---|---|
billing |
Payments, invoices, subscriptions, and refunds |
account |
Login, password, and account access |
technical |
Broken product behavior and integration errors |
other |
Requests outside the three defined categories |
An example conversation looks like this:
{"messages":[{"role":"system","content":"Classify the support ticket as billing, account, technical, or other. Reply with exactly one label and no other text."},{"role":"user","content":"My card was charged twice for the same subscription renewal."},{"role":"assistant","content":"billing"}],"category":"billing"}The final assistant message is the reference answer. During evaluation, Tensorant sends the earlier messages to the model and scores its response against that reference.
Add a labeling rule for ambiguous cases. For example, route an explicit charge dispute to billing even if the ticket also mentions a broken screen. Use the same rule when preparing training labels and test references. Otherwise, disagreements may reflect an unclear policy rather than a model failure.
Step 2: Freeze a useful test set
Use real task variety. Include short requests, long tickets, misspellings, multiple concerns, and cases where a familiar word points to the wrong category. “The billing page crashes” tests product behavior, while “Please resend my invoice” tests billing.
- Import labeled conversations through Sources.
- Import a separately prepared test set with Use these sources for set to Held-out evaluation.
- Review references, resolve ambiguous labels, and approve the examples.
- Open Versions and choose Create version.
- Check category coverage when selecting the version in an experiment.
Tensorant keeps examples sharing a source, an identical normalized prompt, or an image in the same split. It also preserves split reservations across later versions. These checks reduce direct overlap, but paraphrases still need review. See preventing evaluation data leakage.
Aim for at least 10 independent test groups, with useful representation for every label. Ten is the minimum for a comparison interval, not evidence that the test set is large enough for every routing decision. Hundreds of variations from one document still form one source group. See how versions are split.
Step 3: Choose a metric that measures the task
Use Classification with Classification accuracy as the primary metric for the router. Enter the four labels under Allowed labels, one per line. Tensorant treats an answer outside that list as incorrect, including an explanation followed by a label.
For other tasks, use this mapping:
| Task | Useful starting metric | What to inspect alongside it |
|---|---|---|
| Ticket routing | Classification accuracy | Per-label scores and confusion matrix |
| JSON extraction | Selected field accuracy | JSON validity and schema correctness |
| Exact identifiers | Normalized exact match | Case and spacing rules |
| Numeric answers | Numeric correctness | Absolute and relative tolerances |
| Open-ended answers | A carefully defined judge rubric | Human review and factual evidence |
Token F1 measures shared words and punctuation. It cannot establish that an answer is factually correct. Similarly, valid JSON can contain the wrong values. Choose the primary metric for the decision you actually need to make. The evaluation metric reference explains the checks in detail.
Step 4: Run the base-model baseline
- Deploy the project's base model under Deployments and wait for Ready. Use a deployment without FP8 quantization if you will train against this baseline.
- Open Evaluations under Tune and choose Run evaluation.
- Keep Base model selected and choose the frozen version and base-model endpoint.
- Select Classification and Classification accuracy.
- Check the allowed labels and choose No LLM judge.
- Leave System prompt blank to retain each example's system message. If you replace it, use the instruction your application will send.
- Set Maximum output tokens high enough for a complete label. Start with 32 for this short-output task and inspect the returned answers.
- Choose Run baseline evaluation and wait for completion.
Evaluations use temperature 0. Each test prompt plus its output allowance must fit the deployment's context length. Set the deployment's maximum runtime long enough to finish the test set. Evaluation calls prevent idle shutdown, but they do not override that runtime limit. RunPod bills the deployment GPU while it runs.
Open a few answers before interpreting the aggregate score. Reasoning text, a truncated response, or an extra sentence can fail a strict classification check even when the intended label appears somewhere in the response.
Step 5: Evaluate the adapter under the same conditions
Train against the completed baseline, then deploy the verified adapter. On its training run, choose Evaluate adapter. Keep Use ready managed adapter or select an endpoint serving that exact adapter, then choose Run tuned evaluation.
Tensorant reuses the baseline's test data, scoring configuration, system prompt, and output limit. With managed deployments, the model revision, context length, and serving settings also stay aligned. If you use external endpoints, verify their model identity yourself because Tensorant cannot inspect which model they serve.
For a new run, Experiments can perform the baseline, training, adapter deployment, comparison, and cleanup automatically. Choose the same classification configuration, set acceptance thresholds, and use Check readiness before saving the draft. An experiment's evaluations have no LLM judge. Use the manual evaluation workflow when a rubric judge is needed.
Step 6: Read the paired comparison and its uncertainty
Choose Compare when the tuned evaluation completes. Each tuned answer is paired with the baseline's answer to the same test example. This is more informative than comparing two scores from unrelated test sets.
Read these together:
- The tuned model's primary score.
- The improvement in percentage points.
- The 95% range around that improvement.
- The number of independent groups behind the comparison.
A positive improvement with a range entirely above zero is evidence of a gain on this test distribution. A range crossing zero means the available sample does not clearly distinguish the models. A missing range is not proof of stability. Tensorant cannot provide one with fewer than 10 independent groups, or when the grouped changes have no usable variation.
The range uses 2,000 paired resamples, keeping source, prompt, and image groups together. It reflects test-sample uncertainty. It does not measure training randomness, judge bias, or the effect of a changed production workload. See comparison statistics.
Step 7: Inspect regressions before accepting an average gain
- Select Regressed in the comparison's answer filter.
- Open failures and read the ticket, reference, base answer, and tuned answer.
- Use Change by category, recipe, and source to find recurring problems.
- For classification, inspect the confusion matrix. Its rows are reference labels and its columns are predicted labels.
- Record critical error patterns and decide whether they block release.
An improved overall score can hide worse account routing. A large billing category can dominate the average while a small but important account category regresses. Small groups are clues, so collect more independent examples before drawing a strong category-level conclusion.
Choose experiment thresholds before training. For example, a proposed acceptance rule could require at least 90% classification accuracy, a positive improvement, at least 10 independent groups, and the entire interval above zero. These are starting requirements to adapt to your product, not universal quality standards. Add a human release check for critical categories because the overall threshold alone does not enforce a category floor.
Step 8: Extend the evaluation when the output becomes JSON
If the router also returns an escalation flag, change the task to Structured JSON and train references in the same format:
{"category":"billing","needs_human":true}Use this response schema:
{
"type": "object",
"properties": {
"category": {"type": "string", "enum": ["billing", "account", "technical", "other"]},
"needs_human": {"type": "boolean"}
},
"required": ["category", "needs_human"],
"additionalProperties": false
}List /category and /needs_human under Scored JSON fields. Select Selected field accuracy to match field values against the reference, and inspect schema correctness alongside it. Missing fields are wrong, and list order matters when scoring arrays. The schema's required list ensures both keys exist, while additionalProperties: false disallows extra keys. See the JSON Schema object reference.
Define exactly when escalation is required before labeling examples. For open-ended explanations, a manual rubric judge can supplement deterministic checks. Calibrate its scores against human-reviewed answers, and remember that its provider receives each example's full source text. A judge in a vision evaluation sees text and image placeholders, so it cannot verify image details.
Step 9: Turn the result into a release decision
Keep a short record of the frozen version, adapter, model revision, evaluation settings, primary score, paired change, interval, and critical regressions. Tensorant shows inference latency and token counts, which help identify slower or unnecessarily verbose answers. These are telemetry, not billing totals.
If the adapter meets the agreed criteria, continue with deploying a LoRA adapter. If it fails, add new training examples for the observed failure patterns and preserve the original test set. Use fresh held-out cases to confirm the next improvement. The console does not currently offer an evaluation-results download, so capture your decision record separately and use Versions to export the test conversations when needed.