Skip to content
Evaluations

Evaluate

Evaluations

Score the base model and your adapters on a version's test examples, and compare them answer by answer.

An evaluation sends every test example of a frozen version to a model and scores each answer against the example's reference answer. Use evaluations to measure the base model, measure an adapter, and see whether training helped.

There are two kinds:

  • A baseline scores the base model. It fixes the test data, the scoring and the generation settings.
  • A fine-tuned evaluation scores an adapter from a training run. It reuses everything from the run's baseline, so only the model differs.

A comparison then pairs the two evaluations' answers example by example. An experiment does all three for you.

Before you start

Each test example is one call to the model, plus one to the judge if you use one. Your endpoint's provider bills those calls, or RunPod bills your deployment's GPU.

Where the model is served

Deployments

A ready deployment of this project is the simplest choice. Base-model deployments are offered for baselines, and adapter deployments for fine-tuned evaluations.

  • Tensorant checks the token budget before starting.
  • A training run against this baseline uses the same model revision, and its adapter is deployed with the same context length and serving settings, so the comparison stays fair.
  • The deployment does not stop for being idle while the evaluation runs.

Your own endpoints

Any endpoint added in Connections can be used. Tensorant cannot tell what model an endpoint serves: make sure it serves the base model for a baseline, and this run's adapter for a fine-tuned evaluation.

With your own endpoint, the token budget is not checked, each answer may take up to 3 minutes, and the endpoint's provider sees every prompt.

Run a baseline

  1. Open Evaluations under Tune and choose Run evaluation. Base model is selected.
  2. Choose the Dataset version and the Base model endpoint.
  3. Choose the Task preset and the Primary comparison metric. They start from the project's use case.
  4. Fill in what the preset needs: labels, a JSON schema, or tolerances.
  5. Optional: choose a Judge endpoint and edit the LLM judge rubric.
  6. Choose Run baseline evaluation.

For each test example, the model receives the conversation without its last answer. That last answer is the reference.

Settings

Field What it does
Task preset How answers are scored: Q&A, Classification, Structured JSON or Numeric. See Presets and metrics.
Primary comparison metric The metric shown first and used for comparisons. All the preset's metrics are recorded anyway.
Judge endpoint No LLM judge, or a model that scores answers against a rubric. See The LLM judge.
Maximum output tokens The longest answer the model may write. 512 by default. Raise it if answers get cut off.
System prompt Leave blank to keep each example's own system message. Fill it in to replace them all.

The temperature is always 0, so the model gives its most likely answer every time.

Presets and metrics

Each metric scores every answer from 0 to 1. Results show the average as a percentage.

Preset Use it when answers are… Metrics
Q&A Free text Normalized exact match, Token F1
Classification One label from a fixed list Classification accuracy, with per-label results
Structured JSON JSON that follows a schema JSON schema correctness, JSON validity, Selected field accuracy
Numeric A single number Numeric correctness

Every preset also records Normalized exact match and JSON validity. With a judge, every preset also records the LLM judge score.

What each metric means:

  • Normalized exact match: the answer is the same as the reference, ignoring case and spacing.
  • Token F1: how many words the answer shares with the reference. It gives partial credit for partly matching wording, but does not check whether the meaning is right.
  • Classification accuracy: the answer is the reference's label. An answer that is not one of the labels is wrong.
  • JSON validity: the answer is valid JSON.
  • JSON schema correctness: the answer is valid JSON that matches the schema.
  • Selected field accuracy: the share of the fields you listed whose value matches the reference. Missing fields are wrong.
  • Numeric correctness: the answer is a number within your tolerance of the reference. An answer with units or words is wrong.
  • LLM judge score: the judge's score. It is a model's opinion, not a fixed check, and is marked LLM judgment wherever it appears.

Reference answers are checked first

Before the evaluation starts, every reference answer is checked against the preset. A classification reference must be one of the labels, a numeric reference must be a number, and a JSON reference must match the schema. If some don't, the evaluation is refused with "Incompatible reference answers:" and a list of the examples (up to 20).

Classification

Allowed labels is filled in from the test examples when their answers look like labels. Edit the list as you need, one label per line. Results add a table per label and a confusion matrix: rows are the correct labels, columns are what the model answered.

Structured JSON

The JSON schema is filled in when all test examples came from recipes with the same schema. Otherwise, enter it yourself. See Structured JSON.

To score specific fields, list them under Scored JSON fields as paths, one per line, such as /answer or /items/0.

Numeric

An answer is correct when it is within the larger of Absolute tolerance and Relative tolerance × the reference. For example, with an absolute tolerance of 0.5 and a relative tolerance of 0.01 (1%), a reference of 200 accepts 198 to 202, and a reference of 10 accepts 9.5 to 10.5. Both start at 0: only the exact number is correct.

The LLM judge

A judge is a model endpoint that reads each answer and scores it from 0 to 1 against a rubric you write. Use one when the right answer can be worded many ways and word-matching metrics fall short.

  1. Choose a Judge endpoint from your model endpoints.
  2. Edit the LLM judge rubric. It starts with "Score factual correctness and instruction following against the reference, from 0 to 1." Say what a good answer looks like for your task.

For each answer, the judge sees your rubric, the prompt, the reference answer, the full text of the example's source, and the model's answer. It returns a score and a short reason, shown under each answer.

A fine-tuned evaluation uses its baseline's judge and rubric, so both sides are scored the same way.

Note

The judge's provider receives the whole source text of every test example. Choose a judge whose provider may see your data, and whose context is large enough for your sources.

The token budget

With a deployment, each test prompt plus Maximum output tokens must fit in the deployment's context length. If one doesn't, the evaluation is refused and names the first example that is too long.

To fix it, lower Maximum output tokens, shorten the system prompt, or deploy the model with a longer context length. Images count too: see Images and the context length.

Examples with images

In a project with a vision model, test examples can carry images.

  • Examples with images are sent only to a deployment of this project's vision model, never to an outside endpoint. The images stay in your RunPod account.
  • The judge reads text only. It sees [image] in place of each image, and the results say "The judge did not see the images".
  • Results show each prompt's images as thumbnails.

Follow an evaluation

The list shows each evaluation with its kind, model, version, Primary score and status: Queued, Running, Completed, Failed or Cancelled. The score updates as answers come in.

You can also follow it in Activity:

  • Cancel stops it. Answers already scored are kept.
  • Retry continues from the last saved answer.

Each example is tried up to three times. After the third failure, the evaluation fails and shows why.

Evaluate an adapter

  1. On a training run with Adapter ready, choose Evaluate adapter. Or choose Run evaluation and switch to Fine-tuned model.
  2. Choose the Training run.
  3. Keep Use ready managed adapter to use this run's deployment, or choose an endpoint that serves the adapter.
  4. Choose Run tuned evaluation.

The test data, scoring and settings come from the run's baseline. When it completes, the training run shows Completed, and Compare opens the comparison.

Results

Open a completed evaluation with View answers (a baseline) or Compare (a fine-tuned evaluation).

A baseline's answers

  • Tiles: every recorded metric, with the primary one marked.
  • Scores by category, recipe or source. See Group the results.
  • Answers: 25 per page, each with its score. Open one to see the model's answer next to the reference, every score, the judge's reason, and the answer's latency and tokens.
  • Inference telemetry: response times and token counts. These are not billing totals.

Start with the lowest-scoring answers. They show what the model gets wrong, and what data to add.

Compare two evaluations

Compare on a fine-tuned evaluation opens it side by side with its baseline. Both scored the same test examples, so each answer is paired with the other side's answer to the same example.

Read a comparison

  • The change: the adapter's score and how many percentage points it moved, with a 95% range. A change of +8 points means the adapter scored 8 points higher on average.
  • Scores by metric: every metric both sides recorded. Choose a metric's name to compare on it.
  • Change by category, recipe or source: where the gains and losses are. Treat small groups as hints.
  • Answers: filter by Improved, Unchanged or Regressed. Open one to see both answers and the reference.

An answer is improved when it scored higher with the adapter, regressed when it scored lower, and unchanged when the score is the same.

How to read it:

  • A clear improvement has a positive change, a 95% range entirely above zero, and few regressed answers.
  • A range that crosses zero means the test set can't tell whether the adapter is really better. Add more test examples, or more training data.
  • Always read some regressed answers. A good average can hide a new mistake that matters to you.

How the range is computed

The 95% range shows how much the change could move with a different sample of test examples. Tensorant re-samples the test examples 2,000 times, keeping examples from the same source or prompt together. It needs at least 10 independent groups. It does not cover randomness in the model or the judge.

Compare any two evaluations

You can compare any two completed evaluations of the same version, for example two adapters, or the same model with two system prompts.

  1. Open an evaluation.
  2. Under Compare with, choose the other one.

Only metrics both evaluations recorded can be compared.

Group the results

Break down by groups scores by Category, Recipe or Source. Choose a group to filter the answers. Groups with few examples give unreliable scores, so read them as hints.

The console has no download of evaluation results. You can download the test examples from Versions.

Limits

Limit Value
Maximum output tokens 16 to 8,192
System prompt 10,000 characters
LLM judge rubric 5,000 characters
Allowed labels 1 to 256 labels of up to 200 characters, unique ignoring case and spacing
Scored JSON fields 256
Temperature Always 0
Time for one answer from an endpoint 3 minutes
Attempts per test example 3
Independent groups for a range 10
Evaluation name 160 characters

When something goes wrong

Problem What to do
"No model is being served yet…" Deploy the base model in Deployments and wait until it is ready, or add an endpoint in Connections.
"Incompatible reference answers: …" Some reference answers don't fit the preset. Choose another preset, fix the labels, schema or fields, or reject those examples and freeze a new version.
"Reference is not one of the configured labels" Add the missing label to Allowed labels.
A test example exceeds the serving context length See The token budget.
"Images are only sent to this organization's own deployments" Choose a deployment of this project's vision model.
"Deploy this adapter and wait for Ready, or select an external endpoint" Deploy the adapter in Deployments, or choose an endpoint that serves it.
"Restore the frozen judge endpoint before running the tuned evaluation" Change the judge endpoint back to what the baseline used.
"Managed endpoint is not ready; deploy a model and wait for Ready" The deployment stopped. Start a new one and run a new evaluation.
"Endpoint name returned HTTP code" or "did not answer in time" Check the endpoint with its provider, then choose Retry in Activity.
An unexpected error stopped the evaluation Often the judge did not return a score from 0 to 1 with a reason. Check the judge's model and rubric, then choose Retry.
"Your role in this organization cannot make changes" You are a viewer. Ask an owner for the member role.