Skip to content
Evaluation

How to prevent data leakage in LLM evaluation

Prepare held-out LLM evals in Tensorant, separate related sources and prompts, review near-duplicates, and preserve trustworthy tests as your dataset grows.

By Tensorant9 min readIntermediateUpdated

What you will build

Build a held-out evaluation dataset that tests new cases instead of memorized training examples.

Steps in this guide

An LLM evaluation can look excellent while measuring familiarity with training examples. If a ticket, its paraphrase, or its answer key appears on both sides of the split, the score becomes less useful for predicting how the model handles a new request.

Data leakage includes information used during model development that would not be available when predicting new cases. It can produce overly optimistic scores. The scikit-learn guidance on leakage explains the general principle: separate test data before fitting transformations or making model choices with it.

This guide applies that principle to Tensorant's sources, recipes, frozen versions, and evaluations. The example is a support router that returns billing, account, technical, or other. The same preparation matters for extraction, document Q&A, and image-based tasks.

Step 1: Decide what counts as an independent case

Start with the original material, before generating examples or shuffling rows. Ask which records could reveal the answer to another record.

Material Keep this material together
Support tickets The original ticket, follow-ups, copies, and paraphrases
Product incidents Tickets with the same incident-specific clues when testing new incidents
Document Q&A Questions derived from the same source document
Extraction Copies and reformatted versions of the same original record
Vision tasks Reused images and derived image variants

For a router, decide whether the test represents new tickets, new customers, or new incident types. These are different questions. A random ticket split may be reasonable for familiar recurring issues, but it cannot by itself establish performance on an entirely new product incident.

Tensorant links examples when they share a source, an identical prompt after case and spacing normalization, or an image. Links chain: if one example shares a source with another and that example shares a prompt with a third, they belong to one group. A whole group goes to one split.

Semantic relationships such as a shared customer or incident are not automatically inferred. Separate those groups before importing if they define the generalization you want to measure. See version grouping.

Step 2: Reserve original held-out cases before writing training examples

Choose test tickets that were not used to define the prompt, create recipe references, or generate training variations. Keep original records and label them using a written routing policy.

For example, a held-out request might be:

json
{"messages":[{"role":"system","content":"Classify the support ticket as billing, account, technical, or other. Reply with exactly one label and no other text."},{"role":"user","content":"The password reset link says it has expired even when I open the latest email immediately."},{"role":"assistant","content":"account"}],"category":"account"}
  1. Put independently reserved test cases in a separate JSONL file.
  2. Open Sources, choose Add sources, and select Upload files.
  3. Under Use these sources for, select Held-out evaluation.
  4. Import the training collection separately as Training and validation.
  5. In Review, check reference labels and approve both collections.

The purpose applies to the entire file or paste. Held-out examples always go to test. Training and validation sources participate in automatic splitting, which can also place some groups in test. Hugging Face imports use Assign to split instead; existing validation and test split names reserve the corresponding splits. See training or evaluation sources.

A reference answer belongs in the final assistant message. Do not put Correct category: account inside the user prompt. Evaluation removes the final assistant answer, but it still sends every earlier message to the model.

Step 3: Clean source data without adding future information

Use the input your application will actually have. If a support router acts on the initial ticket, training and testing it with the agent's eventual resolution gives it clues that will be absent at routing time.

  1. Remove resolved-team labels, answer keys, internal routing notes, and final outcomes from input messages.
  2. Keep useful facts available at request time, such as the customer's description and the product area.
  3. Apply the same input formatting to training, validation, test, and production requests.
  4. Remove personal data according to your organization's policy before import.
  5. Preserve a separate record of original ticket IDs and grouping decisions outside the model's input.

Tensorant's draft redaction can detect possible emails and phone numbers. It does not remove names or addresses, and it does not change source files or already frozen versions. For generated examples that include their source passage, clean the source before importing and regenerating. Editing the embedded passage can break its schema check. See redaction behavior.

For a time-based test, reserve newer tickets before import and keep related older copies out of training. Tensorant's automatic grouping is not a chronological split. Prepare the date boundary yourself when the purpose is to test future traffic or a policy change.

Step 4: Keep recipe references out of the test set

Recipe references teach a generation endpoint the style and format of examples. They are part of the development process, so choose them from training data.

  1. Write a clear generation instruction and routing policy.
  2. Under Approved references, select only training examples.
  3. Generate training examples from the training collection.
  4. Prepare test cases independently instead of rewriting those references.
  5. Review generated labels against the source and policy before approval.

When you save a recipe, Tensorant reserves its reference examples and related source or prompt groups for training. Held-out evaluation examples cannot be recipe references. This prevents direct reuse, but it cannot make a paraphrase independent of the original incident. See recipe reference reservations.

Choose Training context to match the application. Use Include source passage when production will send the passage with the question. Use Closed book only when the question stands on its own. Evaluating with a helpful passage and deploying without it measures a different task. A document Q&A system with retrieval should also test the actual retrieval path, because a supplied reference passage does not establish that the retriever will find it.

Step 5: Review near-duplicates before freezing

Exact duplicates are only part of the problem. “I was charged twice” and “My renewal produced two card payments” can describe the same original ticket without sharing many words.

  1. Open an example in Review and save any edits.
  2. In PII and similar prompts, choose Review similar prompts.
  3. Inspect each candidate's prompt, source, status, and split.
  4. Choose Next candidates to continue through the collection.
  5. Reject redundant or overlapping examples before creating the version.

The tool checks 200 candidates at a time and reports prompts sharing at least 70% of their words. It compares words rather than meaning, and it changes nothing automatically. An embedded source passage can dominate the overlap score, so review the original records as well. See similar-prompt review.

Use your pre-import ticket or incident record to catch paraphrases the word scan misses. Avoid solving overlap by making a cosmetic prompt edit. That changes the string, but the answer may still be learned from the same underlying case.

Step 6: Freeze a version and resolve split conflicts

Open Versions and choose Create version once quality checks are complete. A frozen version preserves its examples and splits. Later edits affect later versions, so a comparison on the original version continues to use the original data.

If freezing reports Evaluation data overlaps training sources or questions, locate the overlap in Review or your source records. Reject the conflicting example or remove the inappropriate import, then freeze again.

If an edit links previously separate splits, undo it or reject the bridging example. Existing groups retain their reservations across versions, which prevents new data from silently moving an old test question into training. One documented exception can move training groups to validation when validation is missing. It never moves them to test.

Check counts at the group level. The first automatic split needs at least 10 independent groups unless approved held-out evaluation data exists. Later versions, or versions with held-out data, need at least three groups and examples in all three splits. These are minimum dataset requirements. Reliable category-level conclusions usually need more independent test coverage.

Step 7: Check coverage and dataset readiness

Run Check dataset in Versions with the training sequence length you plan to use. Inspect token lengths, empty answers, category balance, and independent groups. Read the inspected-row counts because large datasets may be sampled.

In an experiment's Dataset coverage, confirm each important category has training, validation, and test examples. A billing-heavy set can produce a strong average while leaving account almost untested.

Aim for at least 10 independent test groups to support a paired comparison interval. With automatic splitting, that generally requires roughly 100 total groups. An explicitly held-out set can provide those test groups directly. More examples from one PDF increase row count without increasing source independence. See dataset coverage.

Step 8: Preserve a final check as you iterate

Run the base and tuned models on the same frozen version with the same scoring and generation settings. Use Classification accuracy for the router and inspect regressions by category. Follow the step-by-step LLM evaluation guide.

Repeatedly adapting training data or prompts to one test set turns it into development feedback. Keep its failures for diagnosis, but do not treat the next score on those familiar cases as an independent confirmation.

Before release, collect a fresh acceptance set from new original cases. An existing adapter's Evaluate adapter action uses its training run's original version and baseline. It does not let you select a different version. Test that adapter on fresh cases through the published API with a separate acceptance script:

  1. Prepare a new labeled JSONL collection and keep its original case and group IDs in a separate record. Import it as Held-out evaluation if you want Tensorant to preserve those cases for future versions. Keep the original JSONL file for the acceptance script.
  2. Ask an owner to publish the base model as support-router-base and the verified adapter as support-router. Use the same base-model revision, context length, GPU, and serving settings. Check that your plan has room for both published names. Follow publishing models and the adapter deployment guide.
  3. Create an API key allowed to call both models. Use the API address shown in the console, and keep the key in an environment variable. See the API quickstart.
  4. For each case, remove the final assistant reference from messages. Send the remaining conversation to each model through POST /v1/chat/completions, with the same temperature: 0, max_tokens: 32, and stream: false settings.
  5. Save both returned answers beside the reference and the case ID. For this classification task, compare the complete answer to the reference after normalizing case and spacing. Extra explanations and unknown labels are incorrect.
  6. Handle cold-start and rate-limit responses using Retry-After. Report unresolved request failures as an incomplete acceptance test rather than silently dropping those cases. See API retries.
  7. Calculate each model's accuracy and inspect paired improvements and regressions by category. Record the fresh collection's scope and independent groups with the decision. API requests do not create a console evaluation or its 95% comparison interval automatically.
  8. Stop temporary published models from the API screen when testing is complete. A server remains running while another published model still uses it.

Tensorant's 95% comparison range resamples paired groups. It does not detect hidden incident relationships, eliminate judge bias, or repair a contaminated test set. Review category regressions and read individual answers even when the overall interval supports improvement.

Step 9: Carry the same rules into image and production evaluations

For vision tasks, reusing an image links examples, but a crop, screenshot, or edited duplicate can still need manual grouping before import. Use a deployment of the project's vision model for image evaluation. A text-only judge sees image placeholders and cannot assess visual correctness.

After deployment, collect new failures under the same routing policy and decide which become training examples and which remain future tests. Keep those collections distinct before importing. A useful evaluation record identifies the source collection, frozen version, labeling policy, model revision, primary metric, paired change, uncertainty, and important regressions. This lets the next improvement build on a stable decision process instead of a steadily more familiar benchmark.