Skip to content
Versions

Data

Versions

Freeze approved examples into a numbered version with training, validation and test splits.

A version is a frozen, numbered copy of your approved examples, split three ways:

  • Training examples teach the adapter.
  • Validation examples track its progress while it trains.
  • Test examples are held out. Evaluations score the base model and your tuned model on them.

A version never changes. Later edits and approvals only affect later versions, so scores on the same version always compare the same data.

Before you start

  • You have approved examples in Review, with their quality checks finished. Pending and rejected examples are left out.
  • You have enough independent groups: 10 for the first version, unless you also imported held-out evaluation data. See The 10-group rule.
  • You are a member or an owner. Anyone can download versions.

Create a version

  1. Open Versions under Data and choose Create version. (Review also offers Freeze version when nothing is pending.)
  2. Follow Freeze dataset version in Activity.
  3. The new version appears at the top as vN, with its example counts per split, and how many have images.

If the freeze is refused, the reason shows on the job in Activity. Fix it and choose Retry, or Create version again. To name a version, choose the pencil next to it.

How examples are split

Groups

Examples are linked into one group when they:

  • come from the same source document. Every example generated from one document is in its group.
  • share a user prompt (ignoring case and spacing). The same question about different images does not count.
  • share an image.

Links chain, and a whole group always goes to one split. So no test example shares a source, a prompt or an image with a training example.

Note

Groups count documents, not examples. Three hundred examples generated from four PDFs make at most four groups.

How groups are assigned

  1. Groups that already have a split keep it. See Split reservations.
  2. Groups from Held-out evaluation sources go to test.
  3. About a tenth of all groups go to test and a tenth to validation, at least one each. The rest go to training. The same data always splits the same way.
  4. If nothing ended up in validation (because every row was reserved for train or test at import), about a tenth of the training groups are moved to validation.

Split reservations

Once a group has a split, it keeps it in every later version, so adding data never moves a test question into training. Groups also get a split before any version exists when:

A freeze is refused if held-out evaluation data shares a source or prompt with training data, because the test score would then measure memory, not learning.

The 10-group rule

The first version needs at least 10 groups, unless the project has held-out evaluation data. After that, or with evaluation data, 3 groups are enough. Every version also needs examples in all three splits.

Download a version's splits

Under Export JSONL, choose train, validation or test. You get vN-split.jsonl, one conversation per line:

json
{"messages":[{"content":"How long do refunds take?","role":"user"},{"content":"Refunds reach your card within 5 business days.","role":"assistant"}]}

Categories, evidence and sources are not included. Images appear as {"type": "image", "image": …} parts naming the file on your network volume; the image files are not in the download. See Where images are kept.

Coverage

When you pick a version for an experiment, Dataset coverage shows each category's examples and independent groups per split, and warns about gaps, such as a category with no test examples or fewer than 10 independent test groups. To fill a gap, add examples of that category from new, independent sources and freeze a new version.

Dataset readiness

In Versions, select a frozen version or Current draft · before freezing, set the training sequence length, and choose Check dataset. The report is also available beside a training run's configuration before launch.

The report shows conversation token lengths (median, 95th percentile and maximum), empty assistant answers, category balance, independent source groups, and examples exceeding the smaller of your training sequence length and the model's declared context window. Review beside an issue opens its editable example. Fixing a frozen version requires creating a new version afterward.

Token counts use the project's actual public model tokenizer at the revision shown in the report, including its chat template without truncation. This check downloads tokenizer and configuration files only; it does not launch a GPU or call a paid model. Inaccessible tokenizers, conversations the template cannot render, and image token expansion are reported explicitly; missing counts are never replaced with estimates.

Analysis is bounded to 500 rows per frozen split or 1,500 draft rows, sampled reproducibly with seed 42. Content and tokenizer work also have byte and character limits. Every report shows inspected and total rows, and whether coverage is complete. Category and group counts in a sampled report describe the inspected examples. Unseen draft examples can connect observed groups, so freezing still validates the complete approved set. A draft report includes pending and approved examples; only approved examples enter a version.

Tips

  • More, smaller documents split better. Generate from one file per article or chapter rather than one large PDF. More groups give a bigger, more varied test split.
  • Aim for at least 10 test groups. With fewer, evaluations can't show uncertainty intervals. With automatic splitting that means about 100 groups, or import a held-out evaluation set from 10 or more independent sources.
  • Bring your own test set if you have one. Import it as Held-out evaluation in Sources; it always stays in test.

Limits

Limit Value
Approved examples in a version 50,000
Examples plus their source text 64 MiB
Groups for the first automatic split 10, unless there is held-out evaluation data
Groups for any version 3
Validation and test About a tenth of the groups each, at least one
Version name 160 characters

When something goes wrong

Message What to do
"Approve examples before creating a version" Approve examples in Review first.
"Finish quality checks before freezing a dataset" Wait for checks in Activity, or recheck examples whose checks failed.
"Automatic splits need 10 independent source groups…" Add more documents, or import a separate test set as Held-out evaluation.
"Need at least two independent training sources and one evaluation source" Approve examples from at least three independent sources.
"A version needs train, validation and test examples, and has none for…" Every example is reserved for other splits. Import test data, such as a dataset's test split, or approve examples from new sources.
"Evaluation data overlaps training sources or questions" Find the overlap with search or Review similar prompts in Review, and reject it or delete one of the imports.
"An edit links previously separate splits…" An edited example now matches examples in another split. Reject it, or undo the edit.
"Example id: An image of this example is not in this organization's storage" Reject that example.
"A version supports at most 50,000 approved examples…" or "…exceed the 64 MiB freeze limit…" Reject examples you don't need, especially from very large sources.