Skip to content
Sources

Data

Sources

Add documents and examples to a project from files, pasted text or a Hugging Face dataset.

Sources are what a project learns from. They come in two kinds:

  • Documents, such as manuals, policies or articles. You turn them into examples by generating examples.
  • Examples you already have: question and answer pairs, or whole conversations. They go straight to Review as pending examples.

Each upload, paste or Hugging Face import is listed as one import on the Sources page (under Data in the sidebar). Original files are kept on your RunPod network volume.

Before you start

  • You are a member or an owner. Viewers can only look.
  • Your organization has its storage connected. See Connections.
  • Text files are UTF-8. Scanned PDFs need OCR first: only a PDF's text layer is read.
  • For a private or gated Hugging Face dataset, an owner has added a Hugging Face token that can read it.
  • For images, the project uses a vision model. Images can only come from a Hugging Face dataset.

Add sources

Choose Add sources. The dialog has three tabs.

Upload files

  1. On Upload files, choose up to 20 files.
  2. Under Use these sources for, choose Training and validation or Held-out evaluation. See Training or evaluation.
  3. For files with rows (CSV, JSON Lines, Excel and so on), open Column mapping and worksheet if the columns need mapping. See Map columns.
  4. Choose Import files.

Each file is imported whole or not at all. If a file fails, the dialog stays open with its import report: fix the file and import it again.

Paste content

  1. On Paste content, enter a Source name, choose the Content format and paste into Content.
  2. Choose what to Use these sources for, and the column mapping if needed.
  3. Choose Import pasted content.

The paste is read like an uploaded file of that format.

Import from Hugging Face

  1. On Hugging Face, paste a dataset URL or ID, such as https://huggingface.co/datasets/secemp9/arxiv-complete, and choose Load dataset.
  2. Check the recommended Configuration, Dataset split and Starting row (zero-based). Recommendations prefer usable text or conversations over file inventories; for arXiv, paper_text is recommended instead of files.
  3. Choose Preview rows. Up to five rows are shown, with a suggested mapping.
  4. Check the mapping. See Map columns.
  5. Optionally use Filter and sample below the mapping, then set Maximum rows and Assign to split. With no filters or sampling, All selects every row from the starting row on.
  6. Choose Import up to N rows.

The import runs in the background; follow it in Activity. Rows that arrive stay imported: Cancel keeps them, and Retry on a failed or cancelled import carries on where it stopped.

The preview pins the revision of its rows. Split metadata may have a different processed revision. If the rows or their Parquet export change before or during import, it stops rather than mixing revisions; reload the preview and start a new import.

Filter and sample

  • First matching rows checks rows from the starting row forward. Reproducible random sample draws candidates across the remaining split, using the Random seed. Keep the dataset revision, offset, seed and selection settings the same to reproduce a selection.
  • Add category or language filters and choose the column. Equals any of accepts comma-separated, exact, case-sensitive values; a list column matches if any of its items is listed. Every filter must match.
  • Characters: at least/at most measures the full text in the chosen column, including text shortened in the preview. These are character counts, not model tokens. Numeric bounds work on numeric columns.
  • Maximum source rows to check bounds a filtered selection to at most 50,000 candidates. Random samples without filters choose only the requested number of rows. Filtered samples can find fewer matches than requested; broaden the filters or increase the scan limit when needed.

For example, select paper_text for arXiv, map text as source documents, filter primary_category to cs.AI, cs.LG, choose a random sample with seed 42, and request up to 1,000 rows within a 10,000-row scan. The result may contain fewer than 1,000 papers if too few candidates match.

The five-row preview shows the original source schema and is unfiltered. Selection runs before any rows are imported. Activity shows candidates checked and matches found, then the import progress. Cancel/retry preserves the selection checkpoint. Original source row numbers remain in the provenance and report links.

Filtered and random selections require a complete Parquet conversion and support at most 50,000 selected rows, or 5,000 on shared connections. An ordinary import without filters or sampling retains the existing limits below. A scan limit bounds candidate rows; large rows and many remote shards can still make a scan take time.

The suggested mapping follows the first row: a messages or conversations list becomes conversations; instruction/output, instruction/response, prompt/completion or question/answer become prompt and answer examples; anything else becomes documents from the main text column.

Assign to split decides where the rows go in every version:

  • From a split named train (or any name not below), you choose Train, Validation or Test.
  • A split named validation, valid, val or dev always goes to Validation, and test, eval or evaluation always to Test, also with a suffix such as test_v2.

Rows assigned to Test are held-out evaluation data. Each row keeps its split from then on.

Note

If every example comes from a train split imported as Train, a version has no test examples. Import the dataset's test split too, or another range of rows with Assign to split set to Test.

If a preview row is marked "shortened in this preview", Hugging Face shortened a very long cell for the preview. The import reads the whole cell.

How many rows one import can bring in depends on your organization's connections. With its own connections, there is no limit: the rows go to your own storage. On the platform's connections, one import brings in up to 5,000 rows; start another import at a later row for more.

If Hugging Face converted only part of a very large dataset, rows past that part can't be imported, and the imported rows carry a warning.

Images. With Import as set to Prompt and answer examples or Chat / ShareGPT conversations, an Image column (optional) field appears. It works only in a vision model's project. See Vision models.

File formats

Format Extensions What it becomes
Plain text, Markdown .txt, .md, .markdown One document
HTML .html, .htm One document, visible text only (no scripts, navigation, footers or forms)
PDF .pdf One document, a section per page
Word .docx One document, body text and tables only (no images, headers or footnotes)
CSV, TSV .csv, .tsv One source per row; the first line holds the headings
Excel .xlsx One source per row of every worksheet, or the one named in Excel worksheet (optional); the first row holds the headings
JSON Lines .jsonl, .ndjson One source per line
JSON, YAML .json, .yaml, .yml One source per item of a list, or one for a single record
  • Text formats must be UTF-8.
  • PDFs, Word and Excel files must not be password-protected or encrypted. PDF reading order is best effort, so check tables and columns.
  • CSV, TSV and Excel need unique, nonempty headings. In Excel, replace formulas and error cells with values first.
  • YAML anchors and aliases are not supported.

Map columns

Import as decides how each row is read:

Import as A row becomes Columns you can name
Detect standard text or chat records (default) A conversation if it has messages, else an example if it has question and answer, else a document from text Category
Source documents A document Text, category
Prompt and answer examples One user message and one assistant answer Question, answer, context, system, category, and image for Hugging Face
Chat / ShareGPT conversations A conversation Messages, category, and image for Hugging Face
  • An empty column field uses the name shown in it. Detection only knows the exact names above, so for instruction and output, choose Prompt and answer examples and name them.
  • Nested paths work, such as answers.text.0: dots step into objects and numbers pick list items, from 0.
  • Context is added after the question, separated by a blank line. System becomes a system message.
  • Values must be text, numbers or true/false. For an object or list, name a path to the text inside.
  • A category column is read without mapping. Categories show in Review and in a version's coverage.

A conversation is a list of messages with role (system, user or assistant) and content. ShareGPT's from/value with human/gpt works too. It must have 2 to 100 messages, start with an optional system message, then alternate user and assistant, starting with user and ending with assistant. Messages with tool calls or other fields are skipped. A conversation ending in an unanswered user message is kept up to its last answer and counted as repaired.

Two examples as JSON Lines:

json
{"messages": [{"role": "user", "content": "How do I reset my password?"}, {"role": "assistant", "content": "Open Settings, choose Security, then choose Reset password."}], "category": "account"}
{"question": "How long do refunds take?", "answer": "Refunds reach your card within 5 business days.", "category": "billing"}

The same question and answer as CSV:

csv
question,answer,category
How long do refunds take?,Refunds reach your card within 5 business days.,billing

Rows already in the project are counted as duplicates and not added again, so importing the same file twice adds nothing.

Training or evaluation

Use these sources for sets the purpose of a whole file or paste:

  • Training and validation: examples are split when you freeze a version. Most go to training, some to validation and test.
  • Held-out evaluation: examples always go to test, never to training. Use it for a test set you wrote or collected separately. Examples generated from these sources are held out too.

The same text can't be imported for both purposes. For Hugging Face imports, Assign to split sets the purpose instead.

The import report

Each file, and each Hugging Face import in Activity, reports counts such as "Imported 172 · 14 duplicates · 12 repaired · 2 skipped":

Count Meaning
Imported New sources added
Duplicates Already in the project, not added again
Repaired Conversations whose unanswered last user message was dropped
Skipped Rows that could not be used

Below the counts, each reason is listed with up to 50 of its rows (row, sheet, line or item number). Rows of a Hugging Face import link to the dataset viewer. If no row of a file can be used, the whole file is refused with "No row can be used:" and the first reason.

The source list

The Sources page lists imports, newest first.

  • Dataset health shows how many sources are flagged for possible email or phone numbers, very short text (under 20 characters), empty or long content (over 20,000 characters). Flags are hints for review, not a guarantee.
  • Search names or locations finds sources by file name or location. Press / to jump to it.
  • Choose an import's name to list its sources, and Inspect to read one with its provenance.
  • Tick imports or single sources, then choose Generate examples. See Recipes.

Remove an import

On the import's row, choose the bin icon, then Delete import. This removes its sources and examples, including examples generated from it. Files already on your storage are not erased. You can only remove whole imports, not single sources.

Deleting is refused while a generation, import, quality check or freeze is running, when a version was frozen after the import's examples existed, or when one of its examples is a recipe reference or linked from an experiment. To keep such examples out of later versions, reject them in Review.

Tips

  • Hold out a real test set. Import a dataset's test split, or a set you wrote yourself as Held-out evaluation. Your scores are only as good as the test data.
  • Add categories. A category column lets evaluations break results down by kind of question, and shows gaps in a version's coverage.
  • Many smaller documents split better than one big one. All examples from one document stay in one split, so ten articles give a version more room than one long PDF. See Groups.

Limits

Limit Value
Files per upload 20
Size of a file or paste Shown in the upload dialog
PDF, Word and Excel files 25 MB each; 100 MB and 10,000 parts once unpacked
PDF pages 500
Reading a PDF, Word or Excel file 35 seconds
Text extracted from a file 25 MB
Rows per file 50,000
Excel worksheet 50,000 rows and 1,000 columns
Messages in a conversation 2 to 100, each up to 100,000 characters
Category 80 characters
Column name or path 200 characters
Paste source name 160 characters
Hugging Face rows per import No limit with your own connections; 1 to 5,000 on the platform's connections. 100 by default
Images See Vision models

When something goes wrong

Message What to do
"Exceeds the N MB file limit", "File exceeds 25 MB", "PDF imports support at most 500 pages", "Import at most 50,000 rows" or "Document extraction exceeded 35 seconds…" Split the file into smaller files.
"Text-based files must use UTF-8 encoding" Save the file as UTF-8.
"No extractable text. Scanned PDFs need OCR before import." Run OCR on the PDF, or upload its text.
"row N: wrong number of columns" Fix that line so it has a value for every heading.
"Worksheet name, row N: replace formulas/errors with values before import" Copy the cells and paste them back as values.
"No row can be used: …" The mapping doesn't fit the file. Check Import as and the column names. See Map columns.
"Some rows already exist in this project for a different purpose…" Those rows were already imported for training or for evaluation. Leave them out, or delete the import that holds them.
"Some rows are already assigned to split…" Those rows are reserved for another split. Import them with that split, or leave them out.
"Hugging Face access denied…" Add a token that can read the dataset in Connections, and accept the dataset's terms on Hugging Face with that account.
"No ready splits…" Hugging Face hasn't processed the dataset yet. Try again later.
"The dataset changed during the import" The dataset's files changed on Hugging Face. Choose Load dataset again and start a new import from where this one stopped.
"Organizations on the platform's connections import at most 5,000 rows at a time…" Import up to 5,000 rows, then start another import at a later starting row.
"Dataset viewer revision changed…" The dataset changed on Hugging Face. Choose Load dataset again and start a new import.
"Import exceeds the configured byte limit…" On the platform's connections, import fewer rows at a time.
"Wait for the running generation, import, quality check or freeze to finish" Wait until Activity shows nothing running, then delete the import.
"Version vN was frozen after examples of this import existed…" The import stays. Reject its examples in Review to leave them out of later versions.
"An example of this import is a reference of recipe name" Remove the reference from the recipe and save it, then delete the import.