Data
Vision models
Train and serve models that read images with text and answer in text.
A vision model reads images and text and answers in text. Use one when the questions are about pictures: charts, screenshots, scanned forms, receipts or product photos. The workflow is the same as for a text model, with images carried along: import examples with images, review, freeze a version, train, evaluate and serve. Your application then sends images through the API.
Only images are supported, not audio or video, and answers are always text.
Supported models
| Family | Architecture in config.json |
Example checkpoint |
|---|---|---|
| Qwen2.5-VL | Qwen2_5_VLForConditionalGeneration |
Qwen/Qwen2.5-VL-3B-Instruct |
| Qwen3-VL | Qwen3VLForConditionalGeneration |
Qwen/Qwen3-VL-4B-Instruct |
| Gemma 3, 4B and larger | Gemma3ForConditionalGeneration |
unsloth/gemma-3-4b-it |
| Gemma 4: E2B, E4B, 26B-A4B (mixture of experts) and 31B | Gemma4ForConditionalGeneration |
google/gemma-4-E4B-it, google/gemma-4-31B-it |
| Gemma 4 12B | Gemma4UnifiedForConditionalGeneration |
google/gemma-4-12B-it |
| Qwen3.5, Qwen3.6 and Qwen3.8 | Qwen3_5ForConditionalGeneration |
Qwen/Qwen3.8-27B |
| SmolVLM (version 1) | Idefics3ForConditionalGeneration |
HuggingFaceTB/SmolVLM-Instruct |
Like any base model, it must be public, ungated and have a chat template (see Projects). Its repository must also have a preprocessor_config.json and must not need custom code (no auto_map in its config files). The readiness checks before training or deploying verify this.
- Google's own Gemma 3 repositories are gated. Use an ungated copy such as
unsloth/gemma-3-4b-it. Gemma 3 1B reads text only. - SmolVLM2 is not supported. Use a SmolVLM version 1 model.
- Other vision models, such as Qwen3-VL mixture-of-experts models, Llama 3.2 Vision, LLaVA, Mistral Small 3.1 or Phi-3.5-vision, are not supported.
Before you start
- The Pro plan, for images: importing them, training and evaluating on them, and sending them to a model in the Playground or the API. On any plan, a vision model can be deployed, published and used with text. See Plans and billing.
- A project whose base model is one of the supported models.
- A RunPod network volume as your organization's storage, because images are kept there. RunPod bills the volume for the space. See Connections.
- An image dataset on Hugging Face. Uploaded files and pastes can't carry images.
From dataset to API
- Create a project with a vision base model.
- Import an image dataset from Hugging Face. See below.
- Review the examples. Images show as thumbnails; you edit the text around them. See Review.
- Freeze a version. The same image always stays in one split. See Versions.
- Run a baseline: deploy the base model and evaluate it. Test examples with images can only be sent to your own deployments. See Evaluations.
- Train an adapter with LoRA. See Training runs.
- Deploy and evaluate the adapter against the baseline.
- Try it in the Playground, publish it, and call it from your application.
An experiment runs steps 5 to 7 for you. Choose LoRA as its Training method.
Import images from Hugging Face
- In Sources, choose Add sources, then Hugging Face.
- Enter the dataset ID, choose Load dataset, then the configuration, split and starting row.
- Choose Preview rows. Image cells show as their size, such as "Image · 850 × 600".
- Under Import as, choose Prompt and answer examples or Chat / ShareGPT conversations, and set Image column (optional). In a vision project the suggested mapping fills it in for you.
- Set Maximum rows and Assign to split, and choose Import up to N rows.
The import runs in Activity like any Hugging Face import.
The image column must hold an image or a list of images, up to 4 per row. They are placed like this:
- Prompt and answer examples: the images first, then the question.
- Chat / ShareGPT conversations: the images fill the conversation's placeholders in order, either
{"type": "image"}parts or<image>markers in user messages. If the counts differ, the row is skipped. Without placeholders, the images go at the start of the first user message.
Only user messages can carry images. A row with an empty image cell becomes a text example.
How images are read
The import reads the original image files from Hugging Face's parquet export of the split, not the re-encoded copies in the preview. Each image is checked:
- PNG, JPEG, WebP and GIF are kept as they are. BMP and TIFF are converted to PNG. Other formats are refused.
- An image can be at most 40 megapixels and 20 MB.
A row with an image that fails a check is skipped, with the reason in the import report.
Where images are kept
Images are stored on your organization's RunPod network volume, in the images/ folder, with small thumbnails in thumbnails/. An image used in many examples, imports or projects is stored once.
Images pass through Tensorant when you import them, view them, or send them in the Playground or the API, but are not kept there. Training and evaluations read them straight from your volume, inside your RunPod account. Deleting an import does not delete its images, and Tensorant never deletes them: remove them in RunPod if you need to.
Images and the context length
Each image becomes tokens that count toward the context length, in training (Context length) and when serving (Serving context length). Approximate counts:
| Family | 1024 × 1024 image | 1920 × 1080 image |
|---|---|---|
| Qwen2.5-VL | about 1,371 tokens | about 2,693 tokens |
| Qwen3-VL | about 1,026 tokens | about 2,042 tokens |
| Gemma 3 | 259 tokens at any size | 259 tokens |
| SmolVLM | about 1,548 tokens | about 1,183 tokens |
With Qwen models, the default training Context length of 2,048 is often too short for large photos. Raise it, and the deployment's Serving context length, to fit your longest example with its images and answer.
What is not supported
| What | Notes |
|---|---|
| Images from uploads or pastes | Images come only from Hugging Face datasets. |
| Training a vision model on text only | The version must include examples with images. |
| QLoRA | Vision runs use LoRA. |
| Training the image part of the model | The adapter trains only the language part. |
| Images sent to external model endpoints | Images go only to deployments in your own RunPod account. |
| Images seen by an LLM judge | Judges see [image] in place of each image. See The LLM judge. |
| Images in recipes | Examples with images can't be recipe references, and generated examples are text only. |
| Changing an example's images | Review edits only the text. To drop an image, reject the example. |
| Images in system or assistant messages | Only user messages carry images. |
Limits
| Limit | Value |
|---|---|
| Images per example or API request | 4 |
| Images in one Playground conversation | 4, each up to 10 MB, 15 MB together |
| Stored image | 20 MB and 40 megapixels |
| Rows per Hugging Face import | No limit with your own connections; 1 to 5,000 on the platform's connections |
| New image data per import | 5 GiB (images already stored don't count) |
| API request to a vision model | 20 MiB |
| Large API requests (over 2 MiB) at once | 2 per organization |
| Image link in an API request | https, up to 2,048 characters; the image up to 20 MB, fetched within 10 seconds |
When something goes wrong
Choosing the model
| Message | What to do |
|---|---|
| Warning: "Architecture is not one of the supported model families…" | The model is not on the supported list, and it is treated as a text model. Choose a supported vision model to train with images. For SmolVLM2, use SmolVLM version 1. |
| "V1 supports public, ungated models…" | The model is gated, as Google's Gemma 3 is. Use an ungated copy. |
| "The model has no image processor (preprocessor_config.json)" or "The model needs its own code…" | Choose the family's official checkpoint or another copy that uses standard code. |
Importing
| Message | What to do |
|---|---|
| "Images and vision training are on the Pro plan…" | Ask an owner to upgrade under Settings → Plan and billing. |
| "This project's model reads text only…" | Clear Image column (optional), or switch the project to a vision model in Project settings. |
| "Image data needs RunPod network storage" | Your organization's storage isn't a RunPod network volume. Contact support. |
| "The image column must hold images" | Name a column Hugging Face shows as images. |
| "Images can be imported as prompt-and-answer or chat examples" | Choose Prompt and answer examples or Chat / ShareGPT conversations. |
| "Hugging Face has no parquet export of this split yet" | Try again later. |
| "The row has N image placeholders and M images" (skipped row) | Fix the placeholders in the dataset, or leave those rows out. |
| "The file is not an image Pillow can read" (skipped row) | The image is damaged or in an unsupported format. Convert it to PNG or JPEG. |
| "The image is larger than 40 megapixels" or "…than 20 MB" (skipped row) | Scale the image down in the dataset. |
| "The dataset changed during the import" | Load the dataset again and start a new import. Imported rows stay. |
| "Import exceeds the image byte limit…" | You reached 5 GiB of new images. Start a new import from where this one stopped. |
| "Error: an unexpected error stopped this…" | Often your network volume couldn't be written. Test storage, then choose Retry. |
Training, evaluating and serving
| Message | What to do |
|---|---|
| "Training a vision model on text only isn't supported yet" | Train on a version with image examples. |
| "QLoRA isn't available for vision models yet" | Choose LoRA. |
| "rows are longer than the sequence length of N tokens once images are counted" (in Training logs) | Raise Context length, or reject the longest examples and freeze again. |
| "Images are only sent to this organization's own deployments" | Evaluate on a deployment of this project's vision model. |
| "…exceed the serving context length of C once images are counted" | Deploy with a longer Serving context length, or lower Maximum output tokens. |
| "The model 'name' reads text only" | Send images to a published vision model. |