Skip to content
Vision models

Data

Vision models

Train and serve models that read images with text and answer in text.

A vision model reads images and text and answers in text. Use one when the questions are about pictures: charts, screenshots, scanned forms, receipts or product photos. The workflow is the same as for a text model, with images carried along: import examples with images, review, freeze a version, train, evaluate and serve. Your application then sends images through the API.

Only images are supported, not audio or video, and answers are always text.

Supported models

Family Architecture in config.json Example checkpoint
Qwen2.5-VL Qwen2_5_VLForConditionalGeneration Qwen/Qwen2.5-VL-3B-Instruct
Qwen3-VL Qwen3VLForConditionalGeneration Qwen/Qwen3-VL-4B-Instruct
Gemma 3, 4B and larger Gemma3ForConditionalGeneration unsloth/gemma-3-4b-it
Gemma 4: E2B, E4B, 26B-A4B (mixture of experts) and 31B Gemma4ForConditionalGeneration google/gemma-4-E4B-it, google/gemma-4-31B-it
Gemma 4 12B Gemma4UnifiedForConditionalGeneration google/gemma-4-12B-it
Qwen3.5, Qwen3.6 and Qwen3.8 Qwen3_5ForConditionalGeneration Qwen/Qwen3.8-27B
SmolVLM (version 1) Idefics3ForConditionalGeneration HuggingFaceTB/SmolVLM-Instruct

Like any base model, it must be public, ungated and have a chat template (see Projects). Its repository must also have a preprocessor_config.json and must not need custom code (no auto_map in its config files). The readiness checks before training or deploying verify this.

  • Google's own Gemma 3 repositories are gated. Use an ungated copy such as unsloth/gemma-3-4b-it. Gemma 3 1B reads text only.
  • SmolVLM2 is not supported. Use a SmolVLM version 1 model.
  • Other vision models, such as Qwen3-VL mixture-of-experts models, Llama 3.2 Vision, LLaVA, Mistral Small 3.1 or Phi-3.5-vision, are not supported.

Before you start

  • The Pro plan, for images: importing them, training and evaluating on them, and sending them to a model in the Playground or the API. On any plan, a vision model can be deployed, published and used with text. See Plans and billing.
  • A project whose base model is one of the supported models.
  • A RunPod network volume as your organization's storage, because images are kept there. RunPod bills the volume for the space. See Connections.
  • An image dataset on Hugging Face. Uploaded files and pastes can't carry images.

From dataset to API

  1. Create a project with a vision base model.
  2. Import an image dataset from Hugging Face. See below.
  3. Review the examples. Images show as thumbnails; you edit the text around them. See Review.
  4. Freeze a version. The same image always stays in one split. See Versions.
  5. Run a baseline: deploy the base model and evaluate it. Test examples with images can only be sent to your own deployments. See Evaluations.
  6. Train an adapter with LoRA. See Training runs.
  7. Deploy and evaluate the adapter against the baseline.
  8. Try it in the Playground, publish it, and call it from your application.

An experiment runs steps 5 to 7 for you. Choose LoRA as its Training method.

Import images from Hugging Face

  1. In Sources, choose Add sources, then Hugging Face.
  2. Enter the dataset ID, choose Load dataset, then the configuration, split and starting row.
  3. Choose Preview rows. Image cells show as their size, such as "Image · 850 × 600".
  4. Under Import as, choose Prompt and answer examples or Chat / ShareGPT conversations, and set Image column (optional). In a vision project the suggested mapping fills it in for you.
  5. Set Maximum rows and Assign to split, and choose Import up to N rows.

The import runs in Activity like any Hugging Face import.

The image column must hold an image or a list of images, up to 4 per row. They are placed like this:

  • Prompt and answer examples: the images first, then the question.
  • Chat / ShareGPT conversations: the images fill the conversation's placeholders in order, either {"type": "image"} parts or <image> markers in user messages. If the counts differ, the row is skipped. Without placeholders, the images go at the start of the first user message.

Only user messages can carry images. A row with an empty image cell becomes a text example.

How images are read

The import reads the original image files from Hugging Face's parquet export of the split, not the re-encoded copies in the preview. Each image is checked:

  • PNG, JPEG, WebP and GIF are kept as they are. BMP and TIFF are converted to PNG. Other formats are refused.
  • An image can be at most 40 megapixels and 20 MB.

A row with an image that fails a check is skipped, with the reason in the import report.

Where images are kept

Images are stored on your organization's RunPod network volume, in the images/ folder, with small thumbnails in thumbnails/. An image used in many examples, imports or projects is stored once.

Images pass through Tensorant when you import them, view them, or send them in the Playground or the API, but are not kept there. Training and evaluations read them straight from your volume, inside your RunPod account. Deleting an import does not delete its images, and Tensorant never deletes them: remove them in RunPod if you need to.

Images and the context length

Each image becomes tokens that count toward the context length, in training (Context length) and when serving (Serving context length). Approximate counts:

Family 1024 × 1024 image 1920 × 1080 image
Qwen2.5-VL about 1,371 tokens about 2,693 tokens
Qwen3-VL about 1,026 tokens about 2,042 tokens
Gemma 3 259 tokens at any size 259 tokens
SmolVLM about 1,548 tokens about 1,183 tokens

With Qwen models, the default training Context length of 2,048 is often too short for large photos. Raise it, and the deployment's Serving context length, to fit your longest example with its images and answer.

What is not supported

What Notes
Images from uploads or pastes Images come only from Hugging Face datasets.
Training a vision model on text only The version must include examples with images.
QLoRA Vision runs use LoRA.
Training the image part of the model The adapter trains only the language part.
Images sent to external model endpoints Images go only to deployments in your own RunPod account.
Images seen by an LLM judge Judges see [image] in place of each image. See The LLM judge.
Images in recipes Examples with images can't be recipe references, and generated examples are text only.
Changing an example's images Review edits only the text. To drop an image, reject the example.
Images in system or assistant messages Only user messages carry images.

Limits

Limit Value
Images per example or API request 4
Images in one Playground conversation 4, each up to 10 MB, 15 MB together
Stored image 20 MB and 40 megapixels
Rows per Hugging Face import No limit with your own connections; 1 to 5,000 on the platform's connections
New image data per import 5 GiB (images already stored don't count)
API request to a vision model 20 MiB
Large API requests (over 2 MiB) at once 2 per organization
Image link in an API request https, up to 2,048 characters; the image up to 20 MB, fetched within 10 seconds

When something goes wrong

Choosing the model

Message What to do
Warning: "Architecture is not one of the supported model families…" The model is not on the supported list, and it is treated as a text model. Choose a supported vision model to train with images. For SmolVLM2, use SmolVLM version 1.
"V1 supports public, ungated models…" The model is gated, as Google's Gemma 3 is. Use an ungated copy.
"The model has no image processor (preprocessor_config.json)" or "The model needs its own code…" Choose the family's official checkpoint or another copy that uses standard code.

Importing

Message What to do
"Images and vision training are on the Pro plan…" Ask an owner to upgrade under Settings → Plan and billing.
"This project's model reads text only…" Clear Image column (optional), or switch the project to a vision model in Project settings.
"Image data needs RunPod network storage" Your organization's storage isn't a RunPod network volume. Contact support.
"The image column must hold images" Name a column Hugging Face shows as images.
"Images can be imported as prompt-and-answer or chat examples" Choose Prompt and answer examples or Chat / ShareGPT conversations.
"Hugging Face has no parquet export of this split yet" Try again later.
"The row has N image placeholders and M images" (skipped row) Fix the placeholders in the dataset, or leave those rows out.
"The file is not an image Pillow can read" (skipped row) The image is damaged or in an unsupported format. Convert it to PNG or JPEG.
"The image is larger than 40 megapixels" or "…than 20 MB" (skipped row) Scale the image down in the dataset.
"The dataset changed during the import" Load the dataset again and start a new import. Imported rows stay.
"Import exceeds the image byte limit…" You reached 5 GiB of new images. Start a new import from where this one stopped.
"Error: an unexpected error stopped this…" Often your network volume couldn't be written. Test storage, then choose Retry.

Training, evaluating and serving

Message What to do
"Training a vision model on text only isn't supported yet" Train on a version with image examples.
"QLoRA isn't available for vision models yet" Choose LoRA.
"rows are longer than the sequence length of N tokens once images are counted" (in Training logs) Raise Context length, or reject the longest examples and freeze again.
"Images are only sent to this organization's own deployments" Evaluate on a deployment of this project's vision model.
"…exceed the serving context length of C once images are counted" Deploy with a longer Serving context length, or lower Maximum output tokens.
"The model 'name' reads text only" Send images to a published vision model.