Skip to content
Supported models

Reference

Supported models

The text and vision model families you can train and serve, what a base model needs, and how much GPU memory it takes.

Every project is built on one base model from Hugging Face. You choose it when you create a project. This page lists the model families you can use and how much GPU they need.

What a base model needs

  • Public and not gated on Hugging Face. Gated models such as meta-llama/… and Gemma 3's google/gemma-3-… are refused; use an ungated copy of the same weights, such as those under unsloth/. Gemma 4's google/gemma-4-… repositories are not gated.
  • A chat model: its tokenizer_config.json has a chat_template. Choose the instruct or chat version of a model.
  • Ideally a supported family: see Text models and Vision models. A model of any other family is allowed, with a warning that it may not train or serve. It is treated as a text model.
  • No code of its own (auto_map in its configuration) that Transformers doesn't already have built in: such models are refused. Phi-3's repositories name their own code, but the built-in version is used, so they work.
  • A vision model also needs a preprocessor_config.json.

The model is checked before an experiment, training run, deployment or published model starts, not when you create the project. The revision you choose (main by default) is fixed to one commit, so later changes on Hugging Face do not change your results.

Note

A supported family does not guarantee that every model of that family trains and serves on every GPU.

Text models

Family Architecture in config.json Example
Qwen2 and Qwen2.5 Qwen2ForCausalLM Qwen/Qwen2.5-1.5B-Instruct
Qwen3 Qwen3ForCausalLM Qwen/Qwen3-4B
Llama and models built on it LlamaForCausalLM unsloth/Llama-3.2-1B-Instruct, HuggingFaceTB/SmolLM2-1.7B-Instruct
Mistral MistralForCausalLM mistralai/Mistral-7B-Instruct-v0.3
gpt-oss (mixture of experts) GptOssForCausalLM unsloth/gpt-oss-20b-BF16 to train; openai/gpt-oss-20b to serve
Phi-3 Phi3ForCausalLM microsoft/Phi-3-mini-4k-instruct
GPT-NeoX GPTNeoXForCausalLM Few have a chat template; check yours does

The examples met every rule on 5 October 2026. Models on Hugging Face can change.

gpt-oss: the official checkpoints store their expert weights in MXFP4. They serve, but the trainer can't train MXFP4 weights, so train a BF16 copy such as unsloth/gpt-oss-20b-BF16 (an adapter trained on it serves on the same copy). The readiness checks warn before training an MXFP4 checkpoint. gpt-oss expects its own "harmony" chat format, which its chat template applies. A BF16 copy needs about twice as much GPU memory as its parameter count in billions, in GiB, before training overhead: choose an 80 GB GPU or QLoRA for the 20B model.

Gemma 3 cannot be served in float16: leave Weight precision at Auto or choose bfloat16 in the advanced serving settings.

Vision models

A vision model reads images and text and answers in text. See Vision models.

Family Architecture in config.json Example
Qwen2.5-VL Qwen2_5_VLForConditionalGeneration Qwen/Qwen2.5-VL-3B-Instruct
Qwen3-VL Qwen3VLForConditionalGeneration Qwen/Qwen3-VL-4B-Instruct
Gemma 3, 4B and larger Gemma3ForConditionalGeneration unsloth/gemma-3-4b-it
Gemma 4: E2B, E4B, 26B-A4B (mixture of experts) and 31B Gemma4ForConditionalGeneration google/gemma-4-E4B-it, google/gemma-4-31B-it
Gemma 4 12B Gemma4UnifiedForConditionalGeneration google/gemma-4-12B-it
Qwen3.5, Qwen3.6 and Qwen3.8 Qwen3_5ForConditionalGeneration Qwen/Qwen3.8-27B
SmolVLM (version 1) Idefics3ForConditionalGeneration HuggingFaceTB/SmolVLM-Instruct

Vision models train with LoRA only. Each image uses tokens of the context length.

Gemma 4: use the instruction-tuned -it checkpoints; the base ones have no chat template. Each image takes up to 282 tokens. The 26B-A4B model trains adapters on its attention and shared layers, not on its experts. The 12B model is newer and less proven in serving than the other sizes.

Qwen3.8 27B (and Qwen3.5 and 3.6): most of its layers use linear attention, and adapters train on those layers' feed-forward part only. Its chat template turns on long reasoning by default. It needs a large GPU: choose QLoRA, or an 80 GB GPU for LoRA.

Not supported

Model Use instead
Gated or private models An ungated copy, such as unsloth/Llama-3.2-1B-Instruct or unsloth/gemma-3-4b-it
Base models without a chat template The model's instruct or chat version
SmolVLM2 SmolVLM version 1
Mistral Small 3.1 and 3.2 vision, LLaVA-OneVision, Gemma 3n, Phi-3.5-vision, Llama 3.2 Vision, Llama 4, mixture-of-experts vision models A model from Vision models
Models that run their own code (auto_map), such as Sarvam-30B, Sarvam-105B, InternVL chat, MiniCPM-V and Phi-4-multimodal A model from the tables above

GPU memory

Before a launch, Tensorant estimates the GPU memory the model needs from its parameter count on Hugging Face, and refuses a GPU with less.

Model size Training, LoRA Training, QLoRA Serving Serving with FP8
1.5 billion parameters about 7.6 GiB about 5.0 GiB about 5.6 GiB about 4.2 GiB
4 billion parameters about 13.7 GiB about 6.6 GiB about 11.7 GiB about 8.0 GiB
8 billion parameters about 23.4 GiB about 9.2 GiB about 21.4 GiB about 13.9 GiB

For serving, a lower GPU memory share (%) than the default 90 leaves less of the GPU for the model.

The estimate leaves out context length, batch size and Conversations at once, so a GPU that passes can still run out of memory. Then lower those, choose QLoRA for training, or choose a larger GPU. See Will the model fit? for training and Will the model fit? for serving.

Other GPU rules

  • Training needs an Ampere or newer NVIDIA GPU. On an older GPU, the run fails when training starts.
  • GH200 and GB200 GPUs are not supported.
  • Context length is at most 32,768 tokens, and no more than the model supports.
  • FP8 quantization is for base-model deployments only. See FP8 and adapters.
  • Training runs and deployments use RunPod's Secure Cloud GPUs in your network volume's datacenter. See GPUs.