Reference
Supported models
The text and vision model families you can train and serve, what a base model needs, and how much GPU memory it takes.
Every project is built on one base model from Hugging Face. You choose it when you create a project. This page lists the model families you can use and how much GPU they need.
What a base model needs
- Public and not gated on Hugging Face. Gated models such as
meta-llama/…and Gemma 3'sgoogle/gemma-3-…are refused; use an ungated copy of the same weights, such as those underunsloth/. Gemma 4'sgoogle/gemma-4-…repositories are not gated. - A chat model: its
tokenizer_config.jsonhas achat_template. Choose the instruct or chat version of a model. - Ideally a supported family: see Text models and Vision models. A model of any other family is allowed, with a warning that it may not train or serve. It is treated as a text model.
- No code of its own (
auto_mapin its configuration) that Transformers doesn't already have built in: such models are refused. Phi-3's repositories name their own code, but the built-in version is used, so they work. - A vision model also needs a
preprocessor_config.json.
The model is checked before an experiment, training run, deployment or published model starts, not when you create the project. The revision you choose (main by default) is fixed to one commit, so later changes on Hugging Face do not change your results.
Note
A supported family does not guarantee that every model of that family trains and serves on every GPU.
Text models
| Family | Architecture in config.json |
Example |
|---|---|---|
| Qwen2 and Qwen2.5 | Qwen2ForCausalLM |
Qwen/Qwen2.5-1.5B-Instruct |
| Qwen3 | Qwen3ForCausalLM |
Qwen/Qwen3-4B |
| Llama and models built on it | LlamaForCausalLM |
unsloth/Llama-3.2-1B-Instruct, HuggingFaceTB/SmolLM2-1.7B-Instruct |
| Mistral | MistralForCausalLM |
mistralai/Mistral-7B-Instruct-v0.3 |
| gpt-oss (mixture of experts) | GptOssForCausalLM |
unsloth/gpt-oss-20b-BF16 to train; openai/gpt-oss-20b to serve |
| Phi-3 | Phi3ForCausalLM |
microsoft/Phi-3-mini-4k-instruct |
| GPT-NeoX | GPTNeoXForCausalLM |
Few have a chat template; check yours does |
The examples met every rule on 5 October 2026. Models on Hugging Face can change.
gpt-oss: the official checkpoints store their expert weights in MXFP4. They serve, but the trainer can't train MXFP4 weights, so train a BF16 copy such as unsloth/gpt-oss-20b-BF16 (an adapter trained on it serves on the same copy). The readiness checks warn before training an MXFP4 checkpoint. gpt-oss expects its own "harmony" chat format, which its chat template applies. A BF16 copy needs about twice as much GPU memory as its parameter count in billions, in GiB, before training overhead: choose an 80 GB GPU or QLoRA for the 20B model.
Gemma 3 cannot be served in float16: leave Weight precision at Auto or choose bfloat16 in the advanced serving settings.
Vision models
A vision model reads images and text and answers in text. See Vision models.
| Family | Architecture in config.json |
Example |
|---|---|---|
| Qwen2.5-VL | Qwen2_5_VLForConditionalGeneration |
Qwen/Qwen2.5-VL-3B-Instruct |
| Qwen3-VL | Qwen3VLForConditionalGeneration |
Qwen/Qwen3-VL-4B-Instruct |
| Gemma 3, 4B and larger | Gemma3ForConditionalGeneration |
unsloth/gemma-3-4b-it |
| Gemma 4: E2B, E4B, 26B-A4B (mixture of experts) and 31B | Gemma4ForConditionalGeneration |
google/gemma-4-E4B-it, google/gemma-4-31B-it |
| Gemma 4 12B | Gemma4UnifiedForConditionalGeneration |
google/gemma-4-12B-it |
| Qwen3.5, Qwen3.6 and Qwen3.8 | Qwen3_5ForConditionalGeneration |
Qwen/Qwen3.8-27B |
| SmolVLM (version 1) | Idefics3ForConditionalGeneration |
HuggingFaceTB/SmolVLM-Instruct |
Vision models train with LoRA only. Each image uses tokens of the context length.
Gemma 4: use the instruction-tuned -it checkpoints; the base ones have no chat template. Each image takes up to 282 tokens. The 26B-A4B model trains adapters on its attention and shared layers, not on its experts. The 12B model is newer and less proven in serving than the other sizes.
Qwen3.8 27B (and Qwen3.5 and 3.6): most of its layers use linear attention, and adapters train on those layers' feed-forward part only. Its chat template turns on long reasoning by default. It needs a large GPU: choose QLoRA, or an 80 GB GPU for LoRA.
Not supported
| Model | Use instead |
|---|---|
| Gated or private models | An ungated copy, such as unsloth/Llama-3.2-1B-Instruct or unsloth/gemma-3-4b-it |
| Base models without a chat template | The model's instruct or chat version |
| SmolVLM2 | SmolVLM version 1 |
| Mistral Small 3.1 and 3.2 vision, LLaVA-OneVision, Gemma 3n, Phi-3.5-vision, Llama 3.2 Vision, Llama 4, mixture-of-experts vision models | A model from Vision models |
Models that run their own code (auto_map), such as Sarvam-30B, Sarvam-105B, InternVL chat, MiniCPM-V and Phi-4-multimodal |
A model from the tables above |
GPU memory
Before a launch, Tensorant estimates the GPU memory the model needs from its parameter count on Hugging Face, and refuses a GPU with less.
| Model size | Training, LoRA | Training, QLoRA | Serving | Serving with FP8 |
|---|---|---|---|---|
| 1.5 billion parameters | about 7.6 GiB | about 5.0 GiB | about 5.6 GiB | about 4.2 GiB |
| 4 billion parameters | about 13.7 GiB | about 6.6 GiB | about 11.7 GiB | about 8.0 GiB |
| 8 billion parameters | about 23.4 GiB | about 9.2 GiB | about 21.4 GiB | about 13.9 GiB |
For serving, a lower GPU memory share (%) than the default 90 leaves less of the GPU for the model.
The estimate leaves out context length, batch size and Conversations at once, so a GPU that passes can still run out of memory. Then lower those, choose QLoRA for training, or choose a larger GPU. See Will the model fit? for training and Will the model fit? for serving.
Other GPU rules
- Training needs an Ampere or newer NVIDIA GPU. On an older GPU, the run fails when training starts.
- GH200 and GB200 GPUs are not supported.
- Context length is at most 32,768 tokens, and no more than the model supports.
- FP8 quantization is for base-model deployments only. See FP8 and adapters.
- Training runs and deployments use RunPod's Secure Cloud GPUs in your network volume's datacenter. See GPUs.