DPO preference training
Train with preferred and less preferred answers, using text or image prompts with LoRA or QLoRA.
Direct Preference Optimization (DPO) teaches an adapter to prefer one answer over another for the same prompt. Training objective chooses SFT or DPO; Training method separately chooses LoRA or QLoRA. SFT remains the default.
DPO works with text models and supported vision models. Vision DPO accepts text-only pairs, image-conditioned pairs, or a mixture. It starts from the project's pinned base model, not an existing Tensorant adapter.
Prepare preference pairs
Import JSONL, JSON, CSV or a Hugging Face dataset with prompt, chosen and rejected columns. For other column names, select Preference pairs for DPO under Import as and map the three columns. See Sources.
{"prompt":"What is the refund window?","chosen":"You can request a refund within thirty days.","rejected":"Refunds are never available."}Strings become chat messages. For conversation history, use message arrays:
{"prompt":[{"role":"system","content":"Answer using the support policy."},{"role":"user","content":"What is the refund window?"}],"chosen":[{"role":"assistant","content":"You can request a refund within thirty days."}],"rejected":[{"role":"assistant","content":"Refunds are never available."}]}The shared prompt starts with an optional system message, then alternates user and assistant turns and ends with a user turn. It can contain up to 99 messages. Each candidate must be one nonempty assistant text response, and the candidates must differ. Message text is limited to 100,000 characters; the training window is usually the tighter limit. If a dataset stores full conversations in chosen and rejected, first separate their shared prompt from the two final answers.
Shared image prompts
For a vision model, a shared prompt can contain up to four images in its user turns. Both answers see the same images in the same order; the answers themselves are text.
Use Image column (optional) when importing from Hugging Face. Images fill the prompt's image placeholders in order. Without placeholders, they go at the start of the first user turn. JSON can also refer to images already imported into your organization's storage with {"type":"image","image":"<stored image name>"} parts. Arbitrary image URLs and local file paths are not accepted in dataset rows.
The existing image plan and storage requirements apply. Text-only DPO on a vision model does not require the image capability; later image imports, evaluations or requests still do. A text model cannot read image prompts.
Review and create a version
Review shows both candidates. Approve or reject applies to the whole pair: a rejected review status does not create a negative answer. Editing either candidate returns the pair to pending. You can edit prompt text while its images stay fixed.
Personal-data flags include rejected text. Redaction previews cover only the prompt and chosen answer: check and edit the rejected answer manually before saving.
Create a version. Both training and validation must be nonempty and contain only preference pairs. Test examples may be ordinary conversations or pairs. Source, prompt and image grouping keeps related examples together. Versions preserve both candidates and stay unchanged after later edits. Only training and validation enter the training bundle; test data remains held out.
Configure DPO
- In Configure a run, or an experiment's Setup, choose DPO under Training objective.
- Choose a compatible version. The form explains missing preference pairs and prevents selecting an incompatible version.
- Choose LoRA or QLoRA under Training method, beside Training objective in a run or in an experiment's Compute step.
- Under Advanced settings, check Learning rate, DPO beta and Training window (tokens).
- Start with a training trial. A full run needs a completed baseline of the same version; an experiment can run the baseline for you.
| Setting | DPO default | Meaning |
|---|---|---|
| Learning rate | 5e-7 |
How far each update moves the adapter. Switching objectives resets this default; a saved run keeps its chosen value. |
| DPO beta | 0.1 |
Controls how strongly training stays close to the reference model. Greater than zero and at most one. |
| Training method | LoRA | QLoRA loads a 4-bit model to reduce memory use. Switching objectives keeps the selected method. |
API requests use training.objective: "dpo", training.method: "lora" or "qlora", and optionally training.dpo_beta. Omitting the learning rate uses 5e-7; an explicit rate is preserved. The MCP start_training tool exposes the same choices.
Readiness and trials
DPO scores both candidates and the reference policy, so an SFT trial does not predict DPO memory or speed. LoRA and QLoRA use the base model with its adapter disabled as the reference, without loading a second full model.
The free readiness report is a bounded diagnostic. Vision-processor token counts remain incomplete there, with a launch warning. Before loading model weights, the runner checks every bundled training and validation pair with the pinned tokenizer or processor. The prompt plus either response must fit the training window, including image expansion; oversized pairs are refused rather than shortened.
Model fit, quantization support and training speed depend on the checkpoint, GPU and software runtime. Run a DPO trial on your chosen model, method and GPU before a full run.
Evaluate and use the adapter
Evaluations generate answers and compare them with the chosen response. They do not measure pairwise preference accuracy. DPO reward and margin metrics are training diagnostics, not a substitute for the held-out baseline comparison.
The result is a regular LoRA adapter with the existing deployment, recovery and publishing workflow. Training another run from an existing adapter and resuming from a training checkpoint are not supported. A retry starts from the base model again.
Need a hand? Visit troubleshooting