Serve
Advanced serving settings
Tune memory, throughput and precision of the serving engine for deployments, experiments and published models.
The defaults suit most models. Advanced serving settings let you tune the serving engine when a model does not fit, when you need more throughput, or when you want reproducible answers. It is a fixed list of settings, each a whole number in a range or one choice from a list.
Where they apply
| Where | Settings |
|---|---|
| Deploy model in Deployments | All except Adapters |
| Compute limits of an experiment | All except Adapters, for both deployments. No FP8 quantization. |
| Publish a model and Change in Publishing models | All, including Adapters. No FP8 quantization. |
The settings belong to the GPU that serves the model, so they cannot be set per request in the Playground or the API.
Change the settings
- In the form, open Advanced serving settings. Its header shows "Platform defaults" or a summary of your changes.
- Change the fields you need. Leave a number empty, or a choice at its first option, for the default.
- Fix any problem shown in the section: the form cannot be sent until you do. See Checks.
- To undo everything, choose Reset to defaults.
When an adapter's baseline was evaluated on a deployment, the adapter is deployed with the baseline's settings, shown as "Inherited:" and read-only, so the comparison stays fair.
The settings
Memory
| Setting | Values | Default | What it does |
|---|---|---|---|
| GPU memory share (%) | 50 to 95 | 90 | How much GPU memory the engine may take for weights and cache. |
| KV cache precision | Same as the model, FP8 | Same as the model | FP8 holds about twice as many tokens in the cache. Answers can change slightly. |
| Block size | Default, 32 or 64 tokens | Default | Tokens per cache block. Larger blocks waste more cache on short answers. |
| CUDA graphs | On, Off | On | Off uses less GPU memory but answers more slowly. |
Throughput
| Setting | Values | Default | What it does |
|---|---|---|---|
| Conversations at once | 1 to 256 | 8 | Requests answered together. More raises throughput and memory use. |
| Tokens per step | 256 to 65,536 | Chosen for the GPU | Prompt tokens processed in one step. Leave empty unless you have a reason. |
| Prefix caching | On, Off | On | Reuses work for prompts that start the same way, such as a shared system prompt. |
Precision
| Setting | Values | Default | What it does |
|---|---|---|---|
| Weight precision | Auto, bfloat16, float16 | Auto | Auto uses the checkpoint's own precision. Gemma 2, Gemma 3 and GLM-4 cannot run in float16. |
| Quantization | None, FP8, computed at start | None | About half the memory for weights. Experimental, base-model deployments only, and needs a 16-bit checkpoint. |
Reproducibility
| Setting | Values | Default | What it does |
|---|---|---|---|
| Seed | 1 to 2,147,483,647 | The engine's default | Sampling seed for the whole server. An API request can also send its own seed. |
Adapters
Published models only. A published model's server serves the base model and every adapter on it, loading adapters as needed.
| Setting | Values | Default | What it does |
|---|---|---|---|
| Adapters on the GPU at once | 1 to 8 | 4 | Each slot reserves GPU memory for the largest adapter rank. |
| Adapters kept in memory | 1 to 64 | 24 | Adapters kept ready in the machine's memory. At least Adapters on the GPU at once. |
Tips
- Out of GPU memory? Lower GPU memory share (%) or Conversations at once, or turn CUDA graphs off. A larger GPU is often the simplest fix.
- Need a longer context length? Try KV cache precision FP8, which fits about twice as many tokens. It works with adapters.
- Many requests at once? Raise Conversations at once, if the GPU has memory to spare.
- Want repeatable answers? Set a Seed, or send
seedwith each API request. - Published models share a server only when their settings match. Changing a published model's settings moves it to a server with the new ones, which may start a new GPU. See Change a model.
Checks
The form checks your settings as you type, and Tensorant checks them again before anything starts, so a setting the engine would refuse does not cost a GPU start.
- Numbers are whole and in their range.
- Tokens per step is at least Conversations at once, and at most Conversations at once × context length.
- Adapters kept in memory is at least Adapters on the GPU at once.
- FP8 quantization is not used where adapters can be served. See FP8 and adapters.
- Gemma 2, Gemma 3 and GLM-4 do not use float16.
- FP8 quantization needs a checkpoint that is not already quantized another way.
The last two rules read the model's config.json from Hugging Face. The messages and fixes are under When something goes wrong.
For a deployment, these checks run with the readiness checks as serving_settings. Passing them does not guarantee the model fits: context length, Conversations at once and the cache all use memory the estimate leaves out.
FP8 and adapters
FP8 quantization (Quantization: FP8, computed at start) is not yet verified with adapters, so it is refused wherever an adapter can be served:
- when deploying an adapter, and in experiments;
- for every published model, including the base model, because a published model's server can load adapters.
Use it on a deployment of the base model. KV cache precision FP8 is a different setting and works with adapters.
When the engine does not start
If the serving engine stops before it ever answers, Tensorant shows the reason.
| The reason says | What to do |
|---|---|
| …did not accept a serving setting | Choose Reset to defaults. |
| …ran out of GPU memory with these settings | Lower GPU memory share (%) or Conversations at once, or choose a larger GPU. |
| Too little GPU memory is left for the context length | Lower the context length, raise GPU memory share (%), try KV cache precision FP8, or choose a larger GPU. |
| The context length is longer than this model supports | Lower the context length to the model's maximum. |
| The model cannot be split across this many GPUs | Choose another GPU. |
| This checkpoint is quantized another way | Set Quantization to None. |
| This model cannot run in float16 | Set Weight precision to bfloat16 or Auto. |
| …does not support adapters for this model's architecture | Serve the base model, or choose another base model. |
| …stopped before it was ready | Read the log on your network volume (see Deployments), and reset to defaults. |
What happens next:
- A deployment stops with the reason and releases its GPU. Deploy again with new settings. With default settings and no clear reason, it keeps trying until its 30-minute limit to become ready.
- A published model on a server that has never served shows Cannot start and is not retried by itself, so a failure that would repeat is not billed again. Change its settings, or choose Try again. See When a start fails.
When something goes wrong
| Message | What to do |
|---|---|
| "Setting must be a whole number from min to max." | Enter a whole number in range, or leave it empty for the default. |
| "Tokens per step must be at least…" or "…can be at most…" | Change Tokens per step, or leave it empty. |
| "Adapters kept in memory must be at least Adapters on the GPU at once." | Raise Adapters kept in memory or lower Adapters on the GPU at once. |
| "FP8 quantization is not yet verified with adapters…" | Set Quantization to None. See FP8 and adapters. |
| "Family models cannot run in float16…" | Set Weight precision to bfloat16 or Auto. |
| "This checkpoint is already quantized…" | Set Quantization to None. |
| "The model's config.json could not be read from Hugging Face; try again." | Try again in a moment. |
| "This adapter's baseline was evaluated with these advanced serving settings…" | Deploy the adapter with the inherited settings the form shows. |