Skip to content
Playground

Serve

Playground

Try a running model by hand, or compare two models side by side.

The Playground is where you try a model by hand. You type a prompt and one or two models answer side by side, with their speed and token counts under each answer. Use it to compare the base model with your trained adapter, or to try a prompt before you run an evaluation.

The Playground only talks to models that are already running: it never starts a GPU. Nothing is saved: prompts and answers stay in your browser tab, and they never change your dataset.

Before you start

  • A running model, one of:
    • a deployment of this project that shows Ready, of the base model or an adapter (this includes the tuned model of an experiment you kept running);
    • a model endpoint your organization added in Connections.
  • The member or owner role. Viewers can open the Playground, but cannot send. See roles.

Send a prompt

  1. Open the project, then Playground under Use in the sidebar.
  2. Under Model A, choose a model. The list shows this project's ready deployments, base model first, then your organization's endpoints.
  3. Optional: under Model B, choose a second model to compare. Leave it at None to ask one model.
  4. Optional: change the Settings on the right. See below.
  5. Type in Prompt and choose Send, or press Enter. Shift+Enter starts a new line.

Send another prompt to continue the conversation: each side sends its whole conversation so far. Choose Clear to start over. Choosing another model for a side empties that side.

Good to know:

  • Asking a deployment counts as use, so it does not stop for being idle. It still stops at its maximum runtime.
  • An endpoint's provider receives every prompt you send to it.
  • Published models are not listed: call them through the API. The GPU of an always-on published model may appear as Shared server · base model; it answers as the base model, not as a published adapter.

Attach images

A deployment of a vision model reads images with your prompt. Attach image appears beside Send only when every chosen model is a deployment of this project's vision model. Images are never sent to your own endpoints.

  1. Choose Attach image and pick one or more files, or paste an image into Prompt.
  2. To remove one, choose the cross on its thumbnail.
  3. Type a prompt and choose Send. A prompt needs text: an image alone cannot be sent.

Images are sent again with every later prompt of the conversation, since each side sends the whole conversation. Formats, sizes and counts are under Limits.

Settings

The Settings panel on the right applies to both models, from the next prompt you send. Hover over or focus the ⓘ beside a setting to see what it does. Choose Reset to go back to the defaults, or the panel button to hide it; Show settings above the answers shows it again. Your browser remembers the settings.

Setting What it does Default
Stream tokens On: the answer appears as the model writes it. Off: the model is asked for the whole answer, which appears when it is done. Off
System prompt Optional. Sent to both models before the conversation. Blank
Temperature How varied the answer is, from 0 to 2. 0 picks the most likely words every time. 0.7
Output length The longest answer, from 16 to 32,768 tokens. Max lets the model answer as long as its context allows. Max
Reasoning effort Low, medium or high: how long a reasoning model, such as gpt-oss, thinks before answering. Other models may ignore it or refuse the prompt. Model default
Response format JSON makes the model answer with one JSON object. Ask for JSON in the prompt too. Text
Top P Only the most likely words whose chances add up to this share are considered, from 0.01 to 1. 1
Frequency penalty From −2 to 2. Above 0, a word is less likely the more often it already appeared. 0
Presence penalty From −2 to 2. Above 0, a word that already appeared at all is less likely. 0
Stop sequences Up to 4 pieces of text, each up to 100 characters. The answer ends where the model writes one. Press Enter to add one. None
Seed A whole number. The same seed, settings and prompt give the same answer on most models. Blank: a new one each time

Settings left at their defaults are not sent, so the model uses its own. A model or endpoint that does not accept a setting refuses the prompt: see When something goes wrong.

Timing and token counts

Under each answer, on one line: hover over or focus a number to see what it means.

Measure What it is
Time to first token How long the model took to start answering. Shown only when Stream tokens is on.
Tokens per second With streaming, how fast the model wrote once it had started; "≈" means the model reported no token counts, so they are estimated. Without streaming, the answer's tokens over the total time, which includes reading the prompt.
Total time From sending the prompt to the end of the answer.
in and out Prompt tokens (including the conversation so far) and answer tokens.

Times are measured from Tensorant to the model, so they leave out your own connection.

Your organization's activity log records that you used the Playground, but never your prompts.

Limits

Limit Value
Prompt 20,000 characters
System prompt 10,000 characters
One conversation 200,000 characters of text, and 20 prompts
Output length 16 to 32,768, or Max
Images PNG, JPEG, WebP or GIF; up to 10 MB each; 4 images and 15 MB together per conversation; only for this project's vision deployments
Time for one answer 5 minutes
Prompts 30 a minute per person. A prompt sent to two models counts twice.
Answers being written at once 2 per person, 4 per organization

When something goes wrong

An error appears on the side it belongs to. If part of the answer had arrived, it starts with "Stopped:".

Problem What to do
"Nothing is running to talk to" Deploy a model in Deployments and wait for Ready, or add an endpoint in Connections.
"…returned HTTP 400. The answer length may not fit the model's context…" The conversation plus Output length is longer than the model's context. Lower Output length, set it to Max, or choose Clear.
"…returned HTTP 400. The model may not accept one of the settings…" Choose Reset in Settings, then change one setting at a time to find the one it refuses.
"…did not answer in time" or "The answer was cut off before it finished." The answer took over 5 minutes. Lower Output length and try again.
"You are sending prompts too fast…" or "The playground is busy…" Wait a moment, then send again.
An image is refused Use PNG, JPEG, WebP or GIF under 10 MB. After 4 images, choose Clear.
"Keep the conversation under 200,000 characters" or "…List should have at most 40 items" The conversation is full. Choose Clear.
"Choose a running deployment of this project" The deployment stopped, for example at its maximum runtime. Choose another model, or deploy again.
"Endpoint host does not resolve to a public address" Your endpoint must be on the public internet. See Address rules.
Any other "Endpoint name …" error The endpoint refused or failed. Use Check connection in Connections, or ask its provider.