Skip to content
Chat completions

Use the API

Chat completions

The reference for the chat completions and models routes, their fields, responses, limits and errors.

The API answers chat completions from your organization's published models, in OpenAI's format. For a first request, see the API quickstart.

Route What it does
POST /v1/chat/completions Answers a conversation, whole or streamed.
GET /v1/models Lists the published models the key may call.

Both live under the address on the API screen, which ends in /v1.

Authentication

Send your API key in the Authorization header of every request:

text
Authorization: Bearer YOUR_API_KEY

A key reaches only its organization's published models, and only those it may call. A missing, revoked or expired key gets 401 invalid_api_key.

The request

bash
curl https://YOUR_API_ADDRESS/v1/chat/completions \
  -H "Authorization: Bearer $TUNE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "your-model-name",
    "messages": [
      {"role": "system", "content": "Answer in one sentence."},
      {"role": "user", "content": "What is a LoRA adapter?"}
    ],
    "max_tokens": 200,
    "temperature": 0.7
  }'

A field set to null counts as not sent.

Fields

Field Default What it does
model Required The name of a published model.
messages Required The conversation, 1 to 101 messages. See Messages.
max_tokens 512 The longest answer in tokens, an integer from 1 to 8,192.
max_completion_tokens None OpenAI's newer name for max_tokens. Wins when both are sent.
temperature 0 From 0 to 2. At 0 the model gives its most likely answer; raise it for varied answers.
top_p Model default Above 0 and at most 1.
stop None A string, or up to 4 strings of 1 to 64 characters. The model stops when it writes one.
seed None An integer. The same seed and request can give the same answer.
n 1 Must be 1. Send several requests for several answers.
stream false true sends the answer as it is written. See Streaming.
stream_options None Only with stream: true. {"include_usage": true} adds a last chunk with the token counts.
user None Accepted for compatibility, and ignored. Must be a string.

Send whole numbers as integers: 100.0 is refused for max_tokens.

Fields that are refused

Any other field gets 400 unsupported_parameter, naming it, so you never assume something was applied when it was not. This includes OpenAI's tools, tool_choice, functions, response_format, logprobs, logit_bias, presence_penalty and frequency_penalty. Remove the field and send again.

Messages

Each message has exactly two fields:

  • role: system, user or assistant.
  • content: a string, or a list of parts:
    • {"type": "text", "text": "..."} in any message.
    • {"type": "image_url", "image_url": {"url": "..."}} in user messages, for a vision model. See Images.

A name field, a tool or developer role, or an empty list is refused with 400 invalid_request.

The conversation, its images and max_tokens together must fit the model's context length. If they do not, you get 400 model_refused.

Images

A published vision model reads images in user messages. Send each image as an image_url part:

json
{"role": "user", "content": [
  {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
  {"type": "text", "text": "What does this chart show?"}
]}
  • A link must start with https:// and be at most 2,048 characters, with no spaces. It must be reachable from the internet. The image may be up to 20 MB and must download within 10 seconds.
  • Inline data is a base64 data URL of a PNG, JPEG, WebP or GIF, such as data:image/jpeg;base64,.... Use the file's real type.
  • At most 4 images per request, counting every message.
  • detail inside image_url is accepted and ignored.

Images are not stored. Each image uses tokens of the context length: see Images and the context length. A text model refuses images with 400 model_text_only. Images need the Pro plan: on Free and Starter they are refused with 403 plan_upgrade_required, and text to the same model is answered.

Request size

Model Largest request
Text model 2 MiB
Vision model 20 MiB

Base64 makes an image about a third larger. Your organization can send 2 requests over 2 MiB at a time; more get 429 with Retry-After: 1. Images sent as links keep requests small.

Send an image with Python

python
import base64
import os

from openai import OpenAI

client = OpenAI(base_url="https://YOUR_API_ADDRESS/v1", api_key=os.environ["TUNE_API_KEY"])

with open("receipt.jpg", "rb") as file:
    image = base64.b64encode(file.read()).decode()

answer = client.chat.completions.create(
    model="your-model-name",
    messages=[{
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image}"}},
            {"type": "text", "text": "What is the total on this receipt?"},
        ],
    }],
    max_tokens=300,
)
print(answer.choices[0].message.content)

For a link, put the https:// address in url instead.

The response

json
{
  "id": "chatcmpl-8f2c1a9e4b7d40a1",
  "object": "chat.completion",
  "created": 1791158400,
  "model": "your-model-name",
  "choices": [
    {
      "index": 0,
      "message": {"role": "assistant", "content": "A LoRA adapter is a small set of weights trained on top of a base model."},
      "logprobs": null,
      "finish_reason": "stop"
    }
  ],
  "usage": {"prompt_tokens": 31, "completion_tokens": 17, "total_tokens": 48}
}
  • model is the published name you sent.
  • choices holds one choice. finish_reason is stop, or length when the answer reached max_tokens.
  • usage holds the token counts, or null if the model reported none.
  • logprobs is always null. Only OpenAI's fields are returned.

Streaming

With "stream": true, the answer comes as server-sent events. Each event is data: and a JSON chunk; the last is data: [DONE].

text
data: {"id":"chatcmpl-8f2c1a9e4b7d40a1","object":"chat.completion.chunk","created":1791158400,"model":"your-model-name","choices":[{"index":0,"delta":{"role":"assistant","content":""},"logprobs":null,"finish_reason":null}]}

data: {"id":"chatcmpl-8f2c1a9e4b7d40a1","object":"chat.completion.chunk","created":1791158400,"model":"your-model-name","choices":[{"index":0,"delta":{"content":"A LoRA adapter is a small set of weights."},"logprobs":null,"finish_reason":"stop"}]}

data: [DONE]
  • Each chunk's delta.content holds the next piece of the answer.
  • With "stream_options": {"include_usage": true}, one more chunk comes before [DONE], with empty choices and the usage.
  • Errors found before the answer starts come as normal JSON errors.
  • A stream that stops without data: [DONE] is incomplete: it hit a limit or the model failed. Retry it, or shorten the answer.

List models

bash
curl https://YOUR_API_ADDRESS/v1/models -H "Authorization: Bearer $TUNE_API_KEY"
json
{"object": "list", "data": [{"id": "your-model-name", "object": "model", "created": 1791158400, "owned_by": "organization"}]}

data lists every published model the key may call, stopped ones included. created is when it was published. Listing counts toward the key's requests per minute and the organization's.

Cold starts and retries

A wake-on-request model that is asleep starts when a request arrives. The API holds your request for up to a minute while it starts. If it is still not ready, you get 503 model_starting with Retry-After: 15. An always-on model that is starting answers 503 model_starting at once, with Retry-After: 30.

Answer Retry?
429, 503 model_starting Yes, after the seconds in the Retry-After header.
502 upstream_error, 500 Yes. It may work next time.
400, 401, 404, 413, 503 model_unavailable No. Change the request, or fix the model first.

Limits

Limit Value
Request body 2 MiB, or 20 MiB for a vision model, sent within 30 seconds
Requests over 2 MiB at once 2 per organization
Messages 1 to 101
max_tokens 1 to 8,192 (512 when not sent)
Images 4 per request, for a vision model
Image link https, up to 2,048 characters, image up to 20 MB
Whole answer (not streamed) 2 minutes and 2 MiB
Stream 10 minutes from an always-on model, 5 minutes from a wake-on-request model
Pause in a stream 60 seconds
Stream size 8 MiB in all, 1 MiB per event
Cold start wait Up to 1 minute
Requests per key The key's Requests per minute and Requests at once. See Limits per key.
Requests at once per organization Your plan's limit, across all its keys. See Plan limits. The API may also ask you to retry when it's busy.
Requests a minute per organization Your plan's limit, across all its keys. See Plan limits.
Requests with an invalid key 20 a minute from one address
Requests and tokens a month Your plan's limit. See Plan limits.

A stream holds one of the key's Requests at once, and one of the organization's, until it ends.

Errors

Every error has this shape:

json
{"error": {"message": "max_tokens must be a whole number from 1 to 8192", "type": "invalid_request_error", "code": "invalid_request"}}

Act on code. The message says what was wrong.

400 Bad request

Type invalid_request_error.

Code When What to do
invalid_json The body is not valid JSON. Send valid JSON.
invalid_request A field is missing or out of range. The message names it and the allowed values. Fix the field. See Fields and Messages.
unsupported_parameter A field the API does not take. Remove it. See Fields that are refused.
model_text_only Images sent to a text model. Send images only to a vision model.
model_refused The model refused the request, usually because the conversation and max_tokens do not fit its context length. The message gives the model's reason. Shorten the conversation or lower max_tokens. If the message says to restart the model, an owner stops and starts it on the API screen.

401, 403, 404, 405, 408 and 413

Status Code What to do
401 invalid_api_key Send Authorization: Bearer with an active key.
403 plan_upgrade_required "Images are on the Pro plan." Your organization's plan does not include images. Send text only, or ask an owner to upgrade. Type permission_error.
404 model_not_found The model does not exist, or the key may not call it. Check the name with GET /v1/models.
404 not_found Use /v1/chat/completions or /v1/models.
405 method_not_allowed Use POST for chat completions and GET for models.
408 request_timeout Send the whole body within 30 seconds.
413 request_too_large Over 2 MiB to a text model: shorten the conversation. Over 20 MiB: send fewer or smaller images, or links.

429 Too many requests

Each has a Retry-After header with the seconds to wait.

Code Message Meaning
rate_limit_exceeded "This key sent too many requests. Slow down." The key reached its Requests per minute.
rate_limit_exceeded "This organization sent too many requests this minute. Its plan (plan) allows N a minute. …" Your organization reached its plan's requests a minute, across all its keys. An owner can change the plan.
rate_limit_exceeded "This key has too many requests in flight" The key reached its Requests at once.
rate_limit_exceeded "This organization has N requests in flight, as many as its plan (plan) allows. …" Your organization reached its plan's requests at once, across all its keys. Retry after a moment. An owner can change the plan.
rate_limit_exceeded "The API is busy. Try again in a moment." The API is at capacity. Retry after a moment.
rate_limit_exceeded "The API is busy with large requests. Try again in a moment." Too many requests over 2 MiB at once. See Request size.
rate_limit_exceeded "Too many requests with an invalid API key. Wait a minute." Your address sent 20 requests with an invalid key in a minute. For a minute, all its requests are refused, even with a valid key.
quota_exceeded "This organization used the API requests a month its plan (plan) allows. …" Your organization reached a monthly plan limit. Retry-After runs to the start of next month (UTC). See Plan limits.

The type is insufficient_quota for quota_exceeded and rate_limit_error for the others.

500, 502 and 503

Status Code Message What to do
500 internal_error "The request could not be handled" Try again. If it keeps happening, contact support.
502 upstream_error "The model did not answer in time" The answer took over 2 minutes. Lower max_tokens, or stream.
502 upstream_error "The model's answer was too large" Lower max_tokens.
502 upstream_error Other messages The model failed this time. Try again.
503 model_starting "The model 'name' is starting. Try again shortly." Wait the Retry-After seconds and send it again.
503 model_unavailable "The model 'name' is stopped. …" An owner starts the model on the API screen.
503 model_unavailable "The model 'name' could not be started. …" See When a start fails.
503 model_unavailable "The model 'name' is not available right now." Try again in a moment.

The type is api_error for 500 and 502, and service_unavailable for 503.