Use the API
Chat completions
The reference for the chat completions and models routes, their fields, responses, limits and errors.
The API answers chat completions from your organization's published models, in OpenAI's format. For a first request, see the API quickstart.
| Route | What it does |
|---|---|
POST /v1/chat/completions |
Answers a conversation, whole or streamed. |
GET /v1/models |
Lists the published models the key may call. |
Both live under the address on the API screen, which ends in /v1.
Authentication
Send your API key in the Authorization header of every request:
Authorization: Bearer YOUR_API_KEYA key reaches only its organization's published models, and only those it may call. A missing, revoked or expired key gets 401 invalid_api_key.
The request
curl https://YOUR_API_ADDRESS/v1/chat/completions \
-H "Authorization: Bearer $TUNE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "your-model-name",
"messages": [
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "What is a LoRA adapter?"}
],
"max_tokens": 200,
"temperature": 0.7
}'A field set to null counts as not sent.
Fields
| Field | Default | What it does |
|---|---|---|
model |
Required | The name of a published model. |
messages |
Required | The conversation, 1 to 101 messages. See Messages. |
max_tokens |
512 | The longest answer in tokens, an integer from 1 to 8,192. |
max_completion_tokens |
None | OpenAI's newer name for max_tokens. Wins when both are sent. |
temperature |
0 | From 0 to 2. At 0 the model gives its most likely answer; raise it for varied answers. |
top_p |
Model default | Above 0 and at most 1. |
stop |
None | A string, or up to 4 strings of 1 to 64 characters. The model stops when it writes one. |
seed |
None | An integer. The same seed and request can give the same answer. |
n |
1 | Must be 1. Send several requests for several answers. |
stream |
false |
true sends the answer as it is written. See Streaming. |
stream_options |
None | Only with stream: true. {"include_usage": true} adds a last chunk with the token counts. |
user |
None | Accepted for compatibility, and ignored. Must be a string. |
Send whole numbers as integers: 100.0 is refused for max_tokens.
Fields that are refused
Any other field gets 400 unsupported_parameter, naming it, so you never assume something was applied when it was not. This includes OpenAI's tools, tool_choice, functions, response_format, logprobs, logit_bias, presence_penalty and frequency_penalty. Remove the field and send again.
Messages
Each message has exactly two fields:
role:system,userorassistant.content: a string, or a list of parts:{"type": "text", "text": "..."}in any message.{"type": "image_url", "image_url": {"url": "..."}}in user messages, for a vision model. See Images.
A name field, a tool or developer role, or an empty list is refused with 400 invalid_request.
The conversation, its images and max_tokens together must fit the model's context length. If they do not, you get 400 model_refused.
Images
A published vision model reads images in user messages. Send each image as an image_url part:
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
{"type": "text", "text": "What does this chart show?"}
]}- A link must start with
https://and be at most 2,048 characters, with no spaces. It must be reachable from the internet. The image may be up to 20 MB and must download within 10 seconds. - Inline data is a base64 data URL of a PNG, JPEG, WebP or GIF, such as
data:image/jpeg;base64,.... Use the file's real type. - At most 4 images per request, counting every message.
detailinsideimage_urlis accepted and ignored.
Images are not stored. Each image uses tokens of the context length: see Images and the context length. A text model refuses images with 400 model_text_only. Images need the Pro plan: on Free and Starter they are refused with 403 plan_upgrade_required, and text to the same model is answered.
Request size
| Model | Largest request |
|---|---|
| Text model | 2 MiB |
| Vision model | 20 MiB |
Base64 makes an image about a third larger. Your organization can send 2 requests over 2 MiB at a time; more get 429 with Retry-After: 1. Images sent as links keep requests small.
Send an image with Python
import base64
import os
from openai import OpenAI
client = OpenAI(base_url="https://YOUR_API_ADDRESS/v1", api_key=os.environ["TUNE_API_KEY"])
with open("receipt.jpg", "rb") as file:
image = base64.b64encode(file.read()).decode()
answer = client.chat.completions.create(
model="your-model-name",
messages=[{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image}"}},
{"type": "text", "text": "What is the total on this receipt?"},
],
}],
max_tokens=300,
)
print(answer.choices[0].message.content)For a link, put the https:// address in url instead.
The response
{
"id": "chatcmpl-8f2c1a9e4b7d40a1",
"object": "chat.completion",
"created": 1791158400,
"model": "your-model-name",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "A LoRA adapter is a small set of weights trained on top of a base model."},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {"prompt_tokens": 31, "completion_tokens": 17, "total_tokens": 48}
}modelis the published name you sent.choicesholds one choice.finish_reasonisstop, orlengthwhen the answer reachedmax_tokens.usageholds the token counts, ornullif the model reported none.logprobsis alwaysnull. Only OpenAI's fields are returned.
Streaming
With "stream": true, the answer comes as server-sent events. Each event is data: and a JSON chunk; the last is data: [DONE].
data: {"id":"chatcmpl-8f2c1a9e4b7d40a1","object":"chat.completion.chunk","created":1791158400,"model":"your-model-name","choices":[{"index":0,"delta":{"role":"assistant","content":""},"logprobs":null,"finish_reason":null}]}
data: {"id":"chatcmpl-8f2c1a9e4b7d40a1","object":"chat.completion.chunk","created":1791158400,"model":"your-model-name","choices":[{"index":0,"delta":{"content":"A LoRA adapter is a small set of weights."},"logprobs":null,"finish_reason":"stop"}]}
data: [DONE]- Each chunk's
delta.contentholds the next piece of the answer. - With
"stream_options": {"include_usage": true}, one more chunk comes before[DONE], with emptychoicesand theusage. - Errors found before the answer starts come as normal JSON errors.
- A stream that stops without
data: [DONE]is incomplete: it hit a limit or the model failed. Retry it, or shorten the answer.
List models
curl https://YOUR_API_ADDRESS/v1/models -H "Authorization: Bearer $TUNE_API_KEY"{"object": "list", "data": [{"id": "your-model-name", "object": "model", "created": 1791158400, "owned_by": "organization"}]}data lists every published model the key may call, stopped ones included. created is when it was published. Listing counts toward the key's requests per minute and the organization's.
Cold starts and retries
A wake-on-request model that is asleep starts when a request arrives. The API holds your request for up to a minute while it starts. If it is still not ready, you get 503 model_starting with Retry-After: 15. An always-on model that is starting answers 503 model_starting at once, with Retry-After: 30.
| Answer | Retry? |
|---|---|
429, 503 model_starting |
Yes, after the seconds in the Retry-After header. |
502 upstream_error, 500 |
Yes. It may work next time. |
400, 401, 404, 413, 503 model_unavailable |
No. Change the request, or fix the model first. |
Limits
| Limit | Value |
|---|---|
| Request body | 2 MiB, or 20 MiB for a vision model, sent within 30 seconds |
| Requests over 2 MiB at once | 2 per organization |
| Messages | 1 to 101 |
max_tokens |
1 to 8,192 (512 when not sent) |
| Images | 4 per request, for a vision model |
| Image link | https, up to 2,048 characters, image up to 20 MB |
| Whole answer (not streamed) | 2 minutes and 2 MiB |
| Stream | 10 minutes from an always-on model, 5 minutes from a wake-on-request model |
| Pause in a stream | 60 seconds |
| Stream size | 8 MiB in all, 1 MiB per event |
| Cold start wait | Up to 1 minute |
| Requests per key | The key's Requests per minute and Requests at once. See Limits per key. |
| Requests at once per organization | Your plan's limit, across all its keys. See Plan limits. The API may also ask you to retry when it's busy. |
| Requests a minute per organization | Your plan's limit, across all its keys. See Plan limits. |
| Requests with an invalid key | 20 a minute from one address |
| Requests and tokens a month | Your plan's limit. See Plan limits. |
A stream holds one of the key's Requests at once, and one of the organization's, until it ends.
Errors
Every error has this shape:
{"error": {"message": "max_tokens must be a whole number from 1 to 8192", "type": "invalid_request_error", "code": "invalid_request"}}Act on code. The message says what was wrong.
400 Bad request
Type invalid_request_error.
| Code | When | What to do |
|---|---|---|
invalid_json |
The body is not valid JSON. | Send valid JSON. |
invalid_request |
A field is missing or out of range. The message names it and the allowed values. | Fix the field. See Fields and Messages. |
unsupported_parameter |
A field the API does not take. | Remove it. See Fields that are refused. |
model_text_only |
Images sent to a text model. | Send images only to a vision model. |
model_refused |
The model refused the request, usually because the conversation and max_tokens do not fit its context length. The message gives the model's reason. |
Shorten the conversation or lower max_tokens. If the message says to restart the model, an owner stops and starts it on the API screen. |
401, 403, 404, 405, 408 and 413
| Status | Code | What to do |
|---|---|---|
| 401 | invalid_api_key |
Send Authorization: Bearer with an active key. |
| 403 | plan_upgrade_required |
"Images are on the Pro plan." Your organization's plan does not include images. Send text only, or ask an owner to upgrade. Type permission_error. |
| 404 | model_not_found |
The model does not exist, or the key may not call it. Check the name with GET /v1/models. |
| 404 | not_found |
Use /v1/chat/completions or /v1/models. |
| 405 | method_not_allowed |
Use POST for chat completions and GET for models. |
| 408 | request_timeout |
Send the whole body within 30 seconds. |
| 413 | request_too_large |
Over 2 MiB to a text model: shorten the conversation. Over 20 MiB: send fewer or smaller images, or links. |
429 Too many requests
Each has a Retry-After header with the seconds to wait.
| Code | Message | Meaning |
|---|---|---|
rate_limit_exceeded |
"This key sent too many requests. Slow down." | The key reached its Requests per minute. |
rate_limit_exceeded |
"This organization sent too many requests this minute. Its plan (plan) allows N a minute. …" | Your organization reached its plan's requests a minute, across all its keys. An owner can change the plan. |
rate_limit_exceeded |
"This key has too many requests in flight" | The key reached its Requests at once. |
rate_limit_exceeded |
"This organization has N requests in flight, as many as its plan (plan) allows. …" | Your organization reached its plan's requests at once, across all its keys. Retry after a moment. An owner can change the plan. |
rate_limit_exceeded |
"The API is busy. Try again in a moment." | The API is at capacity. Retry after a moment. |
rate_limit_exceeded |
"The API is busy with large requests. Try again in a moment." | Too many requests over 2 MiB at once. See Request size. |
rate_limit_exceeded |
"Too many requests with an invalid API key. Wait a minute." | Your address sent 20 requests with an invalid key in a minute. For a minute, all its requests are refused, even with a valid key. |
quota_exceeded |
"This organization used the API requests a month its plan (plan) allows. …" | Your organization reached a monthly plan limit. Retry-After runs to the start of next month (UTC). See Plan limits. |
The type is insufficient_quota for quota_exceeded and rate_limit_error for the others.
500, 502 and 503
| Status | Code | Message | What to do |
|---|---|---|---|
| 500 | internal_error |
"The request could not be handled" | Try again. If it keeps happening, contact support. |
| 502 | upstream_error |
"The model did not answer in time" | The answer took over 2 minutes. Lower max_tokens, or stream. |
| 502 | upstream_error |
"The model's answer was too large" | Lower max_tokens. |
| 502 | upstream_error |
Other messages | The model failed this time. Try again. |
| 503 | model_starting |
"The model 'name' is starting. Try again shortly." | Wait the Retry-After seconds and send it again. |
| 503 | model_unavailable |
"The model 'name' is stopped. …" | An owner starts the model on the API screen. |
| 503 | model_unavailable |
"The model 'name' could not be started. …" | See When a start fails. |
| 503 | model_unavailable |
"The model 'name' is not available right now." | Try again in a moment. |
The type is api_error for 500 and 502, and service_unavailable for 503.