Use the API
API quickstart
Create an API key and call a published model with curl, Python or JavaScript.
Your applications call your published models through Tensorant's API. It follows the OpenAI chat completions format, so the official OpenAI SDKs work with it: you change the base URL, the key and the model name.
This page gets you to a first answer. Chat completions is the full reference.
Before you start
- A published model. Its name, such as
support, is what you send asmodel. - An owner of the organization to create the key.
Step 1: Copy the API address
Open API under Use in the sidebar and choose Copy address at the top. The address ends in /v1. In the examples below, replace https://YOUR_API_ADDRESS/v1 with it.
Step 2: Create a key
- Under API keys, choose Create key.
- Enter a Name, such as "Production website". Keep the other settings for now.
- Choose Create key.
- Choose Copy key, store the key with your application's secrets, then choose I copied it. The key is shown only this once.
Keys start with ti_. Keep them on your servers, never in a web page or an app you ship. The examples below read the key from the TUNE_API_KEY environment variable:
export TUNE_API_KEY="ti_..."Step 3: Send a request
Replace your-model-name with the name of your published model.
curl
curl https://YOUR_API_ADDRESS/v1/chat/completions \
-H "Authorization: Bearer $TUNE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "your-model-name", "messages": [{"role": "user", "content": "Hello"}]}'Python
Install the SDK with pip install openai, then:
import os
from openai import OpenAI
client = OpenAI(base_url="https://YOUR_API_ADDRESS/v1", api_key=os.environ["TUNE_API_KEY"])
answer = client.chat.completions.create(
model="your-model-name",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=512,
)
print(answer.choices[0].message.content)JavaScript
Install the SDK with npm install openai, then, in an ES module on Node.js:
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://YOUR_API_ADDRESS/v1", apiKey: process.env.TUNE_API_KEY });
const answer = await client.chat.completions.create({
model: "your-model-name",
messages: [{ role: "user", content: "Hello" }],
max_tokens: 512,
});
console.log(answer.choices[0].message.content);The Quickstart panel on the API screen shows the same request with your address and model filled in.
Stream the answer
Add stream: true to receive the answer as it is written. Add stream_options with include_usage to get the token counts in a last chunk.
stream = client.chat.completions.create(
model="your-model-name",
messages=[{"role": "user", "content": "Write a haiku about the sea."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage:
print(f"\n{chunk.usage.prompt_tokens} in, {chunk.usage.completion_tokens} out")In JavaScript, pass the same stream: true and loop with for await (const chunk of stream). See Streaming for the format.
List your models
GET /v1/models lists the published models the key may call, including stopped ones. In the SDKs: client.models.list().
curl https://YOUR_API_ADDRESS/v1/models -H "Authorization: Bearer $TUNE_API_KEY"When the model is starting
A wake-on-request model that is asleep starts on the first request, and the API waits up to a minute for it. If the model is still starting, the API answers 503 with the code model_starting and a Retry-After header. Wait that many seconds and send the request again. See Cold starts and retries.
Next steps
- Chat completions: every field, limit and error.
- Keys and usage: key limits, revoking keys, and usage.
- Publishing models: always on or wake on request.