Skip to content
API quickstart

Use the API

API quickstart

Create an API key and call a published model with curl, Python or JavaScript.

Your applications call your published models through Tensorant's API. It follows the OpenAI chat completions format, so the official OpenAI SDKs work with it: you change the base URL, the key and the model name.

This page gets you to a first answer. Chat completions is the full reference.

Before you start

  • A published model. Its name, such as support, is what you send as model.
  • An owner of the organization to create the key.

Step 1: Copy the API address

Open API under Use in the sidebar and choose Copy address at the top. The address ends in /v1. In the examples below, replace https://YOUR_API_ADDRESS/v1 with it.

Step 2: Create a key

  1. Under API keys, choose Create key.
  2. Enter a Name, such as "Production website". Keep the other settings for now.
  3. Choose Create key.
  4. Choose Copy key, store the key with your application's secrets, then choose I copied it. The key is shown only this once.

Keys start with ti_. Keep them on your servers, never in a web page or an app you ship. The examples below read the key from the TUNE_API_KEY environment variable:

bash
export TUNE_API_KEY="ti_..."

Step 3: Send a request

Replace your-model-name with the name of your published model.

curl

bash
curl https://YOUR_API_ADDRESS/v1/chat/completions \
  -H "Authorization: Bearer $TUNE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "your-model-name", "messages": [{"role": "user", "content": "Hello"}]}'

Python

Install the SDK with pip install openai, then:

python
import os

from openai import OpenAI

client = OpenAI(base_url="https://YOUR_API_ADDRESS/v1", api_key=os.environ["TUNE_API_KEY"])

answer = client.chat.completions.create(
    model="your-model-name",
    messages=[{"role": "user", "content": "Hello"}],
    max_tokens=512,
)
print(answer.choices[0].message.content)

JavaScript

Install the SDK with npm install openai, then, in an ES module on Node.js:

javascript
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://YOUR_API_ADDRESS/v1", apiKey: process.env.TUNE_API_KEY });

const answer = await client.chat.completions.create({
  model: "your-model-name",
  messages: [{ role: "user", content: "Hello" }],
  max_tokens: 512,
});
console.log(answer.choices[0].message.content);

The Quickstart panel on the API screen shows the same request with your address and model filled in.

Stream the answer

Add stream: true to receive the answer as it is written. Add stream_options with include_usage to get the token counts in a last chunk.

python
stream = client.chat.completions.create(
    model="your-model-name",
    messages=[{"role": "user", "content": "Write a haiku about the sea."}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:
        print(f"\n{chunk.usage.prompt_tokens} in, {chunk.usage.completion_tokens} out")

In JavaScript, pass the same stream: true and loop with for await (const chunk of stream). See Streaming for the format.

List your models

GET /v1/models lists the published models the key may call, including stopped ones. In the SDKs: client.models.list().

bash
curl https://YOUR_API_ADDRESS/v1/models -H "Authorization: Bearer $TUNE_API_KEY"

When the model is starting

A wake-on-request model that is asleep starts on the first request, and the API waits up to a minute for it. If the model is still starting, the API answers 503 with the code model_starting and a Retry-After header. Wait that many seconds and send the request again. See Cold starts and retries.

Next steps