Skip to content
Keys and usage

Use the API

Keys and usage

Create, limit and revoke API keys, and see how much each key used each model.

An application calls your published models with an API key. Each key may call all your published models or only some, and has its own rate limits. The Usage panel shows what each key used.

Both are on the API screen, under Use in the sidebar.

Before you start

  • Only owners create, revoke and delete keys. Everyone in the organization sees the key list (never the keys themselves) and the usage. See Roles.

Create a key

Create one key per application, so you can revoke each alone.

  1. Under API keys, choose Create key.
  2. Enter a Name that says which application uses it.
  3. Under May call, keep All published models of this organization, or choose Only the models I choose and tick the models.
  4. Set Requests per minute and Requests at once, or keep 60 and 4.
  5. Optional: set Expires on.
  6. Choose Create key.
  7. Choose Copy key, store the key with your application's secrets, then choose I copied it.

The key starts with ti_ and is shown only this once. Tensorant does not keep it, so a lost key cannot be shown again: revoke it and create another.

Key settings

Setting Default What it does
Name None Up to 80 characters. Labels the key in the list and in usage.
May call All published models The models the key may call.
Requests per minute 60 1 to 6,000. The most requests the key sends in any 60 seconds.
Requests at once 4 1 to 256. The most requests of the key in progress at the same time.
Expires on Never The key stops working at the end of that day (UTC).

You cannot change a key after you create it. To change its models or limits, create a new key, move your application to it, then revoke the old one.

The models a key may call

  • All published models of this organization includes models you publish later.
  • Only the models I choose limits the key to the ticked models. Any other model gets 404 model_not_found, the same answer as for a model that does not exist. GET /v1/models lists only the models the key may call.

A limited key is tied to the models themselves, not their names. If you delete a model and publish a new one with the same name, the key does not reach the new one. The key list shows the deleted model as "1 deleted".

Limits per key

  • Requests per minute counts every request, GET /v1/models included. Beyond it, the API answers 429 rate_limit_exceeded with a Retry-After header.
  • Requests at once counts requests in progress. A stream counts until it ends. Beyond it, the API answers 429 rate_limit_exceeded with Retry-After: 1.

Your organization also has its plan's requests a minute and requests at once, shared by all its keys, whatever its keys allow. Past either, the API answers 429 rate_limit_exceeded with a Retry-After header. Past the requests at once, the message is "This organization has N requests in flight, as many as its plan (Plan) allows." The API may also ask you to retry when it's busy. A refused request doesn't count toward any limit. See Chat completions for every API limit.

The key list

Column What it shows
Name The key's name and its first characters, such as ti_AbCdEfGh…
State Active, Expired or Revoked
May call "All models", or the model names
Limits For example "60/min · 4 at once"
Last used When it last called the API, or "Never"
Expires When it stops working, or "Never"

Revoke a key

  1. Choose Revoke on the key's row.
  2. Choose Revoke key to confirm.

The key is refused from the next request, with 401 invalid_api_key. Revoking cannot be undone. The key stays in the list as Revoked.

To remove a revoked key from the list, choose Delete, then Delete key. Its usage stays, shown as "A deleted key".

Usage

The Usage panel covers the last 30 days (UTC):

  • Totals for Requests, Errors, Tokens (in and out) and Average time.
  • A chart of requests per day, answered and errors.
  • A table by key and model, busiest first.

A request counts when the gateway reserves capacity immediately before calling the model. Requests refused before admission are not counted, such as an invalid key, a rate limit, an invalid field, an unknown model or a stopped model. An error is a counted request that did not end with a complete answer, including a stream that ended without data: [DONE]. Tokens are counted for streamed answers too.

Only these counts are kept. Prompts and answers are never stored.

Plan limits

Your organization's plan can limit:

  • API keys: keys that are not revoked.
  • API requests a month: counted requests, errors included.
  • API requests a minute: every request in any 60 seconds, across all keys. Past it, the API answers 429 rate_limit_exceeded.
  • API requests at once: requests being answered at the same time, across all keys; a stream counts until it ends. Past it, the API answers 429 rate_limit_exceeded with Retry-After: 1.
  • API tokens a month: prompt and answer tokens together.

Months are calendar months (UTC), and caps are strict. Before calling a model, the gateway reserves one request and its full serving context window in tokens, including prompt, images and answer. A chat completion is refused with 429 quota_exceeded if this reservation will not fit. This can happen even when a shorter answer would fit in the remaining allowance.

When a complete answer reports valid usage, unused token capacity is released. Missing usage, interrupted responses and failed accounting keep the full token reservation charged; usage shows these tokens as reserved or unreported. They stay charged to the admission month, even if the request ends in another month. Capacity becomes available when an active request settles, the plan increases or a new month begins. An owner can change the plan under Settings → Plan and billing.

Limits

Limit Value
Keys per organization 50 that are not revoked, or fewer on your plan
Name 80 characters
Requests per minute 1 to 6,000
Requests at once 1 to 256
Models a key may be limited to 50
Requests at once per organization Your plan's limit, across all keys
Requests a minute per organization Your plan's limit, across all keys
Usage shown The last 30 days

When something goes wrong

Problem What to do
Create key is greyed out Only owners create keys. Ask an owner.
"Your plan (plan) allows N API keys…" or "An organization can have 1,000 keys. Revoke one first." Revoke a key you no longer use, or change your plan.
"Your plan (plan) allows N API keys. …" Revoke a key, or ask an owner to change the plan.
"Choose published models of this organization" A model you ticked was deleted meanwhile. Close the dialog and try again.
401 invalid_api_key The key is mistyped, revoked or expired. Check it, or create a new one.
404 model_not_found The model does not exist, or the key may not call it. Check May call and GET /v1/models.
429 rate_limit_exceeded Wait the Retry-After seconds, or create a key with higher limits. If the message says the organization sent too many requests this minute, spread requests out or ask an owner to change the plan.
429 quota_exceeded Your organization reached a monthly plan limit. An owner can change the plan.