Skip to content
Limits

Reference

Limits

The limits you can run into, by area, with links to the pages that explain them.

The limits you can run into, grouped by area. Each section links to the page that explains them.

Compute

Tensorant sets these, and can set different values for an organization. Your organization's values for training runs and deployments show on the Compute card in Connections.

Limit Usual value
Longest training run 24 hours
Training runs at a time 1
Longest deployment 6 hours
Deployments at a time, always-on published models included 1
Serverless workers across all wake-on-request models 3
Size of an uploaded file or paste 25 MB (the upload dialog shows yours)
New image data per Hugging Face import 5 GiB

Accounts and organization

Limit Value
Requests for a new organization One per person
Organization name 1 to 160 characters
Owners At least one
Invitation link Works once, expires after 7 days
Invitations sent 30 an hour

See Accounts and Members.

Plans

Every organization starts on Free. A limit the plan does not set is unlimited. Enterprise has custom limits: see Enterprise.

Limit Free Starter Pro What counts
Dataset imports a month 1 Unlimited Unlimited Each upload batch or Hugging Face import started, unless it fails before importing anything
Training runs a month 1 5 Unlimited Each run started, experiment and trial runs included, unless it fails or is cancelled before its GPU starts
Published models 1 5 Unlimited Every published model until it is deleted
API requests a month 1,000 50,000 5,000,000 Requests that reached a model in the calendar month (UTC), errors included
API requests a minute 30 600 2,000 Requests in any 60 seconds, across all the organization's API keys, listing models included. More get 429 with a wait time.
API requests at once 4 16 64 Requests being answered at the same time, across all the organization's API keys; a stream counts until it ends. More get 429 and should retry.
Projects 1 20 Unlimited Every project
Members 1 5 10, plus extra seats at $30 a month each Members and pending invitations
API keys 2 10 Unlimited Keys that are not revoked
Images (vision) No No Included Importing images, training and evaluating on images, and images sent to a model in the API or the Playground. Text use of any model is on every plan
API tokens a month Unlimited Unlimited Unlimited Prompt and answer tokens, plus reserved or unreported token capacity

"Unlimited" published models and API keys stop at 1,000 each: see Published models and API keys. See Plans and billing.

Connections

Limit Value
Model endpoints 20 per organization
Endpoint name Up to 60 characters
Endpoint address https on port 443 with a public host name, up to 500 characters
Endpoint answer time 20 seconds for Check connection, 3 minutes otherwise
Endpoint answer size 2 MiB
Connection tests at once 2 per organization
Network volume and datacenter Fixed once the organization has projects

See Connections.

Projects

Limit Value
Project name 1 to 160 characters
Goal 5,000 characters
Generation instructions 10,000 characters
Base model ID 200 characters, as organization/model-name
Base model Fixed once the project has an evaluation, training run, deployment or experiment

See Projects.

Sources and imports

Limit Value
Files per upload 20
File or paste 25 MB unless your organization's upload limit differs
PDF, Word and Excel files 25 MB each; PDFs up to 500 pages; not encrypted
Rows per file 50,000
Excel worksheet 50,000 rows and 1,000 columns
Messages in a conversation 2 to 100, each up to 100,000 characters
Hugging Face rows per import No limit with your own connections; 1 to 5,000 on the platform's connections (100 by default)

See Sources.

Images and vision models

Limit Value
Images per example 4, in user messages only
Image size 20 MB and 40 megapixels
Image formats PNG, JPEG, WebP and GIF; BMP and TIFF are converted to PNG
New image data per Hugging Face import 5 GiB, not counting images already stored
Training method LoRA only
Plan Pro, for images. Text use of a vision model is on every plan. See Plans.

See Vision models. Playground and API image limits are in their own sections below.

Recipes and generation

Limit Value
Examples per passage 1 to 10 (3 by default)
Sources in one generation 2,000, and 25 MB of text
Selection in one generation 50 whole imports and 500 single sources
Reference examples 5 per recipe, 20,000 characters together, no images
Recipe name 160 characters
Generation instructions 10,000 characters
Response JSON schema 20,000 characters, 20 levels of nesting
Endpoint answer time 3 minutes per request

See Recipes.

Review

Limit Value
Examples per page 100
Examples selected at once One page
Messages in an example 2 to 100, each up to 100,000 characters
Category 80 characters
Supporting quotes 100 per example, each up to 6,000 characters

See Review.

Versions

Limit Value
Approved examples in a version 50,000
Examples and their sources' text 64 MiB
Independent source groups for an automatic split 10, unless you import held-out evaluation data
Version name 160 characters

See Versions.

Training runs

Setting Range Default
Epochs 0.1 to 20 2
Learning rate 0.000001 to 0.01 0.0002
Adapter size (LoRA rank) 4 to 256 16
Adapter scaling (LoRA alpha) 1 to 512 32
Context length 128 to 32,768 tokens, in steps of 128 2,048
Micro batch size 1 to 32 1
Gradient accumulation 1 to 128 4
Maximum runtime (hours) 0.1 to your organization's limit 2
GPU list-price limit ($/hour) $0.01 to $100 $2
Container disk (GB) 30 to 500 50
Output volume (GB) 30 to 1,000 80

One GPU per run. Run names are up to 160 characters. See Training runs.

Experiments

Setting Range Default
Training maximum hours 0.1 to your organization's limit 2
Hours per inference deployment 0.1 to your organization's limit 1
Training maximum $/hour, Inference maximum $/hour $0.01 to $100 $2
Inference context length 512 to 32,768 tokens 4,096
Inference idle shutdown (minutes) 5 to 120 15
Maximum answer tokens 16 to 8,192, below the inference context length 512
Minimum tuned score (%) 0 to 100 50
Minimum improvement (percentage points) -100 to 100 1
Minimum independent test groups 1 to 100,000 10

Up to 100 follow-ups per experiment. See Experiments.

Evaluations

Limit Value
Maximum output tokens 16 to 8,192, below the deployment's context length (512 by default)
System prompt 10,000 characters
LLM judge rubric 5,000 characters
Allowed labels 1 to 256, each up to 200 characters
Scored JSON fields 256 fields
Endpoint answer time 3 minutes
Evaluation name 160 characters

See Evaluations.

Playground

Limit Value
Prompt 20,000 characters
System prompt 10,000 characters
One conversation 200,000 characters and 20 prompts
Output length 16 to 32,768, or Max for as long as the model allows
Stop sequences 4, each up to 100 characters
Images 4 per conversation, each up to 10 MB, 15 MB together
Time for one answer 5 minutes
Prompts 30 a minute per person; a prompt sent to two models counts twice
Answers being written at once 2 per person, 4 per organization

See Playground.

Deployments

Setting Range Default
Maximum runtime (hours) 0.1 to your organization's limit 1
Idle shutdown (minutes) 5 to 120 15
GPU list-price limit ($/hour) $0.01 to $100 $2
Serving context length 512 to 32,768 tokens 4,096

A deployment that is not ready within 30 minutes stops. One GPU per deployment. See Deployments.

Advanced serving settings

Setting Range Default
GPU memory share (%) 50 to 95 90
Conversations at once 1 to 256 8
Tokens per step 256 to 65,536, at least Conversations at once Automatic
Seed 1 to 2,147,483,647 Not set
Adapters on the GPU at once (published models) 1 to 8 4
Adapters kept in memory (published models) 1 to 64, at least Adapters on the GPU at once 24

FP8 quantization is for base-model deployments only. See Advanced serving settings.

Published models

Limit Value
Published models Your plan's limit, and at most 1,000 per organization
Name Up to 63 lowercase letters, digits, dots, underscores or hyphens, starting with a letter or digit; cannot be changed
Run for (days) 1 to 30 (7 by default)
Starts allowed in 24 hours 1 to 96 (24 by default)
Max workers 1 to your organization's limit across all wake-on-request models (1 by default)
Keep a worker warm (seconds) 5 to 3,600 (60 by default)
GPU list-price limit ($/hour) $0.01 to $100 ($2 by default)
Context length (tokens) 512 to 32,768 (4,096 by default)

See Publishing models.

Publishing adapters to Hugging Face

Limit Value
Repository owner/name, 3 to 200 characters, each part up to 96
Uploads of one run at a time 1

See Publishing adapters to Hugging Face.

The API

Limit Value
Request body 2 MiB, or 20 MiB for a vision model, sent within 30 seconds
Requests over 2 MiB at once 2 per organization
Messages 1 to 101
max_tokens 1 to 8,192 (512 when not sent)
Images 4 per request; links https and up to 2,048 characters; each image up to 20 MB
Whole answer (not streamed) 2 minutes and 2 MiB
Stream 10 minutes from an always-on model, 5 from a wake-on-request model; 8 MiB
Cold start wait Up to 1 minute, then 503 model_starting
Requests at once per organization Your plan's limit, across all its keys. See Plans. The API may also ask you to retry when it's busy.
Requests a minute per organization Your plan's limit, across all its keys. See Plans.
Requests with an invalid key 20 a minute from one address

See Chat completions.

API keys

Limit Value
Keys Your plan's limit of keys that are not revoked, and at most 1,000
Name 80 characters
Requests per minute 1 to 6,000 (60 by default), within your plan's requests a minute for the whole organization
Requests at once 1 to 256 (4 by default), within your plan's requests at once for the whole organization
Models a key may be limited to 50
Usage shown The last 30 days

See Keys and usage.