Reference
Limits
The limits you can run into, by area, with links to the pages that explain them.
The limits you can run into, grouped by area. Each section links to the page that explains them.
Compute
Tensorant sets these, and can set different values for an organization. Your organization's values for training runs and deployments show on the Compute card in Connections.
| Limit | Usual value |
|---|---|
| Longest training run | 24 hours |
| Training runs at a time | 1 |
| Longest deployment | 6 hours |
| Deployments at a time, always-on published models included | 1 |
| Serverless workers across all wake-on-request models | 3 |
| Size of an uploaded file or paste | 25 MB (the upload dialog shows yours) |
| New image data per Hugging Face import | 5 GiB |
Accounts and organization
| Limit | Value |
|---|---|
| Requests for a new organization | One per person |
| Organization name | 1 to 160 characters |
| Owners | At least one |
| Invitation link | Works once, expires after 7 days |
| Invitations sent | 30 an hour |
Plans
Every organization starts on Free. A limit the plan does not set is unlimited. Enterprise has custom limits: see Enterprise.
| Limit | Free | Starter | Pro | What counts |
|---|---|---|---|---|
| Dataset imports a month | 1 | Unlimited | Unlimited | Each upload batch or Hugging Face import started, unless it fails before importing anything |
| Training runs a month | 1 | 5 | Unlimited | Each run started, experiment and trial runs included, unless it fails or is cancelled before its GPU starts |
| Published models | 1 | 5 | Unlimited | Every published model until it is deleted |
| API requests a month | 1,000 | 50,000 | 5,000,000 | Requests that reached a model in the calendar month (UTC), errors included |
| API requests a minute | 30 | 600 | 2,000 | Requests in any 60 seconds, across all the organization's API keys, listing models included. More get 429 with a wait time. |
| API requests at once | 4 | 16 | 64 | Requests being answered at the same time, across all the organization's API keys; a stream counts until it ends. More get 429 and should retry. |
| Projects | 1 | 20 | Unlimited | Every project |
| Members | 1 | 5 | 10, plus extra seats at $30 a month each | Members and pending invitations |
| API keys | 2 | 10 | Unlimited | Keys that are not revoked |
| Images (vision) | No | No | Included | Importing images, training and evaluating on images, and images sent to a model in the API or the Playground. Text use of any model is on every plan |
| API tokens a month | Unlimited | Unlimited | Unlimited | Prompt and answer tokens, plus reserved or unreported token capacity |
"Unlimited" published models and API keys stop at 1,000 each: see Published models and API keys. See Plans and billing.
Connections
| Limit | Value |
|---|---|
| Model endpoints | 20 per organization |
| Endpoint name | Up to 60 characters |
| Endpoint address | https on port 443 with a public host name, up to 500 characters |
| Endpoint answer time | 20 seconds for Check connection, 3 minutes otherwise |
| Endpoint answer size | 2 MiB |
| Connection tests at once | 2 per organization |
| Network volume and datacenter | Fixed once the organization has projects |
See Connections.
Projects
| Limit | Value |
|---|---|
| Project name | 1 to 160 characters |
| Goal | 5,000 characters |
| Generation instructions | 10,000 characters |
| Base model ID | 200 characters, as organization/model-name |
| Base model | Fixed once the project has an evaluation, training run, deployment or experiment |
See Projects.
Sources and imports
| Limit | Value |
|---|---|
| Files per upload | 20 |
| File or paste | 25 MB unless your organization's upload limit differs |
| PDF, Word and Excel files | 25 MB each; PDFs up to 500 pages; not encrypted |
| Rows per file | 50,000 |
| Excel worksheet | 50,000 rows and 1,000 columns |
| Messages in a conversation | 2 to 100, each up to 100,000 characters |
| Hugging Face rows per import | No limit with your own connections; 1 to 5,000 on the platform's connections (100 by default) |
See Sources.
Images and vision models
| Limit | Value |
|---|---|
| Images per example | 4, in user messages only |
| Image size | 20 MB and 40 megapixels |
| Image formats | PNG, JPEG, WebP and GIF; BMP and TIFF are converted to PNG |
| New image data per Hugging Face import | 5 GiB, not counting images already stored |
| Training method | LoRA only |
| Plan | Pro, for images. Text use of a vision model is on every plan. See Plans. |
See Vision models. Playground and API image limits are in their own sections below.
Recipes and generation
| Limit | Value |
|---|---|
| Examples per passage | 1 to 10 (3 by default) |
| Sources in one generation | 2,000, and 25 MB of text |
| Selection in one generation | 50 whole imports and 500 single sources |
| Reference examples | 5 per recipe, 20,000 characters together, no images |
| Recipe name | 160 characters |
| Generation instructions | 10,000 characters |
| Response JSON schema | 20,000 characters, 20 levels of nesting |
| Endpoint answer time | 3 minutes per request |
See Recipes.
Review
| Limit | Value |
|---|---|
| Examples per page | 100 |
| Examples selected at once | One page |
| Messages in an example | 2 to 100, each up to 100,000 characters |
| Category | 80 characters |
| Supporting quotes | 100 per example, each up to 6,000 characters |
See Review.
Versions
| Limit | Value |
|---|---|
| Approved examples in a version | 50,000 |
| Examples and their sources' text | 64 MiB |
| Independent source groups for an automatic split | 10, unless you import held-out evaluation data |
| Version name | 160 characters |
See Versions.
Training runs
| Setting | Range | Default |
|---|---|---|
| Epochs | 0.1 to 20 | 2 |
| Learning rate | 0.000001 to 0.01 | 0.0002 |
| Adapter size (LoRA rank) | 4 to 256 | 16 |
| Adapter scaling (LoRA alpha) | 1 to 512 | 32 |
| Context length | 128 to 32,768 tokens, in steps of 128 | 2,048 |
| Micro batch size | 1 to 32 | 1 |
| Gradient accumulation | 1 to 128 | 4 |
| Maximum runtime (hours) | 0.1 to your organization's limit | 2 |
| GPU list-price limit ($/hour) | $0.01 to $100 | $2 |
| Container disk (GB) | 30 to 500 | 50 |
| Output volume (GB) | 30 to 1,000 | 80 |
One GPU per run. Run names are up to 160 characters. See Training runs.
Experiments
| Setting | Range | Default |
|---|---|---|
| Training maximum hours | 0.1 to your organization's limit | 2 |
| Hours per inference deployment | 0.1 to your organization's limit | 1 |
| Training maximum $/hour, Inference maximum $/hour | $0.01 to $100 | $2 |
| Inference context length | 512 to 32,768 tokens | 4,096 |
| Inference idle shutdown (minutes) | 5 to 120 | 15 |
| Maximum answer tokens | 16 to 8,192, below the inference context length | 512 |
| Minimum tuned score (%) | 0 to 100 | 50 |
| Minimum improvement (percentage points) | -100 to 100 | 1 |
| Minimum independent test groups | 1 to 100,000 | 10 |
Up to 100 follow-ups per experiment. See Experiments.
Evaluations
| Limit | Value |
|---|---|
| Maximum output tokens | 16 to 8,192, below the deployment's context length (512 by default) |
| System prompt | 10,000 characters |
| LLM judge rubric | 5,000 characters |
| Allowed labels | 1 to 256, each up to 200 characters |
| Scored JSON fields | 256 fields |
| Endpoint answer time | 3 minutes |
| Evaluation name | 160 characters |
See Evaluations.
Playground
| Limit | Value |
|---|---|
| Prompt | 20,000 characters |
| System prompt | 10,000 characters |
| One conversation | 200,000 characters and 20 prompts |
| Output length | 16 to 32,768, or Max for as long as the model allows |
| Stop sequences | 4, each up to 100 characters |
| Images | 4 per conversation, each up to 10 MB, 15 MB together |
| Time for one answer | 5 minutes |
| Prompts | 30 a minute per person; a prompt sent to two models counts twice |
| Answers being written at once | 2 per person, 4 per organization |
See Playground.
Deployments
| Setting | Range | Default |
|---|---|---|
| Maximum runtime (hours) | 0.1 to your organization's limit | 1 |
| Idle shutdown (minutes) | 5 to 120 | 15 |
| GPU list-price limit ($/hour) | $0.01 to $100 | $2 |
| Serving context length | 512 to 32,768 tokens | 4,096 |
A deployment that is not ready within 30 minutes stops. One GPU per deployment. See Deployments.
Advanced serving settings
| Setting | Range | Default |
|---|---|---|
| GPU memory share (%) | 50 to 95 | 90 |
| Conversations at once | 1 to 256 | 8 |
| Tokens per step | 256 to 65,536, at least Conversations at once | Automatic |
| Seed | 1 to 2,147,483,647 | Not set |
| Adapters on the GPU at once (published models) | 1 to 8 | 4 |
| Adapters kept in memory (published models) | 1 to 64, at least Adapters on the GPU at once | 24 |
FP8 quantization is for base-model deployments only. See Advanced serving settings.
Published models
| Limit | Value |
|---|---|
| Published models | Your plan's limit, and at most 1,000 per organization |
| Name | Up to 63 lowercase letters, digits, dots, underscores or hyphens, starting with a letter or digit; cannot be changed |
| Run for (days) | 1 to 30 (7 by default) |
| Starts allowed in 24 hours | 1 to 96 (24 by default) |
| Max workers | 1 to your organization's limit across all wake-on-request models (1 by default) |
| Keep a worker warm (seconds) | 5 to 3,600 (60 by default) |
| GPU list-price limit ($/hour) | $0.01 to $100 ($2 by default) |
| Context length (tokens) | 512 to 32,768 (4,096 by default) |
See Publishing models.
Publishing adapters to Hugging Face
| Limit | Value |
|---|---|
| Repository | owner/name, 3 to 200 characters, each part up to 96 |
| Uploads of one run at a time | 1 |
See Publishing adapters to Hugging Face.
The API
| Limit | Value |
|---|---|
| Request body | 2 MiB, or 20 MiB for a vision model, sent within 30 seconds |
| Requests over 2 MiB at once | 2 per organization |
| Messages | 1 to 101 |
max_tokens |
1 to 8,192 (512 when not sent) |
| Images | 4 per request; links https and up to 2,048 characters; each image up to 20 MB |
| Whole answer (not streamed) | 2 minutes and 2 MiB |
| Stream | 10 minutes from an always-on model, 5 from a wake-on-request model; 8 MiB |
| Cold start wait | Up to 1 minute, then 503 model_starting |
| Requests at once per organization | Your plan's limit, across all its keys. See Plans. The API may also ask you to retry when it's busy. |
| Requests a minute per organization | Your plan's limit, across all its keys. See Plans. |
| Requests with an invalid key | 20 a minute from one address |
See Chat completions.
API keys
| Limit | Value |
|---|---|
| Keys | Your plan's limit of keys that are not revoked, and at most 1,000 |
| Name | 80 characters |
| Requests per minute | 1 to 6,000 (60 by default), within your plan's requests a minute for the whole organization |
| Requests at once | 1 to 256 (4 by default), within your plan's requests at once for the whole organization |
| Models a key may be limited to | 50 |
| Usage shown | The last 30 days |
See Keys and usage.