Tools
The tools an agent gets at each access level, how spending works, how files get in, and what stays in the console.
This page is the reference for what an agent can do in Tensorant. To connect one first, see the MCP quickstart.
An agent works through tools. There is roughly one tool for each thing you do in the console, not one for each button. You do not call them yourself: you ask your agent in plain words, such as "import the first 500 rows of this dataset into my support project" or "compare the trained model with its baseline", and it picks the tools.
An agent sees only the tools its access level allows. If it calls one of the others by name, it is told why not and which level would allow it. Results come with a link to the matching console page where there is one. Long lists are cut and say when there is more.
An agent follows your role, your plan's limits and the separation between organizations exactly as the console does. Tensorant never gives an agent something the console would refuse you.
A typical run
Agents are told to work in this order:
- Add sources: import a Hugging Face dataset, add a small file, or send a larger one through an upload link.
- Write examples (optional): generate examples from your documents with one of your endpoints.
- Review: approve or reject examples. Only approved examples go into a version.
- Create a version: freezes the approved examples into train, validation and test splits.
- Run a baseline: evaluate the base model on the version's test examples, against an endpoint or a test deployment of the base model. A full training run needs a completed baseline.
- Train: a short trial first, which shows whether the settings work and estimates a full run, then the full run.
- Test the result: deploy the finished run for testing, evaluate it and compare it with the baseline.
- Stop test deployments when done, because they cost money every hour.
- Publish (owners): serve the model through the API. You create the API keys in the console.
Read tools
Every agent can use these. They change nothing.
| Tool | What it does |
|---|---|
get_account |
Shows who the agent acts for: you, the organization, the role the agent acts with (your role, lowered to what its access level allows), the agent's access level, your plan with its limits and this month's usage, and which connections are set up (never their values). Agents are told to call it first. |
list_projects |
Lists the organization's projects, newest first, with each one's base model and goal. |
get_project |
Shows one project: base model, goal, instructions, counts of sources, examples, versions and training runs, and the next step to take. |
list_sources |
Lists a project's imports, uploads and Hugging Face imports alike, newest first, with how many sources and examples each brought. |
search_huggingface_dataset |
Looks up a Hugging Face dataset: its subsets and splits, and the one worth starting with. |
preview_huggingface_dataset |
Shows a few rows of a Hugging Face dataset, its columns, and how they would become examples. |
list_recipes |
Lists a project's recipes (how examples are written from documents) and the templates to start from. |
list_examples |
Lists examples with their messages, review status, category and quality flags. Can filter by status, flag, category, text or import. Long messages are shortened. |
list_versions |
Lists a project's versions with the examples in each split, and says whether a new version can be made now and why not. |
get_version |
Shows how a version's data looks before training: examples per category and split with warnings, and a check for examples that are too long or cannot be formatted. |
list_training_runs |
Lists a project's trials and full runs with status, settings and what they cost. |
get_training_run |
Shows one run: status, the error if it failed, settings, cost and progress. For a trial it also gives the trial report: whether it finished, the estimate for a full run and recommendations. |
list_gpus |
Lists the GPUs available to the organization with memory, price per hour and availability, cheapest first, at most 40. Can list GPUs for training, test deployments or wake-on-request models. |
check_model |
Checks whether your organization can use a Hugging Face model: that it exists and is open to you, has a chat template, its size, and the GPU memory it needs. |
list_evaluations |
Lists a project's evaluations, baselines and trained models, with their scores. |
get_evaluation |
Shows one evaluation: status, scores per metric, scores by category, recipe and source, and a page of the model's answers with their scores. |
compare_evaluations |
Compares a trained model with its baseline, or two evaluations of one version: how many examples improved, stayed the same or got worse, how sure the difference is, and answers side by side. |
list_deployments |
Lists deployments, models running on a rented GPU, with status, model, GPU, price per hour, when the time limit ends and what each cost. |
list_published_models |
Lists published models: name, state, always on or wake on request, the model behind it, GPU and price while it runs, and the address apps call. |
list_activity |
Shows the GPUs running now across the organization with what each costs per hour and when it stops. With a project, also its recent background jobs and their progress. |
Build tools
Build agents also use all the Read tools, and these.
| Tool | What it does |
|---|---|
create_project |
Creates a project tied to one base model, with an optional goal, instructions and starting use case. |
update_project |
Changes a project's name, goal, instructions, use case or base model. The base model is fixed once the project has an evaluation, a training run or a deployment. |
import_huggingface_dataset |
Imports rows of a Hugging Face dataset as examples, in the background. Counts toward your plan's imports. Examples wait for review unless the agent approves them as they arrive. |
upload_file |
Adds a file of up to 1 MB from content the agent holds. See Files. |
create_upload_link |
Makes a one-time link for a larger file, sent with curl. See Files. |
create_recipe |
Saves a recipe: the kind of example, whether the source passage is shown, extra instructions, and for structured JSON the schema every answer must match. |
update_recipe |
Replaces a recipe's settings. Its version goes up; examples already written keep the version they were written with. |
generate_examples |
Writes examples from the sources of imports with one of your endpoints, in the background. Your endpoint's provider bills the requests, so the agent gets a preview first. See Spending. |
review_examples |
Approves, rejects or puts back to pending up to 500 examples at once. Says why for each one that cannot be approved. |
approve_matching |
Approves the examples that match a filter, a batch at a time, without listing them first. |
edit_example |
Changes one example's messages, category or answer type. Changing the messages or the answer type puts it back to pending. If a judge already checked it, it is sent to your judge endpoint again, which your provider bills. |
create_version |
Freezes the approved examples into a new version, split into train, validation and test, in the background. |
start_evaluation |
Scores a model on a version's test examples, in the background: a base model as a baseline against an endpoint, or a trained run on its ready deployment. Your endpoint's provider bills the requests, so the agent gets a preview first. |
rename |
Gives a training run, evaluation, deployment or version a name you can recognize. |
cancel_job |
Stops a queued or running background job: an import, generation, evaluation or new version. What it already stored stays. It does not stop a GPU. |
ask_model |
Asks a model a question or a conversation, as the Playground does: one of your endpoints, a ready deployment or a published model. It uses only what is already running and starts no deployment. Asking a published model counts toward your plan's API requests and can wake one that is asleep, as a request from your apps does. |
Full tools
Full agents also use all the Read and Build tools, and these. GPUs start only after the price is shown: see Spending.
| Tool | What it does |
|---|---|
start_training |
Trains a model on a version, on a GPU in your own GPU account: a trial first, then a full run, which needs a completed baseline of the same version. Training can use any connected provider. |
retry_training |
Starts a failed or cancelled run again with the same settings, as a new run. |
cancel_training |
Stops a training run and releases its GPU. A cancelled run is not resumed, only started again from the beginning. |
push_to_huggingface |
Publishes a finished run's adapter to a Hugging Face repository, private by default, and agents are told to ask you before making it public. It uses your organization's Hugging Face token from Connections and creates the repository if it does not exist. |
start_deployment |
Starts a test deployment, the project's base model or a finished run's adapter, on a rented RunPod GPU, so it can be asked or evaluated. It bills by the hour until it stops. |
stop_deployment |
Stops a test deployment and releases its GPU. |
publish_model |
Owners only. Publishes a model so apps can call it by name through the API, always on for 1 to 30 days or on demand. See Publishing models. |
start_published_model |
Owners only. Starts a stopped published model again. |
extend_published_model |
Owners only. Keeps a running always-on published model running for more days, at most 30 days from now. |
stop_published_model |
Owners only. Stops a published model. Apps calling it get errors until it is started again. |
Deployments and published models run on RunPod. Agents cannot create API keys, so you create the keys your apps use in the console: see Keys and usage.
Spending
Anything that starts a GPU, or sends many requests to an endpoint your provider bills, takes two steps:
- Preview. The agent calls the tool, and nothing starts. Tensorant answers with what would start, the GPU and its price per hour, the time limit and the most it can cost, or for an always-on model the days it runs. For a test deployment it gives both what the time limit comes to at today's price ("about") and at the price limit you set ("at most"). For training, it also lists any warnings from the readiness check. For work on your endpoints it says how many passages or test rows will be sent. Where the work counts toward a plan limit, such as training runs a month or published models, it says so. It also gives a confirmation code. The prices are the same ones the console shows.
- Confirm. The agent shows you the preview. If you agree, it calls the same tool again with the same settings and the code.
The code works once, within 10 minutes, for those exact settings and that agent. If anything changes, such as a pricier GPU than the one you saw, or the code is used or too old, Tensorant refuses and the agent has to ask for a new preview.
The tools that take two steps:
| What starts | Tools |
|---|---|
| A GPU on your GPU account | start_training, retry_training, start_deployment, publish_model, start_published_model, extend_published_model |
| Many requests to your endpoints | generate_examples, start_evaluation |
Agents are told to show you every preview and wait for your go-ahead, and never to confirm on their own. Tensorant cannot see whether an agent asked you first. To keep an agent from starting training or deployments at all, connect it with Build. The one exception is asking a published model that is asleep, which wakes it as a request from your apps does.
A few smaller things are not previewed: ask_model sends one question to one model, and edit_example sends one example to your judge endpoint when a judge had checked it. The provider of the endpoint or the GPU bills these, not Tensorant.
Test deployments stop by themselves after their time limit or when they sit idle, and agents are told to stop them when done. You can see every GPU running now, with its cost, under Running now in the sidebar.
Files
An agent can add files to a project in two ways. Either way, the file's extension decides how it is read, the same file types as the console: see File formats. The agent chooses whether the file is training data or held-out test data and whether to approve examples that pass the checks as they arrive.
Small files, up to 1 MB. The agent sends the content itself with
upload_file: text for text formats, or the file's bytes base64-encoded for PDF, Word and Excel. Larger files are refused with "Files over 1 MB go through create_upload_link."Larger files, up to 25 MB. The agent calls
create_upload_link, which answers a one-time address and acurlcommand:curl -F 'file=@<path to your file>' https://mcp.tensorant.io/uploads/<code>The agent runs it with the file's path, or, if it cannot run commands, gives it to you to run. The file goes straight to Tensorant, not through the agent's model.
curlprints the result: the examples added, the duplicates and any problems.
The link is valid for 15 minutes and works once. If a request fails before Tensorant has your file, the link still works, and so it does when Tensorant is briefly unavailable: run the same command again. Once Tensorant has the file, the link is used up, even if the import is then refused, for example by a plan limit. If the upload takes too long to be answered, the link is used up too, since the file may have arrived: ask the agent to check the project's sources before you send it again with a new link. Each agent sends one file through a link at a time, at most 4 uploads run at once across Tensorant, and an upload counts toward your plan's imports like one in the console. The console's other limits apply to the file itself, such as pages and rows: see Sources.
What stays in the console
An agent cannot do these. Ask it for the steps and do them yourself in the console:
- Connect or change connections and their keys.
- Change the plan, or anything in billing.
- Invite or remove members, or change roles.
- Create or change API keys.
- Delete anything.
- Download weights, bundles, datasets or dataset splits.
- Start experiments, which run a baseline, training and deployments on their own. See Experiments.
- Change a published model's settings.
- Retry a background job, or recover a training run.
- Change organization settings or email preferences.
- Create, widen or revoke agents.
If an agent reaches for one of these directly, it is refused with "Agents cannot do this. Use the Tensorant console."
Plan limits
Your plan's limits apply to agents as they do to you. Imports, training runs, published models, projects and API requests count whether an agent or a person started them. When a limit stops something, the agent gets the sentence the console would show, such as "Your plan (Free) allows 1 …", and can tell you. An owner can change the plan. See What a plan limits.
See what an agent did
- Imports, generation, evaluations and new versions an agent starts run as background jobs, and appear in Activity like any other.
- Owners can open the audit trail, where a change made by an agent shows the person's name followed by "via" and the agent's name. It also shows an app connected by browser sign-in ("Connected an agent"), an app that disconnects itself ("An agent disconnected itself"), and an agent's use of
ask_model("Used the playground"). Like the rest of the trail, it never records your content. - Settings → Agents shows when each agent was created and last used.
Limits
| Limit | Value |
|---|---|
| File sent by an agent in one call | 1 MB |
| File sent through an upload link | 25 MB, the console's limit |
| Upload link | Valid for 15 minutes, works once |
| Uploads at once across Tensorant | 4 |
| Uploads through links at once, for one agent | 1 |
| Spending preview | Confirm within 10 minutes, once |
| Examples reviewed in one call | 500 |
| GPUs listed in one call | 40, the cheapest |
Calls per agent and the other limits on agents themselves are in Access and tokens. See Limits for all of them.
When something goes wrong
The messages an agent reports, such as an expired confirmation, a plan limit or an upload that was refused, are explained in When something goes wrong.
Next steps
- Connect your app: the setup for Claude Code, Cursor, Codex, VS Code, Claude and ChatGPT.
- Access and tokens: access levels, personal tokens, revoking and limits.
- Plans and billing: what each plan limits.
- Publishing models: what happens after an owner's agent publishes a model.
Need a hand? Visit troubleshooting