AWS
Connect an AWS account, prepare EC2 networking and a GPU image, and verify training and recovery.
Connect AWS to train on one NVIDIA GPU in your own EC2 account. You provide the region, network, image and credentials; Tensorant creates an on-demand instance for each run and releases it after collecting the adapter.
AWS compute and project storage are separate connections. Set up custom S3 storage before creating your first project, or keep your existing RunPod network storage. Managed inference, publishing and managed experiment evaluation still require RunPod and a RunPod network volume.
1. Choose a region and request GPU quota
Use a commercial AWS region such as us-east-1. The AMI, subnet, security group and optional EC2 key pair must exist in that region. This connector does not support AWS China or GovCloud.
In Service Quotas → AWS services → Amazon Elastic Compute Cloud, select your region and check the on-demand quota for the instance family you intend to use. G instances and P instances have separate quotas. Request enough vCPUs for the instance, including other running instances in that quota group; a quota of one does not mean one GPU. New accounts can have zero GPU-family quota. AWS EC2 instance quotas
Tensorant supports single-GPU, x86_64 instances with a GPU compatible with its trainer. Start with an Ampere-or-newer option that fits your model, such as an A10G, L4 or H100. Multi-GPU machines, ARM machines and Spot requests are not supported. A regional listing does not reserve capacity or increase your quota.
2. Prepare a private subnet with outbound access
Use an existing VPC, or have your AWS administrator create one:
- Create a private subnet in an availability zone that offers your chosen instance type.
- Give that subnet an outbound internet route. A typical setup uses a public NAT gateway in a public subnet, an internet gateway on the VPC, and a
0.0.0.0/0route from the private subnet to the NAT gateway. - Create a security group in the same VPC. Permit outbound HTTPS and the DNS access your VPC uses. If you restrict destinations, include the training container registry, model downloads, Tensorant's storage endpoint and their download hosts.
- Record the private subnet's
subnet-…ID and the security group'ssg-…ID.
Tensorant explicitly disables public-IP assignment. Placing the instance in a subnet with an internet-gateway route alone does not give it internet access. The application does not create a NAT gateway, VPC, routes or security-group rules. NAT and data-transfer charges are separate from training estimates. AWS NAT gateways
Training needs no inbound port. For SSH recovery, allow port 22 only from your VPN, private administration network or bastion security group. Keep the connection private; an SSH key by itself does not provide a network route to the instance.
3. Prepare and verify a GPU AMI
The image must be an available x86_64 Linux/UNIX AMI with an EBS root disk. Tensorant rejects Marketplace product codes. The operating system needs:
- cloud-init and Python 3, with cloud-init enabled for new instances;
- Docker Engine and NVIDIA Container Toolkit, configured for Docker;
- a working NVIDIA driver 580 or newer, compatible with CUDA 13 and the chosen GPU;
- enough root-disk space for the system, training container, model and run output.
A useful starting point is the AWS Deep Learning Base OSS Nvidia Driver GPU AMI (Ubuntu 24.04). Resolve its AMI ID in your chosen region using your administrator's AWS CLI credentials:
aws ssm get-parameter --region us-east-1 \
--name /aws/service/deeplearning/ami/x86_64/base-oss-nvidia-driver-gpu-ubuntu-24.04/latest/ami-id \
--query Parameter.Value --output textReplace us-east-1 with your region. The lookup needs ssm:GetParameter for the person running it; Tensorant does not perform this lookup. Check that release's software list and supported instance families, then retain the returned ami-… ID rather than a moving "latest" reference. AWS DLAMI image lookup and release notes
Before using an image for production, launch a temporary instance yourself on the intended GPU and private subnet. This is a paid EC2 instance. Connect through your recovery route and check:
uname -m
python3 --version
cloud-init --version
nvidia-smi --query-gpu=name,driver_version --format=csv,noheader
sudo docker version
sudo docker run --rm --gpus all ubuntu:24.04 nvidia-smiExpect x86_64, a driver version of at least 580, and a GPU visible inside the container. If packages are missing, follow Docker's Ubuntu installation and NVIDIA's toolkit setup. After installing the toolkit, configure Docker:
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart dockerTest again. If you customized the instance, remove build credentials and saved data, prepare cloud-init for cloning, and create an EBS-backed AMI using EC2 → Instances → Actions → Image and templates → Create image. Wait until it is available and use the new AMI ID. Stop or terminate the temporary builder when finished; Tensorant does not manage it. Creating an EBS-backed AMI
Tensorant uses an encrypted gp3 root disk, increases its size to accommodate the image and requested workspace, and suppresses secondary disks inherited from the AMI. Do not put required software or files only on a secondary AMI disk.
4. Create a dedicated control-plane identity
In IAM → Users, create a dedicated user for this connection without console access. Attach a policy scoped to the approved training resources, then open Security credentials → Access keys → Create access key. Choose the use case appropriate for an application outside AWS. Copy the access key ID and secret access key into a password manager; AWS shows the secret only when it is created. Do not use root-account keys. AWS access-key setup
The following policy is a starting template for this connector. Replace REGION, ACCOUNT_ID, AMI_ID, SUBNET_ID, SECURITY_GROUP_ID and KEY_NAME. Remove the key-pair ARN if you will use only Systems Manager recovery. Your administrator should review permission boundaries, service-control policies and encryption-key policies before applying it.
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "DiscoverAndPrice",
"Effect": "Allow",
"Action": [
"ec2:DescribeImages", "ec2:DescribeSubnets",
"ec2:DescribeSecurityGroups", "ec2:DescribeInstanceTypeOfferings",
"ec2:DescribeInstanceTypes", "ec2:DescribeInstances",
"pricing:GetProducts"
],
"Resource": "*"
},
{
"Sid": "ApprovedLaunchInputs",
"Effect": "Allow",
"Action": "ec2:RunInstances",
"Resource": [
"arn:aws:ec2:REGION::image/AMI_ID",
"arn:aws:ec2:REGION:ACCOUNT_ID:subnet/SUBNET_ID",
"arn:aws:ec2:REGION:ACCOUNT_ID:security-group/SECURITY_GROUP_ID",
"arn:aws:ec2:REGION:ACCOUNT_ID:key-pair/KEY_NAME",
"arn:aws:ec2:REGION:ACCOUNT_ID:network-interface/*"
]
},
{
"Sid": "CreateTaggedComputeAndDisks",
"Effect": "Allow",
"Action": "ec2:RunInstances",
"Resource": [
"arn:aws:ec2:REGION:ACCOUNT_ID:instance/*",
"arn:aws:ec2:REGION:ACCOUNT_ID:volume/*"
],
"Condition": {"StringEquals": {"aws:RequestTag/tensorant:managed": "true"}}
},
{
"Sid": "TagsAtLaunchOnly",
"Effect": "Allow",
"Action": "ec2:CreateTags",
"Resource": [
"arn:aws:ec2:REGION:ACCOUNT_ID:instance/*",
"arn:aws:ec2:REGION:ACCOUNT_ID:volume/*"
],
"Condition": {"StringEquals": {
"ec2:CreateAction": "RunInstances",
"aws:RequestTag/tensorant:managed": "true"
}}
},
{
"Sid": "ManageTaggedTrainingInstances",
"Effect": "Allow",
"Action": ["ec2:StartInstances", "ec2:StopInstances", "ec2:TerminateInstances"],
"Resource": "arn:aws:ec2:REGION:ACCOUNT_ID:instance/*",
"Condition": {"StringEquals": {"ec2:ResourceTag/tensorant:managed": "true"}}
}
]
}Discovery and pricing use permissions that cannot all be restricted to individual instances. The Pricing API is called in us-east-1, even when your GPUs run elsewhere; an account policy that blocks that region can prevent the catalog from loading. The resource and tag restrictions above follow AWS EC2 IAM policy examples.
If your account's default EBS encryption key is customer-managed, grant the required access to that specific KMS key as well. EBS encryption permissions
Arrange a recovery login
Choose at least one way to reach a stopped run after restarting it:
- SSH: create or import an EC2 key pair in the same region, keep the private key, and enter its exact key-pair name. Use the login user supplied by your AMI, such as
ubuntufor the Ubuntu DLAMI. - Systems Manager: supply an existing instance profile with the permissions your SSM setup requires. The AMI needs a running SSM Agent and network access to the SSM endpoints. Enter the instance profile ARN, not the role ARN. SSM instance-profile setup
Training itself does not require an instance profile. If you configure one, add this statement to the control-plane user's policy, replacing ROLE_NAME with the role inside that profile:
{
"Effect": "Allow",
"Action": "iam:PassRole",
"Resource": "arn:aws:iam::ACCOUNT_ID:role/ROLE_NAME",
"Condition": {"StringEquals": {"iam:PassedToService": "ec2.amazonaws.com"}}
}The role can have a path; use its complete ARN. Keep the profile's permissions limited because code on the VM can use them. Passing an EC2 instance role
5. Enter the connection in Tensorant
An organization owner opens Settings → Connections, chooses AWS under compute, and fills in:
| Tensorant field | What to enter |
|---|---|
| Access key ID | The dedicated IAM user's access key ID |
| Secret access key | Its secret access key |
| Session token (optional) | Required only when the key pair is temporary STS credentials |
| Region | Region code, for example us-east-1 |
| Subnet ID | The prepared private subnet's subnet-… ID |
| Security group ID | The same-VPC security group's sg-… ID |
| GPU AMI ID | The verified regional ami-… ID |
| SSH key name (optional) | An existing EC2 key-pair name, not a key file or public-key text |
| Instance profile ARN (optional) | The existing SSM profile's full ARN, if used |
Save the connection. Tensorant checks credentials, image metadata, subnet/security-group compatibility and the available GPU/pricing catalog. Saving does not launch a GPU or prove that the image boots, Docker works, quota is sufficient or outbound networking is functional.
Temporary credentials must remain valid until all runs and retained resources are cleaned up. Tensorant does not renew STS sessions. Rotate credentials before expiry, using an identity that can still see and manage the existing instances.
6. Verify a short training run
- Connect storage, create a project and prepare a small dataset version.
- Open Configure a run, choose Training provider → AWS, then choose an eligible Training GPU.
- Choose Training trial · a small sample. Set an affordable Price limit ($/hour) and enough time for the first container/model download.
- Save the draft, review the launch confirmation, and launch. This starts a paid EC2 instance.
- Confirm the run progresses through preparation, training and adapter verification. Check both GPU released in Tensorant and terminated in EC2, with the job's root EBS volume removed.
The displayed hourly estimate includes the on-demand Linux instance and gp3 capacity, using the monthly storage rate divided by 730. It excludes NAT, transfer, taxes and account-specific discounts. The launch limit is not an AWS billing cap. See training runs for trial settings and cost controls.
Recovery, cleanup and disconnection
If export fails or the run times out, Tensorant stops the EC2 instance and retains its root disk for recovery. Instance compute charges stop when EC2 confirms stopped; EBS and your network infrastructure still incur charges. EC2 lifecycle and billing
Open Recover your adapter on the run. If you do not already have its completed archive, restart the retained instance in EC2 and connect using SSH or SSM. With custom S3, download /var/lib/tensorant/data/adapter.zip; with RunPod network storage, the host path is /var/lib/tensorant/data/runs/RUN_ID/adapter.zip. The host's data directory is /workspace/data inside the container. Select the downloaded archive and choose Verify and recover. Restarting resumes instance billing but does not automatically resume training. An incomplete checkpoint is not a completed adapter archive. See failure and recovery.
Stop the instance again while investigating. Once you have recovered the output, or explicitly decided to discard it, use the run's recovery/cancellation controls to release it. Termination deletes the job's root disk and private network interface. It does not delete your AMI, subnet, security group, key pair, NAT gateway or VPC.
Keep the connection and its permissions valid until Tensorant confirms cleanup. If release is pending, inspect the instance in EC2 and use Retry GPU release. Do not remove tags or replace credentials with an identity from another account while a run is retained.
Troubleshooting
| Symptom | Check |
|---|---|
Connection fails with AccessDenied |
The policy covers all discovery calls and pricing:GetProducts; account policies permit the pricing call in us-east-1. |
| GPU catalog is empty | The subnet's availability zone offers a supported single-GPU type, the region is correct, and AWS returns public Linux pricing. |
| Launch reports a vCPU limit | Increase the relevant G or P on-demand quota in that region. |
| Launch reports insufficient capacity | Try later or configure an approved subnet in another supported zone after existing resources are released. |
| Image validation fails | Check the AMI is available, x86_64, Linux/UNIX, EBS-backed and has no Marketplace product code. |
| Startup fails during host validation | Boot the configured image independently and verify its GPU driver is at least 580. |
| Docker setup, image download or model download fails | Verify Docker, NVIDIA Container Toolkit, DNS and the private subnet's NAT/firewall route. |
| Cannot reach the retained VM | A stopped instance must be restarted; confirm the SSH private route/key/user or SSM Agent/profile access. |
| GPU release fails | Restore credentials and lifecycle permissions, check EC2 state, then retry release. A stopped VM still retains billable storage. |
Need a hand? Visit troubleshooting