RunPod
Connect a RunPod account, create API credentials, and choose network storage or S3 for training.
Connect RunPod to train on GPUs billed to your own account. RunPod also powers Tensorant's managed serving. For an organization using its own connections, managed serving needs both RunPod compute and a RunPod network volume.
You need an Owner role to save connections. Open your name at the foot of the Tensorant sidebar, then Settings → Connections under Organization. If the page says Provided by Tensorant, contact your platform administrator about using your own connections.
1. Prepare your account
- Create an account or sign in to the RunPod console.
- Open Billing, choose an amount, and complete payment. Ordinary RunPod accounts use prepaid credits. Keep enough for compute and any storage you retain; see RunPod billing.
- Choose storage before creating projects. A RunPod network volume supports training and managed serving. Custom S3 supports training across all connected GPU providers.
Saving a compute connection does not rent a GPU. Creating a network volume does create billable storage, even before training starts.
2. Create the compute API key
- Open Credentials in RunPod and select API Keys.
- Choose Create API Key and name it, for example
tensorant-training. - Use Restricted permissions with read/write access for creating, listing, reading, starting, stopping, and deleting Pods. Allow network-volume discovery if using RunPod storage. Add Serverless/template permissions if using managed serving. A read-only key cannot launch or clean up resources.
- Create the key and save its value securely. It is shown once. See RunPod's credential instructions for the current permission controls.
Keep the key valid through recovery and cleanup. An account check cannot prove every write permission because it does not create or delete a Pod.
3. Save the compute connection
- In Tensorant, open Settings → Connections.
- Choose Connect RunPod on the RunPod card.
- Paste the key into RunPod API key, then choose Connect RunPod in the dialog.
- Wait for the check to succeed and the card to show Connected. Use Replace key when rotating an existing key.
The compute API key manages Pods. The separate S3 access key and secret below access your network volume; they are not interchangeable.
Connect storage
Choose one storage backend for the organization. Once it contains projects, its storage location cannot be changed through Connections. Rotating credentials for the same location is supported; moving existing data requires a migration.
Option A: RunPod network volume
- In RunPod, open Storage → New Network Volume.
- Choose a datacenter with the GPUs you need and S3 API support. A volume's location determines where RunPod training can run.
- Enter a name and capacity in GB, choose the storage tier, and select Create Network Volume. Allow room for data, model files, and adapters. Capacity can grow but cannot shrink; see network-volume setup.
- Record the volume's unique ID and datacenter ID, not just its friendly name. RunPod uses the volume ID as the S3 bucket and the datacenter ID as the signing region.
- In Credentials → S3 API Keys, create an S3 API key and save both its access key and secret. See S3 access and endpoints.
- In Tensorant's Data storage card, choose Use RunPod volume. The main Connect S3 storage button opens custom S3 setup.
- Select your Network volume. If needed, choose Enter the volume ID instead and supply its ID and datacenter manually.
- Enter the S3 credentials and choose Connect storage. Wait for the storage check to succeed.
| Tensorant field | Value from RunPod |
|---|---|
| Network volume | The volume you created |
| Network volume ID | Its unique ID, when entering it manually |
| Datacenter | The volume's exact datacenter code, such as EU-RO-1 |
| S3 access key | Your S3 API access key |
| S3 secret key | The corresponding S3 API secret |
Tensorant selects the S3 endpoint from the datacenter. You do not need to rent a Pod or mount the volume manually. Other GPU providers transfer files to this volume over S3 and receive its S3 credentials for transfers; only RunPod mounts it directly.
Option B: Custom S3
Follow the S3 setup guide. RunPod training then uses local working storage and transfers data and adapters to your bucket. No RunPod network volume is required for this training setup. Own-account managed serving currently requires the network-volume option above.
4. Verify a first training trial
- Create a project and dataset version using the quickstart.
- Open Tune → Runs → Configure a run and choose Training trial · a small sample. A trial needs no baseline.
- Select the dataset version, set Training provider to RunPod, and choose an available Ampere-or-newer Training GPU with enough memory. T4 and other pre-Ampere GPUs cannot run the trainer. With network storage, the GPU must be in the volume's datacenter.
- Review Price limit ($/hour) and the trial runtime. Tensorant supplies the training container and requests a CUDA 13-compatible host; you do not choose a Pod template.
- Choose Save trial draft. This saves configuration without renting compute.
- When ready for a billable test, choose Launch trial, review the confirmation, and confirm.
- Watch Training logs, then confirm adapter verification and release of the training Pod.
Availability can change between saving and launching. See training runs for full runs and memory settings.
Recovery and billing
Tensorant verifies the adapter before terminating the Pod. If export fails or a time limit leaves recoverable output, it attempts to stop the Pod and retain files. Storage can still be billed while stopped. Your organization network volume remains after successful runs too; compute deletion does not delete shared storage. See RunPod storage behavior.
With RunPod network storage, choose Recover your adapter → Verify the saved adapter. With custom S3, retrieve /workspace/data/adapter.zip from the Pod and upload it using Verify and recover. Restart a stopped Pod before accessing its local files; restarting resumes GPU billing and depends on capacity.
For manual shell access, register your public key in RunPod's SSH Public Keys tab using the credential guide, then follow the Pod's SSH connection instructions. Keep the private key on your computer. Basic SSH does not support SCP or SFTP; those transfers need a full SSH service and exposed port, which Tensorant does not configure automatically.
Check cleanup status in Tensorant and RunPod after recovery or cancellation. Retrieve any needed files before deleting their only copy. Retained storage requires sufficient account credit.
Common problems
| Problem | What to check |
|---|---|
| API key rejected or HTTP 401/403 | Use the compute key, check its permissions, and replace revoked credentials. S3 credentials cannot manage Pods. |
| Storage check fails | Check the volume ID, exact datacenter, S3 support, and separate S3 key pair. |
| No suitable GPU | Check memory, capacity, location, and hourly ceiling. A network volume constrains RunPod placement to its datacenter. |
| Training fails after startup | Read Training logs for model-access, dataset, memory, or storage errors before launching again. |
| Cleanup pending or failed | Restore account access and retry cleanup. Verify the Pod state in RunPod; closing the browser does not release it. |
Need a hand? Visit troubleshooting