Lambda
Connect Lambda Cloud with an API key, registered SSH key, region and compatible GPU image.
Tensorant launches single-GPU training instances in your Lambda Cloud account and exports adapters to your organization's storage. Setup needs an API key, registered SSH key, region, and an available image compatible with Tensorant's CUDA 13 container.
Lambda cannot stop or suspend these instances. A failed run retained for file recovery continues to incur GPU charges until termination. Configure SSH access before your first run.
An Owner can save connections in Tensorant. Open your name at the foot of the sidebar, then Settings → Connections under Organization. If it says Provided by Tensorant, ask your platform administrator about using your own account.
1. Prepare your account and billing
- Sign up or sign in to the Lambda Cloud console.
- Select the workspace you intend to use. Keep the API identity, SSH key, and compute access consistent with that workspace.
- Open Billing and add a valid credit card. Lambda performs a card pre-authorization; see its billing setup guide for current requirements.
- Confirm that your account can launch an on-demand GPU and has sufficient quota. New accounts can have lower limits; see instance creation.
You do not need to rent an instance or create a Lambda filesystem during setup. Tensorant uses the root disk as working space and exports data to your connected storage.
2. Create a Cloud API key
- In Lambda, open Cloud API keys.
- Choose Generate API Key, enter a name such as
tensorant-training, and generate it. - Copy or download the key to a secure location. It cannot be retrieved after the dialog closes; see Lambda's console instructions.
Your API identity must be allowed to list images, SSH keys, and instance types; launch and inspect instances; and terminate them. Keep the key valid throughout recovery and cleanup. Tensorant's read-only connection check cannot prove launch or termination permissions.
3. Register an SSH key
- In the selected Lambda workspace, open SSH keys → Add SSH Key.
- Add an existing public key or use Lambda's key-generation option. If generating one, download and securely retain the private key.
- Name the key, for example
tensorant-recovery, and save it. Record its exact name for Tensorant. - Keep the private key on your computer. Tensorant takes the key's name, not its contents. See Lambda SSH setup.
To generate a key locally, use ssh-keygen -t ed25519, choose an unused file name, and register the .pub file. Do not overwrite an existing private key.
4. Find a compatible image and region
GPU image ID is required. It must identify an image currently returned by Lambda's image API for your region, with architecture x86_64.
- Open the official Lambda API browser and authenticate with your API key.
- Use List available instance types to find a suitable single-GPU type and region in
regions_with_capacity_available. Record the region'sname, such asus-east-1, rather than its display label. - Use List available images, the read-only
GET /api/v1/imagesendpoint. - In the returned
datalist, find an image with your chosenregion.nameandarchitectureofx86_64. Copy itsid. - Confirm that this exact image includes every host requirement below. Ask Lambda support if its image description does not establish driver and container-runtime versions.
These lookups do not launch compute. Lambda offers Lambda Stack, GPU Base, and plain Ubuntu image families; see available base images.
| Host requirement | Purpose |
|---|---|
| x86_64 Linux, cloud-init, systemd, and Python 3 | Runs launch configuration and supervision |
| Docker | Runs the training container |
| NVIDIA Container Toolkit | Gives the container GPU access |
| NVIDIA driver 580 or newer, compatible with CUDA 13 | Required and checked by Tensorant's training bootstrap |
An image name or the presence of Lambda Stack alone does not establish CUDA 13 support. The connection check validates image identity, region, and architecture; it cannot inspect installed packages without a running VM. Tensorant does not install or upgrade missing host drivers or Docker during startup.
If no listed regional image meets these requirements, ask Lambda for a suitable image or use another provider. An arbitrary custom-image name, AWS AMI ID, or image from another region will not work. This integration selects images exposed by Lambda's API; it does not import custom VM images.
5. Save the connection
In Tensorant's Settings → Connections, choose Connect Lambda and enter:
| Tensorant field | Value |
|---|---|
| API key | Your Lambda Cloud API key |
| Region | Exact API region name from step 4 |
| SSH key name | The existing key name, including case |
| GPU image ID | The image's id, not its family or display name |
Choose Connect Lambda and wait for Connected. The check reads the image, key, and GPU catalogs without launching a VM. Updates require reentering the key and choosing Save connection.
6. Connect storage
Follow the S3 guide for a dedicated bucket or use RunPod network storage. Lambda transfers data to either backend; a RunPod volume is not attached to its machine. Tensorant does not attach Lambda filesystems.
Your organization's own storage credentials are supplied to the training instance for transfers. Scope them appropriately and allow for cross-cloud transfer costs. This integration covers training; managed serving currently uses RunPod.
7. Verify a training trial
- Create a project and dataset version using the quickstart.
- Open Tune → Runs → Configure a run and choose Training trial · a small sample.
- Select the dataset, set Training provider to Lambda, and select an available Ampere-or-newer Training GPU with enough memory. T4 and other pre-Ampere GPUs cannot run the trainer.
- Review Price limit ($/hour) and trial runtime. Tensorant lists compatible single-GPU types with sufficient local disk and capacity in your region; multi-GPU and ARM instances are excluded.
- Choose Save trial draft. Saving does not rent compute.
- When ready for a billable test, choose Launch trial, review the confirmation, and confirm.
- Watch Training logs, then verify that the adapter is saved and the instance terminated. This first run checks image startup as well as training and storage transfers.
Price is checked again at launch. The hourly limit is not a cap on every cloud charge, and listed capacity is not a reservation.
Recovery and billing
After successful export, Tensorant verifies the adapter and requests termination. A recoverable failure keeps the instance alive for file retrieval. Its GPU stays billable until termination. Lambda supports launch, restart, and termination but no stop/suspend action; shutting down Linux does not end billing. See instance lifecycle and billing rules.
- Find the run's instance in Lambda and connect with its SSH instructions and your private key.
- Retrieve
/var/lib/tensorant/data/adapter.zipfrom the host. With RunPod storage, use/var/lib/tensorant/data/runs/<run-id>/adapter.zip. Inside the container, the corresponding root is/workspace/data. - Upload the archive under Recover your adapter in Tensorant and choose Verify and recover.
- Wait for verification and termination, then confirm that the instance is no longer running in Lambda.
If the output is not needed, use Cancel run and confirm termination. Termination destroys local files, so retrieve needed data first. Do not leave failed runs expecting an automatic suspension: a timeout or browser close cannot stop a Lambda GPU.
Common problems
| Problem | What to check |
|---|---|
| API key rejected | Check the key, account/workspace access, and revocation status. |
| SSH key not found | Enter the exact key name, not its ID or public-key text. Check the workspace. |
| Image unavailable | Use a current image ID with matching region and x86_64 architecture. |
| Driver or Docker startup error | Check every image prerequisite above. Changing the GPU does not repair the image. |
| No matching instance | Check capacity, memory, fixed root-disk size, hourly ceiling, quota, and billing setup. |
| HTTP 429 | Wait before rechecking. Lambda rate-limits requests; avoid repeated clicks and parallel polling from other tools. |
| Cleanup pending or failed | Restore API access and retry cleanup. Check Lambda directly because billing continues until termination is confirmed. |
Need a hand? Visit troubleshooting