Azure
Prepare Azure GPU images and networking, create a scoped service principal, and connect training and recovery.
Connect Azure to train on one NVIDIA GPU in your subscription. You supply a service principal, a resource group, a subnet and a prepared GPU image. Tensorant creates a private VM for each run, transfers data through your storage connection, and removes the VM after verifying the adapter.
Connect custom S3 storage separately, or keep existing RunPod network storage. An Azure subscription does not make Azure Blob Storage an S3 endpoint. Managed inference, publishing and managed experiment evaluation continue to require RunPod and a RunPod network volume.
1. Choose a subscription, region and GPU quota
Use an Azure public-cloud subscription. This connector uses public Azure's management and identity endpoints; sovereign-cloud subscriptions are not supported.
Create a dedicated resource group for Tensorant training resources. Choose a region with a supported single-GPU VM size and record its location code, such as eastus, rather than the display name "East US".
In Subscriptions → your subscription → Usage + quotas, check both Total Regional vCPUs and the chosen VM-family quota in that region. Request enough vCPUs for the whole VM and any other concurrent usage. Quota approval does not reserve GPU capacity. Azure vCPU quotas
The catalog is limited to explicitly supported full single-GPU sizes, further filtered by your subscription's regional restrictions and available prices. A100 and H100 sizes are useful starting points; the selected GPU must fit your model. Multi-GPU, fractional-GPU and ARM machines are not supported. Use a GPU and image compatible with the current Ampere-or-newer trainer.
Ask your subscription administrator to register Microsoft.Compute and Microsoft.Network under Subscriptions → Resource providers if they are not already registered. Registration and quota changes are administrator setup tasks, not permissions you need to give Tensorant.
2. Prepare the subnet and a recovery route
- Create or choose a VNet and subnet in the same subscription and region as the training VMs. They may be in a different resource group.
- Configure explicit outbound internet access. For example, create a NAT gateway with a public IP address and associate it with the training subnet.
- Check the subnet's network security group and route table. Permit the required DNS access and outbound HTTPS for container registries, model downloads and your S3 service.
- Record the subnet's complete resource ID. You can copy it from the subnet's properties or retrieve it with the Azure CLI:
az network vnet subnet show \
--resource-group NETWORK_RESOURCE_GROUP \
--vnet-name VNET_NAME --name SUBNET_NAME \
--query id --output tsvReplace the uppercase names with your existing resources. Tensorant creates a private network interface without a public IP. It does not create a VNet, subnet, NAT gateway, NSG rules or firewall routes. Azure NAT gateway setup, NIC subscription and region requirements
Training requires no inbound ports. Arrange private administration access for recovery, such as a VPN or an existing Azure Bastion deployment. If you use SSH, allow port 22 only from that administration route. Create an Ed25519 or RSA key pair and keep the private key outside Tensorant:
ssh-keygen -t ed25519 -f ~/.ssh/tensorant-azureThe connection takes the contents of tensorant-azure.pub. RSA keys must be at least 2048 bits. Run VMs use the Linux account tensorant.
3. Prepare a reusable GPU image
Tensorant accepts either a managed image resource ID or a specific Azure Compute Gallery image-version resource ID. A Marketplace URN, VM ID, disk ID, image-definition ID without a version, or latest is not accepted.
The image must have all of these properties:
| Requirement | What to prepare |
|---|---|
| Operating system | Generalized Linux, x64, Generation 2 |
| Security compatibility | A Standard-security image; do not require Trusted Launch or Confidential VM settings |
| Disks | One OS disk, at most 1023 GiB, without inherited data disks |
| Licensing | No Marketplace purchase plan |
| Provisioning | cloud-init enabled, Python 3 and working Azure provisioning |
| GPU runtime | Docker, NVIDIA Container Toolkit and a working driver 580 or newer for CUDA 13 |
| Region | Managed image in the training region, or gallery version replicated there |
If your organization already publishes an image meeting these requirements, test that image on the chosen GPU and use its full resource ID. For a new setup, the following managed-image workflow gives you a concrete starting point. Azure also supports Compute Gallery image versions for a managed image pipeline.
Build and test the source VM
Create a temporary builder VM in Azure using an official Ubuntu 24.04 x64 Gen2 image, Standard security, a supported GPU size and the prepared subnet. Attach only its OS disk. This VM incurs Azure charges and is managed by you, separately from Tensorant.
Install the GPU driver appropriate for its family using Microsoft's N-series Linux driver instructions. The resulting driver must be version 580 or newer. For NVads A10, use a supported Azure GRID driver that meets that requirement; an older GRID image is not sufficient. Do not assume a generic CUDA driver or a preinstalled driver works for every Azure GPU family.
Install Docker Engine for Ubuntu and NVIDIA Container Toolkit, then configure the runtime:
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
uname -m
python3 --version
cloud-init --version
nvidia-smi --query-gpu=name,driver_version --format=csv,noheader
sudo docker run --rm --gpus all ubuntu:24.04 nvidia-smiConfirm x86_64 architecture and a GPU visible both on the host and inside Docker. Also confirm downloads work from the private subnet. Preserve the Azure guest agent and cloud-init configuration from the official Ubuntu image. Preparing Ubuntu images for Azure
Generalize and capture the image
Only do this to the temporary builder after your tests pass. Generalization removes the provisioning account and makes the VM unsuitable for restarting. Remove any build keys, credentials, shell history and application data first; deprovisioning does not guarantee every secret is removed.
On the builder VM, prepare cloud-init for a new instance and deprovision it:
sudo cloud-init clean --logs --seed
sudo waagent -deprovision+user
exitThen run these commands from your administrator workstation or Azure Cloud Shell, replacing the example resource names:
az vm deallocate --resource-group IMAGE_RESOURCE_GROUP --name BUILDER_VM
az vm generalize --resource-group IMAGE_RESOURCE_GROUP --name BUILDER_VM
az image create --resource-group IMAGE_RESOURCE_GROUP \
--name tensorant-gpu-v1 --source BUILDER_VM --hyper-v-generation V2
az image show --resource-group IMAGE_RESOURCE_GROUP \
--name tensorant-gpu-v1 --query id --output tsvSave the returned ID. Wait for image creation to finish, then remove the temporary builder and its unused resources through Azure. Keep the image. This workflow requires a Standard-security source VM: legacy managed images cannot be captured from Trusted Launch VMs. Generalizing a Linux VM, capturing a Gen2 managed image
For a Compute Gallery, publish a generalized x64 Gen2 version with the same contents and no data disks, replicate it to the training region, and use the full version ID ending in /versions/1.0.0 or another explicit version. Do not supply only the gallery image definition. Azure Compute Gallery image versions
4. Register an application and create its secret
- Open Microsoft Entra ID → App registrations → New registration.
- Name it, for example,
Tensorant training, and choose accounts in your organizational directory only. No redirect URI is needed for this connector. - On the application overview, copy Application (client) ID and Directory (tenant) ID.
- Open Certificates & secrets → Client secrets → New client secret. Choose an expiry your team can rotate before it is reached.
- Copy the secret's Value, not its Secret ID. The value is shown only when created.
- Copy the subscription ID from the subscription overview.
The application registration creates a service principal. Azure resource access comes from the role assignments in the next step; adding Microsoft Graph API permissions does not grant VM access. Microsoft's service-principal setup
5. Assign scoped Azure roles
Have an administrator create custom roles and assign them to this application's service principal. Tensorant needs no Owner, role-assignment or subscription-administration permission.
Save this template as tensorant-training-role.json, replacing SUBSCRIPTION_ID. Create the role with az role definition create --role-definition tensorant-training-role.json, then assign Tensorant Training Resources to the service principal at the dedicated training resource group, not the whole subscription.
{
"Name": "Tensorant Training Resources",
"IsCustom": true,
"Description": "Manage GPU training VMs and their disks and private interfaces.",
"Actions": [
"Microsoft.Compute/virtualMachines/read",
"Microsoft.Compute/virtualMachines/write",
"Microsoft.Compute/virtualMachines/delete",
"Microsoft.Compute/virtualMachines/instanceView/read",
"Microsoft.Compute/virtualMachines/start/action",
"Microsoft.Compute/virtualMachines/deallocate/action",
"Microsoft.Compute/disks/read",
"Microsoft.Compute/disks/write",
"Microsoft.Compute/disks/delete",
"Microsoft.Network/networkInterfaces/read",
"Microsoft.Network/networkInterfaces/write",
"Microsoft.Network/networkInterfaces/delete",
"Microsoft.Network/networkInterfaces/join/action"
],
"NotActions": [],
"DataActions": [],
"NotDataActions": [],
"AssignableScopes": ["/subscriptions/SUBSCRIPTION_ID"]
}AssignableScopes controls where the role is available; it does not assign the role. In the training resource group's Access control (IAM) → Add role assignment, select the custom role, choose User, group, or service principal, and select your registered application. The resource-group scope lets this identity manage resources throughout that group, which is why it should be dedicated to training. Azure custom-role setup
Create the following additional narrow roles using the same JSON structure with a distinct Name and the indicated Actions. Assign each at the stated scope:
| Role | Actions | Assignment scope |
|---|---|---|
| Tensorant GPU Catalog | Microsoft.Compute/skus/read |
The compute subscription |
| Tensorant Training Subnet | Microsoft.Network/virtualNetworks/subnets/read, Microsoft.Network/virtualNetworks/subnets/join/action |
The configured subnet |
| Tensorant Managed Image Reader | Microsoft.Compute/images/read |
The configured managed image |
| Tensorant Gallery Image Reader, instead of Managed Image Reader | Microsoft.Compute/galleries/images/read, Microsoft.Compute/galleries/images/versions/read |
The configured gallery image definition, which also covers its versions |
Keep the subscription in AssignableScopes for these role definitions. Assign access explicitly when the image or network is outside the training resource group. Both image-definition and version reads are required for a gallery. These permissions correspond to Azure's compute actions and network actions.
If you use the CLI to assign roles, --assignee-object-id takes the service principal's object ID, not the application's client ID. Your administrator can obtain it with:
az ad sp show --id APPLICATION_CLIENT_ID --query id --output tsvThe service principal used by Tensorant does not need directory-read access to perform training. The lookup above is an administrator setup command. Allow Azure role assignments time to propagate before checking the connection.
6. Enter the connection in Tensorant
An organization owner opens Settings → Connections, chooses Azure under compute, and enters:
| Tensorant field | What to enter |
|---|---|
| Tenant ID | Directory (tenant) ID from the application registration |
| Application (client) ID | Application ID, not either object's Object ID |
| Client secret | The secret's Value |
| Subscription ID | Subscription that will own the run VMs |
| Resource group | Existing dedicated training resource-group name |
| Location | Region code, for example eastus |
| Subnet resource ID | Full ID ending in /virtualNetworks/VNET/subnets/SUBNET |
| GPU image resource ID | Full managed-image ID or gallery image-version ID |
| SSH public key | Entire single-line ssh-ed25519 … or ssh-rsa … public key |
Save the connection. Tensorant authenticates, reads image and subnet metadata, and checks the subscription's supported GPU sizes and current regional prices. It does not start a VM, install drivers, configure networking or verify usable GPU quota when you save.
7. Verify training, export and release
- Connect S3 storage, then create a project and a small dataset version.
- In Configure a run, choose Training provider → Azure and an eligible Training GPU with enough memory.
- Choose Training trial · a small sample, set a Price limit ($/hour), and allow time for initial container/model downloads.
- Save the draft, review the cost confirmation and launch. Azure starts charging for the VM.
- Confirm training and adapter verification succeed. Then verify GPU released in Tensorant and removal of the job's VM, OS disk and network interface in Azure.
The estimate includes the on-demand VM and its Premium SSD OS disk, using the public monthly disk price divided by 730. Network services, transfer, taxes and discounts are outside that estimate. See training runs for the trial flow and launch limits.
Recovery, cleanup and secret rotation
On a timeout or failed export, Tensorant requests deallocation and keeps the OS disk. GPU compute billing stops when Azure confirms Stopped (deallocated). Stopped without deallocation can still incur compute charges. Managed disks and existing networking continue to bill. Azure VM states and billing
Open Recover your adapter on the run. If you need to retrieve the completed archive, start the retained VM in Azure and connect as tensorant through your private route. With custom S3, download /var/lib/tensorant/data/adapter.zip; with RunPod network storage, use /var/lib/tensorant/data/runs/RUN_ID/adapter.zip. The host's data directory is /workspace/data in the container. Select the archive on the run and choose Verify and recover. Starting the VM resumes compute billing; it does not automatically resume training. Deallocate it while resolving any remaining issue. See failure and recovery.
An Azure administrator can also recover from a copy of the retained disk using a helper VM. Preserve the original attachments, and remove helper resources afterward. Tensorant refuses automatic termination if the run VM's disk/NIC attachments were replaced or extra disks were attached, to avoid deleting unrelated data.
After recovery succeeds, or when you explicitly cancel and discard the local output, Tensorant deletes its VM and the original job-created OS disk and NIC. It leaves your resource group, image, subnet, VNet, NSG and NAT gateway intact. Wait for confirmed cleanup before removing the connection or changing accounts.
Rotate the client secret before its expiry: create a new value for the same application, replace the connection, check it, then retire the old secret. Keep access to existing VMs and disks throughout rotation.
Troubleshooting
| Symptom | Check |
|---|---|
| Authentication fails | Tenant and client IDs are correct; the secret is its Value, belongs to that app and has not expired. |
| Catalog returns authorization failure | Microsoft.Compute/skus/read is assigned at subscription scope. |
| Image or subnet returns forbidden | The reader/join role is assigned on the actual resource, including resources outside the training group. |
| Image is rejected | It is generalized x64 Linux Gen2, uses one OS disk, has no Marketplace plan, and the ID names a managed image or explicit gallery version. |
| VM cannot use the image in this region | The managed image is in that region, or the gallery version's replication there is complete. |
| Quota or allocation failure | Both total regional and family vCPU quotas suffice; retry later if the region lacks capacity. |
| Host validation or container startup fails | Verify driver 580+, Docker, toolkit configuration and the image's cloud-init behavior on the selected GPU. |
| Preparing never reaches training | Check DNS, NAT/firewall routing and access to container/model/storage hosts from the private subnet. |
| GPU release fails | Check role/secret validity and Azure activity logs, restore the original disk/NIC attachments, then retry release. |
Need a hand? Visit troubleshooting