Skip to content
Quickstart

Get started

Quickstart

From sign-up to calling your own fine-tuned model, in seven short steps.

This is the shortest path from a new account to a model your application calls. Each step takes a few minutes of your time and links to the full page if you want more.

Before you start

  • A RunPod account with a payment method. Your GPUs and storage are billed there.
  • Some data: documents to learn from, or example questions and answers, as files, pasted text or a Hugging Face dataset.
  • A base model from the supported list, such as a Qwen or Llama chat model.
  • Optional: an OpenAI-compatible model endpoint, if you want to generate examples from documents.

1. Sign up

  1. Choose Get Started at the top of this page and sign up with GitHub, Google, or your email address and a password.
  2. Enter your Organization name and choose Create organization.
  3. The console opens right away. You are the owner of your new organization.

Joining a colleague's organization instead? Open the invitation link they sent you. See Accounts, organizations and roles.

2. Connect RunPod and storage

An owner does this once for the organization.

  1. In RunPod, create an API key with read and write access. In the console, choose Open Connections, then Set up compute, paste the key and choose Connect RunPod.
  2. In RunPod, create a network volume in a datacenter that offers the S3 API, and create an S3 API key.
  3. Choose Set up storage, enter the volume ID, datacenter and S3 keys, and choose Connect storage.

Both are tested before they are saved. See Connections for where to find each value.

3. Create a project

  1. Choose Create project.
  2. Pick a Use case, such as Support assistant, or Start blank.
  3. Enter a Project name and the Hugging Face base model as organization/model-name, then choose Create project.

See Projects.

4. Add and review your data

  1. Open Sources, choose Add sources, and upload files, paste text or import from Hugging Face.
  2. Working from documents? Select them and choose Generate examples to have a model write questions and answers. See Recipes.
  3. Open Review and approve or reject each example. The A and R keys make this quick.
  4. Open Versions and choose Create version to freeze the approved examples.

See Sources, Review and Versions.

5. Train and compare

An experiment scores the base model, trains your adapter, scores it again and shows you the difference.

  1. Open Experiments and choose New experiment.
  2. Pick your version, the GPU and its hour and price limits, and how answers are scored.
  3. Tick Keep the tuned model running after the comparison if you want to try it in the Playground.
  4. Choose Check readiness, then Save draft, then Review and launch.
  5. Check the cost estimate and choose Approve and start.

When it completes, Compare answers shows the base and tuned answers side by side. See Experiments.

6. Try it in the Playground

Open Playground, choose your tuned model, type a prompt and choose Send. You can pick a second model to compare side by side. See Playground.

7. Publish and call your model

Only owners can publish models and create API keys.

  1. Open API and choose Publish a model.
  2. Enter a Name. Your application sends this as model.
  3. Choose your trained adapter, then Wake on request (pay only while it runs) or Always on (a dedicated GPU). Choose Review, then Publish (Publish and start for an always-on model).
  4. Under API keys, choose Create key. Copy the key now: it is shown only once.
  5. Call your model. Use the address shown at the top of the API screen:
bash
curl https://YOUR_API_ADDRESS/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "your-model-name", "messages": [{"role": "user", "content": "Hello"}]}'

A wake-on-request model that is asleep may first answer 503 with a Retry-After header while its GPU starts. Wait that many seconds and try again. See Publishing models and the API quickstart for Python and JavaScript.

What's next

  • Invite your team from Settings → Members → Invite member. See Members and invitations.
  • Run more experiments with different data or settings, and compare the results.

When something goes wrong

  • A button is greyed out. Your role may not allow it. See roles.
  • "This organization has no connections yet". An owner needs to finish step 2.
  • A readiness check fails. Its message names the problem, such as no GPU capacity in your volume's datacenter or a price above your limit. Choose another GPU or raise the limit.
  • Anything else: see Troubleshooting.