Reference
Troubleshooting
Common problems by area, with the message you see and how to fix it.
Most messages in the console say what to do. This page covers the common ones by area. A word in italics stands for a name or number the real message fills in.
Where messages appear
- A refused action shows "Action could not be completed" with the reason, at the bottom right of the screen. In a dialog, the reason also appears in the dialog.
- A background job shows its reason under the job in Activity.
- A training run, deployment, experiment or published model shows its reason on its row or page.
- The API answers with a JSON error. See The API.
If "Workspace could not refresh." appears at the top of a page, reload the page.
Anywhere in the console
| Problem | What to do |
|---|---|
| "You have view-only access to this organization…" or "Your role in this organization cannot make changes" | You are a viewer. Ask an owner for the member role. See Roles. |
| "Only owners can do this" | Ask an owner. Owners manage connections, published models, API keys, members and the plan. |
| "This organization has no connections yet" | An owner connects RunPod and storage in Connections. |
| "Storage unavailable…" | An owner chooses Test storage connection in Connections, and Replace keys if it fails. |
| "Maintenance is active…" | Try again in a few minutes. Reading still works. |
| "Record not found" | It was deleted, or belongs to another organization. Switch to that organization first. |
| "This record overlaps an existing record. Refresh and retry." | Reload the page and try again. |
| "Your plan (Plan) allows N projects…", or the same for members, published models or API keys | Delete or remove one you no longer need, or ask an owner to change the plan. See Plans and billing. |
A job says an unexpected error stopped this
Kind: an unexpected error stopped this. Its details are not shown, because they may contain your data. Retry, or contact support if it happens again.
The details are hidden to protect your data. Check the usual causes, then choose Retry:
| Job | What to try |
|---|---|
| Generate examples | Lower Examples per passage, or use a more capable model. See Recipes. |
| Check example quality | Recheck with another judge, or with Rule checks only. |
| Evaluate model | Check the judge's model and rubric. See Evaluations. |
| Import Hugging Face dataset | Start a new import with fewer rows. See Sources. |
| Any job | Choose Test storage connection in Connections. |
If it happens again, contact support with the time and the job's name.
Signing in and invitations
| Problem | What to do |
|---|---|
| Forgot your password | Use Forgot password? on the sign-in page. |
| "Verify your email address in your account settings, then sign in again." | Add or verify an email address in Settings → Profile, then sign in again. |
| "Another sign-in already uses this email address…" | Sign in the way you first signed up with this address. |
| "This account was disabled…" | Contact support. |
| "You already belong to an organization or asked for one" | Each person can create one organization. Ask an owner of another for an invitation. |
| "Many people are asking for access right now…" | Many organizations were created in the last hour. Try again in an hour. |
| "This link is no longer valid. Ask for a new one." | Ask the owner for a new invitation. |
| "Sign in as email to accept this invitation" | Sign in as the invited address, then open the link again. |
| No organization yet | Ask an owner to invite you. |
| "An organization needs at least one owner" | Make another member an owner first. |
More: Accounts, organizations and roles.
Connections
| Problem | What to do |
|---|---|
| "RunPod did not accept this key…" | Create a RunPod API key with read and write access, and paste it whole. |
| "Connect RunPod compute in Connections first" | Connect compute before storage. |
| "RunPod network volume is unavailable or belongs to a different datacenter" | Check the volume ID and datacenter, and that the volume is in the same RunPod account. |
| "The S3 keys could not reach this volume…" | Create an S3 API key in RunPod and use its access key and secret. |
| "Use an https:// address on port 443…" | Enter an endpoint address such as https://api.example.com/v1. |
| "Endpoint name returned HTTP code" | Check the endpoint's address, model name and key with its provider. |
| "Endpoint name did not answer in time" | Try again, or use a faster model. |
| "This connection cannot be read. Enter it again in Connections" | An owner enters the key, token or endpoint again. |
More: Connections.
Sources and imports
| Problem | What to do |
|---|---|
| "Select between 1 and 20 files per batch" | Upload at most 20 files at a time. |
| "Exceeds the N MB file limit" | Split the file. |
| "Text-based files must use UTF-8 encoding" | Save the file as UTF-8. |
| "No extractable text. Scanned PDFs need OCR before import." | Run OCR on the PDF, or upload its text. |
| "No row can be used: …" | Change the column mapping to fit the file. |
| "Hugging Face access denied…" | Add a Hugging Face token in Connections, and accept the dataset's terms on Hugging Face. |
| "This project's model reads text only…" | Clear Image column (optional), or use a project with a vision model. |
| "Import exceeds the image byte limit…" | Start a new import from the row where it stopped. |
More: Sources.
Generation, review and versions
| Problem | What to do |
|---|---|
| "Select at most 2,000 sources at a time…" | Generate in several rounds. |
| "Finish or rerun quality checks before approving this example." | Wait for the check in Activity, or Recheck with Rule checks only. |
| "Fix schema errors and provide source evidence…" | Fix the answer and add a quote for every assistant turn, then Save & recheck. |
| "Images can't be changed; remove the example instead" | Edit only the text, or reject the example. |
| "Approve examples before creating a version" | Approve examples in Review first. |
| "Automatic splits need 10 independent source groups…" | Add sources from more documents, or import a separate test set. |
| "A version needs train, validation and test examples…" | Import test data, or approve examples from more sources. |
| "Evaluation data overlaps training sources or questions" | Reject the overlapping examples, or delete one of the imports. |
Choosing a model and GPU
These appear when you start an experiment, training run, deployment or published model.
| Problem | What to do |
|---|---|
| "V1 supports public, ungated models…" | The model is private or gated. Use an ungated copy. See Supported models. |
| "Choose a model with a tokenizer chat template" | Choose the model's instruct or chat version. |
| Warning: "Architecture is not one of the supported model families…" | The model may still work. If a run fails, choose a supported model. |
| "Unable to resolve the public model revision and tokenizer…" | Check the model ID and revision on Hugging Face. |
| "Selected GPU has no reported capacity…" | Try later, or choose another GPU. |
| "Selected GPU exceeds the configured hourly-price limit" | Raise the price limit, or choose a cheaper GPU. |
| "GPU memory is below a conservative N GiB estimate…" | Choose a larger GPU, or QLoRA for training. See GPU memory. |
| "Compute duration exceeds the server limit" | Lower the hours to your limit, shown on the Compute card in Connections. |
| "This GPU runs on arm64 machines…" | Choose a GPU other than GH200 or GB200. |
Training runs
| Problem | What to do |
|---|---|
| "Run a baseline evaluation on this version first." | Evaluate the base model on this version. |
| "An active run or unresolved cleanup already uses the compute slot" | Wait for the other run to finish, or cancel it. |
| The run fails with "training stopped; inspect logs" | Read Training logs. Common causes: the GPU ran out of memory or is older than Ampere, or a vision example is longer than Context length. |
| "QLoRA isn't available for vision models yet" | Train vision models with LoRA. |
| "Maximum runtime reached…" | Choose Verify adapter on RunPod. See Needs recovery. |
| "Cleanup failed: …" | Choose Retry cleanup, and check the Pod in RunPod. |
| Spend still says "billing pending" | RunPod's billing can take a while to update. If it still says so a day after the GPU was released, check that your RunPod key has read and write access. See Estimated and actual spend. |
More: Training runs.
Experiments and evaluations
| Problem | What to do |
|---|---|
| "Review and acknowledge launch and dataset coverage warnings" | Tick I reviewed these launch and dataset warnings. |
| "Launch readiness changed…" | Usually a GPU has no capacity now. Choose Review and continue later, or use another GPU. |
| "No model is being served yet…" | Deploy the model and wait for Ready, or add an endpoint in Connections. |
| "Example id: N prompt tokens + M output tokens exceed the serving context length…" | Lower Maximum output tokens, or deploy with a longer context length. |
| "Restore the frozen judge endpoint before running the tuned evaluation" | Set the judge endpoint back to the one the baseline used. |
More: Experiments and Evaluations.
Deployments and the Playground
| Problem | What to do |
|---|---|
| "Stop the existing inference deployment…" | No deployment slot is free. Stop a deployment. |
| "Model did not become ready within 30 minutes…" | Read the deployment's log under inference/ on your volume. Often the model needs a larger GPU or a shorter context length. |
| "…ran out of GPU memory with these settings…" | Lower GPU memory share or Conversations at once, or choose a larger GPU. |
| "Too little GPU memory is left for the context length…" | Lower the context length, or choose a larger GPU. |
| "The context length is longer than this model supports. Lower it." | Lower the context length. |
| "Idle timeout reached" or "Runtime limit reached" | The deployment stopped at its limits. Deploy again when you need it. |
| "Nothing is running to talk to" | Deploy a model and wait for Ready, or add an endpoint. |
| "You are sending prompts too fast…" | Wait a minute. |
| "Keep the conversation under 200,000 characters" | Choose Clear and start again. |
More: Deployments and Playground.
Published models
| Problem | What to do |
|---|---|
| "This organization already publishes a model with that name" | Choose another name. |
| "No free deployment slot…" | An always-on model needs a deployment slot. Stop a deployment. |
| "Serverless worker limit reached…" | Lower Max workers, or stop another wake-on-request model. |
| "Start limit reached…" | Find out why the starts failed, then raise Starts allowed in 24 hours or wait. |
| "Stop the model before deleting it" | Choose Stop first. |
| "Stopped because the organization was disabled or lost its connections." | Reconnect RunPod and storage, then Start the model. |
More: Publishing models.
Publishing adapters to Hugging Face
| Problem | What to do |
|---|---|
| "Add a Hugging Face token with write access…" or "This Hugging Face token can only read…" | An owner adds a token with write access in Connections. |
| "repository is a public repository…" | Choose another repository, or untick Private repository. |
More: Publishing adapters to Hugging Face.
The API
| Error | What to do |
|---|---|
401 invalid_api_key |
Check the key, or create a new one. |
404 model_not_found |
Check the name with GET /v1/models, and the key's May call. |
400 unsupported_parameter |
Remove the field the message names. |
400 model_refused |
Shorten the conversation or lower max_tokens. |
400 model_text_only |
Send images only to a vision model. |
413 request_too_large |
Shorten the request, or send images as links. |
429 rate_limit_exceeded |
Wait the Retry-After seconds, or use a key with higher limits. |
429 quota_exceeded |
Wait for next month (UTC), or ask an owner to change the plan. |
502 upstream_error |
Try again. If the answer timed out, lower max_tokens or stream. |
503 model_starting |
Wait the Retry-After seconds and send it again. |
503 model_unavailable |
An owner starts the model on the API screen. |
Every error: Chat completions.
Plans and billing
| Problem | What to do |
|---|---|
| "This organization has a subscription. Change or renew it under Manage billing." | Use Manage billing. |
| "This organization has no billing account yet. Choose a plan first." | Choose a plan first. |
| "The platform administrator set this organization's plan…" | Contact support to change it. |
More: Plans and billing.
Still stuck
Write to [email protected] with the time, the page, the exact message and any ID the console shows. Never send API keys, tokens or passwords.