Alias
You can usesls as a shorthand for serverless:
Subcommands
List endpoints
List all your Serverless endpoints:List flags
bool
Include template information in the output.
bool
Include workers information in the output.
Get endpoint details
Get detailed information about a specific endpoint:Get flags
bool
Include template information in the output.
bool
Include workers information in the output.
Create an endpoint
Create a new Serverless endpoint from a template or from a Hub repo:--hub-id, GPU IDs and container disk size are automatically pulled from the Hub release config. You can override the GPU type with --gpu-id. Environment variables from the Hub release are included automatically, and you can override or add to them with --env.
Serverless templates vs Pod templates: Serverless endpoints require a Serverless-specific template. Pod templates (like
runpod-torch-v21) cannot be used because they include configuration, which Serverless does not support. When creating a template with runpodctl template create, use the --serverless flag to create a Serverless template.Each Serverless template can only be bound to one endpoint at a time. To create multiple endpoints with the same configuration, create separate templates for each.Create flags
string
Name for the endpoint. Must be at least 3 characters. If omitted, a name is auto-generated in the format
endpoint-XXXXXXXX.string
Template ID to use (required if
--hub-id is not specified). Use runpodctl template search to find templates.string
Hub listing ID to deploy from (alternative to
--template-id). Use runpodctl hub search to find repos.string
GPU type for workers. Accepts either a GPU type ID (e.g.,
NVIDIA A40, NVIDIA GeForce RTX 4090) or a GPU pool ID (e.g., ADA_24, AMPERE_48). Use runpodctl gpu list to see available GPUs.int
default:"1"
Number of GPUs per worker.
string
default:"GPU"
Compute type (
GPU or CPU). For CPU endpoints, use --instance-id to specify the CPU instance type.string
default:"cpu3g-4-16"
CPU instance ID when using
--compute-type CPU. If omitted, defaults to cpu3g-4-16. Only valid with --compute-type CPU.int
default:"0"
Minimum number of workers.
int
default:"3"
Maximum number of workers.
string
Comma-separated list of preferred datacenter IDs. Use
runpodctl datacenter list to see available datacenters.string
Network volume ID to attach for single-region deployments. Use
runpodctl network-volume list to see available network volumes. Mutually exclusive with --network-volume-ids.string
Comma-separated list of network volume IDs for multi-region deployments. Mutually exclusive with
--network-volume-id.string
Minimum CUDA version required for workers (e.g.,
12.4). Workers will only be scheduled on machines that meet this CUDA version requirement.string
Autoscaling strategy:
delay (scales based on queue wait time in seconds) or requests (scales based on pending request count).int
Trigger point for the autoscaler. For
delay, this is the target queue wait time in seconds. For requests, this is the pending request count that triggers scaling.int
Idle timeout in seconds. Workers shut down after being idle for this duration. Valid range: 1-3600 seconds.
bool
Enable or disable flash boot for faster worker startup. When enabled, workers start from cached container images.
int
Execution timeout in seconds. Jobs that exceed this duration are terminated. The CLI accepts seconds but converts to milliseconds internally.
string
Environment variable in
KEY=VALUE format. Use multiple --env flags to set multiple variables. These values only apply when deploying from --hub-id, where they override the Hub release defaults. With --template-id, environment variables come from the template, so --env is ignored and the CLI prints a note to that effect.string
Model reference URL to attach to the endpoint. Use multiple
--model-reference flags to attach multiple models. Works with both --template-id and --hub-id, and requires GPU compute type.Update an endpoint
Update endpoint configuration:Update flags
string
New name for the endpoint.
string
New template ID to swap to. Use this to change the template attached to an existing endpoint without recreating it.
int
New minimum number of workers.
int
New maximum number of workers.
int
New idle timeout in seconds.
string
Scaler type (
QUEUE_DELAY or REQUEST_COUNT).int
Scaler value.
bool
Enable or disable flash boot for faster worker startup.
int
Execution timeout in seconds. Jobs that exceed this duration are terminated.
Delete an endpoint
Delete an endpoint:Check endpoint health
Get worker counts by state and job counts by outcome for an endpoint. This wrapsGET /v2/<endpoint-id>/health and prints the response verbatim, so new fields returned by the invoke API appear without a CLI update.
Invoke an endpoint
Submit a job to an endpoint and wait for it to finish. The payload must be a JSON object and is sent as{"input": <your JSON>}; pass only the handler payload.
/run and then polled on /status until it reaches a terminal status. The CLI never uses /runsync: /runsync releases the connection after roughly 90 seconds while the job continues running server-side, and until it answers there is no job ID to poll.
The payload is validated as JSON locally before it is sent. Payloads over the invoke API’s 10 MiB /run limit fail as a usage_error without a round trip; the size checked is the body the CLI actually sends (payload compacted and JSON-escaped inside {"input": ...}), so whitespace in an input file does not count against the limit and escaped characters do. If the top-level payload contains a curl-style envelope with an input field alongside policy, webhook, or s3Config, the CLI prints a warning naming those keys because they are ignored when nested inside input.
The job payload is printed on stdout even when the job ends in a FAILED state, because the worker’s own error message is typically the useful artifact. Progress messages and error objects (including the CLI’s JSON error envelope) go to stderr.
Exit codes:
0when the job isCOMPLETED, or when--wait 0/--no-waitsubmitted the job successfully.1when the request fails, when the wait budget runs out, or when the job endsFAILED,CANCELLED, orTIMED_OUT. In every case the last job payload is still printed on stdout.
--wait runs out, the job is still running server-side. The timeout error on stderr names the serverless status command to poll it with.
Run flags
string
JSON payload for the handler. Pass
- to read the payload from stdin. Mutually exclusive with --input-file.string
Path to a file containing the JSON payload. Pass
- to read from stdin. Mutually exclusive with --input.duration
default:"5m"
How long to wait for a terminal job status (for example
90s, 10m). A single API call inside the wait is never given less than one second, so a --wait below one second may overshoot by up to that much. 0 submits and returns without waiting.bool
Submit and print the job ID without waiting. Equivalent to
--wait 0; cannot be combined with an explicit --wait.Check job status
Get the status of a job that was submitted earlier, either byserverless run --no-wait or by a serverless run that hit its --wait budget. By default this checks once and returns; pass --wait to keep polling until the job is terminal.
0when the job isCOMPLETED, or when the job is still queued or running (after a single check with--wait 0).1when the job endsFAILED,CANCELLED, orTIMED_OUT, or when--waitruns out. The job payload is printed on stdout either way.
Status flags
duration
default:"0"
Keep polling until the job is terminal, up to this long.
0 checks once and returns.Errors and exit codes
runpodctl serverless run and runpodctl serverless status print a machine-readable JSON error envelope on stderr with a stable code field. The code values these commands can emit are:
The
timeout code is also emitted by runpodctl model add --wait-for-hash when its wait budget runs out. Previously that condition reported cli_error; the exit code and message text are unchanged, but scripts that branch on the stderr JSON code field need to match timeout instead.