The identity and technical contract details declared by the specification.
LargeShmRequest
booleannull
A large /dev/shm device to mount into a container running the created workload. An shm is a shared file system mounted on RAM.
SecretFieldsUpdatable
objectnull
2 properties
EnvironmentVariableConfigMap
objectnull
Details of the configMap and key use to populate the environment variable
2 properties
ConfigMapInstance
objectnull
WorkloadId2
string
A unique ID of the workload.
DefaultMode
stringnull
File permission mode in octal string format. This value must be a 4-digit octal number, representing the default file mode when mounting a Secret or ConfigMap…
CpuMemoryRequest
stringnull
The amount of CPU memory to allocate for this workload (1G, 20M, .etc). The workload will receive at least this amount of memory. Note that the workload will n…
GpuDevicesRequest
integernull
Requested number of GPU devices. Currently if more than one device is requested, it is not possible to provide values for gpuMemory or gpuPortion.
PriorityClass
stringnull
Specifies the priority class for the workload, which determines its scheduling behavior. Valid values are: very-low, low, medium-low, medium, medium-high, high…
Label
objectnull
Label details to be populated into the container.
3 properties
DepartmentId2
string
The id of the department.
DistributedInferenceStartupPolicyField
objectnull
1 property
NodeSelectorTerm
objectnull
A null or empty node selector term matches no objects. The requirements of them are ANDed.
1 property
ImagePullSecrets
arraynull
A list of references to Kubernetes secrets in the same namespace used for pulling container images.
PvcVolumeMode
stringnull
Default volume mode for the PVC. Choose between Filesystem (default) or Block.
TolerationEffect
stringnull
The taint effect to match. (mandatory)
Annotation
objectnull
Annotation details to be populated into the container.
3 properties
NodeType3
stringnull
Nodes (machines), or a group of nodes on which the workload will run. To use this feature, your Administrator will need to label nodes. For more information, s…
PvcClaimSize
stringnull
Requested size for the PVC. Mandatory when existingPvc is false. Recommended sizes: TB/GB/MB/TIB/GIB/MIB
EnvironmentVariablePodFieldReference
objectnull
Details of the field-reference and key use to populate the environment variable
1 property
Labels
arraynull
Set of labels to populate into the container running the workload.
PvcFieldsNonUpdatable
objectnull
6 properties
DistributedInferenceSpecSpec
Preemptibility
stringnull
Specifies whether the workload can be preempted by higher-priority workloads. Valid values are preemptible and non-preemptible. If explicitly set, this value t…
PvcAddedAttrValues
array
an optional array of key-values pairs that are written as annotations on the created PVC. the allowed attributes are determined according to the storage class…
RunAsUid
integernull
The user id to run the entrypoint of the container which executes the workspace. Default to the value specified in the environment asset runAsUid field (option…
StorageInstanceName
objectnull
1 property
WorkloadCreationMeta
object
4 properties
3 required
WorkloadName
string
The name of the workload.
Image
stringnull
Docker image name. For more information, see [Images](https://kubernetes.io/docs/concepts/containers/images). The image name is mandatory for creating a worklo…
SeccompProfileType
stringnull
Indicates which kind of seccomp profile will be applied to the container. The options are a. RuntimeDefault - the container runtime default profile should be u…
EnvironmentVariable
objectnull
Details of an environment variable which is populated into the container.
8 properties
DistributedInferenceStartupPolicy
stringnull
Determines when the worker pods should start during workload initialization. - LeaderCreated: Workers start after the leader pod is created. - LeaderReady: Wor…
PvcAccessModes
objectnull
Default access mode(s) applied to newly created PVCs unless explicitly overridden.
3 properties
EnvironmentVariables
arraynull
Set of environment variables to populate into the container running the workload.
Probe
objectnull
6 properties
PodAffinity
objectnull
Pod affinity scheduling rules (e.g. co-locate this workload in the same node, zone, etc. as some other workloads).
2 properties
CreateHomeDir
booleannull
When set to true, creates a home directory for the container.
UpdateSpec
object
The specifications of the inference to be updated.
1 property
ProjectId
string
The id of the project.
DistributedInferenceReplicasField
object
1 property
Tolerations
arraynull
Set of tolerations to apply to the workload.
ExtendedResource
objectnull
Quantity of an extended resource.
3 properties
DistributedInferenceServingPort
objectnull
Defines the configuration for the inference serving endpoint. This determines how applications or services can send inference requests to the workload.
DistributedInferenceServingPortProtocol
stringnull
The protocol used to access the port.
PvcItems
arraynull
Set of pvc persistent volume claims to use in the workload.
DistributedInferenceLeaderWorkerSpec1
object
18 properties
Args
stringnull
Arguments to the command that the container running the workload executes.
DistributedInferenceServingPortAccess
object
5 properties
EnvironmentVariableSecret
objectnull
Details of the secret and key use to populate the environment variable
2 properties
Command
stringnull
A command to the server as the entry point of the container running the workload.
Category
stringnull
Specify the workload category assigned to the workload. Categories are used to classify and monitor different types of workloads within the NVIDIA Run:ai platf…
SecretInstance2
objectnull
GpuMemoryRequest
stringnull
Required if and only if gpuRequestType is memory. States the GPU memory to allocate for the created workload, per GPU device. Note that the workload will not b…
DistributedInferenceServingPortContainerAndProtocol
object
2 properties
ProbeHandlerScheme
stringnull
Scheme to use for connecting to the host, defaults to HTTP.
EmptyDir
objectnull
3 properties
UidGidSource
stringnull
Indicate the way to determine the user and group ids of the container. The options are a. fromTheImage - user and group ids are determined by the docker image…
MatchExpression
objectnull
A selector that contains values, a key, and an operator that relates the key and values.
3 properties
2 required
ExtendedResources
arraynull
Extended resources and their quantity.
EmptyDirItems
arraynull
A list of emptyDir volumes to mount in the workload.
NodePools
arraynull
A prioritized list of node pools for the scheduler to run the workload on. The scheduler will always try to use the first node pool before moving to the next o…
RunAsGid
integernull
The group id to run the entrypoint of the container which executes the workspace. Default to the value specified in the environment asset runAsGid field (optio…
SecretItems1
arraynull
Set of secret volumes to use in the workload
MatchExpressionOperator
string
Represents a key's relationship to a set of values (mandatory).
ProbeHandler
objectnull
The action taken to determine the health of the container. (mandatory)
1 property
SecretFieldsNonUpdatable
objectnull
1 property
WorkloadDesiredPhase
string
The desired phase of the workload.
ClaimInfo
objectnull
Claim information for the newly created PVC. The information should not be provided when attempting to use existing PVC.
5 properties
DistributedInferenceLeaderSpecFields
object
1 property
DistributedInferenceCreationRequest
PvcAddedAttrValue
object
2 properties
1 required
GpuRequestType
stringnull
Sets the unit type for GPU resources requests. Stated in terms of portion or memory. Sets the unit type for other GPU request fields. If gpuDevicesRequest 1, o…
TolerationOperator
stringnull
A key's relationship to the value. Equal uses key and value. Exists is equivalent to wildcard for value, so that a workload can tolerate all taints of a partic…
DistributedInferenceServingPortAccessAuthorizationTypeEnum
stringnull
Specifies who can send inference requests to the serving endpoint: Possible values: - public: No authorization is required. (Default) - authenticatedUsers: Any…
SubmitWithTemplateId
object
1 property
ConfigMapItems
arraynull
Set of config map volumes to use in the workload
DistributedInferenceWorkerSpecFields
object
1 property
HttpResponse
object
2 properties
2 required
EnvironmentVariableUserCredential
objectnull
Defines a reference to a user-created credential and a specific key within that credential whose value will populate the environment variable. User credentials…
2 properties
Annotations
arraynull
Set of annotations to populate into the container running the workload.
ExcludeField
objectnull
1 property
ReadOnlyRootFileSystem
booleannull
If true, mounts the container's root filesystem as read-only.
CpuCoreRequest
numbernull
CPU units to allocate for the created workload (0.5, 1, .etc). The workload will receive at least this amount of CPU. Note that the workload will not be schedu…
ClusterId
string
The id of the cluster.
CpuMemoryLimit
stringnull
Limitations on the CPU memory to allocate for this workload (1G, 20M, .etc). The system guarantees that this workload will not be able to consume more than thi…
DistributedInferenceWorkersField
object
1 property
WorkingDir
stringnull
Container's working directory. If not specified, the container runtime default will be used. This may be configured in the container image.
ImagePullPolicy
stringnull
Image pull policy. Defaults to Always if :latest tag is specified, otherwise it is IfNotPresent.
NodeAffinityRequired
objectnull
If the affinity requirements specified by this field are not met at scheduling time, the pod will not be scheduled onto the node. If the affinity requirements…
1 property
DistributedInferenceCommonSpec
objectnull
RunAsNonRoot
booleannull
Force the container to run as a non-root user.
Toleration
objectnull
Toleration details.
7 properties
EmptyDirInstance
objectnull
GpuPortionLimit
numbernull
Limitations on the portion consumed by the workload, per GPU device. The system guarantees The gpuPotionLimit must be no less than the gpuPortionRequest.
Capabilities
arraynull
Add POSIX capabilities to running containers. Defaults to the default set of capabilities granted by the container runtime.
Probes
objectnull
Probes are used to determine if the container is healthy and ready to accept traffic.
1 property
PodAffinityType
stringnull
The affinity type, required or preferred. (mandatory)
CpuCoreLimit
numbernull
Limitations on the number of CPUs consumed by the workload (0.5, 1, .etc). The system guarantees that this workload will not be able to consume more than this…
WorkloadMeta1
object
11 properties
8 required
ImagePullSecret
objectnull
A reference to a secret in the same namespace used to pull container images.
3 properties
PvcFieldsUpdatable
objectnull
1 property
DistributedInferenceSpec
object
The specifications of the distributed inference to be created.
1 property
GpuPortionRequest
numbernull
Required if and only if gpuRequestType is portion. States the portion of the GPU to allocate for the created workload, per GPU device, between 0 and 1. The def…
ConfigMap
objectnull
4 properties
DistributedInferenceRestartPolicy
stringnull
Determines the behavior when a pod fails. - RecreateGroupOnPodRestart: Restarts all pods in the group if any pod fails. - None: No automatic restart behavior i…
GpuMemoryLimit
stringnull
Limitation on the memory consumed by the workload, per GPU device. The system guarantees The gpuMemoryLimit must be no less than gpuMemoryRequest.
SupplementalGroups
stringnull
Comma separated list of groups that the user running the container belongs to, in addition to the group indicated by runAsGid. Use only when the source uid/gid…
Error
object
3 properties
2 required
The full machine-readable OpenAPI contract behind this narrative.
Other APIs NVIDIA Run:ai publishes across the network.