How AI is applied across API Evangelist and APIs.io. Read my AI disclosure →
API Evangelist API Evangelist
Discovery
Learnings
Guidance
Toolbox
Alignment
API Evangelist LLC

Triton Inference Server CUDA Shared Memory API

CUDA shared memory region management

Triton Inference Server CUDA Shared Memory API is one of 12 APIs that Triton Inference Server publishes on the APIs.io network, described by a machine-readable OpenAPI specification.

Tagged areas include CUDA Shared Memory. The published artifact set on APIs.io includes an OpenAPI specification and API documentation.

This API exposes 4 operations across 4 paths, and defines 2 schemas. It is described by OpenAPI 3.2.0, at version 2.0.

Requests are made against 2 base URLs: http://localhost:8000, http://{host}:{port}.

4 operations 4 paths 2 schemas 1 GET3 POST

Metadata

The identity and technical contract details declared by the specification.

Specification
OpenAPI 3.2.0
API Version
2.0
Base URL
http://localhost:8000
Resource Areas
1

Paths & Operations 4

Across 4 paths, the API surfaces 4 operations — 1 GET, 3 POST. Each is listed below with its method, path, parameters, and response codes.

CUDA Shared Memory 4

CUDA shared memory region management

GET
/v2/cudasharedmemory/status
Triton Inference Server Get CUDA Shared Memory Status
cudaSharedMemoryStatus → 200400
POST
/v2/cudasharedmemory/region/{region_name}/register
Triton Inference Server Register a CUDA Shared Memory Region
cudaSharedMemoryRegister 1 param body → 200400
POST
/v2/cudasharedmemory/region/{region_name}/unregister
Triton Inference Server Unregister a CUDA Shared Memory Region
cudaSharedMemoryUnregister 1 param → 200400
POST
/v2/cudasharedmemory/unregister
Triton Inference Server Unregister All CUDA Shared Memory Regions
cudaSharedMemoryUnregisterAll → 200400

Schemas 2

The contract defines 2 schemas that model the data the API accepts and returns. The most detailed are CudaSharedMemoryRegion (3 properties), ErrorResponse (1 property). Each schema is shown below with its type and property counts.

CudaSharedMemoryRegion
object
3 properties
ErrorResponse
object
1 property

Specification

The full machine-readable OpenAPI contract behind this narrative.

Source

triton-cuda-shared-memory-api-openapi.yml Raw ↑

Other APIs Triton Inference Server publishes across the network.

Triton GRPC API
Triton Inference Server Health API
Triton Inference Server Inference API
Triton Inference Server Logging API
Triton Inference Server Metrics API
Triton Inference Server Model Metadata API
Triton Inference Server Model Repository API
Triton Inference Server Server Metadata API
Triton Inference Server Statistics API
Triton Inference Server System Shared Memory API
Triton Inference Server Trace API
Where this information came from

This is an independent, third-party profile of Triton Inference Server CUDA Shared Memory API, published by API Evangelist. We do not operate, host, resell, or support these APIs, and we are not affiliated with or endorsed by the company unless stated above. Everything here is built from publicly available information — the company's own site, developer portal, documentation, public repositories, and the specifications it publishes for public use. Nothing is obtained by breaching a system, defeating an access control, or using credentials.

The Kin Score and Agent Readiness rating are independently calculated assessments of a company's public API artifacts, scored against a published rubric. They are not certifications, endorsements, security assessments, or audits.

Corrections, re-scores, and removal are free — no partnership or purchase required, and you do not need to justify the request. A removed company is recorded as unrated, never scored zero for having asked. Acknowledgement within one business day; removal within two.

info@apievangelist.com · Read the full data-sourcing policy →
On a security or compliance team? Put security in the subject line and you will get a person, not a form — we will tell you exactly which public URLs this profile was built from.