NVIDIA NIM Vision Language Models API
Vision-language model inference through the standard /v1/chat/completions surface with image inputs (base64 or URL) in the messages payload. Supports NVIDIA NeVA, microsoft/kosmos-2, Phi-3-vision, llama-3.2-90b-vision-instruct, and other VLMs hosted in the NIM catalog.
NVIDIA NIM Vision Language Models API is one of 11 APIs that NVIDIA NIM publishes on the APIs.io network, described by a machine-readable OpenAPI specification.
Tagged areas include AI, Artificial Intelligence, Vision, Multimodal, and VLM. The published artifact set on APIs.io includes API documentation and an OpenAPI specification.
This API exposes 1 operation across 1 path, and defines 2 schemas. It is described by OpenAPI 3.1.0, at version 2026-05-25.
Requests are made against 2 base URLs: https://integrate.api.nvidia.com, http://localhost:8000.
Metadata
The identity and technical contract details declared by the specification.
Authentication & Security 1
NVIDIA NIM Vision Language Models API declares
1 security scheme
for authenticating requests.
It accepts HTTP bearer tokens (nvapi-...) (BearerAuth).
By default, every request must be authenticated.
Paths & Operations 1
Across 1 path, the API surfaces 1 operation — 1 POST. Each is listed below with its method, path, parameters, and response codes.
Multimodal vision-language operations
Schemas 2
The contract defines 2 schemas that model the data the API accepts and returns. The most detailed are VisionChatRequest (5 properties), VisionChatResponse (4 properties). Each schema is shown below with its type and property counts.
Specification
The full machine-readable OpenAPI contract behind this narrative.
Source
More from NVIDIA NIM 10
Other APIs NVIDIA NIM publishes across the network.