Cartesia Text-to-Speech API
Single-shot and streamed speech generation over plain HTTP.
Cartesia Text-to-Speech API is one of 21 APIs that Cartesia publishes on the APIs.io network, described by a machine-readable OpenAPI specification.
Tagged areas include Text-to-Speech. The published artifact set on APIs.io includes an OpenAPI specification, API documentation, and an API reference.
This API exposes 2 operations across 2 paths, and defines 5 schemas. It is described by OpenAPI 3.0.3, at version 2026-03-01.
Requests are made against a single base URL, https://api.cartesia.ai.
Metadata
The identity and technical contract details declared by the specification.
Authentication & Security 1
Cartesia Text-to-Speech API declares
1 security scheme
for authenticating requests.
It accepts HTTP bearer tokens (sk_car_... API key or short-lived access token) (bearerAuth).
By default, every request must be authenticated.
bearerAuth— Cartesia API key (skcar...) or a short-lived access token minted via POST /access-token, passed as Authorization: Bearer . Every request also requires the Cart…
Paths & Operations 2
Across 2 paths, the API surfaces 2 operations — 2 POST. Each is listed below with its method, path, parameters, and response codes.
Single-shot and streamed speech generation over plain HTTP.
Schemas 5
The contract defines 5 schemas that model the data the API accepts and returns. The most detailed are TtsRequest (7 properties), Error (6 properties), OutputFormat (4 properties), GenerationConfig (3 properties). Each schema is shown below with its type and property counts.
Specification
The full machine-readable OpenAPI contract behind this narrative.
Source
More from Cartesia 12
Other APIs Cartesia publishes across the network.