Cartesia Speech-to-Text API
Batch transcription of an audio file of any length.
Cartesia Speech-to-Text API is one of 21 APIs that Cartesia publishes on the APIs.io network, described by a machine-readable OpenAPI specification.
Tagged areas include Speech-to-Text. The published artifact set on APIs.io includes an OpenAPI specification, API documentation, and an API reference.
This API exposes 1 operation across 1 path, and defines 3 schemas. It is described by OpenAPI 3.0.3, at version 2026-03-01.
Requests are made against a single base URL, https://api.cartesia.ai.
Metadata
The identity and technical contract details declared by the specification.
Authentication & Security 1
Cartesia Speech-to-Text API declares
1 security scheme
for authenticating requests.
It accepts HTTP bearer tokens (sk_car_... API key or short-lived access token) (bearerAuth).
By default, every request must be authenticated.
bearerAuth— Cartesia API key (skcar...) or a short-lived access token minted via POST /access-token, passed as Authorization: Bearer . Every request also requires the Cart…
Paths & Operations 1
Across 1 path, the API surfaces 1 operation — 1 POST. Each is listed below with its method, path, parameters, and response codes.
Batch transcription of an audio file of any length.
Schemas 3
The contract defines 3 schemas that model the data the API accepts and returns. The most detailed are Error (6 properties), TranscriptResponse (6 properties), WordTiming (3 properties). Each schema is shown below with its type and property counts.
Specification
The full machine-readable OpenAPI contract behind this narrative.
Source
More from Cartesia 12
Other APIs Cartesia publishes across the network.