Need help with your APIs? I offer API discovery, governance & evangelism services. Explore services →
API Evangelist API Evangelist
Discovery
Learnings
Guidance
Toolbox
Alignment
API Evangelist LLC

Kensho Extract API

Transforms unstructured PDF and image documents into machine-readable JSON, identifying titles, subtitles, paragraphs, tables, and footers in natural reading order. Optional OCR and Figure Extraction (FigEx). REST API at extract.kensho.com with asynchronous extractions and presigned upload/download URLs.

Kensho Extract API is one of 9 APIs that S&P Global publishes on the APIs.io network, described by a machine-readable OpenAPI specification.

Tagged areas include Document Extraction, OCR, PDF, Tables, and Unstructured Data. The published artifact set on APIs.io includes an OpenAPI specification, API documentation, an API reference, a getting-started guide, authentication docs, and a JSON-LD context.

This API exposes 5 operations across 5 paths, and defines 3 schemas. It is described by OpenAPI 3.0.2, at version 3.0.0.

Requests are made against a single base URL, https://extract.kensho.com/.

5 operations 5 paths 3 schemas 2 GET2 POST1 PUT

Metadata

The identity and technical contract details declared by the specification.

Specification
OpenAPI 3.0.2
API Version
3.0.0
Base URL
https://extract.kensho.com
Authentication
HTTP Bearer
Resource Areas
1

Authentication & Security 1

Kensho Extract API declares 1 security scheme for authenticating requests. It accepts HTTP bearer tokens (JWT) (bearerAuth).

Paths & Operations 5

Across 5 paths, the API surfaces 5 operations — 2 GET, 2 POST, 1 PUT. Each is listed below with its method, path, parameters, and response codes.

Extractions 5
POST
/v3/extractions
Submit a document for extraction
body → 200400401default
POST
/v3/extractions/upload-url
Upload URL To Submit A Document For Extraction.
body → 200400401default
PUT
/v3/extractions/upload-complete
Mark The Upload As Complete To Start Extraction
body → 204400401404default
GET
/v3/extractions/{request_id}
Retrieve the extracted document
2 params → 200400401404405default
GET
/v3/extractions/download-url/{request_id}
Retrieve The Extracted Document's Download URL
1 param → 200400401404405default

Schemas 3

The contract defines 3 schemas that model the data the API accepts and returns. The most detailed are ContentTree (4 properties), Output (2 properties). Each schema is shown below with its type and property counts.

Annotations
array
Additional data about structure of the document that references text content nodes by their UIDs
ContentTree
object
4 properties 4 required
Output
object
2 properties 2 required

Specification

The full machine-readable OpenAPI contract behind this narrative.

Source

sp-global-extractions-api-openapi.yml Raw ↑

Other APIs S&P Global publishes across the network.

S&P Global LLM-Ready API (kFinance)
Kensho NERD API
Kensho Scribe Batch API v2
Kensho Scribe Real Time API
Kensho Scribe Batch API v1
Kensho Grounding Agent (Alpha)
S&P Capital IQ Pro
S&P Global Marketplace