Retrieve parsed document content
Required capabilities:
filesAcl:READ
Returns parsed document content for the file.
Each document that is uploaded to CDF is OCR’ed and run through layout analysis, which detects titles, paragraphs and tables in the document. The textual content of the document is returned as a hierarchy of layout elements.
The high-level layout elements are again divided into lines and words. Each word has a bounding box for identifying which page it is found on, and where on the page.
Bounding boxes for paragraphs or lines are not provided, but can be easily calculated from the bounding boxes on the individual words.
Note that this endpoint only supports PDF files for now. Support for more file types will be added in the future.
This endpoint is in alpha and is available only when a cdf-version: YYYYMMDD-alpha header
is provided.
Authorizations
Access token issued by the CDF project's configured identity provider. Access token must be an OpenID Connect token, and the project must be configured to accept OpenID Connect tokens. Use a header key of 'Authorization' with a value of 'Bearer $accesstoken'. The token can be obtained through any flow supported by the identity provider.
Headers
cdf version header. Use this to specify the requested CDF release.
"alpha"
Body
Fields to be set for the content request.
Response
Parsed textual content for the document
A title in a document. A title can consist of one or more lines of text.