Input requirements
The input for the document parsing job must meet these criteria:- PDF documents with English text and up to 100 pages. Smaller files usually give better results.
- Embedded text or scanned documents.
- Documents that describe a single asset or piece of equipment.
- Key-value pair data representation.
Before you start
- Ingest the documents into CDF.
- Set up access capabilities.
- Create a view in a data model with properties that reflect the key-value data.
Parse documents
1
Navigate to document parsing
Navigate to Data fusion > Contextualize > Document parsing.
2
Create parsing task
Select Create parsing task, and then select the documents you want to parse.
You can parse several documents at the same time, but data from each document is ingested into a separate data model view.
3
Continue to view selection
Select Next to continue.
4
Select views and run
Select the views you want to populate parsed data into and select Run.
5
Review the parsed data
Review the parsed data.If many properties show low scores, your view property names may not align with field names in the document, for example, abbreviations, different wording, or spelling. Rename properties in the view so they resemble the field names more closely, then run parsing again.
- Select a property in the Parsed data sidebar to zoom into a field in the document.
- Hover over a field to update the value.
- Enable Confidence score to view the confidence level for each extracted property. Use the confidence score to decide how much to trust each value and where to focus your review before you approve:
Exact score bands may vary. Focus review on lower-confidence properties; properties with higher scores usually require less review and can be trusted for automated workflows.
6
Approve or reject parsing
Approve or reject the parsing. The approved data is stored as a data model instance.
Further reading
- About document parsing – Overview of document parsing, what the confidence score means, and how it is calculated