OWLxtract is the process of extracting the PDF and Image Files data and storing the data within the OWLdocs module. During the OWLdocs product search process, these extracted files will be analyzed, and if any matches are found, they will be shown on the result screen.
Steps to Perform OWLxtract:
1. Click OWLdocs.
2. This action is performed on the PDF, handwritten PDF, and Image files (jpg, jpeg, png, tiff).
3. For an existing file, click the action menu under the Action column and click OWLxtract.
4. The pop-up will display two options to extract the file: Text Detection and Text Analysis.
5. Text Detection: During this process, the OWL system detects the text in a document along with the following information:
• The lines and words of detected text
• The relationships between the lines and words of the detected text
• The page that the detected text appears on
• The location of the lines and words of text on the document page
• PDF Documents in English, French, German, Italian, Portuguese, Spanish
• Handwritten Documents in English
6. Text Analysis: During this process, OWL detects text in a document, analyzes documents, forms relationships among the detected text, and performs the following:
• Text Extraction- The raw text extraction from a document
• Form Extraction- Form data extraction from a document in the form of a key-value pair
• Table Extraction- Extracts tables, table cells, and the items within table cells
7. Once you have selected our extraction method, click Process to complete the action.
8. Once the OWLxtract process has started, you will receive an email notification that the text extraction is in progress.
9. Once the text extraction is completed, you will receive another email about the extraction completion.
10. After text extraction is completed, the file status will be changed to Text Extraction Processed.