Text Extractor page

Prev Next

Use this page to set text extractor options.

Note

This page is applicable only to Trellix DLP Discover.

Option definitions

Category

Option

Definition

Text Extractor

Use the following fallback ANSI code page

Allows the administrator to select the fallback character set. The text extractor uses this character set to read input files when there is a problem identifying the correct code page. The default is to use the native language of the endpoint computer operating system.

Maximum input file size to scan (MB)

The maximum file size the text extractor can handle. Default: 50

Maximum output file size (MB)

The maximum file size the text extractor generates to be used by Trellix DLP Discover. Default: 50

Execution timeout per file (seconds)

Maximum time for processing a file. Default: 300

Inactivity timeout per file (seconds)

Time the text extractor waits with no input or export before rejecting the file. Default: 60

Use OCR to extract text from images and scanned PDF files

When selected, enables OCR in classification and remediation scans of file repositories.

Note

To use OCR, you must install the OCR package on the Trellix DLP Discover server. See KB91046 for more information.