The text extractor parses the file content when files are opened or copied and compares it to text patterns and dictionary definitions in the classification rules. When a match occurs, the criteria are applied to the content.
Trellix DLP supports accented characters. When an ASCII text file contains a mix of accented characters, such as French and Spanish, as well as some regular Latin characters, the text extractor might not correctly identify the character set. This issue occurs in all text extraction programs. There is no known method or technique to identify the ANSI code page in this case. When the text extractor cannot identify the code page, text patterns and content fingerprint signatures are not recognized. The document cannot be properly classified, and the correct blocking or monitoring action cannot be taken. To work around this issue, Trellix DLP uses a fallback code page. The fallback is either the default language of the computer or a different language set by the administrator.
Text extraction with Trellix DLP Endpoint
Text extraction is supported on Microsoft Windows and Apple OS X computers.
The text extractor can run multiple processes depending on the number of cores in the processor.
A single core processor runs only one process.
Dual-core processors run up to two processes.
Multi-core processors run up to three simultaneous processes.
If multiple users are logged on, each user has their own set of processes. Thus, the number of text extractors depends on the number of cores and the number of user sessions. The multiple processes can be viewed in the Windows Task Manager. Maximum memory usage for the text extractor is configurable. The default is 75 MB.