Text extraction process to apply classification rules

Prev Next

The text extractor parses the file content when files are opened or copied and compares it to text patterns and dictionary definitions in the classification rules. When a match occurs, the criteria are applied to the content.

Trellix DLP – SaaS supports accented characters. When an ASCII text file contains a mix of accented characters, such as French and Spanish, and some regular Latin characters, the text extractor might not correctly identify the character set. This issue occurs in all text extraction programs. There is no known method or technique to identify the ANSI code page in this case. When the text extractor can't identify the code page, text patterns and content fingerprint signatures are not recognized. The document can't be properly classified, and the correct blocking or monitoring action cannot be taken. To work around this issue, Trellix DLP – SaaS uses a fallback code page. The fallback is either the default language of the computer or a different language set by the administrator.

Text extraction with Trellix DLP Endpoint - SaaS

Text extraction is supported on Microsoft Windows and macOS computers.

The text extractor can run multiple processes depending on the number of cores in the processor.

  • A single-core processor runs only one process.

  • Dual-core processors run up to two processes.

  • Multi-core processors run up to three simultaneous processes.

If multiple users are logged on, each user has their own set of processes. Thus, the number of text extractors depends on the number of cores and the number of user sessions. The multiple processes can be viewed in the Windows Task Manager. Maximum memory usage for the text extractor is configurable. The default is 75 MB.