To protect sensitive content, start by defining and classifying sensitive information that needs to be protected.
Content is classified by defining classifications and classification criteria. Classification criteria defines the conditions on how data is classified. Methods to define criteria include:
Advanced patterns — Regular expressions combined with validation algorithms, used to match patterns such as credit card numbers. Advanced patterns are ranked according to a score, meaning, the number of times the sensitive expressions need to appear in the content for the rule to be triggered. The Classifications editor includes several built-in advanced patterns for ensuring compliance with government regulations and simplifying detection of personal information. You can also create your own advanced patterns.
Dictionaries — Lists of specific words or terms, such as medical terms for detecting possible HIPAA violations.
Keywords — A string value that defines sensitive data. You can add multiple keywords for content classifications. Keywords are not consistent across classifications. If you need to use consistent keywords across classifications, use a dictionary.
File size — The size of the file to detect the sensitive data. You can also define a file size range.
True file types — The true file type to determine which files to identify the sensitive data. True file type helps detect attachment violations when file extensions are renamed and sent as attachments. For example, a .cpp file saved as a .txt file can be detected using the true file type classification criteria.
File extension — The file types to detect the sensitive data, such as MP3 and PDF.
Source or destination location — URLs, network shares, or the application or user that created or received the content.
Location in file — The section of the file to look for the sensitive content; Header, Footer, Body or within the first characters. Specifying the number of characters for the within first (characters) option in a classification looks for the sensitive content in the Header, that is, in the first part of the first page in a document.
Microsoft Word documents - Header, body and footer is identified.
PowerPoint documents - WordArt is considered Header, everything else is identified as Body.
Other documents - Only Body is applicable.
Trellix DLP Endpoint supports third-party classification software. You can classify email using Boldon James Email Classifier. You can classify email or other files using Titus classification clients – Titus Message Classification, Titus Classification for Desktop, and Titus Classification Suite. To implement Titus support, the Titus SDK must be installed on the endpoint computers.