The registered documents feature is an extension of location-based content fingerprinting. It gives administrators another way to define sensitive information, to protect it from being distributed in unauthorized ways. Ignored text is text that Trellix DLP – SaaS ignores when processing file content.
To create registered documents, Trellix DLP – SaaS categorizes and fingerprints the contents of files predefined as sensitive. For example, sales estimate spreadsheets for the upcoming quarter. It uses the fingerprints to create signatures that are stored as registered documents. The signatures created are language-agnostic, that is, the process works for all languages.
Manual registration
Create manually registered documents by uploading files in the Classification module on the Register Documents page. Then create packages from the uploaded files, to create signatures. These signatures are made available to and downloaded by the endpoints from the S3 bucket, and used in rules enforced on the endpoints.
When you create a package, Trellix DLP – SaaS processes all files on the list, and loads the fingerprints (signatures) to ePO - SaaS. When you add or delete documents, you must re-create a package. The software makes no attempt to calculate whether some of the files have already been fingerprinted. It always processes the entire list.
Trellix DLP – SaaS appliances also use manual registration. Signatures of the files are uploaded to ePO - SaaS from Trellix DLP – SaaS when you manually upload files and create a package. These signatures are made available to and downloaded by the appliances from the S3 bucket. The appliance is then able to track any content copied from one of these documents and classify it according to the classification of the registered document signature.
Setting the confidence threshold — Trellix DLP – SaaS allows you to configure the number of fingerprints that must be matched in a manually fingerprinted document to trigger a violation. This helps in increasing the detection confidence as it minimizes false positives by triggering more accurate detections and reduces the analysis time. An incident is triggered when the number of matches is equal to or higher than the set confidence threshold. You can set the Confidence Threshold percentage between 10 to 100 percentage. For example, if a fingerprinted document generates 100 signatures, and if you select 10%, then 10 signatures are matched at random in the scanned document. To set the percentage go to, Classification → Registered Documents → Manual Registration → Confidence Threshold.
Ignored Text — Upload ignored text files on the Ignored Text page. Ignored text does not cause content to be classified, even if parts of it match content classification or content fingerprinting criteria. ignored text that commonly appear in files, such as boilerplates, legal disclaimers, and copyright information. ignored text packages are created separately from the registered documents packages and are distributed to the endpoints in a similar manner.
Files that must be ignored must contain at least 400 characters.
If a file contains both classified and ignored data, the system does not ignore it. All relevant content classification and content fingerprinting criteria associated with the content remain in effect.
Viewing registered documents data
The default Statistics view displays totals for number of files, file size, number of signatures, and so forth, in the left pane, and statistics per file in the right pane. Use this data to remove less important packages if the signature limit is approached.
The Group by view for manual registration allows grouping by classification or type/extension. It displays uploaded files per classification or type. You can filter the data by classification or with a custom filter. Information about last package creation and changes to the file list are displayed in the upper right.