OCR enables accurate text extraction from image files, helping to detect and manage sensitive information.
Goal
You can use the OCR feature for extracting text from image files. The extracted text is matched with classification definitions to classify or remediate files.
In Trellix Data Loss Prevention Discover – SaaS, you can use the OCR feature when scanning any file-based repository. Database scans are not supported.
In Trellix Data Loss Prevention – SaaS Network, you can use the OCR feature when scanning images attached to emails, uploaded in web posts, or found in other network traffic.
In Trellix DLP Endpoint for Windows, you can use the OCR feature to scan image files through various data protection rules. However, OCR is not supported for Clipboard Protection and Screen Capture Protection rules.
Concept
OCR is part of the text extraction feature. When the text extractor comes across an image file, a second pass is made with OCR to extract text and classify, remediate, or register the file according to the relevant rules. The feature also works with images saved as a .pdf file. If a .pdf file contains both text and images, it is scanned as a text file using standard method and applies OCR to extract text from the images as well.
For information about installing and updating the OCR package in Trellix DLP Discover and Trellix DLP Endpointfor Windows, see the Trellix Data Loss Prevention Discover and Trellix DLP Endpoint Installation Guideand KB91046.
No additional software installation is needed in Trellix DLP Network to run the OCR feature.
Example
Trellix Data Loss Prevention Discover Scans a file share containing scanned documents to identify sensitive data.
Trellix DLP Network monitors a hardcopy sensitive document scanned and sent in an email as a .pdf attachment.
Trellix DLP Endpoint for Windows prevents:
Sensitive data within image files from being sent via web uploads.
Access to sensitive information within image files accessed by specific applications.
Sensitive data in image files transferred to removable storage devices.
Supported image formats
Trellix DLP Endpoint for Windows, Trellix DLP Discover and Trellix DLP Network supports scanning of images of these formats:
BMP*
GIF (not supported for Trellix DLP Endpoint for Windows)
JPEG
PCX (not supported for Trellix DLP Endpoint for Windows)
PDF
PNG
TIFF
* Although BMP files can be scanned, 32-bit BMP files (8 bits per color channel and 8-bit alpha channel) are not supported.
The following image formats are not supported in Trellix DLP Network but are supported in Trellix DLP Discover:
TIFF-FX (Fax eXtended)
WMP (Windows Media Photo)
XPS (XML Paper Specification)
Operational conditions
For good image recognition, the image should be of good quality, and it should include at least one text line that includes machine-printed text of at least 25–30 characters (possibly mixed upper and lowercase characters).
Supported angle of rotation
Automatic rotation is the default rotation which is supported by the OCR. If any specific rotation is not configured the default angle of rotation is used to detect the images.
Note
Automatic rotation is not supported on Hebrew and Thai languages.
Supported image size
The available memory, the computing environment, and properties of individual image files all influence the size limits for successful image processing.
By default, the Engine handles images up to 8400 pixels in height and width. The OCR rejects images exceeding this limit, and the loading function shows error.
Supported angle limits
The OCR detects the skew effectively only on images with lower than 15-degree skew. Images which are skewed to a specific angle are not detected by the OCR.
Supported languages
OCR scanning works with all Trellix DLP-supported languages, and most Western and Asian languages. These are the languages supported:
Trellix DLP product | Supported languages |
|---|---|
Trellix DLP Endpoint for Windows | Arabic, Russian, Catalan, Chinese - Simplified, Chinese - Traditional, Czech, Danish, Dutch, English, Esperanto, Finnish, French, German, Hungarian, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Portuguese - Brazilian, Slovenian, Spanish, Swedish, Turkish. |
Trellix DLP Discover | Arabic, Russian, Catalan, Chinese - Simplified, Chinese - Traditional, Czech, Danish, Dutch, English, Esperanto, Finnish, French, German, Hungarian, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Portuguese - Brazilian, Slovenian, Spanish, Swedish, Turkish. |
Trellix DLP Network | Afrikaans, Albanian, Arabic, Aymara, Basque, Bemba, Blackfoot, Breton, Brazilian Portuguese, Bugotu, Byelorussian, Catalan, Chamorro, Chechen, Chinese - Simplified, Chinese - Traditional, Corsican, Crow, Czech, Danish, Dutch, English, Esperanto, Estonian, Faroese, Fijian, Finnish, French, Frisian, Friulian, Galician, Ganda, German, Greek, Guarani, Hani, Hawaiian, Hebrew, Hungarian, Icelandic, Ido, Indonesian, Interlingua, Irish Gaelic, Italian, Japanese, Kashubian, Kawa, Kikuyu, Kongo, Korean, Kpelle, Kurdish (Latin), Latin, Latvian, Lithuanian, Luba, Luxembourgian, Macedonian, Malagasy, Malay, Malinke, Maltese, Maori, Mayan, Miao (Hmong), Minankabaw, Mohawk, Moldavian, Nahuatl, Norwegian, Nyanja, Occidental, Ojibway, Papiamento, Pidgin English, Polish, Portuguese, Portuguese - Brazilian, Provencal (Occitan), Quechua, Rhaetic, Romany, Romanian, Russian, Rundi, Samoan, Sami, Sardinian, Scottish Gaelic, Serbian, Shona, Simplified Chinese, Sioux, Slovak, Slovenian, Somali, Sorbian, Southern Sami, Spanish, Swahili, Swazi, Swedish, Tagalog, Tahitian, Thai, Traditional Chinese, Tswana, Turkish, Tun, Ukrainian, Vietnamese, Visayan, Welsh, Wolof, Xhosa, Zapotec, Zulu. |
Unscannable images with Trellix DLP Network
OCR scanning might fail on certain images because of:
Image size greater than 8400 x 8400 pixels
Resolution is less than 75 dpi or more than 2400 dpi.
OCR scanning time exceeding the timeout period of 5 minutes for an individual image
File corruption that renders the file unreadable
On Trellix DLP Network Prevent, OCR scan failure results in the entire email or web post being treated as unscannable. If there is no higher priority action set, such as BLOCK, the UNSCANNABLE action is executed. The Trellix DLP Network Prevent appliance adds X-RCIS-Action: SCANFAIL header to such unscannable emails received by the Smart Host.
For web posts, the web proxy receives a 400 Bad Request ICAP status. Any other detections in the email or web post still result in DLP incidents.
Limitations
OCR is resource-intensive, and significantly increases the scan time if there are many image files in the repository.
For this reason, in Trellix DLP Discover, it can be disabled with a checkbox on the Text Extractor page of the Server Configuration if it is not needed.