NAME
Mail::SpamAssassin::Handler::PDF - a MIME-part handler for PDFs
SYNOPSIS
loadhandler Mail::SpamAssassin::Handler::PDF
DESCRIPTION
A MIME-part handler that registers itself for application/pdf parts, parses each PDF with the pure-Perl Mail::SpamAssassin::PDF::Parser, and exposes the extracted metadata (page, image and link counts, ratios, encryption flags, JavaScript/OpenAction flags, document details) to rules via the pdf2_* eval rules and _PDF2*_ tags below. URLs found inside the PDF are added to the URI detail list with type pdf.
The pdf2_ prefix for tags and eval rules was chosen so as not to conflict with the PDFInfo plugin.
RETURNS
The handler returns embedded images extracted from the PDF as sub-parts, each a { type => '<mediatype>', data => $bytes } spec that the handler framework dispatches to the image handler (e.g. for OCR). Images are emitted as image/jpeg (for streams already in JPEG form), image/tiff (for CCITT fax data, which is left encoded and wrapped in a TIFF header the image reader decodes natively), or image/png (raw image data re-wrapped as PNG).
This extraction is gated by the pdf_extract_images, pdf_max_images, and pdf_max_image_pixels settings (see "CONFIGURATION"), and is skipped entirely for password-protected PDFs. When image extraction is disabled or the PDF contains no usable images, the handler returns an empty list.
REQUIREMENTS
The following non-core perl modules are required for full-functionality:
- Crypt::RC4 - decrypt RC4-encrypted PDFs (and the RC4 step of AES variants)
- Crypt::Mode::CBC - decrypt AES-encrypted PDFs
- Digest::SHA - decrypt AES-256 (R5/R6) PDFs
- Convert::Ascii85 - decode ASCII85-encoded streams
If a required module is missing the handler still loads (with a warning at startup) and PDFs that do not need it parse normally; however PDFs that require the missing module are skipped.
Additionally, to analyze text from PDF's you need pdftotext (from poppler) on the system. The handler runs it directly; set pdf_pdftotext_path if it is not on the PATH. Without it the extracted text is empty, so pdf2_word_count, the pdftext rule type, and body rules matching PDF text silently produce no hits.
CONFIGURATION
- pdf_pdftotext_path /path/to/pdftotext
-
Full path to the
pdftotextexecutable. If unset, the handler looks forpdftotexton thePATH. If it cannot be found, text extraction is disabled with a debug message, so--lintnever fails merely because the binary is absent. - pdf_text_max_pages N (default: 4)
-
Extract text from only the first N pages of each PDF (passed to
pdftotext -l). This bounds the work done on large documents and is a countermeasure against content stuffing. - pdf_extract_images ( 0 | 1 ) (default: 1)
-
Extract embedded images from PDFs and emit them as sub-parts, which the handler framework dispatches to the image handler (e.g. for OCR). Set to 0 to disable.
- pdf_max_images N (default: 4)
-
Extract at most N images per PDF, to bound the OCR work done downstream.
- pdf_max_image_pixels N (default: 25000000)
-
Skip images larger than N pixels (width times height). Set to 0 to disable the limit entirely and extract every image regardless of size.
- pdf_max_uris N (default: 30)
-
Retain at most N distinct URIs per PDF. Link annotations are read from every page, so a hostile PDF stuffed with links could otherwise flood the URI list and the URIBL lookups that follow. Additional links pointing at an already-seen URI do not count toward the limit, and the
LinkCountmetric still counts every link. Set to 0 to disable the limit entirely.
EVAL RULES
This handler defines the following eval rules:
pdf2_count()
body RULENAME eval:pdf2_count(<min>,[max])
min: required, message contains at least x PDF attachments
max: optional, if specified, must not contain more than x PDF attachments
pdf2_page_count()
body RULENAME eval:pdf2_page_count(<min>,[max])
min: required, message contains at least x pages in PDF attachments.
max: optional, if specified, must not contain more than x PDF pages
pdf2_link_count()
body RULENAME eval:pdf2_link_count(<min>,[max])
min: required, message contains at least x links in PDF attachments.
max: optional, if specified, must not contain more than x PDF links
Note: Multiple links to the same URL are counted multiple times
pdf2_word_count()
body RULENAME eval:pdf2_word_count(<min>,[max])
min: required, message contains at least x words in PDF attachments.
max: optional, if specified, must not contain more than x PDF words
Note: pdf2_word_count requires pdftotext (see REQUIREMENTS). Text is
not extracted from password-protected PDFs.
pdf2_match_details()
body RULENAME eval:pdf2_match_details(<detail>,<regex>);
detail: Any standard PDF attribute: Author, Creator, Producer, Title, CreationDate, ModDate, etc..
regex: regular expression
Fires if any PDF attachment has the given attribute and it's value
matches the given regular expression
pdf2_is_encrypted()
body RULENAME eval:pdf2_is_encrypted()
Fires if any PDF attachment is encrypted
pdf2_is_encrypted_blank_pw()
body RULENAME eval:pdf2_is_encrypted_blank_pw()
Fires if any PDF attachment is encrypted with a blank password
pdf2_is_protected()
body RULENAME eval:pdf2_is_protected()
Fires if any PDF attachment is encrypted with a non-blank password
pdf2_has_javascript()
body RULENAME eval:pdf2_has_javascript()
Fires if any PDF attachment has JavaScript
pdf2_has_open_action()
body RULENAME eval:pdf2_has_open_action()
Fires if any PDF attachment has an OpenAction
The following rules only inspect the first page of each document
pdf2_image_count()
body RULENAME eval:pdf2_image_count(<min>,[max])
min: required, message contains at least x images on page 1 (all attachments combined).
max: optional, if specified, must not contain more than x images on page 1
pdf2_image_ratio()
body RULENAME eval:pdf2_image_ratio(<min>,[max])
min: required, images consume at least x percent of page 1 on any PDF attachment
max: optional, if specified, images do not consume more than x percent of page 1
Note: Percent values range from 0-100
pdf2_click_ratio()
body RULENAME eval:pdf2_click_ratio(<min>,[max])
min: required, at least x percent of page 1 is clickable on any PDF attachment
max: optional, if specified, not more than x percent of page 1 is clickable on any PDF attachment
Note: Percent values range from 0-100
TEXT RULES
To match against text extracted from PDF's, use the following syntax:
pdftext RULENAME /regex/
score RULENAME 1.0
describe RULENAME PDF contains text matching /regex/
These rules behave like body rules and support the multiple and maxhits=N tflags. By default a rule stops at its first match. Text is extracted with pdftotext (see REQUIREMENTS); if it is unavailable these rules produce no hits.
TAGS
The following tags can be defined in an add_header line:
_PDF2COUNT_ - total number of pdf mime parts in the email
_PDF2PAGECOUNT_ - total number of pages in all pdf attachments
_PDF2WORDCOUNT_ - total number of words in all pdf attachments
_PDF2LINKCOUNT_ - total number of links in all pdf attachments
_PDF2IMAGECOUNT_ - total number of images found on page 1 inside all pdf attachments
_PDF2VERSION_ - PDF Version, space seperated if there are > 1 pdf attachments
_PDF2IMAGERATIO_ - Percent of first page that is consumed by images - per attachment, space separated
_PDF2CLICKRATIO_ - Percent of first page that is clickable - per attachment, space separated
_PDF2NAME_ - Filenames as found in the mime headers of PDF parts
_PDF2PRODUCER_ - Producer/Application that created the PDF(s)
_PDF2AUTHOR_ - Author of the PDF
_PDF2CREATOR_ - Creator/Program that created the PDF(s)
_PDF2TITLE_ - Title of the PDF File, if available
_PDF2ERRORS_ - number of PDF attachments that failed to parse
URI DETAILS
This handler creates a new "pdf" URI type. You can detect URI's in PDF's using the URIDetail plugin. For example:
uri-detail RULENAME type =~ /^pdf$/ raw =~ /^https?:\/\/bit\.ly\//