NAME

Mail::SpamAssassin::Handler::PDF - a MIME-part handler for PDFs

SYNOPSIS

loadhandler     Mail::SpamAssassin::Handler::PDF

DESCRIPTION

A MIME-part handler that registers itself for application/pdf parts, parses each PDF with the pure-Perl Mail::SpamAssassin::PDF::Parser, and exposes the extracted metadata (page, image and link counts, ratios, encryption flags, JavaScript/OpenAction flags, document details) to rules via the pdf2_* eval rules and _PDF2*_ tags below. URLs found inside the PDF are added to the URI detail list with type pdf.

The pdf2_ prefix for tags and eval rules was chosen so as not to conflict with the PDFInfo plugin.

RETURNS

The handler returns embedded images extracted from the PDF as sub-parts, each a { type => '<mediatype>', data => $bytes } spec that the handler framework dispatches to the image handler (e.g. for OCR). Images are emitted as image/jpeg (for streams already in JPEG form), image/tiff (for CCITT fax data, which is left encoded and wrapped in a TIFF header the image reader decodes natively), or image/png (raw image data re-wrapped as PNG).

This extraction is gated by the pdf_extract_images, pdf_max_images, and pdf_max_image_pixels settings (see "CONFIGURATION"), and is skipped entirely for password-protected PDFs. When image extraction is disabled or the PDF contains no usable images, the handler returns an empty list.

REQUIREMENTS

The following non-core perl modules are required for full-functionality:

Crypt::RC4 - decrypt RC4-encrypted PDFs (and the RC4 step of AES variants)
Crypt::Mode::CBC - decrypt AES-encrypted PDFs
Digest::SHA - decrypt AES-256 (R5/R6) PDFs
Convert::Ascii85 - decode ASCII85-encoded streams

If a required module is missing the handler still loads (with a warning at startup) and PDFs that do not need it parse normally; however PDFs that require the missing module are skipped.

Additionally, to analyze text from PDF's you need pdftotext (from poppler) on the system. The handler runs it directly; set pdf_pdftotext_path if it is not on the PATH. Without it the extracted text is empty, so pdf2_word_count, the pdftext rule type, and body rules matching PDF text silently produce no hits.

CONFIGURATION

pdf_pdftotext_path /path/to/pdftotext

Full path to the pdftotext executable. If unset, the handler looks for pdftotext on the PATH. If it cannot be found, text extraction is disabled with a debug message, so --lint never fails merely because the binary is absent.

pdf_text_max_pages N (default: 4)

Extract text from only the first N pages of each PDF (passed to pdftotext -l). This bounds the work done on large documents and is a countermeasure against content stuffing.

pdf_extract_images ( 0 | 1 ) (default: 1)

Extract embedded images from PDFs and emit them as sub-parts, which the handler framework dispatches to the image handler (e.g. for OCR). Set to 0 to disable.

pdf_max_images N (default: 4)

Extract at most N images per PDF, to bound the OCR work done downstream.

pdf_max_image_pixels N (default: 25000000)

Skip images larger than N pixels (width times height). Set to 0 to disable the limit entirely and extract every image regardless of size.

pdf_max_uris N (default: 30)

Retain at most N distinct URIs per PDF. Link annotations are read from every page, so a hostile PDF stuffed with links could otherwise flood the URI list and the URIBL lookups that follow. Additional links pointing at an already-seen URI do not count toward the limit, and the LinkCount metric still counts every link. Set to 0 to disable the limit entirely.

EVAL RULES

This handler defines the following eval rules:

pdf2_count()

   body RULENAME  eval:pdf2_count(<min>,[max])
      min: required, message contains at least x PDF attachments
      max: optional, if specified, must not contain more than x PDF attachments

pdf2_page_count()

   body RULENAME  eval:pdf2_page_count(<min>,[max])
      min: required, message contains at least x pages in PDF attachments.
      max: optional, if specified, must not contain more than x PDF pages

pdf2_link_count()

   body RULENAME  eval:pdf2_link_count(<min>,[max])
      min: required, message contains at least x links in PDF attachments.
      max: optional, if specified, must not contain more than x PDF links

      Note: Multiple links to the same URL are counted multiple times

pdf2_word_count()

   body RULENAME  eval:pdf2_word_count(<min>,[max])
      min: required, message contains at least x words in PDF attachments.
      max: optional, if specified, must not contain more than x PDF words

      Note: pdf2_word_count requires pdftotext (see REQUIREMENTS).  Text is
      not extracted from password-protected PDFs.

pdf2_match_details()

   body RULENAME  eval:pdf2_match_details(<detail>,<regex>);
      detail: Any standard PDF attribute: Author, Creator, Producer, Title, CreationDate, ModDate, etc..
      regex: regular expression

      Fires if any PDF attachment has the given attribute and it's value
      matches the given regular expression

pdf2_is_encrypted()

   body RULENAME eval:pdf2_is_encrypted()

      Fires if any PDF attachment is encrypted

pdf2_is_encrypted_blank_pw()

   body RULENAME eval:pdf2_is_encrypted_blank_pw()

      Fires if any PDF attachment is encrypted with a blank password

pdf2_is_protected()

   body RULENAME eval:pdf2_is_protected()

      Fires if any PDF attachment is encrypted with a non-blank password

pdf2_has_javascript()

   body RULENAME eval:pdf2_has_javascript()

      Fires if any PDF attachment has JavaScript

pdf2_has_open_action()

   body RULENAME eval:pdf2_has_open_action()

      Fires if any PDF attachment has an OpenAction

The following rules only inspect the first page of each document

pdf2_image_count()

   body RULENAME  eval:pdf2_image_count(<min>,[max])
      min: required, message contains at least x images on page 1 (all attachments combined).
      max: optional, if specified, must not contain more than x images on page 1

pdf2_image_ratio()

   body RULENAME  eval:pdf2_image_ratio(<min>,[max])
      min: required, images consume at least x percent of page 1 on any PDF attachment
      max: optional, if specified, images do not consume more than x percent of page 1

      Note: Percent values range from 0-100

pdf2_click_ratio()

   body RULENAME  eval:pdf2_click_ratio(<min>,[max])
      min: required, at least x percent of page 1 is clickable on any PDF attachment
      max: optional, if specified, not more than x percent of page 1 is clickable on any PDF attachment

      Note: Percent values range from 0-100

TEXT RULES

To match against text extracted from PDF's, use the following syntax:

pdftext  RULENAME   /regex/
score    RULENAME   1.0
describe RULENAME   PDF contains text matching /regex/

These rules behave like body rules and support the multiple and maxhits=N tflags. By default a rule stops at its first match. Text is extracted with pdftotext (see REQUIREMENTS); if it is unavailable these rules produce no hits.

TAGS

The following tags can be defined in an add_header line:

_PDF2COUNT_      - total number of pdf mime parts in the email
_PDF2PAGECOUNT_   - total number of pages in all pdf attachments
_PDF2WORDCOUNT_   - total number of words in all pdf attachments
_PDF2LINKCOUNT_   - total number of links in all pdf attachments
_PDF2IMAGECOUNT_  - total number of images found on page 1 inside all pdf attachments
_PDF2VERSION_     - PDF Version, space seperated if there are > 1 pdf attachments
_PDF2IMAGERATIO_  - Percent of first page that is consumed by images - per attachment, space separated
_PDF2CLICKRATIO_  - Percent of first page that is clickable - per attachment, space separated
_PDF2NAME_        - Filenames as found in the mime headers of PDF parts
_PDF2PRODUCER_    - Producer/Application that created the PDF(s)
_PDF2AUTHOR_      - Author of the PDF
_PDF2CREATOR_     - Creator/Program that created the PDF(s)
_PDF2TITLE_       - Title of the PDF File, if available
_PDF2ERRORS_      - number of PDF attachments that failed to parse

URI DETAILS

This handler creates a new "pdf" URI type. You can detect URI's in PDF's using the URIDetail plugin. For example:

uri-detail RULENAME  type =~ /^pdf$/  raw =~ /^https?:\/\/bit\.ly\//