NAME

Mail::SpamAssassin::URI - parse a URI into its components

SYNOPSIS

use Mail::SpamAssassin::URI;

my $uri = Mail::SpamAssassin::URI->new('https://example.com/a?url=http%3A%2F%2Fevil.com');

$uri->scheme;                 # https
$uri->host;                   # example.com
$uri->path;                   # /a
$uri->param('url');           # http://evil.com

for my $p ($uri->params) {
  printf "%s = %s\n", $p->{name}, $p->{value};
}

DESCRIPTION

Splits a URI into its RFC 3986 components and percent-decodes each one separately.

The order matters. Percent-encoding exists so that a reserved character can appear as data inside a component, so the component boundaries have to be found in the raw string and only then may each component be decoded. Decoding the whole URI first would invent boundaries that no server or MUA ever sees: in https://host/a%3Fx=1 the %3F is part of the path, not the start of a query string, and in ?url=a%26b=c the %26 belongs to the value of url rather than separating two parameters.

Accessors therefore come in two flavours. authority, raw_path, query and fragment return the component exactly as it appeared, while host, path and the parameter accessors return decoded values.

Nothing here validates a URI. Every string parses, including the empty one, so a caller that needs to know whether it has something usable should ask a question it actually cares about, such as is_http or whether host is defined. Note this differs from Mail::SpamAssassin::HTML::Color, which throws on malformed input; a URI arrives from a hostile message rather than from a configuration file, so refusing to parse it is not helpful.

METHODS

new($string)

Parses $string and returns an object. Never returns undef and never dies; an undefined argument is treated as an empty string.

raw()

The string the object was built from. An object also stringifies to this, so it can be interpolated, compared or used as a hash key wherever the original string would have been; an object is always true, even when that string is empty.

scheme()

The scheme, lowercased and without the trailing colon, or undef.

authority()

The authority exactly as it appeared, including any userinfo and port, or undef when the URI has none. See host for the decoded hostname.

path()

The path, percent-decoded. Always defined, but may be empty.

An encoded separator is not distinguished from a real one: /tr%2Fop and /tr/op both come back as /tr/op. A caller that needs to tell them apart wants raw_path.

raw_path()

The path exactly as it appeared, still percent-encoded.

query()

The query string exactly as it appeared, without the leading ?, or undef. See params for the decoded parameters.

fragment()

The fragment exactly as it appeared, without the leading #, or undef.

host()

The hostname, decoded, lowercased and folded to ASCII (an internationalised name is returned in its xn-- form), with any userinfo and port removed, or undef when the URI has no authority.

The authority is decoded but never re-split, so a %2F inside it stays part of the hostname rather than starting a path. https://host.example%2Fa/b therefore has the (unresolvable) host host.example/a and the path /b, which is what the CPAN URI module reports.

The folding goes further than URI does, deliberately. URI folds an internationalised name given as a character string but leaves a percent-encoded or a mixed-case one alone, so b%C3%BCcher.de, BUECHER.de spelled with an umlaut, and xn--bcher-kva.de are three different hosts to it. Here they are one: a configured hostname has to match however the message spelled it, and mail carries octets rather than character strings.

userinfo()

The userinfo from the authority, undecoded and without the @, or undef.

port()

The port from the authority, or undef.

without_fragment()

The URI as a string with any fragment removed. RFC 3986 defines the fragment as client-side only, so this is the form to request from a server, to use as a cache key, or to compare against a Location header.

is_http()

True when the scheme is http or https.

params()

The query parameters as a list of hashrefs with name and value keys, in the order they appeared. Both are decoded; value is undef for a parameter that had no = at all.

This is the complete form: unlike param_hash it keeps every occurrence of a repeated name.

param($name)

In list context every value for $name, in order. In scalar context the first, or undef when there is none.

param_hash()

The query parameters as a hashref of name to value.

Lossy by nature: a repeated name keeps its first value, matching param in scalar context. Use params when every occurrence matters.

NOTES

Decoding uses Mail::SpamAssassin::Util::url_decode, which resolves both single and double encoding in one pass (%253A and %3A both yield :) but stops one layer short of triple encoding.

A + is not converted to a space. That is a convention of HTML form submission rather than of URI syntax, and this class parses URIs.

Only & (and &) separates query parameters. A ; is left alone, so values such as data:text/html;base64,... survive intact.