Changes for version 0.3211 - 2026-09-27

  • min(), max(), sum() and the other reductions: magical arguments
    • An argument whose value exists only once its get magic has run was read as undef. That covers `$#array`, a tied scalar whose FETCH had not yet been called, and a `substr()` lvalue: none has any value flags before `mg_get()`, and the reductions asked `SvOK()` first, so `sum(0, $#list)` croaked "undefined value at argument index 1" where List::Util's `sum()` returns the index. A tied scalar holding an array ref was likewise not seen as one. This affected `min`, `max`, `sum`, `mean`, `sd`, `var`, `median`, `mode`, `uniq`, `scale`, `skew` and `kurtosis`. Each now runs every argument's get magic once, on entry, and reads the fetched value from then on.
    • The same fault reached a tied element of an ordinary array in `median`, `skew`, `kurtosis`, `mode` and `scale`, which read the array's cells without running their magic. The first four croaked on it; `scale`, which skips undef, silently returned one value fewer than it was given. `sum`, `mean` and the rest already read such an element correctly.
    • A value is now fetched once. Reading one after its magic had run could run FETCH again, on every perl before 5.18 and in `mode` and `uniq` on all of them, and a tie that computes its value then answers differently the second time. `sd`, `var` and `scale` still read a tied array, or a tied element, once per pass, as they are documented to.
    • `median` on a tied array now croaks on a non-numeric element, as it already did on an ordinary one, instead of converting it to a number. `scale` croaks if a tied array answers differently between its counting pass and its value pass, which could previously write past the buffer the first pass sized.
  • Tests
    • t/min.max.sum.ListUtil.t carries List::Util 1.70's own t/min.t, t/max.t and t/sum.t, from the copy bundled with perl 5.44.0. Their GETMAGIC block is what found the fault above. Where the two modules differ by design the header says so: `sum()` of nothing croaks rather than returning undef; a Math::BigInt is reduced as the NV its 0+ overload yields; and a sum is carried in an NV, not an IV.
    • An array reference is read as data whether or not it is blessed and whether or not its class overloads 0+, where List::Util would treat an overloaded object as a single number. This is a decision, not an oversight, so that an array-based object keeps being read as the values it holds, and the port now asserts it in place of upstream's three `example` cases.
    • t/get.magic.args.t checks all twelve functions with six kinds of magical input against the same call with plain values, and counts how often FETCH runs.
  • write_table(): .gz and .bz2 on Windows
    • A compressed file is meant to hold exactly the text the plain file would, and on Windows it did not. The plain file is opened with perl's default layers, which there translate "\n" to CRLF; the :raw that makes way for the compressing layer removed that translation, so in 0.321 a .gz or .bz2 held LF lines where the same call to a plain file wrote CRLF. Where perl is a CRLF shop (PERLIO_USING_CRLF), the buffer above the compressor is now :crlf instead of :perlio, and the two files match again. Other platforms are unchanged. read_table reads either line end from a compressed file, as it does from a plain one.
    • t/write_table.compressed.pandas.t failed 16 of its 88 tests on an MSWin32 smoker for this. Its expected text now uses the line end a plain file actually gets, and the check of the "wrote" line reads its pipe in binary mode and ignores a CRLF from STDOUT's layers, which is perl's doing rather than write_table's.
  • lm()
    • A non-integer exponent inside I() was truncated: the exponent was read with atoi(), so I(hp^0.5) became hp^0, a column of ones that was then aliased against the intercept and reported as NaN, with no sign that anything had gone wrong. I(x^1.5) became x^1. The exponent is now read as a number at the build's NV width, and one that is not a number, such as I(x^abc), is missing rather than 0. glm(), aov() and the other model functions read I() the same way and are fixed with it.
    • An infinite value in the data is refused with R's own message, "lm: NA/NaN/Inf in 'y'" (or 'x'), naming the row, as glm() already refuses it. It used to go into the fit, and every coefficient came back NaN, each reported with t = -Inf and p = 0: a test on the standard error being positive sent a NaN one to the infinite branch. NaN, which is missing, still drops its row.
    • t values, R^2, adjusted R^2 and F are the plain ratios summary.lm() takes, so 0/0 is NaN. A response with no variation at all reported t = Inf, p = 0 and R^2 = 0 for every coefficient; it now reports NaN, as R does, and a nonzero estimate over a standard error of exactly 0 is +-Inf with p = 0.
    • The residual degrees of freedom are now tested against the rank, not the column count. y ~ x + z on three rows with z = 2x has one residual degree of freedom, and R fits it with z NA; lm() refused it as "0 degrees of freedom". A fit with none at all is still refused.
    • Everything lm() allocates is on the save stack, so the new croaks free it all.
    • t/lm.edge.R.t covers it, from R 4.6.1 on its own mtcars and from SciPy 1.18.0's test_regressZEROX (Wilkinson's W.IV.D); t/lm.edge.R.R regenerates the reference values. Where R's QR leaves rounding in an exactly zero residual sum of squares and so reports t = 9.0e15 instead of Inf, the test pins the definition and says why.
  • Row names that are not ASCII
    • fitted.values, residuals and the other hashes keyed by row name, from lm(), glm(), zerotrunc(), hurdle(), svyglm(), ivreg(), lmer(), aov() and predict(), dropped the UTF-8 flag of a name, so "\x{65e5}\x{672c}" came back as its six bytes and the caller's own key did not find it. A HoH key that fits in Latin-1 happened to survive, but the same name given in row.names or _row did not. Row names are now held as UTF-8 and stored as UTF-8 keys, which perl downgrades where it can, so ASCII and Latin-1 byte names come back exactly as they were given.
    • t/rownames.utf8.t checks every one of those functions with each data shape that carries names.
  • min(), max(), sum(), mean(), sd() and var(): faster on plain arrays
    • The loops that walk an ordinary array of numbers now keep four running values instead of one, and combine them at the end. With one, every add or compare waited on the result for the previous element, and for min() and max() that wait was most of the cost: they ran at 1.06 ns/element on 1e4 values where sum() ran at 0.67, and a plain C loop over a flat array of doubles was no faster. sd() and var() share the first pass of sum(), so they gain too, by less.
    • Times are nanoseconds per element, 0.321 -> 0.3211, on perl 5.44.0 at -O2, one pinned core, best of 120 timings taken in alternating runs of the two releases, on the same normal data. At n = 1e6 the data no longer fits in cache, so that column is memory-bound and its cells moved by up to 10% between runs; the others moved by under 1%.
    • | function | n = 1e3 | n = 1e4 | n = 1e5 | n = 1e6 | |---|---|---|---|---| | `max` | 1.094 -> 0.905 (-17%) | 1.063 -> 0.855 (-20%) | 1.060 -> 0.850 (-20%) | 1.461 -> 1.160 (-21%) | | `min` | 1.094 -> 0.906 (-17%) | 1.063 -> 0.854 (-20%) | 1.064 -> 0.850 (-20%) | 1.392 -> 1.176 (-16%) | | `sum` | 0.712 -> 0.617 (-13%) | 0.670 -> 0.572 (-15%) | 0.721 -> 0.578 (-20%) | 1.295 -> 1.064 (-18%) | | `mean` | 0.712 -> 0.617 (-13%) | 0.670 -> 0.571 (-15%) | 0.721 -> 0.583 (-19%) | 1.325 -> 1.048 (-21%) | | `sd` | 1.436 -> 1.352 (-6%) | 1.384 -> 1.286 (-7%) | 1.488 -> 1.332 (-10%) | 2.700 -> 2.596 (-4%) | | `var` | 1.436 -> 1.378 (-4%) | 1.383 -> 1.293 (-6%) | 1.483 -> 1.357 (-9%) | 2.739 -> 2.542 (-7%) |
    • sum(), mean(), sd() and var() now add the elements in a different order, four interleaved partial sums added pairwise, so a result can differ from 0.321's in the last bits. The worst-case rounding error is smaller, since each partial sum is a quarter as long.
    • min() and max() now order -0 below +0, as IEEE 754-2019's minimum() and maximum() do: max(-0.0, 0.0) and max(0.0, -0.0) are both +0, and both min()s are -0. They used to return the first zero they met, as R does, and with four lanes that would have come to depend on which lane each zero fell in. NumPy and List::Util return the last zero they meet, so none of the three agrees with this, and t/min.max.signed_zero.t records all of them. The perl literal -0 is the integer 0, which has no sign, so max(-0, 0) was always 0. The table above includes the cost of tracking the sign, which is 0-6% on normal and integer data and up to 9% on data that is mostly zeros.
    • An array holding a numeric string, a tied element or a hole is still read exactly as before: the fast loop stops at that element and the ordinary one finishes from there.

Modules

Get basic statistical functions, like in R, but with Perl using XS for performance