Changes for version 0.321 - 2026-09-25
- read_table(): warnings about stray quotes
- A '"' that opens a quoted field takes every newline into that field up to the next '"'. <=0.320 nothing said when that had happened: a file with one unclosed '"' came back with the rest of the file in a single cell, and a '"' used as text in the middle of a field (an inch mark, 5'10") ran rows together or ended in an "Alignment error" that said nothing about quotes. read_table cannot tell a stray '"' from CSV quoting by the bytes alone, and neither can R or pandas, so it now says what it saw and names the line the quote opened on.
- A file that ends inside a quoted field is warned about, and what was read is kept, as R's scan() does ("EOF within quoted string"); pandas raises instead. That cell no longer gains a newline the file did not have when the file has no final one.
- A '"' in the middle of a field that opens a quoted field running past the end of its line is warned about, once per file. It is still read as a quote, as R reads it (tests/reg-tests-1d.R pins '="Total' opening one); pandas keeps such a '"' as text. A quoted cell that starts at the start of its field and holds a line break is ordinary CSV and is not warned about.
- An alignment error on a row that a quoted field ran across lines in now says where that quote opened and whether it was ever closed, on the XS fast path and through a filter alike.
- quote => '' warns when every field on the first line is wrapped in '"', as R's write.csv() writes them, since the quote marks would then stay in every name. An .xlsx is not checked.
- t/read_table.quote.R.pandas.t covers all of it, from pandas' test_quoting.py, test_skiprows.py, test_eof_states and GH 62739 and R 4.6.1's reading of the same inputs; t/read_table.quote.R and t/read_table.quote.pandas.py regenerate the reference values.
- write_table(): .gz and .bz2 output
- A file name ending in .gz or .bz2 is written gzip- or bzip2-compressed, holding exactly the text the plain file would. The rest of the name picks the default sep as before, so x.tsv.gz is tab-separated. It is streamed through a PerlIO::via layer at gzip's and bzip2's default levels, and a gzip header carries no name or time, so the same table always makes the same bytes. Up to 0.320 such a name got plain text.
- A write that croaks partway leaves a truncated file, which read_table refuses, never a whole-looking one; for bzip2 the first block is written at once so that even a croak on the first row leaves one. A compressed write that cannot reach the disk croaks; a plain write's errors are still not checked.
- Only delimited text is compressed: .tex.gz, .xlsx.bz2, or tex/xlsx with a compressed name, is an error. So is a .bgz name, which promises bgzip's BGZF, which this does not write.
- t/write_table.compressed.pandas.t covers it, from pandas' test_compression.py and test_to_csv.py.
- Minimum perl
- perl 5.10.1 is now the oldest supported, and Compress::Raw::Bzip2, core from that release, is a declared prerequisite.
- read_table(): gzip and bzip2 input
- A gzip or bzip2 file is read as the text inside it, with nothing to ask for. It is recognised by its first bytes, not its name, as R's file() recognises one for read.table, so a compressed file without a .gz suffix is read, and a plain one that has one is still text. The extension still picks the default sep, from the name inside: x.tsv.gz is tab-separated. Up to 0.320 a compressed file was parsed as if its bytes were text, and came back as garbage rows or an alignment error.
- The file is inflated as it is read, 64 KB at a time, through a PerlIO::via layer over Compress::Raw::Zlib or Compress::Raw::Bzip2, so a large .tsv.gz takes no more memory than the plain file would; the gzip read of a 110 MB CSV costs about what zcat does on top of reading it plain.
- Every member is read: bgzip's BGZF (every .vcf.gz), R's gzfile(, "a") and `cat a.gz b.gz' all write several, and stopping at the first would have dropped all but the first block without a word. A truncated file, a bad CRC, and data after the last member that is not NUL padding are errors that name the file, never a short read.
- Both need only core modules. Compress::Raw::Bzip2 is loaded only when a bzip2 file is read, and a read without it says what is missing.
- t/read_table.compressed.R.t covers it, from R 4.6.1's tests/reg-tests-1b.R compressed read.table and append-mode cases and pandas' test_compression.py; t/read_table.compressed.R writes the fixtures.
- read_table(): rows of empty fields, and two header fixes
- A line holding nothing but separators is a row of empty fields, as R 4.6.1's read.table and pandas 3.0.4's read_csv both read it. Up to 0.320 a tab-separated "\t" was taken for a blank line and skipped, so a TSV row with every cell empty went missing without a word and the row count no longer matched R's or pandas'. A line of blanks none of which is the separator is still skipped, and so is any line of blanks under qr/\s+/. t/read_table.blank_lines.R.pandas.t covers it, from pandas' test_empty_lines and test_whitespace_lines; t/read_table.blank_lines.R and t/read_table.blank_lines.pandas.py regenerate the reference values.
- auto.row.names with a commented-out header ("# a\tb") refused the header, because the data rows are one field wider than it -- the very shape auto.row.names looks for -- and made the first data row the header instead. The header is now kept and the rows named.
- A worksheet name in an .xlsx is decoded by the one-pass decoder the cells use. It went through five substitutions in turn before, so a reference one of them produced was decoded again by a later one ("&lt;" became "<"), and a numeric reference too large for chr() died.
- read_table(): speed
- Once the XS fast path is building the rows, an empty field is parsed straight to the undef it becomes, instead of to an empty string that was then freed and replaced. On a 300,000 x 10 CSV with nine cells in ten empty, an aoh read goes from 0.185 s to 0.106 s, a hoa from 0.164 s to 0.097 s and an aoa from 0.152 s to 0.073 s; a file with no empty cells is unchanged. A filter is still handed '' for an empty field, as before, and an .xlsx gap is treated the same way.
- na.strings is checked by comparing the few strings given rather than by hashing every cell: with three of them, an aoa read of a 300,000 x 10 CSV goes from 0.155 s to 0.134 s. Past eight strings the hash is used as before.
- Tests
- t/read_table.aoa.t failed 3 subtests on Windows (Strawberry Perl 5.42.2): write_table writes its file in text mode, so its lines end "\r\n" there, as R's write.csv() does, and the test read the result back raw. It now reads it through the platform's text layer. Nothing in the module changed.
Modules
Get basic statistical functions, like in R, but with Perl using XS for performance