Record format
The record format puts \x1F (ASCII unit separator) between fields and \x1E
(record separator) after each row, with no quoting. A field can hold anything
else, tabs, quotes, and newlines included, and a program in any language reads
it by splitting twice.
import { enumerate } from "@j50n/proc";
import { toRecord } from "@j50n/proc/transforms";
const notes = [
["1", "first line\nsecond line"],
["2", "a\ttab"],
];
// awk splits records on \036 (\x1E) and fields on \037 (\x1F).
const summary = await enumerate(notes)
.transform(toRecord())
.run(
"awk",
'BEGIN { RS = "\\036"; FS = "\\037" } { print $1 ": " length($2) }',
)
.lines
.collect();
console.log(summary);
[ "1: 22", "2: 5" ]
toRecord() writes it and fromRecordToRows() reads it. Reach for it when you
hand rows to another program, or take them from one, and the fields might hold
characters that would break TSV. Unlike CSV, the reading side needs no parser:
split the input on \x1E, dropping the empty piece after the last one, then
each record on \x1F. In awk, set RS and FS as above. The separators are
exported as RECORD_SEPARATOR and FIELD_SEPARATOR.
Reading
import { run } from "@j50n/proc";
import { fromRecordToRows } from "@j50n/proc/transforms";
// printf writes \037 between fields and \036 after each record.
const rows = await run(
"printf",
"Ada\\037line one\\nline two\\036Grace\\037\\036",
)
.transform(fromRecordToRows())
.flatten()
.collect();
console.log(rows);
// A newline after the last record, as echo adds, is a record of its own.
const withNewline = await run("sh", "-c", "printf 'a\\037b\\036'; echo")
.transform(fromRecordToRows())
.flatten()
.collect();
console.log(withNewline);
[ [ "Ada", "line one\nline two" ], [ "Grace", "" ] ]
[ [ "a", "b" ], [ "\n" ] ]
The input is split on \x1E, then each record on \x1F. Nothing else is
special. Text after the last \x1E is a final record, so a record needs no
\x1E at the end of the input. Every piece between separators counts:
- An empty record (
\x1E\x1E) reads as[""]. - A newline after the last
\x1E, asechoorprintadds, reads as a row["\n"]. Make the writer leave it off, or filter that row out.
A UTF-8 byte order mark at the start is dropped. Batches close at about 128 KiB
of text. fromRecordToLazyRows() yields string-backed LazyRows and does the
same work, so it is no faster; use it only when the code downstream expects
LazyRows.
Writing
toRecord() refuses a field holding \x1E or \x1F, throwing an Error such
as Invalid character (field separator) in record data at row 2, field 2; rows
before it have already been written. It refuses a lone surrogate too, which
UTF-8 can’t hold, and a first field starting with U+FEFF, which
fromRecordToRows() would drop as a byte order mark. Any other text is written
as it is.
Unlike TSV, the record format keeps a row holding one empty field: [""] is
written as \x1E and reads back as [""]. A row with no fields, [] inside a
batch, would be written the same way, so toRecord() refuses it:
Invalid row (no fields) in record data at row 3.
The record format and flatdata
The flatdata CLI converts CSV and TSV to records in a separate
process (flatdata csv2record), so a program in another language can read CSV
without a CSV parser. It writes with toRecord(), so a CSV field holding \x1E
or \x1F stops it with the same error and exit code 1, rather than coming out
as extra records or fields.
See toRecord and
fromRecordToRows
for the reference.