Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Record format

The record format puts \x1F (ASCII unit separator) between fields and \x1E (record separator) after each row, with no quoting. A field can hold anything else, tabs, quotes, and newlines included, and a program in any language reads it by splitting twice.

import { enumerate } from "@j50n/proc";
import { toRecord } from "@j50n/proc/transforms";

const notes = [
  ["1", "first line\nsecond line"],
  ["2", "a\ttab"],
];

// awk splits records on \036 (\x1E) and fields on \037 (\x1F).
const summary = await enumerate(notes)
  .transform(toRecord())
  .run(
    "awk",
    'BEGIN { RS = "\\036"; FS = "\\037" } { print $1 ": " length($2) }',
  )
  .lines
  .collect();

console.log(summary);
[ "1: 22", "2: 5" ]

toRecord() writes it and fromRecordToRows() reads it. Reach for it when you hand rows to another program, or take them from one, and the fields might hold characters that would break TSV. Unlike CSV, the reading side needs no parser: split the input on \x1E, dropping the empty piece after the last one, then each record on \x1F. In awk, set RS and FS as above. The separators are exported as RECORD_SEPARATOR and FIELD_SEPARATOR.

Reading

import { run } from "@j50n/proc";
import { fromRecordToRows } from "@j50n/proc/transforms";

// printf writes \037 between fields and \036 after each record.
const rows = await run(
  "printf",
  "Ada\\037line one\\nline two\\036Grace\\037\\036",
)
  .transform(fromRecordToRows())
  .flatten()
  .collect();

console.log(rows);

// A newline after the last record, as echo adds, is a record of its own.
const withNewline = await run("sh", "-c", "printf 'a\\037b\\036'; echo")
  .transform(fromRecordToRows())
  .flatten()
  .collect();

console.log(withNewline);
[ [ "Ada", "line one\nline two" ], [ "Grace", "" ] ]
[ [ "a", "b" ], [ "\n" ] ]

The input is split on \x1E, then each record on \x1F. Nothing else is special. Text after the last \x1E is a final record, so a record needs no \x1E at the end of the input. Every piece between separators counts:

  • An empty record (\x1E\x1E) reads as [""].
  • A newline after the last \x1E, as echo or print adds, reads as a row ["\n"]. Make the writer leave it off, or filter that row out.

A UTF-8 byte order mark at the start is dropped. Batches close at about 128 KiB of text. fromRecordToLazyRows() yields string-backed LazyRows and does the same work, so it is no faster; use it only when the code downstream expects LazyRows.

Writing

toRecord() refuses a field holding \x1E or \x1F, throwing an Error such as Invalid character (field separator) in record data at row 2, field 2; rows before it have already been written. It refuses a lone surrogate too, which UTF-8 can’t hold, and a first field starting with U+FEFF, which fromRecordToRows() would drop as a byte order mark. Any other text is written as it is.

Unlike TSV, the record format keeps a row holding one empty field: [""] is written as \x1E and reads back as [""]. A row with no fields, [] inside a batch, would be written the same way, so toRecord() refuses it: Invalid row (no fields) in record data at row 3.

The record format and flatdata

The flatdata CLI converts CSV and TSV to records in a separate process (flatdata csv2record), so a program in another language can read CSV without a CSV parser. It writes with toRecord(), so a CSV field holding \x1E or \x1F stops it with the same error and exit code 1, rather than coming out as extra records or fields.

See toRecord and fromRecordToRows for the reference.