Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

TSV

TSV is a tab between fields and a line feed after each row, with no quoting. fromTsvToRows() reads it and toTsv() writes it.

import { read } from "@j50n/proc";
import { fromTsvToRows } from "@j50n/proc/transforms";

const big = await read("data-orders.tsv")
  .transform(fromTsvToRows())
  .flatten()
  .drop(1) // the header
  .filter((row) => Number(row[3]) > 1)
  .map(([, customer, item]) => `${customer}: ${item}`)
  .collect();

console.log(big);
[ "Ada: Widget, large", "Linus: Bolt" ]

The parser is the CSV parser with a tab separator and quoting turned off, in WebAssembly, and yields a batch of rows for about every 128 KiB of input. There are no options.

Because nothing is quoted, TSV suits data you also want to read with cut, awk, or grep, and whose fields never hold a tab or a line break. When they might, use CSV or the record format.

Edge cases

import { enumerate } from "@j50n/proc";
import { fromTsvToRows } from "@j50n/proc/transforms";

async function parse(text: string): Promise<string> {
  try {
    const rows = await enumerate([new TextEncoder().encode(text)])
      .transform(fromTsvToRows())
      .flatten()
      .collect();
    return JSON.stringify(rows);
  } catch (error) {
    return error instanceof Error ? error.message : String(error);
  }
}

const inputs = [
  "a\tb\r\nc\td", // CRLF; no line end on the last row
  "a\n\nb\n", // blank lines are skipped
  "\t\n", // but a lone tab is two empty fields
  '"a\tb"\tc\n', // quotes are text, so they don't protect the tab
  "a\t\tb\t\n", // empty fields
  "a\tb\rc\n", // a CR that isn't part of CRLF is an error
];
for (const text of inputs) {
  console.log(JSON.stringify(text).padEnd(18), await parse(text));
}
"a\tb\r\nc\td"     [["a","b"],["c","d"]]
"a\n\nb\n"         [["a"],["b"]]
"\t\n"             [["",""]]
"\"a\tb\"\tc\n"    [["\"a","b\"","c"]]
"a\t\tb\t\n"       [["a","","b",""]]
"a\tb\rc\n"        Invalid character (CR) in TSV data at row 1, field 2
  • Lines end in LF or CRLF, and the last line needs no LF.
  • Blank lines are skipped. A line holding only spaces, or only a tab, is a row.
  • Quotes are text. A CSV-style quoted field doesn’t protect a tab.
  • A CR anywhere else throws an Error naming the row and field, so a file with CR-only line ends fails at its first line. TSV can’t hold a CR in a field.
  • A UTF-8 byte order mark at the start is dropped.

Writing

toTsv() joins fields with tabs and ends each row with LF. A field holding a tab, CR, or LF throws an Error naming the row and field, as in Invalid character (tab) in TSV data at row 2, field 2; rows before it have already been written (What writers refuse shows one). A row of one empty field, [""], throws too, since it would be a blank line and the parser skips those: Invalid row (one empty field) in TSV data at row 3. So do a lone surrogate, which UTF-8 can’t hold, and a first field starting with U+FEFF, which the parser would drop as a byte order mark. If your fields might hold these and a lossy fix is acceptable, replace them first:

import { enumerate } from "@j50n/proc";
import { toTsv } from "@j50n/proc/transforms";

const rows = [["id", "note"], ["1", "tab\there"], ["2", "two\r\nlines"]];

// TSV can't hold a tab, CR, or LF in a field. If yours might, replace them.
await enumerate(rows)
  .map((row) => row.map((field) => field.replace(/[\t\r\n]+/g, " ")))
  .transform(toTsv())
  .toStdout();
id	note
1	tab here
2	two lines

LazyRows from TSV

fromTsvToLazyRows() yields LazyRows that decode a field only when you read it. To filter on a field or two, it is several times faster than fromTsvToRows().

To convert TSV to CSV, tsvToCsv() skips the rows altogether; see CSV.

See fromTsvToRows and toTsv for the reference.