The flatdata CLI
flatdata converts between CSV, TSV, and the record format on the command line.
Put it in front of a program, and the program reads simple records instead of
CSV, with the parsing done in another process.
cat data.csv | flatdata csv2record | ./process | flatdata record2csv > out.csv
Install it with Deno:
deno install -g --allow-read --allow-write -n flatdata jsr:@j50n/proc@0.29.0/flatdata
The permissions are for -i and -o; without them it still works on stdin and
stdout. To run it without installing,
deno run jsr:@j50n/proc@0.29.0/flatdata csv2tsv < data.csv.
flatdata --help lists the commands, and flatdata <command> --help a
command’s options.
Commands
Commands are named <from>2<to>. Each reads stdin and writes stdout, or
-i <file> and -o <file>.
| Command | Options |
|---|---|
csv2record, csv2tsv | -d <char> CSV separator, default , |
tsv2csv, record2csv | -d <char>, --crlf |
tsv2record, record2tsv | none |
--crlf ends rows with CRLF. The commands are the library’s transforms behind a
command line: csv2tsv is csvToTsv(), csv2record is fromCsvToRows() into
toRecord(), and so on. So they read and refuse what the library does, its
parsers read flatdata’s output, and its writers make flatdata’s input.
In a pipeline
import { read } from "@j50n/proc";
import { fromRecordToRows } from "@j50n/proc/transforms";
// Installed, the command is just "flatdata". Here it runs from the package.
const flatdata = import.meta.resolve("@j50n/proc/flatdata");
await read("data-orders.csv")
.run("deno", "run", flatdata, "csv2record") // parses in the child process
.transform(fromRecordToRows())
.flatten()
.drop(1) // the header
.map(([id, customer, item]) => `${id} ${customer}: ${item}`)
.toStdout();
1 Ada: Widget, large
2 Grace: Gear "XL"
3 Linus: Bolt
With flatdata installed, write .run("flatdata", "csv2record"). The parsing
then runs in its own process, alongside your program instead of in it. Records
end in \x1E, not newlines, so read them with fromRecordToRows(), not
.lines.
Errors
A field the output format can’t hold stops the conversion, as the library’s
writers do: a tab, CR, or LF for TSV, \x1E or \x1F for the record format. So
does a row the output would lose, a CR in CSV or TSV input that isn’t part of a
CRLF, and a CSV quote still open at the end of the input. flatdata prints the
message to stderr as one line naming the row and field, as in
flatdata: Invalid character (tab) in TSV data at row 2, field 1, and exits
with code 1, which proc turns into an ExitCodeError:
import { enumerate, ExitCodeError } from "@j50n/proc";
const flatdata = import.meta.resolve("@j50n/proc/flatdata");
const csv = (text: string) => enumerate([new TextEncoder().encode(text)]);
// A tab inside a quoted CSV field: TSV can't hold it, so csv2tsv fails. It
// writes one line to stderr, naming the row and field, and exits with 1.
try {
await csv('a,b\n"c\td",e\n')
.run(
{
fnStderr: (stderr) => stderr.lines.collect(),
fnError: (error, stderrLines) => {
if (error instanceof ExitCodeError) {
console.log(`exit code ${error.code}: ${stderrLines?.join("; ")}`);
}
if (error) throw error;
},
},
"deno",
"run",
flatdata,
"csv2tsv",
)
.lines
.forEach((line) => console.log(JSON.stringify(line)));
} catch (error) {
if (!(error instanceof ExitCodeError)) throw error;
}
exit code 1: flatdata: Invalid character (tab) in TSV data at row 2, field 1
On stdout, output written before the error stays, and can end partway through a
row; here there was none. -o is safer: flatdata writes a new file beside it
and renames it into place only once the conversion has succeeded, so after an
error the old file is as it was. That also lets -o name the file it reads, by
-i or by <, to convert a file in place. -d must be one ASCII character
other than ", CR, or LF.
When the reader of its output goes away, as with | head, flatdata stops and
exits with 0, as other commands do.
Invalid UTF-8 stops the commands that make rows (csv2record, tsv2record,
record2csv, record2tsv) with a TypeError. csv2tsv and tsv2csv copy
bytes without decoding them, so they pass it through unchanged.