JSON lines
JSON lines is one JSON value per line. fromJsonToRows() parses each line with
JSON.parse, and toJson() writes each value with JSON.stringify.
import { read } from "@j50n/proc";
import { fromJsonToRows } from "@j50n/proc/transforms";
type Event = { id: string; level: string; ms: number };
const slow = await read("data-events.jsonl")
.transform(fromJsonToRows<Event>()) // the type is asserted, not checked
.flatten()
.filter((event) => event.ms > 10)
.map((event) => event.id)
.collect();
console.log(slow);
[ "e1", "e2" ]
Despite its name, fromJsonToRows() yields values, not rows: objects, arrays,
strings, numbers, anything JSON holds. They come in batches of about 128 KiB of
text, so .flatten() before working value by value. Blank lines, and lines of
only whitespace, are skipped. A CR at the end of a line is whitespace to
JSON.parse, so CRLF files read fine.
The type parameter (fromJsonToRows<Event>()) is only an assertion. Nothing
checks the values unless you pass a schema.
Checking values
import { enumerate } from "@j50n/proc";
import { fromJsonToRows } from "@j50n/proc/transforms";
type Event = { id: string; ms: number };
// Anything with a parse() that throws on a bad value works; a Zod schema does.
const EventSchema = {
parse(value: unknown): Event {
const e = value as Partial<Event>;
if (typeof e?.id !== "string" || typeof e.ms !== "number") {
throw new TypeError(`not an event: ${JSON.stringify(value)}`);
}
return e as Event;
},
};
const text = '{"id":"e1","ms":12}\n{"id":"e2","ms":"slow"}\n';
try {
await enumerate([new TextEncoder().encode(text)])
.transform(fromJsonToRows({ schema: EventSchema }))
.flatten()
.forEach((event) => console.log(event.id));
} catch (error) {
if (error instanceof TypeError) console.log(error.message);
}
not an event: {"id":"e2","ms":"slow"}
| Option | Default | Meaning |
|---|---|---|
schema | none | an object whose parse(value) returns the value to yield, or throws |
sampleSize | every value | check only the first sampleSize values |
Whatever schema.parse throws stops the stream and reaches your catch as it
is, so instanceof finds a Zod error. What it returns is the value you get, so
a Zod schema’s defaults and transforms apply, and keys it doesn’t know are
stripped, as with any parse call.
With sampleSize, only the first sampleSize values go through the schema. The
rest come as JSON.parse made them, unchecked and untransformed, though they
are typed as the schema’s output, just as values are with no schema at all. So
use sampleSize only with a schema that checks values without changing them.
Note that e1 never printed. A batch is parsed whole before it is yielded, so
an error stops the stream before any value in the same batch reaches you.
A line that isn’t JSON throws a SyntaxError naming the line, counted from 1
with blank lines included: Invalid JSON at line 4001: Unexpected token ....
Writing
import { enumerate, read } from "@j50n/proc";
import { fromCsvToRows, toJson } from "@j50n/proc/transforms";
// toJson takes one value per item, so flatten the parser's batches first.
await read("data-orders.csv")
.transform(fromCsvToRows())
.flatten()
.drop(1) // the header
.map(([id, customer, , qty]) => ({ id, customer, qty: Number(qty) }))
.transform(toJson())
.toStdout();
// A value with no JSON form is refused, naming the item.
try {
await enumerate([{ total: 13 }, undefined]).transform(toJson()).toStdout();
} catch (error) {
if (error instanceof TypeError) console.log(error.message);
}
{"id":"1","customer":"Ada","qty":2}
{"id":"2","customer":"Grace","qty":1}
{"id":"3","customer":"Linus","qty":10}
{"total":13}
Item 2 can't be written as JSON (undefined)
toJson() takes one value per item and writes it on a line of its own, the
reverse of fromJsonToRows() followed by .flatten(). The parsers yield
batches, so flatten them first; a batch passed as it is would be written as one
JSON array on one line.
Inside a value, JSON.stringify’s rules apply: a property holding undefined
or a function is left out, and in an array it becomes null. NaN and
Infinity become null too, without an error, so check numbers parsed from
text (Number("12,5") is NaN) before they get here. An item with no JSON form
at all (undefined, a function, a symbol) throws a TypeError naming the item,
counted from 1, as the second half of the example shows. So does an item
JSON.stringify throws on, such as a BigInt. Items before it have already
been written.
From JSON to rows
To write JSON values as CSV or TSV, turn each into an array of strings:
.map((e) => [e.id, String(e.ms)]) after .flatten(), before toCsv().
See
fromJsonToRows
and toJson for the
reference.