Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

JSON lines

JSON lines is one JSON value per line. fromJsonToRows() parses each line with JSON.parse, and toJson() writes each value with JSON.stringify.

import { read } from "@j50n/proc";
import { fromJsonToRows } from "@j50n/proc/transforms";

type Event = { id: string; level: string; ms: number };

const slow = await read("data-events.jsonl")
  .transform(fromJsonToRows<Event>()) // the type is asserted, not checked
  .flatten()
  .filter((event) => event.ms > 10)
  .map((event) => event.id)
  .collect();

console.log(slow);
[ "e1", "e2" ]

Despite its name, fromJsonToRows() yields values, not rows: objects, arrays, strings, numbers, anything JSON holds. They come in batches of about 128 KiB of text, so .flatten() before working value by value. Blank lines, and lines of only whitespace, are skipped. A CR at the end of a line is whitespace to JSON.parse, so CRLF files read fine.

The type parameter (fromJsonToRows<Event>()) is only an assertion. Nothing checks the values unless you pass a schema.

Checking values

import { enumerate } from "@j50n/proc";
import { fromJsonToRows } from "@j50n/proc/transforms";

type Event = { id: string; ms: number };

// Anything with a parse() that throws on a bad value works; a Zod schema does.
const EventSchema = {
  parse(value: unknown): Event {
    const e = value as Partial<Event>;
    if (typeof e?.id !== "string" || typeof e.ms !== "number") {
      throw new TypeError(`not an event: ${JSON.stringify(value)}`);
    }
    return e as Event;
  },
};

const text = '{"id":"e1","ms":12}\n{"id":"e2","ms":"slow"}\n';

try {
  await enumerate([new TextEncoder().encode(text)])
    .transform(fromJsonToRows({ schema: EventSchema }))
    .flatten()
    .forEach((event) => console.log(event.id));
} catch (error) {
  if (error instanceof TypeError) console.log(error.message);
}
not an event: {"id":"e2","ms":"slow"}
OptionDefaultMeaning
schemanonean object whose parse(value) returns the value to yield, or throws
sampleSizeevery valuecheck only the first sampleSize values

Whatever schema.parse throws stops the stream and reaches your catch as it is, so instanceof finds a Zod error. What it returns is the value you get, so a Zod schema’s defaults and transforms apply, and keys it doesn’t know are stripped, as with any parse call.

With sampleSize, only the first sampleSize values go through the schema. The rest come as JSON.parse made them, unchecked and untransformed, though they are typed as the schema’s output, just as values are with no schema at all. So use sampleSize only with a schema that checks values without changing them.

Note that e1 never printed. A batch is parsed whole before it is yielded, so an error stops the stream before any value in the same batch reaches you.

A line that isn’t JSON throws a SyntaxError naming the line, counted from 1 with blank lines included: Invalid JSON at line 4001: Unexpected token ....

Writing

import { enumerate, read } from "@j50n/proc";
import { fromCsvToRows, toJson } from "@j50n/proc/transforms";

// toJson takes one value per item, so flatten the parser's batches first.
await read("data-orders.csv")
  .transform(fromCsvToRows())
  .flatten()
  .drop(1) // the header
  .map(([id, customer, , qty]) => ({ id, customer, qty: Number(qty) }))
  .transform(toJson())
  .toStdout();

// A value with no JSON form is refused, naming the item.
try {
  await enumerate([{ total: 13 }, undefined]).transform(toJson()).toStdout();
} catch (error) {
  if (error instanceof TypeError) console.log(error.message);
}
{"id":"1","customer":"Ada","qty":2}
{"id":"2","customer":"Grace","qty":1}
{"id":"3","customer":"Linus","qty":10}
{"total":13}
Item 2 can't be written as JSON (undefined)

toJson() takes one value per item and writes it on a line of its own, the reverse of fromJsonToRows() followed by .flatten(). The parsers yield batches, so flatten them first; a batch passed as it is would be written as one JSON array on one line.

Inside a value, JSON.stringify’s rules apply: a property holding undefined or a function is left out, and in an array it becomes null. NaN and Infinity become null too, without an error, so check numbers parsed from text (Number("12,5") is NaN) before they get here. An item with no JSON form at all (undefined, a function, a symbol) throws a TypeError naming the item, counted from 1, as the second half of the example shows. So does an item JSON.stringify throws on, such as a BigInt. Items before it have already been written.

From JSON to rows

To write JSON values as CSV or TSV, turn each into an array of strings: .map((e) => [e.id, String(e.ms)]) after .flatten(), before toCsv().

See fromJsonToRows and toJson for the reference.