Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

LazyRow

A LazyRow is a row that decodes a field only when you read it. Use one when a pipeline filters on a field or two, or reads a few fields of wide rows.

import { read } from "@j50n/proc";
import { fromCsvToLazyRows, toCsv } from "@j50n/proc/transforms";

await read("data-orders.csv")
  .transform(fromCsvToLazyRows())
  .flatten()
  .drop(1) // the header
  .filter((row) => !row.fieldEquals(3, "1")) // compares bytes; decodes nothing
  .map((row) => {
    const fields = row.toStringArray(); // a LazyRow can't change; its array can
    fields[1] = fields[1].toUpperCase();
    return fields;
  })
  .transform(toCsv())
  .toStdout();
1,ADA,"Widget, large",2
3,LINUS,Bolt,10

fromCsvToLazyRows() and fromTsvToLazyRows() yield rows that are views of the bytes the WebAssembly parser produced. getField(i) decodes field i alone, and fieldEquals(i, value) compares field i with value byte by byte, without making a string, so the filter above decodes nothing. It is the fastest way to filter: about three times the speed of fromCsvToRows(), as How fast shows.

Methods

MemberWhat it does
columnCountthe number of fields (a property)
getField(i)field i, counted from 0; RangeError outside [0, columnCount)
fieldEquals(i, value)whether field i is exactly value; same RangeError
toStringArray()all fields as a new string[]
LazyRow.fromStringArray(arr)wrap a string[], without copying it

A LazyRow can’t be changed. To change a row, take toStringArray() and change the array, as the example does; the writers take plain rows as well. Every row writer (toCsv, toTsv, toRecord) takes LazyRows, singly or in batches, mixed with plain rows if you like.

When it doesn’t help

fromRecordToLazyRows() splits every record into strings first and wraps the array, so it does all the work fromRecordToRows() does. If you read every field anyway, toStringArray() decodes them all, and a plain Row is simpler.

Traps

  • A row from CSV or TSV keeps its whole batch alive, the rows of about 128 KiB of input. To hold on to a few rows from a large stream, keep their toStringArray() instead.
  • Invalid UTF-8 throws a TypeError when it is decoded, not from the parser, so a bad field nobody reads goes unnoticed. toStringArray() decodes the row’s whole batch at once, so it throws if any row in the batch is bad.

See LazyRow for the reference.