LazyRow
A LazyRow is a row that decodes a field only when you read it. Use one when a
pipeline filters on a field or two, or reads a few fields of wide rows.
import { read } from "@j50n/proc";
import { fromCsvToLazyRows, toCsv } from "@j50n/proc/transforms";
await read("data-orders.csv")
.transform(fromCsvToLazyRows())
.flatten()
.drop(1) // the header
.filter((row) => !row.fieldEquals(3, "1")) // compares bytes; decodes nothing
.map((row) => {
const fields = row.toStringArray(); // a LazyRow can't change; its array can
fields[1] = fields[1].toUpperCase();
return fields;
})
.transform(toCsv())
.toStdout();
1,ADA,"Widget, large",2
3,LINUS,Bolt,10
fromCsvToLazyRows() and fromTsvToLazyRows() yield rows that are views of the
bytes the WebAssembly parser produced. getField(i) decodes field i alone,
and fieldEquals(i, value) compares field i with value byte by byte,
without making a string, so the filter above decodes nothing. It is the fastest
way to filter: about three times the speed of fromCsvToRows(), as
How fast shows.
Methods
| Member | What it does |
|---|---|
columnCount | the number of fields (a property) |
getField(i) | field i, counted from 0; RangeError outside [0, columnCount) |
fieldEquals(i, value) | whether field i is exactly value; same RangeError |
toStringArray() | all fields as a new string[] |
LazyRow.fromStringArray(arr) | wrap a string[], without copying it |
A LazyRow can’t be changed. To change a row, take toStringArray() and change
the array, as the example does; the writers take plain rows as well. Every row
writer (toCsv, toTsv, toRecord) takes LazyRows, singly or in batches,
mixed with plain rows if you like.
When it doesn’t help
fromRecordToLazyRows() splits every record into strings first and wraps the
array, so it does all the work fromRecordToRows() does. If you read every
field anyway, toStringArray() decodes them all, and a plain Row is simpler.
Traps
- A row from CSV or TSV keeps its whole batch alive, the rows of about 128 KiB
of input. To hold on to a few rows from a large stream, keep their
toStringArray()instead. - Invalid UTF-8 throws a
TypeErrorwhen it is decoded, not from the parser, so a bad field nobody reads goes unnoticed.toStringArray()decodes the row’s whole batch at once, so it throws if any row in the batch is bad.
See LazyRow for the
reference.