Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Searching logs

The task: find the database errors in an application’s logs, where the current file is plain text and the rotated ones are gzipped (app.log, app.log.1.gz, app.log.2.gz, as logrotate leaves them).

import { enumerate, read } from "@j50n/proc";

/** The lines of a log file, unzipping it first if it ends in `.gz`. */
function logLines(path: string) {
  const bytes = read(path);
  return path.endsWith(".gz")
    ? bytes.transform(new DecompressionStream("gzip")).lines
    : bytes.lines;
}

// app.log, app.log.1.gz, app.log.2.gz, ...
const dir = "recipes-logs";
const files: string[] = [];
for await (const entry of Deno.readDir(dir)) {
  if (entry.name.startsWith("app.log")) files.push(`${dir}/${entry.name}`);
}
// Oldest first: app.log.2.gz, app.log.1.gz, app.log (a plain sort would put
// app.log.10.gz before app.log.2.gz).
const age = (path: string) => Number(path.match(/\.(\d+)\.gz$/)?.[1] ?? 0);
files.sort((a, b) => age(b) - age(a));

await enumerate(files)
  .flatMap(logLines)
  .filter((line) => line.includes(" ERROR db:"))
  .toStdout();
2026-10-03 08:14:02 ERROR db: connection refused
2026-10-03 17:42:30 ERROR db: query timeout after 30s
2026-10-04 06:31:41 ERROR db: query timeout after 30s
2026-10-04 13:48:09 ERROR db: connection refused
2026-10-05 09:01:13 ERROR db: connection refused

logLines() is the one piece of logic: a .gz file goes through DecompressionStream before .lines, a plain one doesn’t. flatMap reads the files one after another as a single stream of lines, so memory stays flat however big the logs are, and toStdout() prints each match as it is found. Deno.readDir lists files in no particular order, hence the sort.

The script needs --allow-read=recipes-logs and nothing else: no command is run.

Summarizing

Counting is a forEach into a Map. Slicing fixed columns is the fastest way to pick a line apart when the format is fixed; use a regular expression when it isn’t.

import { enumerate, read } from "@j50n/proc";

const files = ["app.log.2.gz", "app.log.1.gz", "app.log"]
  .map((name) => `recipes-logs/${name}`);

const perDay = new Map<string, number>();
const perSource = new Map<string, number>();

await enumerate(files)
  .flatMap((path) =>
    path.endsWith(".gz")
      ? read(path).transform(new DecompressionStream("gzip")).lines
      : read(path).lines
  )
  .filter((line) => line.slice(20, 25) === "ERROR")
  .forEach((line) => {
    // 2026-10-05 09:01:13 ERROR db: connection refused
    const day = line.slice(0, 10);
    const source = line.slice(26).split(":")[0];
    perDay.set(day, (perDay.get(day) ?? 0) + 1);
    perSource.set(source, (perSource.get(source) ?? 0) + 1);
  });

console.log("errors per day:", Object.fromEntries(perDay));
const ranked = [...perSource].sort(([, a], [, b]) => b - a);
console.log("by source:", ranked.map(([s, n]) => `${s} ${n}`).join(", "));
errors per day: { "2026-10-03": 3, "2026-10-04": 4, "2026-10-05": 2 }
by source: db 5, http 3, disk 1

Timestamps in this format sort as strings, so a time window is a string comparison: .filter((line) => line >= "2026-10-04 06:00" && line < "2026-10-04 14:00").

Letting grep do the searching

For gigabytes of logs, grep finds matches faster than a JavaScript filter, and zgrep reads plain and gzipped files alike. The trap is that grep exits with code 1 when nothing matches, which proc, like set -e, treats as a failure. An fnError handler that lets exit code 1 through turns “no matches” back into “no lines”:

import { ExitCodeError, run } from "@j50n/proc";

/** grep and zgrep exit 1 when nothing matches; treat that as no lines. */
function noMatchIsFine(error?: Error) {
  if (error instanceof ExitCodeError && error.code === 1) return;
  if (error) throw error;
}

const logs = ["app.log.2.gz", "app.log.1.gz", "app.log"]
  .map((name) => `recipes-logs/${name}`);

for (const pattern of ["timeout", "segfault"]) {
  const hits = await run(
    { fnError: noMatchIsFine },
    "zgrep",
    "-h",
    pattern,
    ...logs,
  ).lines.collect();
  console.log(`${pattern}: ${hits.length}`, hits);
}
timeout: 2 [
  "2026-10-03 17:42:30 ERROR db: query timeout after 30s",
  "2026-10-04 06:31:41 ERROR db: query timeout after 30s"
]
segfault: 0 []

Exit code 2 (a missing file, a bad pattern) still throws ExitCodeError. Errors has more on fnError.

The same applies to grep in the middle of a pipeline: read(path).run({ fnError: noMatchIsFine }, "grep", "ERROR").

Variations

  • Following a log as it grows: read() stops at the end of the file. For tail -f, run it: await run("tail", "-F", "app.log").lines.forEach(...) goes on until you stop it.
  • Many large files at once: wrap the per-file work in concurrentMap, one file per call, and combine the counts at the end.
  • The first match only: .find((line) => ...) stops reading at the match and closes the file.
  • Structured logs (one JSON object per line): .transform(jsonParse) after .lines, or see JSON lines.
  • Writing the matches to a file: end with .writeTo(path) instead of .toStdout(); it writes each line with a newline, as toStdout does.