Searching logs
The task: find the database errors in an application’s logs, where the current
file is plain text and the rotated ones are gzipped (app.log, app.log.1.gz,
app.log.2.gz, as logrotate leaves them).
import { enumerate, read } from "@j50n/proc";
/** The lines of a log file, unzipping it first if it ends in `.gz`. */
function logLines(path: string) {
const bytes = read(path);
return path.endsWith(".gz")
? bytes.transform(new DecompressionStream("gzip")).lines
: bytes.lines;
}
// app.log, app.log.1.gz, app.log.2.gz, ...
const dir = "recipes-logs";
const files: string[] = [];
for await (const entry of Deno.readDir(dir)) {
if (entry.name.startsWith("app.log")) files.push(`${dir}/${entry.name}`);
}
// Oldest first: app.log.2.gz, app.log.1.gz, app.log (a plain sort would put
// app.log.10.gz before app.log.2.gz).
const age = (path: string) => Number(path.match(/\.(\d+)\.gz$/)?.[1] ?? 0);
files.sort((a, b) => age(b) - age(a));
await enumerate(files)
.flatMap(logLines)
.filter((line) => line.includes(" ERROR db:"))
.toStdout();
2026-10-03 08:14:02 ERROR db: connection refused
2026-10-03 17:42:30 ERROR db: query timeout after 30s
2026-10-04 06:31:41 ERROR db: query timeout after 30s
2026-10-04 13:48:09 ERROR db: connection refused
2026-10-05 09:01:13 ERROR db: connection refused
logLines() is the one piece of logic: a .gz file goes through
DecompressionStream before .lines, a plain one doesn’t. flatMap reads the
files one after another as a single stream of lines, so memory stays flat
however big the logs are, and toStdout() prints each match as it is found.
Deno.readDir lists files in no particular order, hence the sort.
The script needs --allow-read=recipes-logs and nothing else: no command is
run.
Summarizing
Counting is a forEach into a Map. Slicing fixed columns is the fastest way
to pick a line apart when the format is fixed; use a regular expression when it
isn’t.
import { enumerate, read } from "@j50n/proc";
const files = ["app.log.2.gz", "app.log.1.gz", "app.log"]
.map((name) => `recipes-logs/${name}`);
const perDay = new Map<string, number>();
const perSource = new Map<string, number>();
await enumerate(files)
.flatMap((path) =>
path.endsWith(".gz")
? read(path).transform(new DecompressionStream("gzip")).lines
: read(path).lines
)
.filter((line) => line.slice(20, 25) === "ERROR")
.forEach((line) => {
// 2026-10-05 09:01:13 ERROR db: connection refused
const day = line.slice(0, 10);
const source = line.slice(26).split(":")[0];
perDay.set(day, (perDay.get(day) ?? 0) + 1);
perSource.set(source, (perSource.get(source) ?? 0) + 1);
});
console.log("errors per day:", Object.fromEntries(perDay));
const ranked = [...perSource].sort(([, a], [, b]) => b - a);
console.log("by source:", ranked.map(([s, n]) => `${s} ${n}`).join(", "));
errors per day: { "2026-10-03": 3, "2026-10-04": 4, "2026-10-05": 2 }
by source: db 5, http 3, disk 1
Timestamps in this format sort as strings, so a time window is a string
comparison:
.filter((line) => line >= "2026-10-04 06:00" && line < "2026-10-04 14:00").
Letting grep do the searching
For gigabytes of logs, grep finds matches faster than a JavaScript filter, and
zgrep reads plain and gzipped files alike. The trap is that grep exits with
code 1 when nothing matches, which proc, like set -e, treats as a failure. An
fnError handler that lets exit code 1 through turns “no matches” back into “no
lines”:
import { ExitCodeError, run } from "@j50n/proc";
/** grep and zgrep exit 1 when nothing matches; treat that as no lines. */
function noMatchIsFine(error?: Error) {
if (error instanceof ExitCodeError && error.code === 1) return;
if (error) throw error;
}
const logs = ["app.log.2.gz", "app.log.1.gz", "app.log"]
.map((name) => `recipes-logs/${name}`);
for (const pattern of ["timeout", "segfault"]) {
const hits = await run(
{ fnError: noMatchIsFine },
"zgrep",
"-h",
pattern,
...logs,
).lines.collect();
console.log(`${pattern}: ${hits.length}`, hits);
}
timeout: 2 [
"2026-10-03 17:42:30 ERROR db: query timeout after 30s",
"2026-10-04 06:31:41 ERROR db: query timeout after 30s"
]
segfault: 0 []
Exit code 2 (a missing file, a bad pattern) still throws ExitCodeError.
Errors has more on fnError.
The same applies to grep in the middle of a pipeline:
read(path).run({ fnError: noMatchIsFine }, "grep", "ERROR").
Variations
- Following a log as it grows:
read()stops at the end of the file. Fortail -f, run it:await run("tail", "-F", "app.log").lines.forEach(...)goes on until you stop it. - Many large files at once: wrap the per-file work in
concurrentMap, one file per call, and combine the counts at the end. - The first match only:
.find((line) => ...)stops reading at the match and closes the file. - Structured logs (one JSON object per line):
.transform(jsonParse)after.lines, or see JSON lines. - Writing the matches to a file: end with
.writeTo(path)instead of.toStdout(); it writes each line with a newline, astoStdoutdoes.