Skip to main content

Command Palette

Search for a command to run...

Unix Philosophy in Node.js: Building CLI Tools That Compose

Published
•13 min read•View as Markdown

Unix Philosophy in Node.js: Building CLI Tools That Compose

The Unix philosophy is deceptively simple: write programs that do one thing well, write programs that work together, and write programs that handle text streams. These principles, articulated by Doug McIlroy in 1978, remain the most powerful software design paradigm ever conceived. Yet most Node.js CLI tools ignore them entirely, building monolithic commands that can't talk to anything else.

After building 35+ CLI tools in Node.js, I've learned that the difference between a toy project and a production utility comes down to one question: can it compose? In this article, we'll build CLI tools that pipe, transform, and chain together like proper Unix citizens.

The Unix Philosophy, Distilled

Ken Thompson summarized it best: "When in doubt, use brute force." But more practically, the Unix philosophy boils down to three rules:

  1. Do one thing well. A tool that parses JSON shouldn't also format tables.
  2. Expect your output to become someone else's input. Always write to stdout in a machine-readable format.
  3. Use text streams as the universal interface. Everything is a string of bytes flowing through a pipe.

The shell pipe operator | is the glue. It connects stdout of one process to stdin of the next. This simple mechanism lets you build arbitrarily complex workflows from small, focused tools:

cat server.log | grep "ERROR" | cut -d' ' -f3 | sort | uniq -c | sort -rn

Six tiny tools. Zero shared state. Infinite composability. Let's bring this to Node.js.

Reading from stdin: The Foundation

Every composable CLI tool starts with stdin. In Node.js, process.stdin is a readable stream that represents piped input:

#!/usr/bin/env node

let input = '';

process.stdin.setEncoding('utf8');

process.stdin.on('data', (chunk) => {
  input += chunk;
});

process.stdin.on('end', () => {
  const result = processData(input);
  process.stdout.write(result);
});

function processData(data) {
  // Your transformation logic here
  return data.toUpperCase();
}

This tool reads everything piped into it, transforms it, and writes the result to stdout. It works with any upstream tool:

echo "hello world" | node uppercase.js
# HELLO WORLD

cat README.md | node uppercase.js
# (entire file in uppercase)

But there's a problem. Run it without piping anything, and it hangs, waiting for input that never comes. We need to detect whether we're receiving piped data.

Detecting TTY vs Piped Input

The process.stdin.isTTY property tells you whether stdin is connected to a terminal (interactive) or a pipe:

#!/usr/bin/env node

if (process.stdin.isTTY) {
  // No piped input — run in standalone mode
  console.log('Usage: echo "data" | mytool');
  console.log('  or: mytool --file input.txt');
  process.exit(0);
}

// Piped input — process it
let input = '';
process.stdin.setEncoding('utf8');
process.stdin.on('data', (chunk) => { input += chunk; });
process.stdin.on('end', () => {
  process.stdout.write(transform(input));
});

This pattern is crucial. A well-behaved CLI tool should:

  • When piped to: read from stdin silently, write results to stdout
  • When run interactively: show help, accept arguments, or start an interactive mode

Here's a more complete pattern that handles both modes:

#!/usr/bin/env node
import { readFileSync } from 'fs';

const args = process.argv.slice(2);

async function main() {
  let input;

  if (!process.stdin.isTTY) {
    // Piped mode: read from stdin
    input = await readStdin();
  } else if (args[0]) {
    // File argument mode
    input = readFileSync(args[0], 'utf8');
  } else {
    console.error('Usage: mytool <file> or echo "data" | mytool');
    process.exit(1);
  }

  const result = transform(input);
  process.stdout.write(result);
}

function readStdin() {
  return new Promise((resolve) => {
    let data = '';
    process.stdin.setEncoding('utf8');
    process.stdin.on('data', (chunk) => { data += chunk; });
    process.stdin.on('end', () => resolve(data));
  });
}

main();

stdout for Data, stderr for Status

This is the rule most developers break first, and it destroys composability. Here it is plainly: stdout is for data. stderr is for everything else.

Progress bars, status messages, warnings, debug info — all of it goes to stderr. If you write a "Processing..." message to stdout, it contaminates the data stream and breaks every downstream pipe.

// WRONG — this breaks piping
console.log('Processing 1,547 records...');  // goes to stdout!
console.log(JSON.stringify(results));

// RIGHT — status to stderr, data to stdout
console.error('Processing 1,547 records...');  // stderr
process.stdout.write(JSON.stringify(results));  // stdout

The user sees both streams in their terminal. But when piped, only stdout flows to the next tool. stderr stays visible to the user:

cat data.csv | mytool transform | mytool format > output.json
# "Processing 1,547 records..." appears on screen (stderr)
# Only the JSON data flows through the pipe (stdout)

I wrap this in a utility across all my tools:

const log = {
  info: (msg) => process.stderr.write(`${msg}\n`),
  warn: (msg) => process.stderr.write(`⚠ ${msg}\n`),
  error: (msg) => process.stderr.write(`ERROR: ${msg}\n`),
  data: (obj) => process.stdout.write(
    typeof obj === 'string' ? obj : JSON.stringify(obj)
  ),
};

Line-by-Line Processing with readline

Buffering all input into memory doesn't scale. For large files or continuous streams, process line by line using Node's readline module:

#!/usr/bin/env node
import { createInterface } from 'readline';

const rl = createInterface({
  input: process.stdin,
  crlfDelay: Infinity,
});

let lineNumber = 0;

rl.on('line', (line) => {
  lineNumber++;

  // Example: filter lines containing "error" (case-insensitive)
  if (line.toLowerCase().includes('error')) {
    process.stdout.write(`${lineNumber}: ${line}\n`);
  }
});

rl.on('close', () => {
  process.stderr.write(`Scanned ${lineNumber} lines\n`);
});

This processes each line as it arrives, keeping memory constant regardless of input size. It also starts producing output immediately — the next tool in the pipe doesn't have to wait for all input to be consumed.

Line-by-line processing is the Unix way. It enables real streaming pipelines where data flows through all tools simultaneously.

Transform Streams as Unix Filters

Node.js Transform streams are the perfect abstraction for Unix-style filters. A Transform stream reads input, modifies it, and writes the result — exactly what a pipe stage does:

#!/usr/bin/env node
import { Transform } from 'stream';
import { pipeline } from 'stream/promises';

const csvToJson = new Transform({
  objectMode: false,
  construct(callback) {
    this.headers = null;
    this.buffer = '';
    callback();
  },
  transform(chunk, encoding, callback) {
    this.buffer += chunk.toString();
    const lines = this.buffer.split('\n');
    this.buffer = lines.pop(); // Keep incomplete last line

    for (const line of lines) {
      if (!line.trim()) continue;

      const values = line.split(',').map((v) => v.trim());

      if (!this.headers) {
        this.headers = values;
        continue;
      }

      const obj = {};
      this.headers.forEach((h, i) => {
        obj[h] = values[i] || '';
      });

      this.push(JSON.stringify(obj) + '\n');
    }
    callback();
  },
  flush(callback) {
    // Process any remaining data in buffer
    if (this.buffer.trim() && this.headers) {
      const values = this.buffer.split(',').map((v) => v.trim());
      const obj = {};
      this.headers.forEach((h, i) => {
        obj[h] = values[i] || '';
      });
      this.push(JSON.stringify(obj) + '\n');
    }
    callback();
  },
});

await pipeline(process.stdin, csvToJson, process.stdout);

Now you have a proper Unix filter:

cat users.csv | node csv-to-json.js | grep "admin" | jq '.email'

The pipeline() function from stream/promises handles backpressure correctly and propagates errors — always use it instead of manually piping.

NDJSON: The Interchange Format

CSV is fragile. Plain text loses structure. JSON is great but can't be streamed line by line (you need the complete document to parse it). The solution is NDJSON — Newline Delimited JSON:

{"name":"alice","role":"admin","active":true}
{"name":"bob","role":"user","active":false}
{"name":"carol","role":"admin","active":true}

Each line is a complete, valid JSON object. This format is perfect for Unix pipes because:

  • Streamable: each line can be parsed independently
  • Structured: preserves types, nesting, and arrays
  • Composable: grep, jq, sed all work on it
  • Appendable: just add another line

Here's a reusable NDJSON filter pattern:

#!/usr/bin/env node
import { createInterface } from 'readline';

const rl = createInterface({
  input: process.stdin,
  crlfDelay: Infinity,
});

const filterField = process.argv[2];  // e.g., "role=admin"
const [key, value] = (filterField || '').split('=');

rl.on('line', (line) => {
  if (!line.trim()) return;

  try {
    const obj = JSON.parse(line);

    if (!filterField || obj[key] === value) {
      process.stdout.write(JSON.stringify(obj) + '\n');
    }
  } catch {
    process.stderr.write(`Skipping invalid JSON: ${line}\n`);
  }
});

Usage:

cat users.ndjson | node ndjson-filter.js "role=admin" | node ndjson-pick.js "name,email"

I use NDJSON as the default output format in all my CLI tools. When a tool produces multiple records, it emits one JSON object per line. This makes every tool instantly composable with every other tool.

Exit Codes as Communication

In a pipeline, exit codes tell the shell whether each stage succeeded. This is how set -o pipefail and && chains work:

#!/usr/bin/env node

try {
  const result = await processInput();

  if (result.length === 0) {
    process.stderr.write('No matches found\n');
    process.exit(1);  // Failure — no results
  }

  result.forEach((r) => {
    process.stdout.write(JSON.stringify(r) + '\n');
  });

  process.exit(0);  // Success
} catch (err) {
  process.stderr.write(`Fatal: ${err.message}\n`);
  process.exit(2);  // Error — something broke
}

Follow the convention:

Exit CodeMeaning
0Success
1General failure (no results, validation error)
2Misuse (bad arguments, missing input)
126Permission denied
127Command not found
130Interrupted (Ctrl+C)

Handle SIGINT gracefully to avoid zombie processes in pipelines:

process.on('SIGINT', () => {
  process.stderr.write('\nInterrupted\n');
  process.exit(130);
});

process.on('SIGPIPE', () => {
  // Downstream consumer closed — exit silently
  process.exit(0);
});

The SIGPIPE handler is especially important. When a downstream tool like head closes early, your tool receives SIGPIPE. Without handling it, Node.js throws an unhandled error. With the handler, it exits cleanly.

Building a Pipeline of Our Own Tools

Let's build three small tools that compose into a powerful pipeline. Imagine we're analyzing npm package data.

Tool 1: pkg-fetch — Fetches package metadata as NDJSON

#!/usr/bin/env node
// pkg-fetch: emit package info as NDJSON
import https from 'https';

const packages = process.argv.slice(2);

if (!packages.length && process.stdin.isTTY) {
  process.stderr.write('Usage: pkg-fetch <pkg1> <pkg2> ...\n');
  process.exit(2);
}

// Read package names from stdin if piped
const names = packages.length
  ? packages
  : (await readStdin()).split('\n').filter(Boolean);

for (const name of names) {
  try {
    const data = await fetchPackage(name);
    const record = {
      name: data.name,
      version: data['dist-tags']?.latest,
      description: data.description,
      downloads: data.time ? Object.keys(data.time).length : 0,
      license: data.license,
      homepage: data.homepage,
    };
    process.stdout.write(JSON.stringify(record) + '\n');
  } catch (err) {
    process.stderr.write(`Failed to fetch ${name}: ${err.message}\n`);
  }
}

Tool 2: pkg-filter — Filters NDJSON by field value

#!/usr/bin/env node
// pkg-filter: filter NDJSON records
import { createInterface } from 'readline';

const [field, op, value] = parseFilter(process.argv[2]);

const rl = createInterface({ input: process.stdin, crlfDelay: Infinity });

rl.on('line', (line) => {
  if (!line.trim()) return;
  const obj = JSON.parse(line);

  if (matchesFilter(obj[field], op, value)) {
    process.stdout.write(line + '\n');
  }
});

function parseFilter(expr) {
  if (!expr) return [null, null, null];
  const match = expr.match(/^(\w+)(>=|<=|!=|=|>|<)(.+)$/);
  return match ? [match[1], match[2], match[3]] : [null, null, null];
}

function matchesFilter(fieldVal, op, target) {
  if (!op) return true;
  const num = Number(target);
  const val = isNaN(num) ? target : num;
  const fv = typeof fieldVal === 'number' ? fieldVal : fieldVal;

  switch (op) {
    case '=':  return fv == val;
    case '!=': return fv != val;
    case '>':  return fv > val;
    case '<':  return fv < val;
    case '>=': return fv >= val;
    case '<=': return fv <= val;
    default:   return true;
  }
}

Tool 3: pkg-format — Formats NDJSON as a table

#!/usr/bin/env node
// pkg-format: render NDJSON as a table
import { createInterface } from 'readline';

const fields = (process.argv[2] || '').split(',').filter(Boolean);
const rows = [];

const rl = createInterface({ input: process.stdin, crlfDelay: Infinity });

rl.on('line', (line) => {
  if (!line.trim()) return;
  rows.push(JSON.parse(line));
});

rl.on('close', () => {
  if (!rows.length) {
    process.stderr.write('No data to format\n');
    process.exit(1);
  }

  const cols = fields.length ? fields : Object.keys(rows[0]);
  const widths = cols.map((c) =>
    Math.max(c.length, ...rows.map((r) => String(r[c] ?? '').length))
  );

  // Header
  const header = cols.map((c, i) => c.padEnd(widths[i])).join('  ');
  const separator = widths.map((w) => '-'.repeat(w)).join('  ');

  process.stdout.write(header + '\n');
  process.stdout.write(separator + '\n');

  for (const row of rows) {
    const line = cols
      .map((c, i) => String(row[c] ?? '').padEnd(widths[i]))
      .join('  ');
    process.stdout.write(line + '\n');
  }
});

Now compose them:

# Fetch packages, filter by license, display as table
echo -e "express\nlodash\nreact\nvue" \
  | pkg-fetch \
  | pkg-filter "license=MIT" \
  | pkg-format "name,version,license,description"
name     version  license  description
-------  -------  -------  -------------------------------------------
express  4.21.1   MIT      Fast, unopinionated web framework for node.
lodash   4.17.21  MIT      Lodash modular utilities.
react    18.3.1   MIT      React is a JavaScript library for UIs.
vue      3.5.12   MIT      The progressive JavaScript framework.

Each tool is under 50 lines. Each does exactly one thing. Together they form a powerful data pipeline.

Real Composition With Existing Unix Tools

The beauty of following Unix conventions is that your Node.js tools instantly work with the entire Unix ecosystem:

# Count packages by license type
cat packages.txt | pkg-fetch | jq -r '.license' | sort | uniq -c | sort -rn

# Find packages without homepages
cat packages.txt | pkg-fetch | pkg-filter "homepage=" | jq -r '.name'

# Monitor for new versions (combine with watch)
watch -n 60 'echo "express" | pkg-fetch | jq .version'

# Feed results into xargs for bulk operations
cat packages.txt | pkg-fetch | pkg-filter "downloads>100" \
  | jq -r '.name' | xargs -I {} npm info {} dist-tags.latest

# Diff package states over time
pkg-fetch express lodash > snapshot-monday.ndjson
# ... later ...
pkg-fetch express lodash > snapshot-friday.ndjson
diff <(jq -S . snapshot-monday.ndjson) <(jq -S . snapshot-friday.ndjson)

With 35+ tools in our toolkit, the composition possibilities multiply. Here are patterns I use daily:

# Audit dependencies, filter vulnerabilities, generate a report
npmdeps list | depcheck-ai scan | pkg-format "name,severity,fix"

# Scan repo stats across multiple projects
ls ~/projects | xargs -I {} gitstats ~/projects/{} \
  | ndjson-filter "commits>100" | pkg-format "repo,commits,contributors"

# Chain URL analysis with metadata extraction
cat urls.txt | urlmeta extract | ndjson-filter "status=200" \
  | jq '{url, title: .meta.title}' | pkg-format "url,title"

The Composability Checklist

Before publishing any CLI tool, I run through this checklist:

  1. Does it read from stdin when piped? Use !process.stdin.isTTY to detect.
  2. Does it write data to stdout only? Status, progress, and errors go to stderr.
  3. Does it output structured data? NDJSON for multiple records, JSON for single objects.
  4. Does it use meaningful exit codes? 0 for success, 1 for failure, 2 for misuse.
  5. Does it handle SIGPIPE? Exit cleanly when downstream closes.
  6. Does it work line-by-line? Avoid buffering entire input when possible.
  7. Does it accept file arguments too? mytool file.txt should work alongside cat file.txt | mytool.
  8. Is it silent on success in pipe mode? No banners, no "Done!" messages to stdout.

A tool that passes all eight checks is a proper Unix citizen. It works with grep, sort, jq, xargs, tee, head, tail, and every other tool in the ecosystem — including your other Node.js tools.

The Template

Here's the minimal template I use to start every new CLI tool:

#!/usr/bin/env node
import { createInterface } from 'readline';
import { readFileSync } from 'fs';

const args = process.argv.slice(2);

// Handle --help
if (args.includes('--help') || args.includes('-h')) {
  process.stderr.write(`Usage: mytool [options] [file]

  Reads from stdin or file. Outputs NDJSON to stdout.

  Options:
    -h, --help    Show this help
    -q, --quiet   Suppress status messages

  Examples:
    cat data.csv | mytool
    mytool data.csv
    mytool data.csv | other-tool | jq '.field'
`);
  process.exit(0);
}

const quiet = args.includes('-q') || args.includes('--quiet');
const log = (msg) => { if (!quiet) process.stderr.write(msg + '\n'); };

async function main() {
  const input = !process.stdin.isTTY
    ? process.stdin
    : args[0]
      ? createReadStreamFromFile(args[0])
      : (log('No input. Use --help for usage.'), process.exit(2));

  const rl = createInterface({ input, crlfDelay: Infinity });
  let count = 0;

  for await (const line of rl) {
    if (!line.trim()) continue;
    const result = processLine(line);
    if (result) {
      process.stdout.write(JSON.stringify(result) + '\n');
      count++;
    }
  }

  log(`Processed ${count} records`);
}

function processLine(line) {
  // Your logic here
  return { raw: line, length: line.length };
}

// Graceful shutdown
process.on('SIGINT', () => process.exit(130));
process.on('SIGPIPE', () => process.exit(0));

main().catch((err) => {
  process.stderr.write(`Fatal: ${err.message}\n`);
  process.exit(1);
});

This template gives you stdin/file input, NDJSON output, proper exit codes, signal handling, and quiet mode — in under 50 lines. Every tool I build starts here.

Closing Thoughts

The Unix philosophy isn't nostalgia. It's engineering pragmatism. Small, composable tools are easier to test, easier to debug, easier to maintain, and infinitely more flexible than monolithic commands.

Node.js gives us everything we need: process.stdin, process.stdout, process.stderr, Transform streams, readline, and excellent process management. The question isn't whether Node.js can build Unix-style tools — it's why so few developers bother.

Build your next CLI tool like it's 1978. Read from stdin. Write to stdout. Do one thing well. Then pipe it into something beautiful.

More from this blog

W

Wilson Xu

108 posts