Summary
A pure function always produces the same output for the same input and has no observable side effects — no database writes, no external API calls, no mutation of shared state. In data pipelines, applying this principle to transformation logic dramatically improves testability, composability, and reproducibility.
Problem
Pipeline transformations that mix computation with I/O are difficult to test, reason about, and reuse. A function that reads from a database, transforms data, and writes results is doing three jobs and cannot be safely re-executed or unit-tested without infrastructure.
- How do you unit-test a transformation without a running Spark cluster or database?
- How do you guarantee that rerunning a pipeline step produces the same result?
- How do you compose small transformations into complex pipelines without coupling them?
Solution
Separate transformation logic from I/O. Transformations take data in, return data out, and touch nothing else.
# Impure — reads and writes mixed with logic
def process_orders(conn):
orders = conn.execute("SELECT * FROM orders").fetchall()
result = [o for o in orders if o['total'] > 100]
conn.execute("INSERT INTO high_value_orders VALUES ...", result)
# Pure — just the transformation
def filter_high_value(orders: list[dict], threshold: float) -> list[dict]:
return [o for o in orders if o['total'] > threshold]
# I/O happens at the boundary, separately
orders = read_orders(conn)
high_value = filter_high_value(orders, threshold=100.0)
write_orders(conn, high_value)
The transformation filter_high_value is trivially unit-testable, composable, and deterministic.
The same separation holds in TypeScript pipelines: the transformation takes plain data and returns plain data, with I/O pushed to the caller.
interface Order {
id: string;
total: number;
}
// Impure — reads and writes mixed with logic
async function processOrders(db: Database): Promise<void> {
const orders = await db.query<Order>("SELECT * FROM orders");
const result = orders.filter((o) => o.total > 100);
await db.execute("INSERT INTO high_value_orders VALUES ...", result);
}
// Pure — just the transformation
function filterHighValue(orders: Order[], threshold: number): Order[] {
return orders.filter((o) => o.total > threshold);
}
// I/O happens at the boundary, separately
const orders = await readOrders(db);
const highValue = filterHighValue(orders, 100.0);
await writeOrders(db, highValue);
filterHighValue takes no db, no Date.now(), no ambient state — every unit test supplies fixtures directly and asserts on the return value, with no mocks or fakes required.
Live Playground
Experiment with the pattern below. The two tabs show an impure transform (before.ts) that mixes a fake “database read,” a running total mutation, and filtering into one function, and a pure equivalent (after.ts) that separates I/O from computation.
When to Use
- Any data transformation logic: filtering, mapping, aggregating, joining, reshaping
- Pipeline steps that need to be unit-tested without infrastructure
- Transformations intended to be reused across multiple pipeline stages
- Anywhere reproducibility and idempotency are requirements
Avoid when:
- The operation is inherently stateful or I/O-bound (e.g., deduplication using a lookup table) — isolate the I/O, keep the comparison logic pure
Trade-offs
| Benefit | Cost |
|---|---|
| Unit-testable with simple data fixtures — no mocks needed | Requires architectural discipline to push I/O to boundaries |
| Deterministic: same input always gives same output | Can feel unnatural when the natural framing is stateful |
| Composable: pipe outputs of one function into another | Large intermediate datasets may have memory implications |
| Parallelizable: pure functions are inherently safe to run concurrently | — |
Related Patterns
- Schema-Driven Validation — Validation is a pure transformation: data in, valid/invalid out
- Medallion Architecture — Silver and Gold transformations should be pure functions over Bronze data
- Batch vs Streaming — pure transformation logic is execution-model agnostic: the same functions run in batch Spark jobs and streaming Flink operators
- Single Responsibility — Pure functions are the function-level expression of single responsibility