Detail chart · Data Pipeline

Schema-Driven Validation

Pattern ◆◆◇◇◇

Define and enforce data contracts at pipeline boundaries to catch structural violations before they propagate downstream.

Summary

Schema-Driven Validation encodes data expectations — types, nullability, value ranges, referential constraints — as explicit schema definitions. Data is validated against the schema automatically, separating the declaration of what is valid from the logic that checks it. Violations surface as structured errors rather than silent corruption.

Problem

Imperative validation (if column is null, if value not in set) is scattered, hard to audit, and silently fails to cover edge cases. As data volumes and sources grow, ad-hoc checks become unmaintainable and incomplete.

Solution

Define a schema object that describes the expected structure and constraints. Pass data through the schema validator; handle pass/fail at the pipeline boundary.

import pandera as pa

order_schema = pa.DataFrameSchema({
    "order_id":   pa.Column(str,   nullable=False, unique=True),
    "total":      pa.Column(float, pa.Check.greater_than(0)),
    "status":     pa.Column(str,   pa.Check.isin(["pending", "complete", "cancelled"])),
    "created_at": pa.Column(pa.DateTime, nullable=False),
})

# Validation is declarative and automatic
validated_df = order_schema.validate(raw_df)  # raises SchemaError on failure

Schemas serve as living documentation of data contracts. They can be versioned, shared with upstream teams, and used to generate synthetic test data.

The same declarative approach applies in TypeScript pipelines, where a schema library like Zod replaces scattered if checks with a single parseable contract:

import { z } from "zod";

const OrderSchema = z.object({
  orderId: z.string().uuid(),
  total: z.number().positive(),
  status: z.enum(["pending", "complete", "cancelled"]),
  createdAt: z.coerce.date(),
});

type Order = z.infer<typeof OrderSchema>;

interface ValidationResult {
  valid: Order[];
  rejected: { record: unknown; reason: string }[];
}

function validateBatch(records: unknown[]): ValidationResult {
  const valid: Order[] = [];
  const rejected: ValidationResult["rejected"] = [];

  for (const record of records) {
    const result = OrderSchema.safeParse(record);
    if (result.success) {
      valid.push(result.data);
    } else {
      rejected.push({ record, reason: result.error.issues[0]?.message ?? "invalid" });
    }
  }

  return { valid, rejected };
}

OrderSchema is the single source of truth: the type Order is inferred from it rather than kept in sync by hand, and safeParse turns validation into a data transformation — bad records are routed to rejected instead of throwing mid-pipeline.

Live Playground

Experiment with the pattern below. The two tabs show an imperative validator (before.ts) with scattered if checks that silently drop context on failure, and a schema-driven equivalent (after.ts) where a single Zod schema is both the type source and the validation contract.

When to Use

Avoid when:

Trade-offs

BenefitCost
Expectations are explicit, auditable, and version-controlledSchema must be maintained as data evolves
Violations surface immediately with structured error detailStrict schemas can break on legitimate upstream changes
Schemas serve as documentation and contractsTeams must align on schema governance process
Enables automated quality reportingAdds a validation step to pipeline runtime