Summary
Schema-Driven Validation encodes data expectations — types, nullability, value ranges, referential constraints — as explicit schema definitions. Data is validated against the schema automatically, separating the declaration of what is valid from the logic that checks it. Violations surface as structured errors rather than silent corruption.
Problem
Imperative validation (if column is null, if value not in set) is scattered, hard to audit, and silently fails to cover edge cases. As data volumes and sources grow, ad-hoc checks become unmaintainable and incomplete.
- How do you ensure every column is validated consistently across all pipeline runs?
- How do you communicate data contracts to upstream producers?
- How do you catch schema drift (source system adds/removes columns) automatically?
Solution
Define a schema object that describes the expected structure and constraints. Pass data through the schema validator; handle pass/fail at the pipeline boundary.
import pandera as pa
order_schema = pa.DataFrameSchema({
"order_id": pa.Column(str, nullable=False, unique=True),
"total": pa.Column(float, pa.Check.greater_than(0)),
"status": pa.Column(str, pa.Check.isin(["pending", "complete", "cancelled"])),
"created_at": pa.Column(pa.DateTime, nullable=False),
})
# Validation is declarative and automatic
validated_df = order_schema.validate(raw_df) # raises SchemaError on failure
Schemas serve as living documentation of data contracts. They can be versioned, shared with upstream teams, and used to generate synthetic test data.
The same declarative approach applies in TypeScript pipelines, where a schema library like Zod replaces scattered if checks with a single parseable contract:
import { z } from "zod";
const OrderSchema = z.object({
orderId: z.string().uuid(),
total: z.number().positive(),
status: z.enum(["pending", "complete", "cancelled"]),
createdAt: z.coerce.date(),
});
type Order = z.infer<typeof OrderSchema>;
interface ValidationResult {
valid: Order[];
rejected: { record: unknown; reason: string }[];
}
function validateBatch(records: unknown[]): ValidationResult {
const valid: Order[] = [];
const rejected: ValidationResult["rejected"] = [];
for (const record of records) {
const result = OrderSchema.safeParse(record);
if (result.success) {
valid.push(result.data);
} else {
rejected.push({ record, reason: result.error.issues[0]?.message ?? "invalid" });
}
}
return { valid, rejected };
}
OrderSchema is the single source of truth: the type Order is inferred from it rather than kept in sync by hand, and safeParse turns validation into a data transformation — bad records are routed to rejected instead of throwing mid-pipeline.
Live Playground
Experiment with the pattern below. The two tabs show an imperative validator (before.ts) with scattered if checks that silently drop context on failure, and a schema-driven equivalent (after.ts) where a single Zod schema is both the type source and the validation contract.
When to Use
- Bronze-to-Silver transitions in a Medallion Architecture — validate before promoting data
- Any pipeline that accepts data from external sources (APIs, upstream teams, user uploads)
- Teams that need to communicate data contracts formally
- Anywhere silent data corruption is more dangerous than a failed pipeline run
Avoid when:
- Schema-free document stores where structure is intentionally variable — validate at the application layer instead
- Extremely high-throughput streaming where per-record validation overhead is prohibitive (validate samples instead)
Trade-offs
| Benefit | Cost |
|---|---|
| Expectations are explicit, auditable, and version-controlled | Schema must be maintained as data evolves |
| Violations surface immediately with structured error detail | Strict schemas can break on legitimate upstream changes |
| Schemas serve as documentation and contracts | Teams must align on schema governance process |
| Enables automated quality reporting | Adds a validation step to pipeline runtime |
Related Patterns
- Pure Functions — Validation logic should be a pure function: data in, result out
- Medallion Architecture — Schema validation is the quality gate at the Bronze-to-Silver boundary
- Batch vs Streaming — schema validation applies at ingestion regardless of execution model; streaming pipelines validate per-event, batch pipelines validate per-partition
- Dependency Injection — Schema objects can be injected to allow environment-specific validation rules
- CQRS — command objects on the write side are validated against a schema before being accepted by command handlers