Summary
Medallion Architecture organizes a data lakehouse into three progressive refinement layers — Bronze (raw ingestion), Silver (cleansed and conformed), and Gold (business-ready aggregates). Each layer has a defined contract for data quality, enabling downstream consumers to trust the data they use without re-engineering upstream pipelines.
Problem
Raw data from source systems is messy: schema drift, nulls, duplicates, inconsistent types, and mixed grain. Processing it repeatedly for different consumers leads to duplicated logic, inconsistent results, and fragile pipelines.
- How do you preserve raw data for reprocessing while serving clean data to consumers?
- How do you enforce quality gates without blocking ingestion?
- How do you build domain-level aggregates that multiple teams can trust?
Solution
Divide the data lake into three named zones with clear contracts:
Source Systems
│
▼
[ Bronze Layer ] — Raw, unmodified ingestion. Schema-on-read. Append-only.
│
▼
[ Silver Layer ] — Cleansed, deduped, typed, joined. Row-level quality enforced.
│
▼
[ Gold Layer ] — Business aggregates, KPIs, domain models. Query-optimized.
- Bronze: Ingest exactly as received. Preserve original records for auditability and reprocessing.
- Silver: Apply schema validation, deduplication, type casting, and referential joins. One record = one business entity.
- Gold: Aggregate, denormalize, and shape data for specific analytical or operational consumers.
// Bronze — raw ingestion, schema-on-read, no transformation
interface BronzeOrderEvent {
raw: Record<string, unknown>;
ingestedAt: Date;
sourceFile: string;
}
async function ingestToBronze(events: unknown[], sourceFile: string): Promise<BronzeOrderEvent[]> {
return events.map((raw) => ({
raw: raw as Record<string, unknown>,
ingestedAt: new Date(),
sourceFile,
}));
}
// Silver — cleansed, typed, deduped. One record = one Order.
interface SilverOrder {
orderId: string;
customerId: string;
total: number;
placedAt: Date;
}
function promoteToSilver(bronze: BronzeOrderEvent[]): SilverOrder[] {
const seen = new Set<string>();
return bronze
.filter((b) => typeof b.raw.order_id === "string" && typeof b.raw.total === "number")
.filter((b) => {
const id = b.raw.order_id as string;
if (seen.has(id)) return false;
seen.add(id);
return true;
})
.map((b) => ({
orderId: b.raw.order_id as string,
customerId: b.raw.customer_id as string,
total: b.raw.total as number,
placedAt: new Date(b.raw.placed_at as string),
}));
}
// Gold — business aggregate, query-optimized for a specific consumer
interface GoldDailyRevenue {
date: string;
orderCount: number;
totalRevenue: number;
}
function aggregateToGold(silver: SilverOrder[]): GoldDailyRevenue[] {
const byDate = new Map<string, GoldDailyRevenue>();
for (const order of silver) {
const date = order.placedAt.toISOString().slice(0, 10);
const bucket = byDate.get(date) ?? { date, orderCount: 0, totalRevenue: 0 };
bucket.orderCount += 1;
bucket.totalRevenue += order.total;
byDate.set(date, bucket);
}
return [...byDate.values()];
}
Each function is pure and stage-specific: promoteToSilver never reaches into a database, and aggregateToGold never re-parses raw fields. A failure in aggregateToGold never corrupts Bronze — the run can be retried from the Silver snapshot without re-ingesting the source.
from dataclasses import dataclass
from datetime import datetime, date
from typing import Any
# Bronze — raw ingestion, schema-on-read, no transformation
@dataclass(frozen=True)
class BronzeOrderEvent:
raw: dict[str, Any]
ingested_at: datetime
source_file: str
def ingest_to_bronze(events: list[dict[str, Any]], source_file: str) -> list[BronzeOrderEvent]:
return [
BronzeOrderEvent(raw=raw, ingested_at=datetime.utcnow(), source_file=source_file)
for raw in events
]
# Silver — cleansed, typed, deduped. One record = one Order.
@dataclass(frozen=True)
class SilverOrder:
order_id: str
customer_id: str
total: float
placed_at: datetime
def promote_to_silver(bronze: list[BronzeOrderEvent]) -> list[SilverOrder]:
seen: set[str] = set()
silver: list[SilverOrder] = []
for b in bronze:
order_id = b.raw.get("order_id")
total = b.raw.get("total")
if not isinstance(order_id, str) or not isinstance(total, (int, float)):
continue
if order_id in seen:
continue
seen.add(order_id)
silver.append(
SilverOrder(
order_id=order_id,
customer_id=b.raw["customer_id"],
total=total,
placed_at=datetime.fromisoformat(b.raw["placed_at"]),
)
)
return silver
# Gold — business aggregate, query-optimized for a specific consumer
@dataclass
class GoldDailyRevenue:
date: date
order_count: int
total_revenue: float
def aggregate_to_gold(silver: list[SilverOrder]) -> list[GoldDailyRevenue]:
by_date: dict[date, GoldDailyRevenue] = {}
for order in silver:
d = order.placed_at.date()
bucket = by_date.setdefault(d, GoldDailyRevenue(date=d, order_count=0, total_revenue=0.0))
bucket.order_count += 1
bucket.total_revenue += order.total
return list(by_date.values())
Live Playground
Experiment with the pattern below. The two tabs show a monolithic transform (before.ts) that cleans and aggregates raw events in one pass, and the Medallion-refactored equivalent (after.ts) with distinct Bronze, Silver, and Gold stages.
When to Use
- Data lakehouses built on Delta Lake, Apache Iceberg, or similar open table formats
- Organizations with multiple downstream consumers needing consistent, trusted data
- Pipelines requiring reprocessability (ability to rerun Silver/Gold from raw Bronze)
- Regulatory environments requiring full audit trails of raw data
Avoid when:
- Small-scale ETL with a single well-understood source and consumer — layers add overhead
- Real-time streaming with sub-second SLAs — the batch layering model may not fit
Trade-offs
| Benefit | Cost |
|---|---|
| Raw data preserved for full reprocessing | Storage costs increase (data stored at multiple stages) |
| Quality enforced at well-defined checkpoints | Latency added by multi-stage processing |
| Gold layer trusted by all consumers without re-derivation | Schema evolution must be managed across all layers |
| Failures isolated to a layer — Bronze intact even if Silver fails | Requires disciplined ownership of layer boundaries |
Related Patterns
- Schema-Driven Validation — Applied at the Bronze-to-Silver transition to enforce quality
- Pure Functions — Silver and Gold transformations benefit from pure, deterministic logic
- Batch vs Streaming — both execution models write into Medallion zones; batch jobs promote Bronze snapshots, streaming pipelines land micro-batch deltas
- CQRS — Gold layer tables are pre-projected read models, echoing CQRS’s query-side optimization philosophy