Detail chart · Data Pipeline

Medallion Architecture

Pattern ◆◆◆◆◆

Progressively refine raw ingested data through Bronze, Silver, and Gold layers, each adding quality and semantic richness.

Summary

Medallion Architecture organizes a data lakehouse into three progressive refinement layers — Bronze (raw ingestion), Silver (cleansed and conformed), and Gold (business-ready aggregates). Each layer has a defined contract for data quality, enabling downstream consumers to trust the data they use without re-engineering upstream pipelines.

Problem

Raw data from source systems is messy: schema drift, nulls, duplicates, inconsistent types, and mixed grain. Processing it repeatedly for different consumers leads to duplicated logic, inconsistent results, and fragile pipelines.

Solution

Divide the data lake into three named zones with clear contracts:

Source Systems
     │
     ▼
[ Bronze Layer ]   — Raw, unmodified ingestion. Schema-on-read. Append-only.
     │
     ▼
[ Silver Layer ]   — Cleansed, deduped, typed, joined. Row-level quality enforced.
     │
     ▼
[ Gold Layer ]     — Business aggregates, KPIs, domain models. Query-optimized.
// Bronze — raw ingestion, schema-on-read, no transformation
interface BronzeOrderEvent {
  raw: Record<string, unknown>;
  ingestedAt: Date;
  sourceFile: string;
}

async function ingestToBronze(events: unknown[], sourceFile: string): Promise<BronzeOrderEvent[]> {
  return events.map((raw) => ({
    raw: raw as Record<string, unknown>,
    ingestedAt: new Date(),
    sourceFile,
  }));
}

// Silver — cleansed, typed, deduped. One record = one Order.
interface SilverOrder {
  orderId: string;
  customerId: string;
  total: number;
  placedAt: Date;
}

function promoteToSilver(bronze: BronzeOrderEvent[]): SilverOrder[] {
  const seen = new Set<string>();
  return bronze
    .filter((b) => typeof b.raw.order_id === "string" && typeof b.raw.total === "number")
    .filter((b) => {
      const id = b.raw.order_id as string;
      if (seen.has(id)) return false;
      seen.add(id);
      return true;
    })
    .map((b) => ({
      orderId: b.raw.order_id as string,
      customerId: b.raw.customer_id as string,
      total: b.raw.total as number,
      placedAt: new Date(b.raw.placed_at as string),
    }));
}

// Gold — business aggregate, query-optimized for a specific consumer
interface GoldDailyRevenue {
  date: string;
  orderCount: number;
  totalRevenue: number;
}

function aggregateToGold(silver: SilverOrder[]): GoldDailyRevenue[] {
  const byDate = new Map<string, GoldDailyRevenue>();
  for (const order of silver) {
    const date = order.placedAt.toISOString().slice(0, 10);
    const bucket = byDate.get(date) ?? { date, orderCount: 0, totalRevenue: 0 };
    bucket.orderCount += 1;
    bucket.totalRevenue += order.total;
    byDate.set(date, bucket);
  }
  return [...byDate.values()];
}

Each function is pure and stage-specific: promoteToSilver never reaches into a database, and aggregateToGold never re-parses raw fields. A failure in aggregateToGold never corrupts Bronze — the run can be retried from the Silver snapshot without re-ingesting the source.

from dataclasses import dataclass
from datetime import datetime, date
from typing import Any

# Bronze — raw ingestion, schema-on-read, no transformation
@dataclass(frozen=True)
class BronzeOrderEvent:
    raw: dict[str, Any]
    ingested_at: datetime
    source_file: str


def ingest_to_bronze(events: list[dict[str, Any]], source_file: str) -> list[BronzeOrderEvent]:
    return [
        BronzeOrderEvent(raw=raw, ingested_at=datetime.utcnow(), source_file=source_file)
        for raw in events
    ]


# Silver — cleansed, typed, deduped. One record = one Order.
@dataclass(frozen=True)
class SilverOrder:
    order_id: str
    customer_id: str
    total: float
    placed_at: datetime


def promote_to_silver(bronze: list[BronzeOrderEvent]) -> list[SilverOrder]:
    seen: set[str] = set()
    silver: list[SilverOrder] = []
    for b in bronze:
        order_id = b.raw.get("order_id")
        total = b.raw.get("total")
        if not isinstance(order_id, str) or not isinstance(total, (int, float)):
            continue
        if order_id in seen:
            continue
        seen.add(order_id)
        silver.append(
            SilverOrder(
                order_id=order_id,
                customer_id=b.raw["customer_id"],
                total=total,
                placed_at=datetime.fromisoformat(b.raw["placed_at"]),
            )
        )
    return silver


# Gold — business aggregate, query-optimized for a specific consumer
@dataclass
class GoldDailyRevenue:
    date: date
    order_count: int
    total_revenue: float


def aggregate_to_gold(silver: list[SilverOrder]) -> list[GoldDailyRevenue]:
    by_date: dict[date, GoldDailyRevenue] = {}
    for order in silver:
        d = order.placed_at.date()
        bucket = by_date.setdefault(d, GoldDailyRevenue(date=d, order_count=0, total_revenue=0.0))
        bucket.order_count += 1
        bucket.total_revenue += order.total
    return list(by_date.values())

Live Playground

Experiment with the pattern below. The two tabs show a monolithic transform (before.ts) that cleans and aggregates raw events in one pass, and the Medallion-refactored equivalent (after.ts) with distinct Bronze, Silver, and Gold stages.

When to Use

Avoid when:

Trade-offs

BenefitCost
Raw data preserved for full reprocessingStorage costs increase (data stored at multiple stages)
Quality enforced at well-defined checkpointsLatency added by multi-stage processing
Gold layer trusted by all consumers without re-derivationSchema evolution must be managed across all layers
Failures isolated to a layer — Bronze intact even if Silver failsRequires disciplined ownership of layer boundaries