The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Databricks medallion architecture organizes lakehouse data into progressively more reliable and consumer-ready layers: bronze for source-faithful raw data, silver for validated and refined records, and gold for business-ready products such as metrics and aggregates. Databricks calls it a recommended best practice, not a requirement, so use it when these boundaries improve your team’s quality controls, reuse, governance, or operations.
What each medallion layer should do
The layers represent increasing data quality and readiness for use, not simply three storage locations. A sound design gives each layer a clear purpose and a known set of consumers.
| Layer | Primary role | Typical contents |
|---|---|---|
| Bronze | Preserve what arrived from source systems so it can be traced and reprocessed. | Incrementally appended raw records, source metadata, and minimally transformed payloads. |
| Silver | Validate, standardize, and integrate records into a reusable trusted form. | Cleaned, typed, deduplicated, non-aggregated records; integrated datasets. |
| Gold | Serve defined business, analytics, machine-learning, or operational needs. | Data marts, dimensional models, metrics, summaries, and aggregates. |
Databricks describes the pattern and its layer roles in its medallion lakehouse architecture documentation.
Recommended Free Tools
Keep bronze faithful to its sources
Use bronze as a replayable foundation rather than a place to make data look clean. Incremental ingestion with minimal transformation helps preserve evidence of source issues, supports downstream rebuilds, and makes schema changes easier to handle without losing the original representation.
#1 Best Overall
- Retain useful provenance, such as source identifiers and arrival information, so records can be traced.
- Preserve most source fields. Databricks suggests flexible representations such as strings, VARIANT, or binary where unexpected schema changes could otherwise break ingestion.
- Apply only the parsing or technical handling needed to land data reliably; put business validation and extensive cleanup downstream.
- Manage access and retention deliberately: raw data can contain sensitive fields and should not become an uncontrolled holding area.
Databricks recommends Unity Catalog managed tables across the medallion layers and Unity Catalog volumes for landing zones and raw unstructured data. External tables can be appropriate when files must remain at specific storage paths. See the medallion guidance and lakehouse governance best practices.
Make silver the trusted record layer
Silver is where teams turn landed data into dependable, reusable records. Build it from bronze or from existing silver tables, and preserve at least one validated, non-aggregated representation of each record. This detail supports varied downstream analysis and machine-learning needs; add aggregates in silver only when a concrete consumer need warrants them, since aggregates typically belong in gold.
Handle the failure-prone data changes here
- Enforce or evolve schemas and cast values to appropriate types.
- Handle nulls, corrupt records, duplicates, and late or out-of-order events.
- Join sources where integration is part of the reusable dataset’s purpose.
- Apply explicit data-quality checks and document the rules and expected freshness.
For most append-only inputs, Databricks recommends reading from bronze rather than writing ingestion output directly to silver: schema changes or corrupt records can otherwise disrupt the curated layer. Its guidance favors streaming reads for most such inputs and batch reads for small datasets, such as small dimensions. Choose according to the source and workload, rather than treating either mode as universal.
Rank #2
Build gold for identifiable consumers
Gold should answer a defined need, not duplicate bronze as another raw store. Shape it around actual consumers and outcomes: a reporting mart, a set of trusted business metrics, a dashboard summary, or a model-ready product. Dimensional models and aggregates can reduce repeated downstream work when their definitions and ownership are clear.
- State who owns each product and who may use it.
- Apply appropriate protection, including anonymization, row-level access, or column masking.
- Publish metrics and summaries with documented meaning and freshness expectations.
- Choose centralized, domain-owned, or hybrid publication based on governance and organizational needs.
For hub-and-spoke organizations, Databricks recommends a shared hub for organization-wide data, domain-specific ingestion and curation, explicit publishing policies, and Unity Catalog catalogs that distinguish hub and domain assets. Its governance guidance discusses these organizational choices.
Choose pipeline patterns by transformation and operations
Lakeflow guidance distinguishes incremental row-level work from transformations that benefit from incremental recomputation. Streaming tables suit raw ingestion and row-level steps such as filtering, cleaning, and parsing. Materialized views suit enrichment joins or complex aggregations that benefit from incremental refresh, including precomputed gold summaries. Check current Lakeflow pipeline documentation before specifying syntax or relying on release-specific capabilities.
When practical, separate ingestion from downstream transformation pipelines. This lets teams schedule, monitor, and troubleshoot each independently; a transformation failure need not prevent new records from landing in bronze. The trade-off is more operational components to own, so keep the boundary where it meaningfully improves recovery or control.
Build quality and governance into every layer
Quality should improve at each hop, with checks at ingestion and stricter expectations in curated layers. Databricks identifies constraints, expectations, primary- and foreign-key metadata, and Lakehouse Monitoring among the available quality-related capabilities. Primary and foreign keys described as informational metadata should not be mistaken for enforced constraints.
- Use Unity Catalog for data discovery and lineage, and choose catalogs and schemas to match the governance model.
- Use managed tables by default when they fit storage and governance requirements; retain external tables where fixed paths are a real requirement.
- Define checks and ownership before delivery pressure makes them easy to omit.
- Avoid unmanaged table sprawl and make publishing policies explicit for shared products.
See Databricks lakehouse data-governance best practices for managed-table, Unity Catalog, and organizational guidance.
Rank #4
Decide whether the pattern fits your workload
Medallion architecture is a design pattern, not a fixed Databricks configuration. Databricks states: “Following the medallion architecture is a recommended best practice but not a requirement.” The right question is whether distinct raw, validated, and consumer-ready boundaries solve a real problem for your data and teams.
- Latency and ingestion: Decide whether the workload calls for batch, streaming, or change data capture.
- Ownership: Choose centralized, domain-based, or hybrid governance and publishing.
- Consumer needs: Keep detailed reusable records where analysts and models need them; publish aggregates or marts for defined use cases.
- Storage control: Prefer managed tables where suitable, but account for requirements to keep data at fixed paths.
- Operational boundaries: Separate ingestion and transformation when independent monitoring and recovery justify the added pipeline management.
Databricks’ medallion documentation makes clear that the pattern is recommended rather than mandatory; its governance guidance and Lakeflow documentation cover related implementation decisions.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

