Lakehouse and Warehouse in One Fabric Architecture
A lakehouse and a warehouse are not mutually exclusive choices in Microsoft Fabric. Both can store analytical data in OneLake using Delta format, but they expose different development experiences and are suited to different patterns of ingestion, transformation, governance, and consumption.
For DP-600 candidates and Fabric Analytics Engineer Associate teams, the architecture question is not ‘Which product wins?’ It is ‘Which workload belongs where, and how do the parts share data without creating another set of silos?’
A strong Fabric design can use a lakehouse for engineering-oriented transformation and a warehouse for governed SQL consumption, while keeping the data platform coherent through OneLake, shortcuts, shared semantic models, and deliberate ownership boundaries.
Start with workload shape, not product preference
Teams often choose technology based on the skills they already have. Spark-oriented engineers naturally reach for a lakehouse; SQL-oriented analytics teams prefer a warehouse. Those preferences matter because an architecture that no one can operate is not sustainable, but they should not be the only criteria.
Consider data type, transformation style, concurrency, governance needs, and downstream consumers. Semi-structured files, notebooks, and large-scale engineering may fit a lakehouse. Relational dimensional models, T-SQL development, and governed reporting may fit a warehouse.
The distinction becomes clearer when the team understands both a data lake and the analytical responsibilities of a warehouse rather than treating either as a marketing label.
OneLake reduces the old duplication pressure
Historically, hybrid architectures often copied data from a data lake into a separate warehouse platform before BI could use it. That duplication created additional pipelines, storage, refresh windows, and reconciliation work.
Fabric changes the physical boundary because lakehouse and warehouse data can live in OneLake and use Delta. The experiences remain different, but the underlying data estate is more interoperable. Shortcuts and cross-item access can reduce the need to create another physical copy just to satisfy a different engine.
This does not mean every layer should point at the same unmanaged tables. Logical ownership, data contracts, and quality gates still matter. Shared storage is useful only when consumers understand which tables are authoritative.
A common pattern is engineering in the lakehouse and serving in the warehouse
One practical architecture places raw and refined engineering data in lakehouses and publishes consumption-ready relational structures into a warehouse. Data engineers can use Spark and notebooks where that flexibility is valuable, while BI developers and analysts work with a familiar SQL surface for curated models.
This resembles the layered responsibilities described in data warehouse architecture: upstream processes absorb source complexity, while the serving layer presents stable, queryable structures.
The handoff should be explicit. Define which layer owns deduplication, history, keys, business mappings, and slowly changing dimensions so the same transformation is not reimplemented independently in lakehouse and warehouse code.
The warehouse can also be the center of a medallion design
A warehouse is not limited to a final gold layer. For strongly relational data and teams that want to remain in SQL, Fabric can support warehouse-oriented bronze, silver, and gold patterns as well. The appropriate design depends on how much semi-structured processing and Spark-specific capability the workload requires.
This is an important correction to simplistic architecture diagrams. Medallion describes progressive refinement and trust; it does not require a particular product for every layer.
The broader difference between a data warehouse and an operational database also remains useful: analytical layers are shaped for historical analysis and aggregation rather than transactional application behavior.
Shared data still needs a clear contract
When a lakehouse table is consumed by a warehouse, semantic model, notebook, or KQL workflow, schema changes have a wider blast radius. A renamed field or changed data type may break several experiences at once.
Treat high-value tables as published interfaces. Document grain, keys, null expectations, historical behavior, refresh cadence, and ownership. Stage breaking changes so downstream teams can move deliberately.
The more interoperable the platform becomes, the more important interface discipline becomes. Ease of access should not be confused with permission to change shared structures casually.
Star schemas still matter above open storage
Open Delta storage does not eliminate dimensional modeling. Business users still need facts with clear grain, dimensions with usable attributes, consistent dates, and measures whose definitions are stable across reports.
A curated star-schema design can live in a warehouse or be represented through a semantic model over lakehouse data. The logical analytical structure remains valuable even when the physical storage layer is more flexible.
Do not let a convenient OneLake table become the public BI contract simply because it is easy to connect. Engineering tables are often optimized for transformation rather than human interpretation.
Security boundaries should follow ownership and consumption
Workspaces, item permissions, OneLake security, SQL permissions, and semantic-model security can all participate in the final access design. Mixing lakehouse and warehouse components without a security plan can create paths that are technically functional but difficult to reason about.
Decide where each audience should enter the system. Engineers may need broad write access to transformation layers, while report consumers should normally reach curated data through governed warehouse objects or semantic models.
Test the actual path a user takes. An administrator’s successful query proves little about how a Viewer, analyst, or external consumer will experience the architecture.
Performance is end-to-end, not item-by-item
A lakehouse can be healthy while the semantic model is slow, or a warehouse can answer SQL quickly while a transformation pipeline becomes the bottleneck. Architecture reviews should trace the entire analytical path from ingestion through transformation, storage, modeling, and user query.
Consider file size and Delta maintenance in the lakehouse, SQL query patterns in the warehouse, semantic-model cardinality and measures, and Fabric capacity contention across all of them.
Local optimization can simply move the bottleneck. The useful question is whether the overall workload meets its freshness, latency, concurrency, and cost targets.
Coexistence is strongest when each layer has a reason to exist
Using both products is not automatically more sophisticated. Every additional layer adds ownership, deployment, monitoring, and troubleshooting responsibility. A small structured workload may be simplest in a warehouse alone; a heavy Spark engineering workload may remain lakehouse-centric.
Good data architecture is about assigning responsibilities clearly, not collecting platform features.
Choose coexistence when the workloads genuinely differ and the interoperability of OneLake reduces the cost of serving them. If two layers perform the same transformation for the same consumers, the architecture has probably become more complicated than the problem.
Shortcuts deserve architectural discipline as well. They can expose data in another location without copying it, which is valuable for reducing duplicate storage and keeping one authoritative source. But a shortcut also creates a dependency on the source item’s availability, schema, security, and ownership. Document that dependency so a downstream team does not assume the data is locally controlled when it is actually governed elsewhere.
A mixed lakehouse-and-warehouse design needs a clear rule for where business transformations stop. If bronze and silver logic lives in Spark while gold dimensional logic lives in SQL, the boundary should be intentional. Otherwise teams may implement customer matching in a notebook, again in a warehouse procedure, and a third time in a semantic model. Duplicate logic is one of the fastest ways to create reconciliation problems.
Operational ownership should follow the layer. Data engineering may own ingestion and Delta maintenance, analytics engineering may own curated relational structures, and BI teams may own semantic measures. Those roles can overlap, but every production table and metric should have a named owner who knows which upstream and downstream contracts it participates in.
Disaster recovery and replay behavior also differ by layer. If raw data can be reingested from a durable source, the lakehouse may be rebuilt differently from a warehouse that contains curated historical transformations. Capture which artifacts are reproducible from code and source data, which require backups or retained history, and how long a full rebuild would take under real capacity limits.
Cost should be evaluated across the whole architecture. Avoiding a copy can save storage and pipeline work, while adding another serving layer can increase deployment and monitoring overhead. Conversely, a curated warehouse can reduce repeated ad hoc transformation in dozens of reports. The cheapest design is not always the one with the fewest items; it is the one that minimizes repeated work while preserving the service level the business needs.
Use architecture reviews to challenge every boundary periodically. As Fabric capabilities evolve, a layer that was once necessary may become redundant, or a workload may grow large enough to deserve separation. Coexistence should remain a deliberate response to workload diversity rather than a permanent monument to the first version of the platform.
The semantic layer can also bridge both stores without forcing every consumer to understand the physical split. A well-designed model can present consistent measures and dimensions while upstream teams choose the storage experience that fits each workload. That abstraction is useful only if lineage remains visible enough for support teams to trace a metric back to its authoritative table and transformation path.
When both items exist, naming matters. Calling every refined table ‘gold’ without identifying whether it is a lakehouse table, warehouse table, shortcut, or published semantic source makes operations harder. Use names and documentation that communicate the item’s role, owner, and expected consumers so the architecture remains understandable after the original designers move on.
Fabric makes lakehouse and warehouse coexistence practical because they can participate in one governed analytical estate instead of behaving like isolated platforms.
The design still requires judgment: put transformations, interfaces, and security where they are easiest to own and operate, then keep the boundaries explicit enough that teams know which data product they can trust.