Capabilities

Updated Date:

A Fabric lakehouse combines the scale and flexibility of a data lake with the querying power of a warehouse. Structured and unstructured data live in one place, on OneLake, and are analyzed with both Spark and SQL without moving data between systems.

Every lakehouse has two areas. Files is the flexible, schema-on-read side: land raw CSVs, JSON, logs, or images without modeling anything up front. Tables is the managed side, where data is stored as Delta tables. When Delta data lands in Tables, Fabric discovers it, reads its metadata, and registers it automatically, so it’s queryable without writing a single CREATE TABLE statement.

Delta Lake is the foundation. It gives lakehouse tables ACID transactions (Atomicity, Consistency, Isolation, and Durability), schema enforcement to keep bad data out, and time travel to query earlier versions of a table.

One copy of data serves everyone. Data engineers, data analysts, and data scientists work against the same storage. With OneLake shortcuts, that extends beyond your workspace: you can reference data in other lakehouses, external clouds, or even other organizations’ tenants through cross-tenant sharing, without making a copy.

Microsoft Fabric OneLake, Workspaces, Data Lake, Lakehouse

Solution

Enable your organization to analyze various data formats across multiple sources by leveraging Microsoft Fabric. Combine unstructured data from social media, log analytics, or partners with structured transactional data, all in one governed location.

Getting data in is flexible. Shortcuts provide live access to data where it already lives. Pipelines handle scheduled, orchestrated ingestion. Dataflows Gen2 give business-facing teams a low-code option. Notebooks and Spark job definitions cover engineering-grade transformation in Python, Scala, SQL, or R.

Getting insight out is just as flexible. From the lakehouse itself you can:

  • Query tables with T-SQL through the automatically created SQL analytics endpoint. It’s read-only and only shows Delta tables, which makes it a good fit for exploration, reporting, and ad-hoc analysis.
  • Run Spark SQL directly in the lakehouse’s query explorer, or open a notebook for deeper work.
  • Open the data in an Eventhouse endpoint to run KQL for real-time analytics.
  • Build a Power BI semantic model on top of the data for reporting. Note that Fabric no longer creates one for you automatically.

Workspaces remain the container for organization, collaboration, and a layer of security.

Lakehouse or Warehouse?

Both store Delta data on OneLake and share the same SQL engine, so this is less about capability and more about how your team works:

  • Choose a lakehouse for Spark-first teams, mixed structured and unstructured data, data science, and medallion (bronze / silver / gold) architectures.
  • Choose a warehouse for T-SQL-first teams, dimensional modeling, BI reporting, and workloads that need multi-table transactions.

In practice, many solutions use both: land and transform data in a lakehouse with Spark, then serve curated data from a warehouse for SQL-based reporting.

Further reading

Leave a comment

Trending