Originally published June 2025.
Updated Date:
Fabric’s Native Execution Engine (NEE)
When I first wrote about Fabric’s Native Execution Engine (NEE), the pitch was simple. It swaps Spark’s row-based JVM execution for a columnar, vectorized C++ path, and your jobs get faster and cheaper. That’s still true. What has changed is how much the engine covers, how easy it is to turn on, and how clearly you can now see what it’s doing. Here’s the refreshed version.

What it is
The NEE is built on two open-source projects. Velox is a C++ execution library that came out of Meta. Apache Gluten (incubating) is a middle layer, started at Intel, that hands Spark’s work off to native engines like Velox.
When an operator is supported, Gluten moves it off the JVM onto Velox. There it runs as columnar, SIMD-accelerated code. SIMD means the processor applies one instruction to many values at once, which is where much of the speedup comes from.
There are three details worth highlighting.
1.) No code changes and no lock-in. You keep writing standard Spark APIs. The engine only changes what runs underneath.
2.) The optimizer still works. Adaptive query execution, cost-based rewrites, column pruning, and predicate pushdown all stay active when operators are offloaded. You lose none of Fabric’s query planning.
3.) Better Delta support. Delta snapshots can load in parallel. Tables organized with Z-ordering or Liquid Clustering get extra gains.
The engine runs on Runtime 1.3 (Spark 3.5 / Delta 3.2) and Runtime 2.0 (Spark 4.1 / Delta 4.1). It reads Parquet, Delta, and CSV. It works wherever your data sits in OneLake, including data reached through shortcuts.
Where it helps most
The engine pays off on compute-heavy work: complex transformations, joins, and aggregations over Parquet or Delta. It helps much less on simple jobs that spend most of their time on I/O.
The biggest factor is fallback. If a query uses something the engine doesn’t support, that part drops back to regular JVM Spark. The more of your plan that stays native, the more you gain. Microsoft keeps the list of supported operators and functions in the Apache Gluten documentation.
Since my original post, the engine has closed some important gaps. It now supports Python UDFs, Scala UDFs, and complex data types. CSV reads use a vectorized parser. On Runtime 2.0, ANSI SQL mode is supported, with error behavior that matches JVM Spark.

Turning it on
For a whole environment. In your Fabric environment, go to Spark compute → Acceleration. Check Enable native execution engine, then Save and Publish. Every notebook and Spark job definition that uses the environment inherits the setting. If you enabled the engine the old way, through Spark properties in the environment, switch to this toggle. The Spark property still works if you prefer it.
For one notebook or Spark job definition. Put this at the very start. In a notebook, that means the first cell.
json
%%configure{ "conf": { "spark.native.enabled": "true" }}
The engine works with live pools, so the setting takes effect right away. You don’t need to start a new session.
For a single query. If you know a query uses an unsupported operator, turn the engine off just for that cell:
sql
%%sqlSET spark.native.enabled=FALSE;
Or in PySpark:
python
spark.conf.set("spark.native.enabled", "false")
Cells run in order, so this setting carries over to later cells. Set it back to TRUE in the next cell.
Checking that it’s actually working
Turning it on is the easy part. Confirming it helps your workload is where the optimization happens.
Physical plan. Run df.explain(), or open the query in the Spark UI or Spark History Server. Look for nodes whose names end in Transformer, NativeFileScan, or VeloxColumnarToRowExec. Examples include ProjectExecTransformer and BroadcastHashJoinExecTransformer. Those nodes ran natively.
Gluten SQL / DataFrame tab. This tab in the Spark UI shows, for each query, how many nodes ran natively and how many fell back to the JVM.
Query execution graph. Green nodes ran on the native engine. Light-blue nodes ran on the JVM. It’s the quickest way to see where a plan drops out of native execution.
Fabric Spark Advisor. Advisor now shows fallback alerts directly in the notebook cell output. You find out about unsupported operators while you work, not after the job finishes.
Fallback is automatic, so your jobs won’t break. But a job that silently falls back most of the time is paying for acceleration it isn’t getting. Make checking these views part of how you tune jobs.
Limitations to know before you flip the switch
These apply on every runtime:
- Structured streaming isn’t supported and falls back to the JVM.
- JSON and XML sources aren’t accelerated.
- Date comparisons need matching types on both sides. Cast explicitly, as in
CAST(order_date AS DATE) = '2024-05-20', rather than comparing a timestamp column to a string. - Managed private endpoints: if your workspace reaches storage through them, you need separate endpoints for the Blob and DFS services, even when both point to the same storage account.
Runtime 1.3 only. ANSI mode isn’t supported. There are also several correctness differences from JVM Spark:
- decimal-to-float casting
- unrecognized timezone settings
round()behavior- duplicate-key checks in
map() - element order and intermediate types in
collect_list()andcollect_set()
All of these are fixed in Runtime 2.0. If your pipelines depend on any of them, that’s a strong reason to test the move to 2.0.
Execution vs. optimization, revisited
My original post argued for balancing speed of delivery against the long-term savings of optimization. The Native Execution Engine is about as close as Fabric gets to optimization without trading away delivery speed. It’s a single toggle, it needs no code rewrite, and fallback protects you if something isn’t supported.
The effort goes into measurement. Enable the engine, run your heaviest jobs, and check the execution graph for fallbacks. Fix the easy ones, such as mismatched date types. Compare run time and capacity consumption before and after. Some workloads will improve a lot and others barely at all, and you’ll only know which is which by measuring.
Reference: Native execution engine for Fabric Data Engineering – Microsoft Learn






Leave a comment