Databricks Data Engineer Associate · Free practice question 12 of 12
Spark lazy evaluation transformations vs actions
An engineer at Dunlin Health runs a notebook cell containing df = spark.read.table('clinic.silver.visits').filter('visit_year = 2025').select('patient_id'). The cell finishes almost instantly and no Spark job appears in the UI. Why?
- A.The result was served from a cached copy of the table
- B.The filter matched no rows, so Spark skipped reading the table
- C.The notebook is not attached to compute, so the code was not executed
- D.Transformations are evaluated lazily, so Spark only runs a job when an action such as count() or display() needs a result
Show answer and explanation
Correct answer: D. Transformations are evaluated lazily, so Spark only runs a job when an action such as count() or display() needs a result
Why: Spark transformations such as filter and select are lazy: they add steps to a query plan without processing data, and execution starts only when an action like count, collect, write, or display requires a result. Fast completion does not imply caching or an empty result, and code cannot run at all without attached compute.
More free Databricks Data Engineer Associate questions
- Delta shallow vs deep clone
- Auto Loader schema hints
- Reading notebook task parameters
- collect_set vs collect_list aggregation
- Triggered vs continuous pipeline mode
- Job maximum concurrent runs
- Job task timeout setting
- SHOW GRANTS on Unity Catalog objects
- Transferring object ownership
- Compute policies for cluster governance
- ALTER TABLE ADD COLUMNS on Delta