CertKeen

Databricks Data Engineer Associate · Free practice question 12 of 12

Spark lazy evaluation transformations vs actions

An engineer at Dunlin Health runs a notebook cell containing df = spark.read.table('clinic.silver.visits').filter('visit_year = 2025').select('patient_id'). The cell finishes almost instantly and no Spark job appears in the UI. Why?

  1. A.The result was served from a cached copy of the table
  2. B.The filter matched no rows, so Spark skipped reading the table
  3. C.The notebook is not attached to compute, so the code was not executed
  4. D.Transformations are evaluated lazily, so Spark only runs a job when an action such as count() or display() needs a result
Show answer and explanation

Correct answer: D. Transformations are evaluated lazily, so Spark only runs a job when an action such as count() or display() needs a result

Why: Spark transformations such as filter and select are lazy: they add steps to a query plan without processing data, and execution starts only when an action like count, collect, write, or display requires a result. Fast completion does not imply caching or an empty result, and code cannot run at all without attached compute.

More free Databricks Data Engineer Associate questions