Databricks Data Engineer Associate Practice Exam
Practice questions for the Databricks Certified Data Engineer Associate certification: the Databricks Intelligence Platform and lakehouse architecture, compute types including serverless, Delta Lake fundamentals, notebooks and Databricks Connect, ingestion with Auto Loader, COPY INTO and Lakeflow Connect, Spark SQL and PySpark transformations, the medallion architecture, Lakeflow Spark Declarative Pipelines (formerly Delta Live Tables) with expectations and CDC, MERGE and higher-order functions, Lakeflow Jobs with task dependencies and repair runs, Declarative Automation Bundles (formerly Databricks Asset Bundles), and Unity Catalog governance including grants, managed and external tables, lineage, OpenSharing (formerly Delta Sharing), Lakehouse Federation and system tables. Every question includes a written explanation.
100 questions · 12 free preview
Studying more than one? All Databricks exams for $29 · every exam for $79
Free sample questions
- Sample · question 1 · Delta shallow vs deep clone
Before testing a risky schema change, an engineer at Tessaly Books wants a copy of the Delta table books.silver.orders that is created almost instantly without copying data files, while leaving the source table untouched. Which statement fits?
- A.CREATE TABLE books.dev.orders_test DEEP CLONE books.silver.orders
- B.CREATE VIEW books.dev.orders_test AS SELECT * FROM books.silver.orders
- C.CREATE TABLE books.dev.orders_test LIKE books.silver.orders
- D.CREATE TABLE books.dev.orders_test SHALLOW CLONE books.silver.orderscorrect
Why: A shallow clone copies only the table metadata and references the source's existing data files, so it is created quickly, and changes made to the clone are written as new files without affecting the source. A deep clone copies all data files, which takes longer and uses more storage. A view cannot be altered independently to test schema changes, and CREATE TABLE LIKE copies only the definition, with no data.
Open this question on its own page → - Sample · question 2 · Auto Loader schema hints
Auto Loader at Ferro Cycles infers the schema of incoming JSON files, but it keeps typing the column order_total as STRING when the team needs DECIMAL(10,2). All other columns should still be inferred. Which option addresses this?
- A.cloudFiles.schemaHints = 'order_total DECIMAL(10,2)'correct
- B.cloudFiles.schemaEvolutionMode = 'rescue'
- C.cloudFiles.useNotifications = 'true'
- D.cloudFiles.maxFilesPerTrigger = '1'
Why: Schema hints let you declare the type of selected columns while Auto Loader continues to infer the rest of the schema. The rescue evolution mode changes how unexpected data is handled rather than setting column types, notification mode changes how new files are discovered, and maxFilesPerTrigger only limits batch size.
Open this question on its own page → - Sample · question 3 · Reading notebook task parameters
A notebook at Glenrock Energy is run by a Lakeflow Job that passes a parameter named run_date to the notebook task. Inside the notebook, how should the code read that value?
- A.spark.conf.get('run_date')
- B.dbutils.widgets.get('run_date')correct
- C.os.environ['run_date']
- D.dbutils.fs.head('run_date')
Why: Parameters passed to a notebook task are exposed to the notebook as widget values, which the code reads with dbutils.widgets.get. They are not set as Spark configuration values or environment variables, and dbutils.fs.head reads the beginning of a file.
Open this question on its own page → - Sample · question 4 · collect_set vs collect_list aggregation
For each customer, an analyst at Wicklow Grocers needs an array of the distinct store IDs where the customer shopped, with no repeated values. Which aggregate function should be used with GROUP BY customer_id?
- A.collect_set(store_id)correct
- B.collect_list(store_id)
- C.array(store_id)
- D.explode(store_id)
Why: collect_set aggregates values into an array without duplicates. collect_list keeps every value, including repeats. array builds an array from its arguments within a single row, and explode turns array elements into rows instead of aggregating.
Open this question on its own page → - Sample · question 5 · Triggered vs continuous pipeline mode
A Lakeflow Spark Declarative Pipeline (formerly a Delta Live Tables pipeline) at Fenmore Transit only needs its tables refreshed once every morning before business hours, and the team wants compute to shut down between refreshes. Which pipeline execution mode fits?
- A.Continuous mode
- B.Triggered mode, run on a schedulecorrect
- C.Continuous mode combined with development mode
- D.Continuous mode with a larger cluster size
Why: In triggered mode, each update processes the available data and then stops, so it can be scheduled, for example from a job, and compute shuts down between runs. Continuous mode keeps the pipeline running to process new data with low latency, which adds cost when a daily refresh is enough. Development mode affects compute reuse and retries rather than how often updates run.
Open this question on its own page → - Sample · question 6 · Job maximum concurrent runs
A job at Keswick Media is scheduled every 10 minutes, and occasionally a run takes longer than 10 minutes. The team wants to prevent a new run from starting while the previous run is still executing. Which job setting controls this?
- A.Task retry count, set to 0
- B.Run if condition, set to All succeeded
- C.Maximum concurrent runs, set to 1correct
- D.Auto termination, set to 10 minutes
Why: The maximum concurrent runs setting limits how many runs of a job can be active at the same time; with a value of 1, a new run does not start while one is already in progress; it is skipped, or queued if queueing is enabled on the job. Retries control reruns of failed tasks, Run if conditions depend on upstream task results, and auto termination applies to idle clusters.
Open this question on its own page → - Sample · question 7 · Job task timeout setting
A task in a job at Lochby Analytics sometimes hangs indefinitely while waiting on an external system. The team wants the task stopped and treated as failed if it runs for longer than 45 minutes. What should they configure?
- A.A retry policy with a 45-minute interval between retries
- B.A cron schedule that starts the job every 45 minutes
- C.A timeout of 45 minutes on the taskcorrect
- D.An expectation with ON VIOLATION FAIL UPDATE
Why: A task timeout sets the maximum time a task may run; when it is exceeded, the task is terminated and its run is marked as timed out rather than succeeded. A retry interval controls the wait between attempts, a schedule controls how often the job starts, and expectations validate data in declarative pipelines rather than limiting task duration.
Open this question on its own page → - Sample · question 8 · SHOW GRANTS on Unity Catalog objects
An auditor at Parnell Bank wants to see which users and groups hold privileges on the Unity Catalog table risk.gold.exposures. Which statement lists them?
- A.SHOW GRANTS ON TABLE risk.gold.exposurescorrect
- B.SHOW TABLES IN risk.gold
- C.SELECT * FROM risk.gold.exposures WITH GRANTS
- D.DESCRIBE HISTORY risk.gold.exposures
Why: SHOW GRANTS ON a securable lists the principals and the privileges they hold on it. SHOW TABLES lists the tables in a schema, WITH GRANTS is not valid query syntax, and DESCRIBE HISTORY lists the table's write operations rather than its permissions.
Open this question on its own page → - Sample · question 9 · Transferring object ownership
The owner of the Unity Catalog table ops.gold.kpis at Tarrant Foods is leaving the company, and the team wants the group platform_admins to become the table's owner. Which statement does this?
- A.GRANT OWNERSHIP ON TABLE ops.gold.kpis TO `platform_admins`
- B.ALTER TABLE ops.gold.kpis OWNER TO `platform_admins`correct
- C.GRANT ALL PRIVILEGES ON TABLE ops.gold.kpis TO `platform_admins`
- D.ALTER TABLE ops.gold.kpis SET TBLPROPERTIES ('owner' = 'platform_admins')
Why: In Unity Catalog, ownership of an object is transferred with ALTER ... OWNER TO followed by the new principal, which can be a group. GRANT OWNERSHIP is not Databricks SQL syntax, ALL PRIVILEGES grants privileges without changing the owner, and ownership is not controlled by a table property.
Open this question on its own page → - Sample · question 10 · Compute policies for cluster governance
Workspace admins at Yelverton Labs want to limit which instance types and maximum sizes users can choose when they create classic clusters, and to require specific cost-center tags. Which feature is designed for this?
- A.Serverless SQL warehouses
- B.Unity Catalog row filters
- C.Compute (cluster) policiescorrect
- D.Databricks Git folders
Why: Compute policies define rules that constrain cluster configuration, such as allowed instance types, maximum sizes, and required tags, and they can be assigned to users and groups. Serverless SQL warehouses are a type of compute rather than a way to constrain it, row filters restrict table rows, and Git folders provide version control.
Open this question on its own page → - Sample · question 11 · ALTER TABLE ADD COLUMNS on Delta
The team at Carlow Marine needs to add a nullable column vessel_class of type STRING to the existing Delta table fleet.silver.vessels without rewriting existing data. Which statement does this?
- A.CREATE OR REPLACE TABLE fleet.silver.vessels (vessel_class STRING)
- B.INSERT INTO fleet.silver.vessels (vessel_class) VALUES (NULL)
- C.UPDATE fleet.silver.vessels SET vessel_class = NULL
- D.ALTER TABLE fleet.silver.vessels ADD COLUMNS (vessel_class STRING)correct
Why: ALTER TABLE ... ADD COLUMNS adds new columns to a Delta table as a metadata change, and existing rows return NULL for the new column. CREATE OR REPLACE with a single column would replace the table definition and its data. INSERT and UPDATE cannot reference a column that does not yet exist.
Open this question on its own page → - Sample · question 12 · Spark lazy evaluation transformations vs actions
An engineer at Dunlin Health runs a notebook cell containing df = spark.read.table('clinic.silver.visits').filter('visit_year = 2025').select('patient_id'). The cell finishes almost instantly and no Spark job appears in the UI. Why?
- A.The result was served from a cached copy of the table
- B.The filter matched no rows, so Spark skipped reading the table
- C.The notebook is not attached to compute, so the code was not executed
- D.Transformations are evaluated lazily, so Spark only runs a job when an action such as count() or display() needs a resultcorrect
Why: Spark transformations such as filter and select are lazy: they add steps to a query plan without processing data, and execution starts only when an action like count, collect, write, or display requires a result. Fast completion does not imply caching or an empty result, and code cannot run at all without attached compute.
Open this question on its own page →
Like the sample?
Other practice exams
- AnthropicClaude Certified Architect — Foundations100 questions · $19
- CompTIACompTIA Security+ (SY0-701)100 questions · $19
- ISC2CISSP100 questions · $19
- DatabricksDatabricks Data Engineer Professional100 questions · $19
- Google CloudGoogle Cloud Associate Cloud Engineer100 questions · $19
- Google CloudGoogle Cloud Professional Cloud Architect100 questions · $19
- Google CloudGoogle Cloud Professional Data Engineer100 questions · $19
- HashiCorpTerraform Associate (004)100 questions · $19
- Microsoft Power BI & FabricPower BI Data Analyst (PL-300)100 questions · $19
- Microsoft Power BI & FabricFabric Analytics Engineer (DP-600)100 questions · $19
- SnowflakeSnowPro Core (COF-C03)250 questions · $19
- SnowflakeSnowPro Advanced: Data Engineer100 questions · $19
- SnowflakeSnowPro Advanced: Architect100 questions · $19
- AWSAWS Cloud Practitioner (CLF-C02)100 questions · $19
- AWSAWS Solutions Architect Associate (SAA-C03)100 questions · $19
- AWSAWS AI Practitioner (AIF-C01)100 questions · $19
- Microsoft AzureAzure Fundamentals (AZ-900)100 questions · $19
- Microsoft AzureAzure Administrator (AZ-104)100 questions · $19
- Microsoft AzureAzure AI Fundamentals (AI-901)100 questions · $19
- Microsoft AzureAzure Solutions Architect Expert (AZ-305)100 questions · $19