Databricks Data Engineer Professional · Free practice question 2 of 12
Partitioning guidance for mid-size tables
A new Delta table at Cragside Analytics will hold about 400 GB, and an engineer proposes partitioning it by customer_id, which has around 2 million distinct values. What does Databricks guidance recommend?
- A.Partition by customer_id, because more partitions always improve data skipping
- B.Partition by customer_id and run OPTIMIZE ZORDER BY on the same column
- C.Do not partition a table of this size, especially on a high-cardinality column; use liquid clustering instead
- D.Partition by a hash of customer_id into 2 million buckets
Show answer and explanation
Correct answer: C. Do not partition a table of this size, especially on a high-cardinality column; use liquid clustering instead
Why: Databricks recommends against partitioning most tables under about 1 TB, and a partition column should leave each partition with at least about 1 GB of data; a high-cardinality column produces millions of tiny partitions and small files. Liquid clustering gives data skipping on customer_id without those problems. Z-ordering cannot be applied to partition columns, and hash buckets create the same small-file problem.
More free Databricks Data Engineer Professional questions
- Databricks SQL alerts with notifications
- VACUUM retention duration safety check
- SQL warehouse sizing versus scaling
- BROWSE privilege for data discovery
- UNDROP TABLE for managed tables
- Workspace-catalog binding isolation
- Bundle generate and deployment bind
- Driver memory and toPandas
- Continuous job trigger for streaming
- Table update triggers for jobs
- Serverless environment dependencies