SnowPro Advanced: Data Engineer · Free practice question 6 of 10
Maintaining externally managed Iceberg tables
An externally-managed Apache Iceberg table backed by S3 is written by a Spark job using merge-on-read with position delete files. Over several weeks, Snowflake queries against the table have become significantly slower and the number of small delete files in S3 keeps growing. What is the correct fix?
- A.Enable Snowflake Automatic Clustering on the Iceberg table so Snowflake will compact the data and delete files.
- B.Set ENABLE_ICEBERG_MERGE_ON_READ = FALSE in Snowflake to have Snowflake rewrite the files into copy-on-write form.
- C.Run regular data-file compaction and delete-file cleanup in the external engine (Spark), then let Snowflake pick up the results via its normal metadata refresh.
- D.Increase the frequency of ALTER ICEBERG TABLE ... REFRESH in Snowflake so old delete files are pruned faster.
Show answer and explanation
Correct answer: C. Run regular data-file compaction and delete-file cleanup in the external engine (Spark), then let Snowflake pick up the results via its normal metadata refresh.
Why: For externally-managed Iceberg tables Snowflake is a read-only consumer — it cannot rewrite data files or clean up delete files owned by the external catalog. Compaction and maintenance must be run in the writing engine (Spark, Trino, etc.). Snowflake's job is to refresh its metadata pointer once the external maintenance completes.
More free SnowPro Advanced: Data Engineer questions
- Snowpark lazy evaluation
- Dynamic tables with TARGET_LAG
- Snowpipe Streaming for sub-10-second latency
- Fan-in task DAGs with AFTER
- Time Travel vs Fail-safe recovery window
- External table partition metadata refresh
- Tag propagation across data movement
- Query Acceleration max scale factor
- Alerting on Cortex AI credit usage